Training method and translation method of multi-language neural machine translation model, equipment and storage medium

By iterative training and sub-network model extraction of multilingual neural machine translation models, combined with gradient-based pruning and joint iterative training, the performance degradation problem of multilingual neural machine translation models on high resource language pairs is solved, achieving more efficient and stable translation performance.

CN120146068APending Publication Date: 2025-06-13MILITARY SCI INFORMATION RES CENT ACAD OF MILITARY SCI OF THE CHINESE PEOPLES LIBERATION ARMY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510219817.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Multilingual neural machine translation models have performance degradation problems in practical applications, especially when dealing with high resource language pairs, the translation quality is not as good as bilingual models specially trained for specific language pairs.

Method used

By iteratively training the pre-created initial network model using multilingual sample data, an intermediate network model was obtained; then iteratively training the intermediate network model using the sample training data of each language pair to obtain the corresponding subnetwork model for each language pair; fuse the subnetwork model with the intermediate network model, and obtain the multilingual neural machine translation model through joint iterative training.

Benefits of technology

The gradient-based pruning criterion removes weights that have little impact on the translation performance of high-resource languages, reduces interference between different language pairs, improves the effectiveness of parameters, maximizes parameter efficiency, and reduces the complexity of the model and reduces the risk of overfitting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146068A_ABST
    Figure CN120146068A_ABST
Patent Text Reader

Abstract

The invention provides a training method and a translation method of a multi-language neural machine translation model, equipment and a storage medium. The method comprises the following steps: training an initial network model by using multi-language sample data to obtain an intermediate network model; respectively training the intermediate network model by using the sample data of each language pair to obtain a sub-network model; in each training step, the training gradient of each weight in the intermediate network model is calculated, the training gradient of each weight is accumulated after each preset number of training steps is completed and multiplied by the weight, an importance score corresponding to each weight is obtained and compared with a pruning threshold value, and a pruning mask is generated; pruning the intermediate network model; obtaining a sub-network model after all training steps are completed; fusing the sub-network model and the intermediate network model, and then carrying out joint training by using multi-language sample data to obtain a multi-language neural machine translation model; according to the method, the problem of performance degradation of a multi-language neural machine translation technology in a practical application process can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of language translation, and particularly to a training method, a translation method, a device and a storage medium for a multi - language neural machine translation model. Background Art

[0002] Multi - language neural machine translation (MNMT) enables mutual translation between multiple languages through a unified model, reducing system complexity and improving the translation quality of low - resource languages. It shares parameters to facilitate cross - language knowledge transfer, but faces problems of performance degradation and negative interference, especially performing worse than bilingual models on high - resource language pairs. To solve this problem, researchers have proposed methods such as language - specific components and model pruning, but still need to balance translation quality and computational efficiency. Therefore, how to improve the efficiency and stability of multi - language neural machine translation models while ensuring translation quality remains an urgent problem to be solved.

[0003] Existing multi - language neural machine translation technologies achieve translation between multiple languages through a unified model, aiming to reduce system complexity and utilize the commonalities of multi - language data to improve the translation quality of low - resource languages. Its core idea is parameter sharing, that is, during training, different language pairs share the parameters of the same model, thereby promoting cross - language knowledge transfer. This method can learn the similarities and differences between different languages, and combine general language structures and semantic representations during translation, while capturing the unique features of each language.

[0004] However, due to the existence of negative interference, that is, the differences in grammar, vocabulary and semantics between different language pairs cause the shared parameters to be unable to find the optimal settings, there are significant performance degradation problems in practical applications. When dealing with high - resource language pairs, the translation quality is often inferior to that of bilingual models trained specifically for specific language pairs. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide a training method, a translation method, a device and a storage medium for a multi - language neural machine translation model to eliminate or improve one or more defects existing in the prior art. It can solve the performance degradation problem that occurs in the process of practical application of multi - language neural machine translation technology.

[0006] One aspect of the present invention provides a training method for a multi - language neural machine translation model, the method comprising the following steps:

[0007] Iteratively train a pre - created initial network model using multi - language sample data to obtain an intermediate network model; the multi - language sample data includes sample training data corresponding to at least three languages, where every two languages form a language pair, and there are sample training data for mutual translation between every two languages;

[0008] Using the sample training data corresponding to each language pair, iteratively train the intermediate network model separately to obtain the sub-network model corresponding to each language pair; wherein, during the iterative training process of the intermediate network model with the sample training data of each language pair, in each training step, calculate the training gradient of each weight in the intermediate network model, and after every preset number of training steps, accumulate the training gradients of each weight and multiply them by the weight magnitude respectively to obtain the importance score corresponding to each weight, compare the importance score with the pruning threshold to generate a pruning mask, and perform a pruning operation on the intermediate network model through the pruning mask; after completing all training steps, obtain the sub-network model;

[0009] Fuse the sub-network model with the intermediate network model to obtain a fusion model;

[0010] Use multi-language sample data to perform joint iterative training on the fusion model to obtain a multi-language neural machine translation model.

[0011] In some embodiments of the present invention, during the iterative training process of the intermediate network model with the sample training data of each language pair, it further includes:

[0012] Determine the target pruning ratio based on the number of steps in the current training step;

[0013] Sort the importance scores corresponding to each weight according to their magnitudes;

[0014] Select the importance score at the corresponding position from the sorted importance score list as the pruning threshold according to the target pruning ratio.

[0015] In some embodiments of the present invention, determining the target pruning ratio based on the number of steps in the current training step includes:

[0016] In the case where the number of steps is less than the first preset number of steps, determine the first preset pruning ratio as the target pruning ratio; or, in the case where the number of steps is greater than or equal to the second preset number of steps, determine the second preset pruning ratio as the target pruning ratio; the first preset pruning ratio is less than the second preset pruning ratio.

[0017] In some embodiments of the present invention, in the case where the number of steps is greater than or equal to the first preset number of steps and less than the second preset number of steps, it further includes:

[0018] Based on the number of steps, the first preset number of steps, the second preset number of steps, and the second preset pruning ratio, obtain a dynamic pruning ratio according to a preset linear growth algorithm as the target pruning ratio; the dynamic pruning ratio is greater than the first preset pruning ratio and less than the second preset pruning ratio.

[0019] In some embodiments of the present invention, a multi - language neural machine translation model is obtained by jointly iteratively training a fusion model using multi - language sample data, including:

[0020] In each process of joint iterative training, randomly select the sample training data corresponding to a language pair from the multi - language sample data as the target sample training data;

[0021] Train the fusion model through the target sample training data and a preset loss function. During the backpropagation process, only update the model parameters of the sub - network model corresponding to the target sample training data;

[0022] Repeat the above steps, sequentially process the sample training data corresponding to all language pairs, and perform joint iterative training on the fusion model to obtain a multi - language neural machine translation model.

[0023] In some embodiments of the present invention, an initial network model created in advance is iteratively trained using multi - language sample data to obtain an intermediate network model, including:

[0024] Randomly select the sample training data corresponding to at least one language pair from the language sample data to form the current batch of sample training data;

[0025] Iteratively train the initial network model through the current batch of sample training data and a preset loss function to obtain an intermediate network model.

[0026] In some embodiments of the present invention, the initial network model adopts a Transformer structure.

[0027] One aspect of the present invention provides a multi - language translation method, which includes:

[0028] Obtain a multi - language neural machine translation model; the multi - language neural machine translation model is trained by using the training method of the above - mentioned multi - language neural machine translation model;

[0029] Input the text to be translated, the language identifier corresponding to the text to be translated, and the target language identifier into the multi - language neural machine translation model together. The multi - language neural machine translation model translates the text to be translated into the target language based on the target language identifier.

[0030] One aspect of the present invention provides an electronic device, including a processor, a memory, and a computer program / instructions stored on the memory. The processor is used to execute the computer program / instructions. When the computer program / instructions are executed, the electronic device implements the steps of the above - mentioned training method of the multi - language neural machine translation model or the multi - language translation method.

[0031] One aspect of the present invention provides a computer-readable storage medium, on which computer programs / instructions are stored. When the computer programs / instructions are executed by a processor, the steps of the above-mentioned training method for a multi-lingual neural machine translation model or the multi-lingual translation method are implemented.

[0032] The beneficial effects of the present invention include:

[0033] The training method, translation method and device of the multi-lingual neural machine translation model of the present invention can solve the problem of performance degradation that occurs in the actual application process of multi-lingual neural machine translation technology; through the pruning criterion based on gradients, the importance score of each weight can be accurately calculated, and the pruning mask can be determined according to these scores, so as to effectively remove the weights that have little impact on the translation performance of high-resource languages, ensuring that the remaining parameters are the most critical parts for the model performance, improving the effectiveness of the parameters while reducing the interference between different language pairs; in this way, while reducing the number of parameters of the model, the translation performance can be maintained or even improved, achieving the maximization of parameter efficiency.

[0034] In addition, the pruning ratio gradually increases from low to high, enabling the model to gradually adapt to the reduction of parameters during the training process, avoiding performance degradation caused by sudden large-scale pruning; at the same time, without affecting the performance, the number of parameters of the model can be reduced, improving the training efficiency and inference speed. When translating complex sentences, the model can utilize the structure after progressive pruning to more efficiently process long-distance dependencies and accurately translate various language components in the sentence, thereby improving the accuracy and fluency of translation.

[0035] In addition, through progressive pruning and sub-network model extraction, the complexity of the model can also be reduced, reducing the risk of overfitting of the model to the training data. During the extraction process of the sub-network model, specific sub-network models are extracted for each language pair. These sub-networks focus on learning the characteristics of specific language pairs, reducing the interference between different language pairs, enabling the model to focus more on learning useful language patterns rather than overlearning the noise and details in the training data, and reducing the risk of overfitting.

[0036] In addition, the joint iterative training adopts a language-aware data batching strategy. During the training process, the sample training data for each batch comes from the same language pair. This strategy enables the model to focus on learning the characteristics of specific language pairs during each training, avoiding interference between different language pairs; at the same time, during the backpropagation process, only the sub-network parameters related to the current language pair are updated. This parameter update strategy can enable the model to optimize each sub-network more specifically, improving the training efficiency.

[0037] In addition, through targeted parameter updates, the application of optimization algorithms, and the use of regularization techniques, joint iterative training can effectively improve the performance and stability of the multilingual neural machine translation model, enabling the model to better adapt to the translation requirements of different language pairs and providing strong guarantees for high-quality multilingual translation.

[0038] Additional advantages, objectives, and features of the present invention will be partially described below and will become partially apparent to those of ordinary skill in the art after studying the following text, or may be learned through the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the specification and the accompanying drawings.

[0039] Those skilled in the art will understand that the objectives and advantages that can be achieved by the present invention are not limited to the above specific descriptions, and the above and other objectives that the present invention can achieve will be more clearly understood according to the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not limit the present invention. In the drawings:

[0041] Figure 1 is a flowchart of a method for training a multilingual neural machine translation model provided by an embodiment of the present invention.

[0042] Figure 2 is a flowchart of a multilingual translation method provided by another embodiment of the present invention.

[0043] Figure 3 is a block diagram of a training device for a multilingual neural machine translation model provided by another embodiment of the present invention.

[0044] Figure 4 is a block diagram of a multilingual translation device provided by another embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the embodiments and the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.

[0046] Here, it should also be noted that in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution of the present invention are shown in the drawings, while other details less related to the present invention are omitted.

[0047] It should be emphasized that when the term "comprising / including" is used herein, it refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.

[0048] Here, it should also be noted that if not otherwise specified, the term "connection" in this text can not only refer to a direct connection, but also represent an indirect connection with an intermediate.

[0049] In the following, embodiments of the present invention will be described with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar components, or the same or similar steps.

[0050] The following provides a detailed introduction to the training method of the multi - language neural machine translation model provided by this application.

[0051] In some embodiments of the present invention, the execution subject of the training method of the multi - language neural machine translation model provided by this application is an electronic device, which can be a terminal such as a computer, a mobile phone, a tablet computer, a camera, etc., or can also be a server. The implementation manner of the electronic device is not limited in this embodiment.

[0052] This embodiment provides a training method of a multi - language neural machine translation model, as Figure 1 shown, the method at least includes step S101 to step S104:

[0053] Step S101, iteratively train a pre - created initial network model using multi - language sample data to obtain an intermediate network model.

[0054] Among them, the multi - language sample data includes sample training data corresponding to at least three languages. Each two languages form a language pair, and there are sample training data for mutual translation corresponding to each pair of languages. For example, sample training data corresponding to English - German language pair, English - French language pair or English - Japanese language pair, etc.

[0055] In order to enable the trained multi - language neural machine translation model to learn the general features and semantic representations between different languages, in some embodiments of the present invention, a large - scale multi - language translation corpus is used as the sample training data to perform preliminary iterative training on the initial network model to obtain an intermediate network model.

[0056] In some embodiments of the present invention, the multilingual translation corpus includes, but is not limited to, the International Workshop on Spoken Language Translation (IWSLT) dataset or the Workshop on Statistical Machine Translation (WMT) dataset. These corpora contain parallel texts in multiple language pairs and serve as the basis for the model to learn language knowledge and translation patterns. For language pairs such as English-French, English-German, and English-Japanese, a large number of bilingual sentence pairs are collected from publicly available multilingual corpora to ensure that the corpus covers a rich variety of language expressions, semantic scenarios, and domain knowledge. The collected corpus is preprocessed, including operations such as text cleaning, tokenization, and tagging, to convert the original text into a format that the model can process.

[0057] Among them, the IWSLT dataset mainly comes from speeches and conversations at international conferences, covering multiple language pairs such as English-German, English-French, and English-Japanese. The characteristic of this dataset is that the language style is relatively colloquial, the sentence structure is relatively simple, and it is closer to daily life and actual communication scenarios. In scenarios such as conference records and interview conversations, the IWSLT dataset can provide rich samples of colloquial expressions and natural language communication, helping the model learn the flexible use of daily language and colloquial translation skills. This enables the model trained based on this dataset to have certain advantages when dealing with translation tasks in scenarios such as daily communication and cross-border tourism, and can generate more natural and fluent translations, meeting the needs of actual communication.

[0058] The WMT dataset mainly comes from fields such as news, government documents, and academic papers, and the language pairs are also rich and diverse. Compared with the IWSLT dataset, the language of the WMT dataset is more formal and standardized, the sentence structure is complex, and it contains a large number of professional terms and complex grammar structures. In news reports, it will involve professional vocabulary and complex sentence patterns in various fields such as politics, economy, and culture; government documents and academic papers require even more accuracy and rigor in language. This enables the model trained based on the WMT dataset to perform excellently when dealing with tasks such as translating formal documents and professional fields, and can accurately understand and translate complex sentence structures and professional terms, meeting the translation needs of fields such as business and academia.

[0059] By selecting these two datasets with different characteristics and application scenarios, the multilingual neural machine translation technology can be evaluated from multiple dimensions. Specifically, through the IWSLT dataset, the translation performance of the model in colloquial and daily communication scenarios can be examined, and its understanding and translation ability of natural language expressions can be evaluated; through experiments on the WMT dataset, the translation ability of the model in formal and professional fields can be tested, including the ability to handle complex sentence structures and professional terms. Such a dataset selection strategy can more comprehensively and objectively verify the effectiveness and superiority of the technology in different scenarios, providing strong support for the further optimization and application of the technology.

[0060] In order to construct an efficient multi - language neural machine translation model, in some embodiments of the present invention, the initial network model adopts the Transformer architecture. Through the powerful self - attention mechanism of the Transformer architecture, it can effectively capture long - distance dependencies in the text and demonstrate excellent performance in the field of machine translation.

[0061] When using the IWSLT dataset as the sample training data, since its sentence structure is relatively simple and the language complexity is relatively low, the Transformer - small model can be used as the initial network model. At this time, the parameter scale and computational complexity of the initial network model are relatively small. For example: the number of attention heads of the initial network model is 4, the number of layers is 6, the dimension of the word vector is 512, and the dimension of the feed - forward neural network is 1024.

[0062] The parameter scale of the Transformer - small model can better adapt to the characteristics of the IWSLT dataset. It can not only ensure that the model has a certain expressive ability, but also effectively control the computational cost and training time, avoid the over - fitting problem caused by too many parameters in the model, and can also converge quickly to improve the training efficiency.

[0063] When using the WMT dataset as the sample training data, since the WMT dataset contains a large amount of formal texts and content in professional fields, with complex sentence structures and rich and diverse language expressions, the Transformer - base model can be used as the initial network model. At this time, the parameter scale and computational complexity of the initial network model are relatively large.

[0064] For example: the number of attention heads of the initial network model is 8, the number of layers is 6, the dimension of the word vector is 512, and the dimension of the feed - forward neural network is 2048.

[0065] By increasing the number of attention heads and the dimension of the feed - forward neural network, the Transformer - base model can better capture long - distance dependencies and complex semantic information in the text, thereby improving the translation ability for complex sentences and professional terms, giving full play to the advantages of the model, adapting to the training requirements of large - scale and complex datasets, and enhancing the performance of the model in professional field translation tasks.

[0066] In addition, in order to enable the initial network model to perform targeted translation according to different language pairs, in some embodiments of the present invention, a language embedding layer is introduced into the Transformer architecture adopted by the initial network model to identify different languages through different identifiers. In this way, when inputting the text to be translated, the identifiers of the source language and the target language are input into the model together, so as to achieve targeted translation from the source language to the target language.

[0067] In actual implementation, in order to enhance the adaptability of the model to different languages, some learnable parameters are added to the initial network model, such as language-specific bias terms, etc. These parameters can be adjusted according to the characteristics of different languages, thereby improving the translation ability of the final multilingual neural machine translation.

[0068] In some embodiments of the present invention, the initial network model is iteratively trained using multilingual sample data, and a preset loss function is used to optimize the parameters of the initial network model. By continuously adjusting the parameters, the performance of the initial network model in the multilingual translation task is gradually improved, and an intermediate network model is obtained through training.

[0069] Specifically, iteratively training the pre-created initial network model using multilingual sample data to obtain an intermediate network model includes: randomly selecting sample training data corresponding to at least one language pair from the language sample data to form the current batch of sample training data; and iteratively training the initial network model through the current batch of sample training data and the preset loss function to obtain the intermediate network model. Among them, the preset loss function includes but is not limited to the Negative Log-Likelihood Loss (NLL Loss) or the Cross-Entropy Loss.

[0070] At the same time, during the initial iterative training process of the initial network model, appropriate hyperparameters can also be set, such as the learning rate, batch size, number of training epochs, etc., to ensure that the initial network model can converge stably. Among them, the learning rate adopts a dynamic adjustment strategy, uses the Adaptive Moment Estimation (Adam) optimizer, and decays the learning rate according to the number of training steps to balance the convergence speed and stability of the initial network model.

[0071] Specifically, in terms of training parameters, the Adam optimizer is adopted, and its parameters are set as β1 = 0.9 and β2 = 0.98. The Adam optimizer adaptively adjusts the learning rate, and can converge quickly during the training process while maintaining good stability. The learning rate adopts a dynamic adjustment strategy, the initial learning rate is set to a certain value, and it decays according to a preset rule during the training process to balance the convergence speed and training effect of the model. In the initial stage of training, a larger learning rate can enable the model to quickly adjust parameters and accelerate the convergence speed; as the training progresses, gradually reducing the learning rate can enable the model to converge more stably and avoid overfitting.

[0072] During the training process, the maximum number of updates and the checkpoint saving strategy are set. For different datasets and models, the maximum number of updates will be adjusted according to the actual situation. When saving checkpoints, the parameters of the model and the training status are saved regularly, so that in case of problems during the training process, it can be restored to the previous state, and it is also convenient to evaluate and analyze the model. At the same time, an early stopping mechanism is also set. When the performance of the model on the validation set does not improve significantly in a certain number of iterations, the training is stopped to avoid overfitting and save computing resources.

[0073] By reasonably configuring the structural parameters and training parameters of the model, the advantages of the Transformer model can be fully utilized, enabling it to better adapt to datasets of different scales and characteristics, providing strong support for the experimental verification of the multi - language neural machine translation technology based on gradient progressive pruning.

[0074] Step S102: Use the sample training data corresponding to each language pair to iteratively train the intermediate network model respectively, and obtain the sub - network model corresponding to each language pair.

[0075] Among them, during the iterative training of the intermediate network model with the sample training data of each language pair, in each training step, calculate the training gradient of each weight in the intermediate network model. After every preset number of training steps, accumulate the training gradients of each weight and multiply them by the weight magnitude respectively to obtain the importance score corresponding to each weight. Compare the importance score with the pruning threshold to generate a pruning mask, and perform pruning operations on the intermediate network model through the pruning mask; the sub - network model is obtained after all training steps are completed.

[0076] Among them, the preset number refers to the number of training steps set in advance, including but not limited to 10 to 50 steps, 50 to 200 steps, 200 to 1000 steps, or 4000 steps, etc. In actual implementation, the number of training steps can be adjusted according to the scale of the sample training data corresponding to each language pair or the complexity of the training task. This embodiment does not limit the value of the preset number.

[0077] In some embodiments of the present invention, the extraction of the sub - network model is a key step to achieve language - pair - specific translation optimization. This step extracts the sub - network model for each language pair from the multi - language machine translation model through a pruning mask. These sub - networks can better capture the unique language features of specific language pairs, thereby reducing the interference between different language pairs and improving the translation performance.

[0078] For each language pair, filter out the retained weights from the intermediate network model according to the corresponding pruning mask, and these retained weights constitute the sub - network model of this language pair.

[0079] For example, in a translation model that contains multiple language pairs such as English-French, English-German, etc., for the English-French language pair, according to its pruning mask, the weights in the model that are related to English-French translation and have an importance score higher than the pruning threshold are extracted to form an English-French sub-network model. The weights in this sub-network model are screened to remove the parts that have less impact on English-French translation, making the sub-network model more focused on learning the language conversion patterns and semantic correspondences between English and French.

[0080] The structure and parameters of the extracted sub-network model are closely related to the original intermediate network model, but have language-pair specific characteristics. The sub-network model inherits some structures of the original model, such as the layer structure and attention mechanism in the Transformer architecture, but the parameters are refined and optimized. Through the pruning operation, the sub-network model removes some parameters that are not important for the translation of a specific language pair, thereby reducing the complexity and computational amount of the model. At the same time, since the sub-network model is extracted based on the language-pair specific pruning mask, it can better adapt to the grammar structure, vocabulary characteristics, and semantic expressions of a specific language pair, improving the pertinence and adaptability to the translation task of that language pair. By extracting the language-pair specific sub-network model, the model structure can be optimized, the translation efficiency and quality can be improved, providing a more targeted solution for multi-language translation.

[0081] In practical applications, the advantages of sub-network model extraction are reflected in multiple aspects. Since the sub-network model focuses on the translation of a specific language pair, it can more effectively utilize limited computational resources and improve translation efficiency.

[0082] For example, when processing a large number of English-French translation tasks, the English-French sub-network model can complete the translation quickly and accurately without having to consider other language pairs like the original multi-language model, thus saving computational time and resources. The sub-network model can better capture the language features of a specific language pair, reduce the interference between different language pairs, and improve translation quality. During the translation process, the sub-network model can accurately handle the selection of vocabulary, the conversion of grammar, and the expression of semantics according to the language characteristics of English-French, generating a more natural, fluent, and accurate translation.

[0083] In the process of iteratively training the intermediate network model using the sample training data corresponding to each language pair respectively, the training gradient of each weight in the intermediate network model is calculated in each training step. After completing the preset number of training steps, the training gradients of each weight are accumulated and multiplied by the weight size respectively to obtain the importance score corresponding to each weight. Specifically, the importance score corresponding to each weight can be expressed by the following formula:

[0084]

[0085] In the formula, W ij represents the weight; represents the importance score of the weight W ij after T times of gradient updates with a preset number; α represents the learning rate, represents the gradient of the loss function L with respect to the weight W ij , which is the partial derivative of the loss function L with respect to the weight W ij ; t represents the number of gradient update times.

[0086] Among them, W ij refers to the learnable parameters in the model, including but not limited to the weights in the self-attention mechanism (such as query weight, key weight, value weight, output projection weight, etc.), the weights in the feed-forward neural network (such as the weights of the first fully connected layer, the weights of the second fully connected layer, etc.), the weights in the embedding layer (such as word embedding weight, position encoding weight, etc.), and other weights (such as normalization parameters, output layer weights, etc.).

[0087] The loss function L includes but not limited to the negative log-likelihood loss function or the cross-entropy loss function. The gradient reflects the degree of influence of the change in the weight W ij on the loss function L. If the absolute value of the gradient is large, it means that a small change in the weight will cause a large change in the loss function L. At this time, this weight has a great influence on the performance of the model; conversely, if the absolute value of the gradient is small, it means that the change in the weight has a small influence on the loss function, that is, this weight has a relatively small influence on the model performance. At the same time, the magnitude of the weight W ij itself also reflects its importance to the model to a certain extent. Based on this, in some embodiments of the present invention, the gradient is multiplied by the weight W ij to comprehensively consider these two factors and more comprehensively measure the importance of each weight in the model.

[0088] By accumulating the training gradients, the importance changes of the weights during the entire training process can be captured. In the initial stage of training, the model parameters may not have converged yet, and the importance scores of the weights may fluctuate greatly; as the training progresses, the model gradually converges, and the importance scores of the weights also tend to be stable. The importance scores calculated in this way can more accurately reflect the actual role of the weights in the model and provide a reliable basis for subsequent pruning operations.

[0089] After calculating the importance scores in a preset number of training steps, the importance scores are compared with the pruning threshold to generate a pruning mask, and the intermediate network model is pruned through the pruning mask to trim the model parameters. Among them, generating the pruning mask is the key step to achieve model parameter pruning.

[0090] In some embodiments of the present invention, the pruning mask is a binary matrix. The element values in the pruning mask indicate whether the corresponding weights are retained or pruned, thus directly affecting the structure and performance of the model.

[0091] During the process of comparing the importance score of each weight with the pruning threshold, if the importance score of a certain weight is lower than the pruning threshold, the element at the corresponding position in the pruning mask is set to 0 to indicate that the weight will be pruned; conversely, if the importance score of a certain weight is greater than or equal to the pruning threshold, the element at the corresponding position in the pruning mask is set to 1, indicating that the weight will be retained.

[0092] In actual implementation, the specific application method of the pruning mask is as follows: during the training process of the model, when the model performs forward propagation, the pruning mask is multiplied by each weight matrix of the model, so that the weights corresponding to the positions where the value in the mask is 0 are ignored in the calculation, thereby achieving the pruning of these weights.

[0093] In a neural network containing multiple layers, each layer has a corresponding weight matrix and pruning mask. When performing forward propagation calculation, the weight matrix of each layer is multiplied by the corresponding pruning mask, and then subsequent calculations are performed, such as matrix multiplication, activation function operations, etc. In this way, during the calculation process of the model, the pruned weights no longer participate in the calculation, thereby reducing the calculation amount and the number of parameters of the model.

[0094] In some embodiments of the present invention, the generation and application of the pruning mask is a dynamic process. During the model training process, the pruning mask is dynamically adjusted according to different training stages and pruning strategies.

[0095] In the initial stage of training, in order to avoid excessive pruning having too great an impact on the model performance, a relatively low pruning threshold is set, and only those weights that have a minimal impact on the model performance are pruned; as the training progresses, the model gradually converges and the performance gradually stabilizes, and the pruning threshold can be gradually increased to further reduce the number of parameters of the model. By this way of dynamically adjusting the pruning mask, it is possible to gradually optimize the structure of the model while ensuring the model performance, and improve the training efficiency and inference speed of the model.

[0096] In some embodiments of the present invention, importance scores are compared with pruning thresholds to generate pruning masks. Among them, the pruning threshold is a preset importance score threshold, and the pruning thresholds corresponding to different stages may be the same or different. For example, the training stage is divided into three stages, including the initial pruning stage, the pruning transition stage, and the pruning completion stage; in the initial pruning stage, the first preset pruning threshold is used; in the pruning transition stage, the second preset pruning threshold is used; in the pruning completion stage, the third preset pruning threshold is used; among them, the first preset pruning threshold is less than the second preset pruning threshold, and the second preset pruning threshold is less than the third preset pruning threshold.

[0097] In some other embodiments of the present invention, the pruning threshold is determined based on the pruning ratio. Among them, the pruning ratio is used to determine an importance score as the pruning threshold among the importance scores corresponding to each calculated weight.

[0098] In each stage, first, the importance scores of each weight are sorted in order of magnitude to obtain a list of importance scores; then, the pruning ratio is used to determine an importance score at the corresponding position in the list of importance scores as the pruning threshold.

[0099] Specifically, during the iterative training of the intermediate network model on the sample training data of each language pair, it further includes: determining the target pruning ratio based on the current iteration number; sorting the importance scores corresponding to each weight in order of magnitude; and selecting the importance score at the corresponding position from the sorted list of importance scores as the pruning threshold according to the target pruning ratio.

[0100] The adjustment of the pruning ratio is a key link, which can directly affect the performance and parameter efficiency of the model. On the premise of ensuring the translation quality of the model, by setting a reasonable pruning ratio, the number of model parameters is gradually reduced, thereby improving the training efficiency and inference speed of the model.

[0101] To achieve this goal, in some embodiments of the present invention, the target pruning ratio is determined by setting different training step intervals to ensure that the model can make a smooth transition during training, while taking into account performance optimization and accuracy maintenance.

[0102] Specifically, the iterative training of the model is divided into three stages, including the initial pruning stage, the pruning transition stage, and the pruning completion stage. Among them, the initial pruning stage refers to the stage with less than the first preset number of steps, the pruning transition stage refers to the stage between the first preset number of steps and the second preset number of steps, and the pruning completion stage refers to the stage with more than the second preset number of steps.

[0103] In the initial stage of pruning, a first preset pruning ratio is used to determine the pruning threshold; in the pruning transition period, a dynamic pruning ratio is calculated using the first preset pruning ratio, the second preset pruning ratio, and the number of steps to determine the pruning threshold; in the pruning completion period, the second preset pruning ratio is used to determine the pruning threshold. Among them, the dynamic pruning ratio is greater than the first preset pruning ratio and less than the second preset pruning ratio.

[0104] Specifically, determining the target pruning ratio based on the number of steps of the current training step includes: in the case where the number of steps is less than the first preset number of steps, determining the first preset pruning ratio as the target pruning ratio; or, in the case where the number of steps is greater than or equal to the second preset number of steps, determining the second preset pruning ratio as the target pruning ratio; the first preset pruning ratio is less than the second preset pruning ratio.

[0105] In the case where the number of steps is greater than or equal to the first preset number of steps and less than the second preset number of steps, it further includes: based on the number of steps, the first preset number of steps, the second preset number of steps, and the second preset pruning ratio, obtaining a dynamic pruning ratio according to a preset linear growth algorithm as the target pruning ratio.

[0106] For example: taking the pruning ratio starting from 0 and gradually increasing until reaching the preset target pruning ratio as an example. At this time, 0 is the first preset pruning ratio. The adjustment of the pruning ratio can be expressed by the following formula:

[0107]

[0108] In the formula, R t represents the pruning ratio at training step t; T 1 and T 2 represent the preset training step thresholds, T 1 is the first preset number of steps, T 1 +T 2 is the second preset number of steps; R max represents the second preset pruning ratio.

[0109] In the initial stage of pruning, that is, when t < T 1 , since at this time in the initial stage of model training, the model is still in the process of learning language features and establishing basic semantic representations, pruning at this time may damage the learning ability of the model, resulting in the model being unable to converge or a significant decline in performance. At this stage, the model mainly focuses on learning the common features and semantic relationships between different language pairs through a large amount of training data, laying a foundation for subsequent pruning operations. Therefore, the pruning ratio R t is set to 0.

[0110] In the pruning transition period, that is, when T 1 ≤t < T1 +T 2 When, the pruning ratio R t begins to gradually increase. Specifically, the pruning ratio linearly increases in the manner of . This linear growth pattern can make the pruning process smoother and avoid the impact on the model performance caused by a large amount of pruning at one time.

[0111] During the pruning transition period, since the model has already learned and understood certain language features to some extent, gradually increasing the pruning ratio can, without affecting the model performance, gradually reduce the number of model parameters and improve the training efficiency of the model. As the pruning ratio increases, the model will gradually adapt to the reduction of parameters and maintain good translation performance by adjusting its own structure and parameters.

[0112] In the pruning completion period, that is, when t ≥ T 1 +T 2 When, the pruning ratio R t reaches the second preset pruning ratio R max . At this time, the model has completed most of the pruning operations, and the subsequent training is mainly to fine-tune on the pruned model structure to further optimize the model performance. At this stage, the number of model parameters has been reduced to the target value, and through continuous training, the model can make better use of the remaining parameters to improve the accuracy and fluency of translation.

[0113] Through this progressive pruning ratio adjustment strategy, the model can gradually adapt to the reduction of parameters during the training process and avoid the performance degradation caused by a sudden large amount of pruning. At the same time, this strategy can also dynamically adjust the pruning ratio according to the progress of training, so that the model can achieve a better balance between training efficiency and translation performance.

[0114] In traditional multi-language neural machine translation models, since different language pairs share the same model parameters, the model faces difficulties in capturing the unique features of high-resource languages and is easily interfered by other language pairs, resulting in performance degradation.

[0115] Based on this, in some embodiments of the present invention, through the gradient-based pruning criterion, the importance score of each weight can be accurately calculated, and the pruning mask is determined according to these scores, so as to effectively remove the weights that have less impact on the translation performance of high-resource languages and reduce the interference between different language pairs. For example, when dealing with the high-resource language pair of English-German, the model can focus on learning the specific language patterns and semantic relationships between English and German, avoiding the interference of parameters of other language pairs, and making the translation result more accurately reflect the characteristics of the two languages.

[0116] Meanwhile, the application of the progressive pruning mechanism further optimizes the performance of the model. The pruning ratio gradually increases from low to high, enabling the model to gradually adapt to the reduction of parameters during training and avoiding performance degradation caused by sudden large-scale pruning. In the initial stage of pruning, the model mainly focuses on learning the basic features and semantic representations of the language. Maintaining a low pruning ratio at this time can ensure the stability and learning ability of the model. As training progresses, during the pruning transition period, as the model's understanding of the language gradually deepens, the pruning ratio is gradually increased. In this way, without affecting performance, the number of model parameters can be reduced, improving the training efficiency and inference speed. When translating complex sentences, the model can utilize the structure after progressive pruning to more efficiently handle long-distance dependencies, accurately translate various language components in the sentence, thereby improving the accuracy and fluency of translation.

[0117] In addition, traditional multilingual neural machine translation models often contain a large number of learnable parameters, some of which contribute less to the model performance but increase the computational complexity and training cost of the model. At the same time, traditional pruning methods include magnitude-based pruning methods, second-derivative-based pruning methods, or random pruning methods, etc. Among them, the magnitude-based pruning method determines the pruning mask based on the absolute value of the weights, assuming that small weights contribute less to the model and can be pruned. Although this method is computationally simple and easy to implement, it ignores the function and role of the weights and may lead to the accidental deletion of key weights, especially affecting performance in complex models. The second-derivative-based pruning method evaluates the impact of the weights on the loss function by calculating the Hessian matrix of the weights, which can more accurately identify important weights, but has a high computational cost and poor numerical stability, and is not suitable for large-scale models; the random pruning method randomly selects weights for pruning without complex calculations, but may accidentally delete key weights, resulting in performance degradation and poor effects in multilingual neural machine translation.

[0118] Based on this, in some embodiments of the present invention, the importance score of each weight in the model is calculated based on gradient information, and the pruning mask is determined according to the importance score and a preset pruning threshold, thereby accurately removing the weights that have less impact on the model performance. In the multilingual model of the Transformer architecture, through the pruning operation, some unimportant connection weights in the attention mechanism and the feed-forward neural network are removed, making the model structure more compact. This not only reduces the number of model parameters and the consumption of computing resources but also significantly improves the training and inference speed of the model.

[0119] In addition, through the progressive pruning mechanism, the number of parameters of the model is gradually reduced, enabling the model to achieve better performance with limited parameters. At the initial stage of training, the number of parameters of the model is relatively large. As training progresses, the pruning ratio gradually increases, and the model continuously adjusts its own structure to adapt to the reduction of parameters. During this process, the model can make more effective use of the remaining parameters, improving the utilization efficiency of the parameters. At the same time, the gradient-based pruning criterion ensures that the retained parameters are the most critical parts for the model performance, further enhancing the effectiveness of the parameters. In this way, while reducing the number of parameters, the model can maintain or even improve the translation performance, achieving the maximization of parameter efficiency.

[0120] In addition, through progressive pruning and sub-network model extraction, the complexity of the model can also be reduced, lowering the risk of overfitting of the model to the training data. During the extraction of the sub-network model, specific sub-network models are extracted for each language pair. These sub-networks focus on learning the features of specific language pairs, reducing the interference between different language pairs, enabling the model to focus more on learning useful language patterns rather than overlearning the noise and details in the training data, and reducing the risk of overfitting.

[0121] Step S103: Integrate the sub-network model with the intermediate network model to obtain a fused model.

[0122] Since the intermediate network model has learned the general features and semantic representations between different languages through multi-language sample data during the pre-training stage and has strong cross-language transfer ability; while the sub-network model is the result of specialized optimization for each language pair (such as English-German, English-French, etc.), which can better capture the unique grammar structures, lexical characteristics, and translation patterns of specific language pairs. Based on this, integrating the sub-network model with the intermediate network model combines the general feature representation ability of the intermediate network model and the language-specific characteristics of each language pair's sub-network model, enabling the construction of an efficient model that can handle both multi-language general tasks and meet the requirements of specific language pairs.

[0123] In some embodiments of the present invention, the sub-network model is integrated with the intermediate network model by parameter sharing and fine-tuning. A part of the parameters in the sub-network model are shared with the intermediate network model to retain its general feature extraction ability, and at the same time, the language-specific part of the parameters is fine-tuned to enhance the support for specific language pairs.

[0124] In some other embodiments of the present invention, the sub-network model is integrated with the intermediate network model by weighted combination; different weights are assigned to the intermediate network model and the sub-network model based on the importance of the language pair or the requirements of the translation task to form a new fused model. For example, for frequently used language pairs, a higher weight can be given to their sub-networks.

[0125] In some other embodiments of the present invention, the sub-network model and the intermediate network model are fused through modular integration; the intermediate network model is used as the basic framework, and the sub-network models of each language pair are embedded therein in a modular manner. The corresponding sub-network module is selected according to the input language pair for translation, and at the same time, the intermediate network model provides global support.

[0126] Step S104, jointly iteratively train the fusion model using multilingual sample data to obtain a multilingual neural machine translation model.

[0127] After obtaining the fusion model, in order to further improve the performance of the multilingual neural machine translation model, it is also necessary to jointly iteratively train the fusion model, co-train the extracted language-pair specific sub-network model and the intermediate network model, and through reasonable training strategies and optimization algorithms, enable the model to better learn and utilize the information of each language pair, thereby improving the translation quality and stability.

[0128] In some embodiments of the present invention, the joint iterative training adopts a language-aware data batching strategy. During the training process, the sample training data of each batch comes from the same language pair. This strategy enables the model to focus on learning the features of a specific language pair during each training, avoiding interference between different language pairs.

[0129] For example: in a training set containing multiple language pairs such as English-French, English-German, etc., the sentence pairs of English-French are grouped into one batch, and the sentence pairs of English-German are grouped into another batch, and they are trained separately; in this way, when the model processes the English-French batch, it can concentrate on learning the language conversion rules between English and French without being affected by the English-German language pair, thereby more effectively capturing the unique features of each language pair.

[0130] At the same time, during the backpropagation process, only the sub-network parameters related to the current language pair are updated. This parameter update strategy enables the model to optimize each sub-network more specifically, improving the training efficiency.

[0131] For example: when processing the English-French training batch, only the parameters in the English-French sub-network are updated, while the sub-network parameters of other language pairs remain unchanged; in this way, the model can quickly adjust the parameters of the sub-network according to the English-French training data, making it better adapt to the English-French translation task, while avoiding unnecessary interference to other sub-networks, reducing the complexity and computational amount of training.

[0132] Specifically, the fusion model is jointly iteratively trained using multilingual sample data to obtain a multilingual neural machine translation model, including: in each joint iterative training process, randomly select the sample training data corresponding to a language pair from the multilingual sample data as the target sample training data; train the fusion model through the target sample training data and a preset loss function, and in the backpropagation process, only update the model parameters of the sub-network model corresponding to the target sample training data; repeat the above steps, sequentially process the sample training data corresponding to all language pairs, and perform joint iterative training on the fusion model to obtain a multilingual neural machine translation model.

[0133] In addition, to further optimize the training process, some optimization algorithms are also adopted, such as Adaptive Gradient Algorithm (Adagrad), Adaptive Delta (Adadelta), Adaptive Moment Estimation (Adam), etc. These algorithms can dynamically adjust the learning rate according to the training situation of the model, enabling the model to converge faster during training and improving the training efficiency. The Adam algorithm can adaptively adjust the learning rate of each learnable parameter, and set different learning rates for different learnable parameters according to the gradient and historical gradient information of the learnable parameters, thereby accelerating the convergence speed of the model. At the same time, in the joint iterative training, using the Adam algorithm can enable the model to flexibly adjust the learning rate according to the training situation of each sub-network model when processing different language pairs, improving the training effect of the model.

[0134] In addition, during the training process, regularization techniques such as L1 regularization (Least Absolute Shrinkage and Selection Operator, LASSO) and L2 regularization (Ridge Regression) can also be adopted to prevent the model from overfitting. The regularization technique adds a regularization term to the loss function to constrain the parameters of the model, enabling the model to pay more attention to the overall characteristics of the data during the learning process rather than overfitting the noise and details in the training data. In the joint iterative training, the parameters of each sub-network model are constrained through L2 regularization to prevent the sub-network model from overfitting when learning the features of a specific language pair, thereby improving the generalization ability of the model and the stability of translation.

[0135] In some embodiments of the present invention, the extraction and joint iterative training of the sub-network models make important contributions to the performance improvement of the model. By extracting specific sub-network models for each language pair, the multi-lingual neural machine translation model can better capture the unique language features of that language pair, further reducing the interference between different language pairs. At the same time, in the joint iterative training stage, a language-aware data batching strategy is adopted, enabling the multi-lingual neural machine translation model to focus on learning the specific features of each language pair, further reducing the interference between different language pairs. During the backpropagation process, only the sub-network parameters related to the current language pair are updated, avoiding the interference of the parameters of other language pairs, enabling the model to more effectively learn the translation knowledge of each language pair, improving the adaptability and translation ability of the model to different language pairs, reducing the occurrence of overfitting, and avoiding the mutual interference of multiple language pairs during the training process, thus generating more accurate and natural translations. When dealing with multiple language pairs, the sub-network models of different language pairs can work together to perform accurate translations according to the characteristics of their respective languages, thereby improving the overall performance of multi-lingual translation.

[0136] At the same time, through targeted parameter updates, the application of optimization algorithms, and the use of regularization techniques, the joint iterative training can effectively improve the performance and stability of the multi-lingual neural machine translation model, enabling the model to better adapt to the translation requirements of different language pairs, providing a strong guarantee for high-quality multi-lingual translation. Through these optimization measures, the model can converge more stably during the training process and exhibit better translation performance on the test data. Therefore, the reliability and practicality of the model are improved.

[0137] In addition, during the entire training process, the model also needs to be evaluated and verified regularly. The performance of the model is evaluated using the reserved validation set, and translation quality metrics such as the Bilingual Evaluation Understudy (BLEU), Recall-Oriented Understudy for Gisting Evaluation (ROUGE), etc. are monitored. According to the evaluation results, the hyperparameters and training strategies of the model are adjusted to ensure that the model can be continuously optimized during the training process and ultimately achieve good translation performance.

[0138] In summary, the training method of the multi - language neural machine translation model provided by this application iteratively trains a pre - created initial network model using multi - language sample data to obtain an intermediate network model; the multi - language sample data includes sample training data corresponding to at least two language pairs; uses the sample training data corresponding to each language pair to iteratively train the intermediate network model respectively to obtain a sub - network model corresponding to each language pair; wherein, during the iterative training process of the intermediate network model with the sample training data of each language pair, in each training step, calculate the training gradient of each weight in the intermediate network model, and after every preset number of training steps, accumulate the training gradients of each weight respectively and multiply them by the weight magnitude to obtain the importance score corresponding to each weight, compare the importance score with the pruning threshold to generate a pruning mask, and perform a pruning operation on the intermediate network model through the pruning mask; obtain the sub - network model after completing all training steps; fuse the sub - network model with the intermediate network model to obtain a fused model; use the multi - language sample data to perform joint iterative training on the fused model to obtain a multi - language neural machine translation model; it can solve the problem of performance degradation that occurs in the actual application process of multi - language neural machine translation technology; through the gradient - based pruning criterion, it can accurately calculate the importance score of each weight, determine the pruning mask according to these scores, thereby effectively removing the weights that have little impact on the translation performance of high - resource languages, ensuring that the retained parameters are the most critical parts of the model performance, improving the effectiveness of the parameters while reducing the interference between different language pairs; in this way, while reducing the number of parameters of the model, it can maintain or even improve the translation performance, achieving the maximization of parameter efficiency.

[0139] In addition, the pruning ratio gradually increases from low to high, enabling the model to gradually adapt to the reduction of parameters during the training process, avoiding performance degradation caused by sudden large - scale pruning; at the same time, without affecting performance, it can reduce the number of parameters of the model, improving the training efficiency and inference speed. When translating complex sentences, the model can utilize the structure after progressive pruning to more efficiently handle long - distance dependencies and accurately translate various language components in the sentence, thereby improving the accuracy and fluency of translation.

[0140] In addition, through progressive pruning and sub - network model extraction, the complexity of the model can also be reduced, reducing the risk of overfitting of the model to the training data. During the extraction process of the sub - network model, specific sub - network models are extracted for each language pair. These sub - networks focus on learning the characteristics of specific language pairs, reducing the interference between different language pairs, enabling the model to focus more on learning useful language patterns rather than over - learning the noise and details in the training data, reducing the risk of overfitting.

[0141] In addition, joint iterative training adopts a language-aware data batching strategy. During the training process, the sample training data for each batch comes from the same language pair. This strategy enables the model to focus on learning the features of a specific language pair during each training session, avoiding interference between different language pairs. At the same time, during the backpropagation process, only the sub-network parameters related to the current language pair are updated. This parameter update strategy can enable the model to optimize each sub-network more specifically and improve the training efficiency.

[0142] In addition, through targeted parameter updates, the application of optimization algorithms, and the use of regularization techniques, joint iterative training can effectively improve the performance and stability of the multi-language neural machine translation model, enabling the model to better adapt to the translation requirements of different language pairs and providing a strong guarantee for high-quality multi-language translation.

[0143] Based on the above embodiments, Figure 2 This is a multi-language translation method provided by an embodiment of the present application. The method at least includes steps S201 to S202:

[0144] Step S201, obtain a multi-language neural machine translation model; the multi-language neural machine translation model is trained using the training method of the multi-language neural machine translation model in the above embodiments. For relevant details, refer to the above embodiments.

[0145] Step S202, input the text to be translated, the language identifier corresponding to the text to be translated, and the target language identifier into the multi-language neural machine translation model. The multi-language neural machine translation model translates the text to be translated into the target language based on the target language identifier.

[0146] The multi-language neural machine translation model training method and the multi-language translation method provided by the present application have broad application prospects in the field of multi-language neural machine translation, covering multiple important application scenarios. At the same time, the application scope and limitations in these fields are clearly defined.

[0147] For example, in the field of international business, the multi - language neural machine translation model training method and the multi - language translation method provided by this application can be applied to scenarios such as the translation of business documents in multinational companies and real - time translation in international conferences. In the translation of business documents, it can accurately translate various types of documents such as contracts, business reports, and market analyses, ensuring the accurate conveyance of business information and promoting the smooth progress of cross - border trade. In real - time translation in international conferences, it can enable barrier - free communication among participants speaking different languages, improving the efficiency and effectiveness of the conference. However, when applying in this field, attention should be paid to the accurate translation of professional terms and the understanding of business cultural backgrounds. There are differences in business cultures in different countries and regions, such as business etiquette and negotiation styles, and these factors may affect the accuracy and appropriateness of translation. Therefore, when applying the multi - language neural machine translation model training method and the multi - language translation method provided by this application, translation needs to be carried out in combination with relevant business knowledge and cultural backgrounds to ensure that the translation results meet the requirements of business scenarios.

[0148] On social media platforms, the multi - language neural machine translation model training method and the multi - language translation method provided by this application can be used for communication translation among multi - language users, achieving real - time language conversion and promoting global user interaction. In scenarios such as when users post updates, comments, and private messages, it can quickly and accurately translate the user's language into other languages, breaking down language barriers and enhancing the user experience. But when applying in this field, the characteristics of social media language need to be considered, such as conciseness, colloquialism, and buzzwords. The language on social media is often more casual and changeable, containing a large number of internet buzzwords, abbreviations, and emojis, which all increase the difficulty of translation. Therefore, when applying the multi - language neural machine translation model training method and the multi - language translation method provided by this application, the translation model needs to be continuously updated and optimized to adapt to the changes in social media language and improve the accuracy and naturalness of translation.

[0149] In the news media industry, the multi - language neural machine translation model training method and the multi - language translation method provided by this application can be used for the multi - language translation of news manuscripts, helping the media quickly spread news content to all over the world. In the translation of news reports, feature articles, comments, etc., it can timely and accurately translate news content into multiple languages, meeting the needs of readers in different regions and expanding the dissemination scope of news. But when applying in this field, attention should be paid to the timeliness and accuracy of news language. News reports often need to be translated and published within a short time, and at the same time, it is necessary to ensure that the translated content is accurate and free of ambiguity. In addition, news content covers multiple fields such as politics, economy, and culture, and certain knowledge of different fields is required to ensure the professionalism of translation.

[0150] In the field of academic research, the multi - language neural machine translation model training method and multi - language translation method provided in this application can be used for the translation of academic literature, helping researchers access global academic resources. In the translation of academic papers, research reports, conference proceedings, etc., it can accurately translate professional terms and complex academic sentences, promoting academic exchanges and cooperation. However, academic literature often has high professionalism and rigor, involving a large number of professional terms and complex theoretical knowledge. When applying the patented technology of this application, it is necessary to combine the knowledge of the professional field and the term library for translation to ensure that the translation results can accurately convey the academic meaning of the original text. At the same time, due to the continuous development of academic research, new terms and concepts are emerging continuously, and it is necessary to update the translation model in a timely manner to adapt to the changes in the academic field.

[0151] In summary, the multi - language neural machine translation model training method and multi - language translation method provided in this application have important application values in multiple application fields of multi - language neural machine translation. However, when applying in different fields, corresponding measures need to be taken according to the characteristics and requirements of each field to ensure the quality and effect of translation. At the same time, the limitations and challenges that may be faced when applying in each field are also clarified.

[0152] Figure 3 is a block diagram of a training device for a multi - language neural machine translation model provided by an embodiment of this application. The device at least includes the following modules: an initial training module 310, a sub - model extraction module 320, a model fusion module 330, and a joint training module 340.

[0153] The initial training module 310 is used to iteratively train a pre - created initial network model using multi - language sample data to obtain an intermediate network model; the multi - language sample data includes sample training data corresponding to at least three languages, where every two languages form a language pair, and there are sample training data for mutual translation corresponding to each pair of languages.

[0154] The sub - model extraction module 320 is used to iteratively train the intermediate network model using the sample training data corresponding to each language pair to obtain a sub - network model corresponding to each language pair; among them, during the iterative training process of the intermediate network model using the sample training data of each language pair, in each training step, calculate the training gradient of each weight in the intermediate network model. After every preset number of training steps, accumulate the training gradients of each weight respectively and multiply them by the weight magnitude to obtain the importance score corresponding to each weight. Compare the importance score with the pruning threshold to generate a pruning mask, and perform a pruning operation on the intermediate network model through the pruning mask; obtain the sub - network model after completing all training steps.

[0155] The model fusion module 330 is used to fuse the sub - network model with the intermediate network model to obtain a fusion model.

[0156] The joint training module 340 is configured to perform joint iterative training on the fusion model using multilingual sample data to obtain a multilingual neural machine translation model.

[0157] For related details, refer to the above embodiments.

[0158] It should be noted that: when training the multilingual neural machine translation model provided in the above embodiments, only the division of the above functional modules is used for illustration. In practical applications, the above functions can be allocated to different functional modules as needed to complete all or part of the functions described above. In addition, the training device of the multilingual neural machine translation model provided in the above embodiments and the embodiments of the multilingual neural machine translation method belong to the same concept. For the specific implementation process, refer to the method embodiments and will not be elaborated here.

[0159] Figure 4 It is a block diagram of a multilingual translation device provided by an embodiment of the present application. The device at least includes the following modules: a model acquisition module 410 and a text translation module 420.

[0160] The model acquisition module 410 is configured to acquire a multilingual neural machine translation model; the multilingual neural machine translation model is trained using the multilingual neural machine translation model training method provided in the above embodiments;

[0161] The text translation module 420 is configured to input the text to be translated, the language identifier corresponding to the text to be translated, and the target language identifier into the multilingual neural machine translation model together. The multilingual neural machine translation model translates the text to be translated into the target language based on the target language identifier.

[0162] For related details, refer to the above embodiments.

[0163] It should be noted that: when performing multilingual translation using the multilingual translation device provided in the above embodiments, only the division of the above functional modules is used for illustration. In practical applications, the above functions can be allocated to different functional modules as needed to complete all or part of the functions described above. In addition, the multilingual translation device provided in the above embodiments and the multilingual translation embodiments belong to the same concept. For the specific implementation process, refer to the method embodiments and will not be elaborated here.

[0164] Correspondingly to the above method, the present invention further provides an electronic device, including a processor, a memory, and a computer program / instructions stored in the memory. The processor is configured to execute the computer program / instructions, and when the computer program / instructions are executed, the electronic device implements the steps of any one of the methods in the above embodiments.

[0165] The present invention further provides a computer-readable storage medium, on which a computer program / instructions are stored, and characterized in that when the computer program / instructions are executed by a processor, the steps of any one of the methods in the above embodiments are implemented.

[0166] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in connection with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present invention are programs or code segments used to execute the required tasks. The programs or code segments can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave on a transmission medium or a communication link.

[0167] It should be clear that the present invention is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, the detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present invention is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.

[0168] In the present invention, the features described and / or illustrated for one embodiment can be used in the same or a similar manner in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.

[0169] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for training a multilingual neural machine translation model, characterized in that: The method comprises the following steps: Iteratively training the pre-created initial network model using multilingual sample data to obtain an intermediate network model; the multilingual sample data includes sample training data corresponding to at least three languages, wherein each two languages ​​form a language pair, and each language corresponds to sample training data translated from each other; Using the sample training data corresponding to each language pair, the intermediate network model is iteratively trained to obtain the sub-network model corresponding to each language pair; wherein, in the iterative training process of the intermediate network model using the sample training data of each language pair, the training gradient of each weight in the intermediate network model is calculated in each training step, and after completing a preset number of training steps, the training gradient of each weight is accumulated and multiplied by the weight to obtain the importance score corresponding to each weight, and the importance score is compared with the pruning threshold to generate a pruning mask, and the intermediate network model is pruned by the pruning mask; the sub-network model is obtained after completing all the training steps; Fusing the sub-network model with the intermediate network model to obtain a fusion model; The fusion model is jointly iteratively trained using the multilingual sample data to obtain a multilingual neural machine translation model.

2. The method for training a multilingual neural machine translation model according to claim 1, characterized in that: The iterative training process of the intermediate network model using sample training data of each language pair also includes: Determine the target pruning ratio based on the number of steps in the current training step; Sort the importance scores corresponding to each weight by size; According to the target pruning ratio, the importance score of the corresponding position is selected from the sorted importance score list as the pruning threshold.

3. The method for training a multilingual neural machine translation model according to claim 2, characterized in that: The step of determining the target pruning ratio based on the number of steps in the current training step includes: When the number of steps is less than a first preset number of steps, a first preset pruning ratio is determined as the target pruning ratio; or, when the number of steps is greater than or equal to a second preset number of steps, a second preset pruning ratio is determined as the target pruning ratio; the first preset pruning ratio is less than the second preset pruning ratio.

4. The method for training a multilingual neural machine translation model according to claim 3, characterized in that: In the case where the number of steps is greater than or equal to the first preset number of steps and less than the second preset number of steps, the method further includes: Based on the number of steps, the first preset number of steps, the second preset number of steps and the second preset pruning ratio, a dynamic pruning ratio is obtained according to a preset linear growth algorithm as the target pruning ratio; the dynamic pruning ratio is greater than the first preset pruning ratio and less than the second preset pruning ratio.

5. The method for training a multilingual neural machine translation model according to claim 1, characterized in that: The fusion model is jointly iteratively trained using the multilingual sample data to obtain a multilingual neural machine translation model, including: In each joint iterative training process, randomly selecting sample training data corresponding to a language pair from the multilingual sample data as target sample training data; The fusion model is trained using the target sample training data and a preset loss function, and in the back propagation process, only the model parameters of the sub-network model corresponding to the target sample training data are updated; Repeat the above steps, process the sample training data corresponding to all language pairs in turn, perform joint iterative training on the fusion model, and obtain the multilingual neural machine translation model.

6. The method for training a multilingual neural machine translation model according to claim 1, characterized in that: The method of iteratively training the pre-created initial network model using multilingual sample data to obtain an intermediate network model includes: Randomly selecting sample training data corresponding to at least one language pair from the pairs of language sample data to form a current batch of sample training data; The initial network model is iteratively trained using the current batch of sample training data and a preset loss function to obtain the intermediate network model.

7. The method for training a multilingual neural machine translation model according to claim 1, characterized in that: The initial network model adopts a Transformer structure.

8. A multilingual translation method, characterized in that: The method includes: Acquire a multilingual neural machine translation model; the multilingual neural machine translation model is trained using a multilingual neural machine translation model training method according to any one of claims 1 to 7; The text to be translated, the language identifier corresponding to the text to be translated, and the target language identifier are input into a multilingual neural machine translation model, and the multilingual neural machine translation model translates the text to be translated into a target language based on the target language identifier.

9. An electronic device comprising a processor, a memory and a computer program / instruction stored in the memory, characterized in that: The processor is used to execute the computer program / instructions. When the computer program / instructions are executed, the electronic device implements the steps of the method according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.