Multi-Domain Adaptive Neural Machine Translation Method Based on Domain-Specific Subnetwork DsCN

The DsCN method addresses parameter interference and catastrophic forgetting in multi-domain NMT by using domain-specific sub-networks that inherit parameters based on inter-domain gradient similarity, enhancing translation accuracy across domains.

CN116542266BActive Publication Date: 2025-07-15KUNMING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310582479.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-23
Publication Date
2025-07-15
Estimated Expiration
2043-05-23

AI Technical Summary

Technical Problem

There are problems of parameter interference and catastrophic forgetting in multi-domain adaptive neural machine translation, and the existing methods are complex and ineffective.

Method used

The domain-specific subnetwork DsCN method is adopted to generate subnetworks by defining inheritance standards for inter-domain gradient similarity, only subnetwork parameters are updated, parameter interference is avoided and common domain performance is maintained.

Benefits of technology

The performance of multi-domain neural machine translation model is improved, parameter interference and catastrophic forgetting are avoided, and translation accuracy is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116542266B_ABST
    Figure CN116542266B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-domain adaptive neural machine translation method based on a domain-specific sub-network DsCN, which relates to the technical field of natural language processing. Multi-domain adaptive neural machine translation aims to use a single model to translate multiple domains, and joint training among multiple domains has proven to be successful. However, this joint training causes a performance degradation in resource-rich domains, namely the general domain, which is attributed to parameter interference. To address these problems, the DsCN of the present invention learns the sub-networks of each domain to counteract parameter interference. A large number of experiments on multi-domain datasets from English to German and from Chinese to English show that the method of the present invention significantly outperforms various baselines, improving 3 BLEU respectively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a multi-domain adaptive neural machine translation method based on a domain-specific sub-network DsCN, and belongs to the technical field of natural language processing. Background Art

[0002] In recent years, neural machine translation (NMT) has shown good performance with the support of large-scale corpora. However, since it is trained on a large amount of general data, its translation performance in specific fields such as law and medicine is often unsatisfactory. The use of professional terms and text styles in specific fields may confuse the model, and the scarcity of professional corpus data makes it difficult for the model to learn appropriate parameters. Multi-domain adaptive neural machine translation adopts a domain adaptation method to map the general information contained in the general domain to the specific domain, making up for the lack of training data in the specific domain and making full use of the language information in the specific domain to achieve the purpose of accurate translation.

[0003] Although multi-domain adaptive neural machine translation shows great potential, it still faces various challenges. One of the most serious problems is catastrophic forgetting caused by fine-tuning. Existing methods usually use domain-specific data to fine-tune a general-domain model trained on a large corpus to improve the translation performance in the target domain. However, this often leads to a decline in the translation performance in the original general domain. To solve this problem, Zhu et al. proposed a method called mixed fine-tuning (MFT), which trains a single model by mixing general-domain data with multiple domain-specific data. However, this leads to parameter interference, that is, the same parameter has a positive or negative impact in different domains. Therefore, Liang et al. proposed a method called sequence pruning and fine-tuning (SPT), which optimizes the general-domain model by discovering and adjusting domain-specific sub-networks. However, this method involves multiple steps and requires manual intervention, which introduces some complexity and subjectivity. Summary of the Invention

[0004] Aiming at the above problems, the present invention proposes a multi-domain adaptive neural machine translation method based on a domain-specific sub-network DsCN. The present invention effectively solves the problem of parameter interference and avoids the catastrophic forgetting problem in the general domain, improving the performance of the multi-domain neural machine translation model.

[0005] The method of the present invention models the general features of language and the specific features of each domain in a single model, avoiding the influence of parameter interference and not degrading the performance in the general domain during fine-tuning; the DsCN of the present invention uses an inheritance criterion defined as the inter-domain gradient similarity to generate a sub-domain network for each specific domain from the parent network in the general domain; each sub-network shares some parameters with other domains while also retaining its own domain-specific parameters. In DsCN, the general domain uses all the parameters of the model, i.e., the parent network, while the specific domain uses the parameters of the sub-network. At the same time, the sub-network itself determines the sharing strategy. Experimental results show that the method proposed by the present invention can effectively improve the performance of the multi-domain neural machine translation model on the English-German and Chinese-English public datasets.

[0006] The technical solution of the present invention is: a multi-domain adaptive neural machine translation method based on a domain-specific sub-network DsCN, and the specific steps of the method are as follows:

[0007] Step1. Download the training dataset on the machine translation competition website, and use tools to perform data preprocessing operations such as word segmentation and byte pair processing on the data;

[0008] Step2. Initialize the multi-domain NMT model, and each domain uses the inheritance criterion to evaluate all the parameters of the model to generate a domain sub-network;

[0009] Step3. Finally, jointly train and obtain a unified multi-domain neural machine translation model by mixing the general domain and specific domain data controlled by the domain label, and use the unified model for translation.

[0010] As a further solution of the present invention, the specific steps of the Step 1 are:

[0011] Download the general dataset and domain datasets on websites such as WMT, CCMT, and UM-Corpus, and use tools such as sentencespiece, StanfordNLP, and mosestokenizer to perform word segmentation and byte pair processing operations on the data to complete the data preprocessing.

[0012] As a further solution of the present invention, the specific steps of the Step2 are:

[0013] First, initialize a fully shared multi-domain NMT model and train for several rounds;

[0014] Secondly, all parameters of the parent network are evaluated using the inheritance criterion for each specific domain; and they are classified as inheritable or non-inheritable; the parameters that meet the inheritance criterion are considered inheritable, and the corresponding values in the mask matrix are set to 1, while the parameters that do not meet the conditions are considered non-inheritable and discarded, and the corresponding values in the mask matrix are set to 0;

[0015] Then, the parent network is multiplied element-wise with the mask matrix of the specific domain to obtain the domain-specific sub-network;

[0016] Finally, in the backpropagation training, only the parameters of the sub-network are updated, and these parameters are continuously iteratively updated until convergence is achieved.

[0017] As a further aspect of the present invention, the inheritance criterion in the step of evaluating all parameters of the parent network using the inheritance criterion for each specific domain is used to determine whether the parameters from the parent network should be inherited in the sub-network of the specific domain, and an evaluation criterion based on the inter-domain gradient cosine similarity is defined, which is called the inheritance criterion.

[0018] As a further aspect of the present invention, the parameter θ i is initially stored in the networks of the specific domains d1 and d2. In order to find the parameters suitable for joint training in the general domain D and the specific domains d1 and d2, the inter-domain gradient cosine similarity is first defined as the inheritance criterion; more formally, in order to determine whether the i-th parameter θ i in the parent network should be retained in the sub-network of the specific domain, the fitness score i of the parameter θ is defined as:

[0019]

[0020] where and are the gradient directions of the general domain D and the specific domain d i on the parameter θ k ;

[0021] Intuitively, the gradient determines the optimization direction, and the gradient represents the global optimal direction in the domain d k ;

[0022] For each specific domain, the gradients of all parameters in the parent network are evaluated on the validation set data, and the fitness i of the parameter θ ki in the specific region d is calculated, which helps to identify the parameters that cause serious interference to the joint training, and the superiority of these parameters is marked by comparing with the fitness threshold F T ;

[0023] Treat each parameter as an individual and assume the model parameter θ i has a fitness score greater than the fitness threshold F of the domain d k indicating that the parameter has strong viability in both the general domain D and the specific domain d T ; thus, the parameter θ k can be inherited into the sub-network of d i and the setting corresponding to this parameter k is set to 1, that is:

[0024]

[0025] In this way, the domain-specific sub-network θ of the domain d k is obtained: k :

[0026] θ k = θ ⊙ M k (3)

[0027] As a further solution of the present invention, each parameter in the parent network is evaluated using the inheritance criterion. Specifically, in the multi-domain NMT model, different granularity levels, including layers and parameters, are used to distinguish the evaluation. The layer granularity represents different layers in the model, and the parameter granularity represents each parameter in the model; at the layer granularity, the parameters within the layer are concatenated into a vector and calculated together using the inheritance criterion.

[0028] As a further solution of the present invention, batches are generated from the multi-domain dataset and it is ensured that each batch contains only samples from one domain; more specifically, first, the training data of the domain d k is sampled to create batches Then, the batches are used to train the domain-specific sub-network.

[0029] The beneficial effects of the present invention are:

[0030] The domain-specific sub-network DsCN proposed by the present invention models the general characteristics of the language and the specific characteristics of each domain, effectively solves the parameter interference problem, and avoids the catastrophic forgetting problem of the general domain, thereby improving the performance of the multi-domain neural machine translation model;

[0031] The method proposed by the present invention is simpler than the traditional multi-domain neural machine translation. Experimental results show that the BLEU values of this method are generally improved compared to the baseline system. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 ​This is the overall flowchart in the present invention;

[0033] Figure 2 This is the diagram of the domain-specific sub-network DsCN model in the present invention;

[0034] Figure 3 This is the inheritance standard diagram in the present invention. Detailed implementation manners

[0035] Example 1: As Figures 1-3 shown, a multi-domain adaptive neural machine translation method based on the domain-specific sub-network DsCN. To test the effectiveness of the proposed method in the multi-domain adaptive neural machine translation task, the present invention conducts experiments on the translations from English to German and from Chinese to English respectively, and compares various baseline models. The specific steps of the multi-domain adaptive neural machine translation method based on the domain-specific sub-network DsCN are as follows:

[0036] Step1. Download the training data set on the machine translation competition website, and perform data preprocessing operations such as word segmentation and byte pair processing on the data using tools; for the English-to-German data set, the present invention constructs data sets in three specific domains: general domain, speech, biomedicine, and fiction. In the general domain, the present invention uses the WMT14 news translation task as the training set, Newstest2013 as the validation set, and newstest2014 as the test set. For the speech domain, the present invention uses IWSLT14 as the training set, Dev2010 as the validation set, and Dev2014 as the test set. In the biomedicine domain, we use EMEA NewsCrawl as the training set, and use Khresmoi medical abstract translation test data 2.0 as the validation and test set. For the fiction domain, the present invention uses the OPUS book data set as the training set, randomly selected chapters from "Jane Eyre" as the validation set, and "The Metamorphosis" as the test set.

[0037] For the Chinese-to-English data set, the present invention constructs data in three specific domains: general domain, papers, spoken language, and education. For the general domain, the present invention uses the WMT14 news translation task as the training set, Newstest2017 as the validation set, and News Dev2017 as the test set. For the specific domains, the present invention uses academic papers, spoken language, and education fields from the UM-Corpus as experimental data.

[0038] The data set statistical information is shown in Table 1.

[0039] Table 1 Data set statistical information

[0040]

[0041]

[0042] The present invention applies the Byte Pair Encoding (BPE) algorithm to all datasets, where the vocabulary sizes of the OPUS and WMT datasets are 64k, and the vocabulary size of the IWSLT dataset is 32k. Since English and German belong to the West Germanic language branch, the present invention uses a shared vocabulary of size 32k for the English-to-German dataset and uses sentencespiece for tokenization. For the Chinese-to-English dataset, the non-shared vocabulary sizes are 44K for Chinese and 33K for English. The present invention uses Stanford NLP Segmenter and tokenizertools to tokenize Chinese and English respectively.

[0043] During the inference process, the beam search for English-German and Chinese-English is set to 4, and the length penalties for English-German and Chinese-English are set to 0.6 and 1.0 respectively. In addition, the present invention uses the Sacre-BLEU tool to evaluate the performance of the translation models for all experiments. For simplicity, the present invention only reports the BLEU from the best multi-domain neural machine translation model.

[0044] Step 2: The experimental environment is the Linux Ubuntu 20.04 system, the PyTorch version is 1.7.1+cu110, and the compilation language is Python 3.7. The present invention conducts experiments using the basic settings of the Transformer architecture with 6 encoder-decoder layers, 1024 / 4096 hidden dimensions, and 16 attention heads. The model is trained using the Adam optimizer during training, and the hyperparameters β1 and β2 are set to 0.9 and 0.98 respectively. The initial learning rate is set to 5x10 -4 , the dropout is 0.1, and all models are trained on an NVIDIA RTX3090ti GPU with a global batch size of 4096 words. The experimental parameter settings are shown in Table 2.

[0045] Table 2 Experimental Parameter Settings

[0046]

[0047]

[0048] The main objective of the present invention is to determine the parameters that contribute to jointly training a multi-domain adaptive NMT model in general and specific domains. To achieve this goal, the present invention introduces a method DsCN for dynamically discovering and learning domain-specific sub-networks for multi-domain adaptive NMT, and its inheritance criterion is defined as the similarity of inter-domain gradients.

[0049] The present invention first initializes a fully shared multi-domain NMT model. After several rounds of training, for each specific domain, inheritance criteria are used to evaluate all parameters in the parent network and classify them as inheritable or non-inheritable. Parameters that meet the inheritance criteria are considered inheritable and will be retained in the sub-network of that specific domain, and the corresponding values in the mask matrix are set to 1, while parameters that do not meet the criteria are considered non-inheritable and discarded, and the corresponding values in the mask matrix are set to 0; then, the parent network is multiplied by the mask matrix of the specific domain to obtain the domain-specific sub-network; in subsequent backpropagation training, only the parameters of the sub-network are updated. The present invention continues to iteratively update these parameters until convergence is achieved.

[0050] A key issue for evaluating parameter quality is to define an evaluation criterion that can help determine whether parameters from the parent network should be inherited in the sub-network of a specific domain. The present invention defines an evaluation criterion based on the cosine similarity of gradients between domains, called the inheritance criterion, where parameters with conflicting gradients are more likely to hinder joint training in the general domain and specific domains. On the contrary, parameters with a consistent gradient optimization direction can help solve catastrophic forgetting and parameter interference.

[0051] As Figure 3 shown, the parameter θ i is initially stored in the networks of specific domains d1 and d2. To find parameters that are suitable for joint training in both the general domain d and specific domains D1 and D2, the present invention first defines the cosine similarity of gradients between domains as the inheritance criterion. More formally, to determine whether the i-th parameter θ i in the parent network should be retained in the sub-network of the specific domain, the fitness score i of the parameter θ is defined as:

[0052]

[0053] where and are the gradient directions of the general domain d and the specific domain D i on the parameter θ k .

[0054] Intuitively, the gradient determines the optimization direction. For example, in Figure 3 , the gradient represents the global optimal direction in the domain D k . This gradient has the largest negative cosine similarity, as and they point in opposite directions, hindering optimization and have been shown to be unfavorable for multi-task learning.

[0055] For each specific domain, the gradients of all parameters in the parent network are evaluated on the validation set data. The present invention calculates the parameter θ i in the specific region d k of the fitness which helps to identify the parameters that cause serious interference to the joint training, and by comparing the fitness threshold f T to mark the superiority of these parameters.

[0056] The present invention treats each parameter as an individual, and assumes that the fitness score i of the model parameter θ is greater than the fitness threshold F k of the domain d T indicating that this parameter has strong viability in both the general domain d and the specific domain D k . Therefore, the parameter θ i can be inherited into the sub-network of D k . In this algorithm, the corresponding of this parameter is set to 1, that is:

[0057]

[0058] In this way, the present invention obtains the domain-specific sub-network θ k of the domain D k :

[0059] θ k = θ ⊙ M k (3)

[0060] The consistent gradient optimization direction helps the model to maintain its performance in the general domain as much as possible, but multi-domain adaptive NMT also needs to preserve domain information. By adjusting different fitness thresholds, this problem can be well solved. Obviously, the smaller the value of F T , the looser the model evaluates the parameters, and the more domain-specific information is preserved.

[0061] Theoretically, each parameter in the parent network can be evaluated using the inheritance criterion. However, in practice, scoring each parameter may be time-consuming and resource-consuming, especially in multi-domain NMT models that may contain millions to billions of parameters. To solve this problem, the present invention uses different granularity levels, such as layers and parameters, to distinguish the evaluation. Layer granularity represents different layers in the model, and parameter granularity represents each parameter in the model. For example, at the layer granularity, the parameters within a layer are concatenated into a vector and calculated together using the inheritance criterion.

[0062] In the present invention, batches are generated from multi-domain datasets, and it is ensured that each batch contains samples from only one domain. This method is different from the training of traditional fully shared multi-domain NMT models, where each batch may contain sentence pairs from different domains. More specifically, the present invention first samples the training data of domain d k to create batches Then, the present invention uses the batches to train domain-specific sub-networks.

[0063] Step 3, Finally, a unified multi-domain neural machine translation model is jointly trained and obtained by mixing general-domain and domain-specific data controlled by domain labels.

[0064] To verify the effectiveness of the proposed multi-domain adaptive neural machine translation method based on the domain-specific sub-network DsCN of the present invention, comparisons are made with some classical methods and the current state-of-the-art (SOTA) methods:

[0065] 1) Multi-domain NMT model (MDNMT): This model is trained only on parallel corpus data from the general domain.

[0066] 2) Fine-tuning (FT): First, it is trained on the general-domain corpus, and then fine-tuned with the domain-specific corpus to continue training.

[0067] 3) Mixed fine-tuning (MFT): Mix data from the general domain and all domain-specific data to train the model.

[0068] 4) Mixed domain tags (MDT): Mix data from the fairy tale domain and all domain-specific data to train a unified multi-domain NMT model, but add domain tags to the data in each domain for differentiation.

[0069] 5) Adapter: The adapter takes the intermediate layer of a pre-trained general model as input and adds a smaller task-specific layer (adapter) on top of it, enabling the model to be fine-tuned for the new domain.

[0070] 6) Sequence pruning tuning (SPT): Find and freeze the most informative parameters on the general-domain model, then prune unnecessary network model parameters, and finally use a mask matrix and domain-specific data to fine-tune the domain-specific sub-network parameters.

[0071] The results of each model on the English-German dataset and the Chinese-English dataset are shown in Tables 3 and 4 respectively.

[0072] Table 3 Results of each model on the English-German dataset

[0073]

[0074] Table 4 Results of each model on Chinese-English datasets

[0075]

[0076] Table 3 and Table 4 respectively present the results of the method of the present invention and the baseline method on English-German and Chinese-English datasets. In the English-to-German translation task, DsCN has always outperformed the baseline model in general, speech, and biology domains. The average BLEU score is 30.22, which is about 3 percentage points higher than the MDNMT model. It is worth noting that DsCN has achieved the most significant improvement in the speech and novel domains. Similarly, on the Chinese-English dataset, DsCN outperforms the baseline model in the general domain and the paper domain, with an average BLEU score of 20.26, which is about 3 percentage points higher than the MDNMT model.

[0077] Compared with the catastrophic forgetting caused by fine-tuning, the performance of DsCN on English-German and Chinese-English datasets is always better than that of all baseline models in the general domain. Compared with the parameter interference problem introduced by MFT, the performance of DsCN in six domains of English-German and Chinese-English respectively outperforms MFT, and the BLEU scores are increased by about 2 and 3 percentage points respectively. It is worth noting that DsCN has achieved better performance than the SOTA method SPT in all domains of English-German.

[0078] In addition, for different granularities in DsCN, the present invention finds that compared with hierarchical granularity, parameter granularity achieves the best results in most domains of English-German and Chinese-English. The present invention speculates that this is because more fine-grained control is performed on the sub-networks at the parameter granularity.

[0079] Therefore, the present invention can know that the method of the present invention models the general features of language and the specific features of each domain by generating domain-specific sub-networks, avoiding the influence of parameter interference, thereby improving the performance of the multi-domain neural machine translation model. Secondly, compared with fine-tuning, DsCN can effectively avoid the catastrophic forgetting problem in the general domain. Although hybrid domain labels, adapters, and pruning followed by expansion can also solve this problem to a certain extent, the method proposed by the present invention has a more obvious effect. The fundamental reason is that DsCN only updates the parameters corresponding to the sub-networks, avoiding the impact of updating all parameters on the performance of other domains. Another advantage of updating the corresponding parameters is its rapid adaptation to specific domains. This alleviates the problem of performance degradation in the general domain and improves the performance in multiple specific domains.

[0080] The specific embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.

Claims

1. A multi-domain adaptive neural machine translation method based on a domain-specific sub-network DsCN, characterized in that: The specific steps of the method are as follows: Step1. Download the training data set from the machine translation competition website, and use tools to perform data preprocessing operations such as word segmentation and byte pair processing on the data; Step2. Initialize the multi-domain NMT model, and each domain uses all the parameters inherited from the standard evaluation model to generate a domain sub-network; Step3. Finally, jointly train by mixing general-domain and specific-domain data controlled by domain labels to obtain a unified multi-domain neural machine translation model, and use the unified model for translation; The specific steps of Step2 are: First, initialize a fully shared multi-domain NMT model and train for several rounds; Secondly, for each specific domain, use all the parameters of the inherited standard evaluation parent network; and classify them as inheritable or non-inheritable; the parameters that meet the inheritance standard are considered inheritable, and the corresponding values in the mask matrix are set to 1, while the parameters that do not meet the conditions are regarded as non-inheritable and discarded, and the corresponding values in the mask matrix are set to 0; Then, multiply the parent network by the mask matrix of the specific domain to obtain a domain-specific sub-network; Finally, in the backpropagation training, only update the parameters of the sub-network, and continue to iteratively update these parameters until convergence is reached; The inheritance standard in the step of using the inheritance standard to evaluate all the parameters of the parent network for each specific domain is used to determine whether the parameters from the parent network should be inherited in the sub-network of the specific domain, and an evaluation criterion based on the cosine similarity of gradients between domains is defined, which is called the inheritance criterion; Parameter θ i Initially stored in the networks of specific domains d1 and d2, to find the parameters suitable for joint training in both the general domain D and specific domains d1 and d2, an inter-domain gradient cosine similarity was first defined as the inheritance criterion; more formally, to determine whether the i-th parameter θ i in the parent network should be retained in the sub-network of the specific domain, the fitness score of the parameter θ i is defined as: was defined as: wherein and are the general domain D and the specific domain d i on the parameter θ k gradient directions; Intuitively, the gradient determines the optimization direction, and the gradient represents the global optimal direction in domain d k ; For each specific domain, the gradients of all parameters in the parent network are evaluated on the validation set data to calculate the parameter θ i in the specific region d k of the fitness helps to identify the parameters that cause serious interference to the joint training, and by comparing the fitness threshold F T to mark the superiority of these parameters; Treat each parameter as an individual and assume the model parameter θ i has a fitness score greater than the domain d k has a fitness threshold F T , indicating that the parameter has strong viability in both the general domain D and the specific domain d k ; therefore, the parameter θ i can be inherited into the sub-network of d k and the corresponding is set to 1, that is: In this way, the domain d is obtained k The domain-specific sub-network θ k : θ k = θ ⊙ M k (3).

2. The multi-domain adaptive neural machine translation method based on the domain-specific sub-network DsCN according to claim 1, characterized in that: The specific steps of Step 1 are: Download the general data set and domain data set from the WMT, CCMT and UM-Corpus websites, and use tools such as sentencespiece, StanfordNLP and mosestokenizer to perform word segmentation and byte pair processing operations on the data to complete data preprocessing.

3. The multi-domain adaptive neural machine translation method based on a domain-specific sub-network DsCN according to claim 1, wherein: Each parameter in the parent network is evaluated using the inheritance standard. Specifically, in the multi-domain NMT model, different granularity levels, including layers and parameters, are used to distinguish the evaluation. The layer granularity represents different layers in the model, and the parameter granularity represents each parameter in the model; at the layer granularity, the parameters within the layer are concatenated into a vector and calculated together using the inheritance standard.

4. The multi-domain adaptive neural machine translation method based on the domain-specific sub-network DsCN according to claim 1, characterized in that: Generate batches from a multi-domain dataset and ensure that each batch contains samples from only one domain; more specifically, first sample the training data for domain d k to create batches Then, use the batches to train the domain-specific sub-network.

Citation Information

Patent Citations

  • Pretreatment-based automatic detection method for vulnerable plaque of IVOCT image

    CN108416769A

  • Mongolian and Chinese neural machine translation method based on transfer learning strategy

    CN108829684A