Data enhancement method based on bidirectional interpretation adversarial network
By generating high-quality adversarial samples and diverse data through a bidirectional back-translation adversarial network, the bias problem of low-resource language translation models in specific domains is solved, and the robustness and translation performance of the model are improved.
Patent Information
- Application Number
- CN202510866925.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-11-04
AI Technical Summary
Machine translation models for low-resource languages perform poorly in specific domains. The lack of large-scale and diverse training data leads to biased translation results, and existing data augmentation methods have failed to effectively improve the robustness of the models.
A data augmentation method based on bidirectional back-translation adversarial networks is adopted to generate high-quality adversarial examples and diverse training data through adversarial training and bidirectional back-translation techniques. The model is then retrained using a pre-trained masked language model to enhance its robustness.
It significantly improves the performance and robustness of machine translation models under low-resource conditions, reduces overfitting of models to specific data, and improves translation accuracy and generalization ability.
Smart Images

Figure CN120893449A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to a data enhancement method based on a bidirectional back-translation adversarial network and belongs to the technical field of natural language processing. BACKGROUND
[0002] In recent years, the application of large language models (LLM) in the field of machine translation has undoubtedly been a major breakthrough, significantly improving the accuracy and fluency of translation. Although more translation data can be generated through data enhancement techniques, if the original data itself has biases or deficiencies, the enhanced data may not fully reflect the true situation of the language. Low-resource languages are not well translated by LLM due to the lack of sufficient training data.
[0003] The problem of large translation bias in the translation process by increasing perturbation and then performing a translation process again indicates that the model is sensitive to certain perturbations. Based on this, data augmentation is performed using adversarial training, and the reliability of adversarial instances is enhanced by combining the bidirectional translation method and the pre-trained masking language model on the basis of the adversarial back-translation framework. Retraining the model with these data can enhance its robustness, thereby effectively improving the overall translation performance. SUMMARY
[0004] The technical problem solved by the application is that the application provides a data enhancement method based on a bidirectional back-translation adversarial network to solve the problem of bias in machine translation results in a specific field due to the lack of large-scale and diverse training corpus in low-resource language machine translation. The application combines adversarial training and bidirectional back-translation technology to generate high-quality adversarial samples and diverse training data using a pre-trained masking language model to enhance the robustness and performance of the model and improve the performance of neural machine translation models under low-resource conditions.
[0005] The technical solution of the application is a data enhancement method based on a bidirectional back-translation adversarial network, which comprises:
[0006] Step 1, using an adversarial back-translation method to back-translate pseudo-parallel data of source language and target language to enhance the instances;
[0007] Step 2, introducing adversarial samples with consistent semantics to provide more extensive training data for model training;
[0008] Step 3, combining bidirectional translation and MLM to generate enhanced data and retraining the model with these data.
[0009] Further, the Step 1 comprises:
[0010] Step1.1, constructing a module that can realize two-way back translation, including a back translation module BT from the source language to the target language and back to the source language s = {M s→t , M t→s}, and a back translation module BT from the target language to the source language and back to the target language t = {M t→s , M s→t}; wherein M s→t refers to a model for translating a source language sentence s to a target language sentence t, and M t→s refers to a model for translating a target language sentence t back to a source language sentence s;
[0011] Step 1.2, a set of sentence pairs D = {s, t} to be attacked is generated through the back translation process BT(·) to produce a new parallel sentence pair, denoted as back translation instance D' = {s', t'}, wherein s is a source language sentence, s' is the translation of s after reconstruction by the BT s process, t is a target language sentence, and t' is the translation of the target language sentence after reconstruction by the BT t process.
[0012] Further, the Step2 includes:
[0013] Step2.1, by constructing a reverse translation instance sentence, the reconstructed translation D' and D are compared to calculate the translation deviation of the model; by analyzing this deviation, the accuracy and reliability of the model's translation ability for sentence pairs are effectively evaluated;
[0014] Step2.2, when the similarity between D' and D is high, D' is considered a good adversarial instance and is selected as the adversarial target, which will be attacked by the adversarial attack module for corpus data enhancement; the similarity calculation standard selected is to calculate the BLEU value.
[0015] Further, in Step3, when the back translation instance D' selected from the corpus is screened and enters the adversarial attack module for attack, an adversarial instance D'' is generated by using the adversarial attack module; the adversarial attack module in data enhancement includes a source language mask language model MLM and a translation language model TLM in the source language to target language translation direction.
[0016] In the generation of adversarial instances, the mask language model MLM is used to determine the part of the instance that is most vulnerable to attack, and then the translation language model TLM in the source language to target language translation direction is used to replace the target vocabulary.
[0017] Further, the Step3 includes:
[0018] Step3.1, the masked language model MLM reversely predicts the words in the masked positions by using the context information after randomly masking some words in random positions of the input sentence; the context understanding ability of the MLM is used to attack the source language sentence s;
[0019] Step3.2, the TLM is used to generate the adversarial instance t’ on the target language side; after a set of replacement words is obtained, the attack position is determined, and the sentence pair D is attacked.
[0020] Further, in Step3.2, the attack includes the following steps:
[0021] Step3.2.1, determining the attack position n of the word x n ,…,x l in the source language sentence s={x1,…,x
[0022] A statistical-based phrase alignment tool is used to find the attacked source language word x n by using the alignment relationship between the source language and the target language sentences in the sentence pair D;
[0023] Step3.2.2, attacking the source language sentence s;
[0024] The source language sentence s is replaced by the new word x″ n predicted by the MLM after the word x n in the attack position n is replaced, and the new source language sentence s” after the attack is formed;
[0025] Step3.2.3, determining the alignment relationship {y1,…,y m ,…,y p ,…,y ; of the attacked source language word x n in the target language vocabulary; the attack position m-p of the target language is determined, and the attack position m-p of the target language corresponding to the attack position n of the source language attack vocabulary is obtained by using the alignment relationship;
[0026] Step3.2.4, attacking the target language sentence t;
[0027] After the target language attack position m-p is obtained, the target language sentence t=
[0028] {y1,…,y m ,…,y p ,…,y l} a mask operation is performed to replace the phrases of the target end with the [MASK] label to obtain the target language sentence t with the mask label m , a new sentence s" of the source language after confrontation is obtained by splicing the new sentence s" of the source language after confrontation and the target language sentence t with the mask label m An input sequence input of a TLM is composed t ={x1,…,x n ,…,x l ,[SEP],y1,…,y m-1 ,[MASK],y p ,…,y l} is obtained; the mask label in the target language is predicted by using the cross-language prediction capability of the TLM to obtain the output of the TLM, that is, a new sentence t' of the target language after confrontation;
[0029] Step3.2.5, a final confrontation instance D" is obtained through the new sentence s" of the source language after confrontation and the new sentence t' of the target language after confrontation.
[0030] After the above steps, the new sentence s" of the source language after confrontation predicted by the MLM and the new sentence t' of the target language after confrontation predicted by the TLM are obtained, and finally the stability of the newly generated sentence pair is evaluated by an iterative back-translation evaluation model, when the stability is poor, the confrontation instance does not meet the similarity principle of the corpus and will be discarded.
[0031] The application also provides a data enhancement system based on a bidirectional back-translation confrontation network, which comprises a module for executing the data enhancement method based on the bidirectional back-translation confrontation network.
[0032] The application also provides an electronic device comprising a memory, a processor and a computer program stored on the memory and executable on the processor, wherein the processor executes the program to implement the data enhancement method based on the bidirectional back-translation confrontation network.
[0033] The application also provides a non-transitory computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the data enhancement method based on the bidirectional back-translation confrontation network.
[0034] The application also provides a computer program product comprising a computer program, wherein the computer program is executed by a processor to implement the data enhancement method based on the bidirectional back-translation confrontation network.
[0035] The application has the following beneficial effects:
[0036] 1. The application enhances the robustness of the model by combining adversarial training and bidirectional back-translation technology, using pre-trained masked language models to generate high-quality adversarial samples and diversified training data.
[0037] 2. The experimental results show that the method reduces the overfitting of the model to specific types of data, and significantly enhances the robustness and performance of the NMT model at the data level. BRIEF DESCRIPTION OF DRAWINGS
[0038] Figure 1 The total model structure in the application. DETAILED DESCRIPTION
[0039] Example 1: As shown, the data augmentation method based on bidirectional back-translation adversarial network includes: Figure 1
[0040] Step 1, the pseudo-parallel data of source language and target language are back-translated by an adversarial back-translation method to enhance the instances;
[0041] Step 2, by introducing adversarial examples with consistent semantics, more extensive training data is provided for model training;
[0042] Step 3, the bidirectional translation and MLM are fused to generate enhanced data, and the model is retrained with these data.
[0043] Further, the Step 1 includes:
[0044] Step 1.1, a module capable of bidirectional back-translation is constructed, including a back-translation module BT s ={M s→t ,M t→s} from source language to target language and back to source language, and a back-translation module BT t ={M t→s ,M s→t} from target language to source language and back to target language; wherein M s→t is a model for translating a source language sentence s to a target language sentence t, and M t→s is a model for translating a target language sentence t back to a source language sentence s; such configuration ensures that the two translation processes of source language to target language and back to source language, and target language to source language and back to target language are covered;
[0045] Step 1.2, a set of sentences to be attacked D={s,t} is generated through the back-translation process BT(·) to produce a new parallel sentence pair, denoted as back-translation instance D’={s′,t′}, wherein s is a source language sentence, s’ is s through BT s The translation after process reconstruction, t is the target language sentence, t' is the target language sentence through BT t The translation after process reconstruction.
[0046] Further, the Step2 includes:
[0047] Step2.1, by constructing the example sentence of back translation, further comparing the reconstructed translation D' and D, to calculate the translation deviation of the model; by analyzing this deviation, effectively evaluating the accuracy and reliability of the model on sentence pair translation ability;
[0048] Step2.2, when the similarity between D' and D is high, D' is considered a good adversarial example, and is selected as the adversarial target, which will be attacked by the adversarial attack module for enhancing the corpus data; the similarity calculation standard is to calculate the BLEU value.
[0049] The similarity between the original data and the back translation is calculated as the standard for judging whether this example can be used as an adversarial sample, as shown in formula 1:
[0050]
[0051] Among them, Similarity s is the similarity between s and the translation s' reconstructed by the translation model in the process of forward translation, and Similarity t is the similarity between the target language and the back translated sentence. This step is used to filter some reconstruction failed examples;
[0052] Table 1 is the algorithm of back translation and screening
[0053]
[0054] Further, in Step3, when the back translated example D' selected from the corpus passes the screening, it will enter the adversarial attack module for attack, and the adversarial example D" is generated by using the adversarial attack module; the adversarial attack module in data enhancement is shown in part b of Figure 1 The adversarial attack module in data enhancement includes the source language masked language model MLM (Masked Language Model) and the translation language model TLM (Translation Language Model) in the source language to target language translation direction.
[0055] In the generation of adversarial examples, attention should be paid to the quality of the generated adversarial examples, and identifying specific positions and vocabulary types in the dataset that are prone to attack is an important step in the attack; a masked language model MLM is used to determine the most vulnerable part of the example, and then a translation language model TLM in the source language to target language translation direction is used to replace the target vocabulary; this method aims to generate examples better than traditional adversarial methods and one-way adversarial methods to improve the accuracy of translation;
[0056] Further, the Step3 comprises:
[0057] Step3.1, the masked language model exhibits good context understanding ability and perception ability, and the application of the language model to some downstream tasks has advanced effect; the masked language model MLM predicts the words in the masked position by randomly masking some input sentence words in the masked position through context information; the context understanding ability of MLM is used to attack the source language sentence s;
[0058] Step3.2, the TLM cross-language characteristics are used to assist in generating target language side adversarial examples t''; after obtaining a set of replacement word tables, the attack position is determined to attack the sentence pair D.
[0059] Further, in the Step3.2, the attack comprises the following steps:
[0060] Step3.2.1, determining the word attack position n in the source language sentence s={x1,…,x n ,…,x ;};
[0061] The alignment between the bilingual sentence pair D is embodied between phrases, and the MLM model is based on word granularity [MASK], so incorrect attack positions will lead to incorrect replacement types;
[0062] A statistical-based phrase alignment tool is used to generate the source language word x n to be attacked by finding the alignment relationship between the source language and the target language sentences in the sentence pair D;
[0063] Step3.2.2, attack the source language sentence s;
[0064] The source language sentence s obtains the word attack position n of the source language word x n to be attacked by using the alignment tool, and the word in the word attack position n is replaced with the new word x'' predicted by the MLM n to form the new source language sentence s'' after the attack;
[0065] Step3.2.3, determining the source language word x attacked n The alignment relationship in the target language vocabulary {y1,…,y m ,…,y p ,…,y ;};Not only need to attack the vocabulary x n in the source language, but also need to determine the attack position m-p of the target language in order to ensure the semantic consistency between the source language sentence and the target language in the sentence pair after the attack, and the attack position m-p of the target language corresponding to the word attack position n of the source language attack vocabulary is obtained through the alignment relationship;
[0066] Step3.2.4, attack the target language sentence t;
[0067] After obtaining the target language attack position m-p, the target language sentence t={y1,…,y m ,…,y p ,…,y l} is subjected to a mask operation, and the phrase of the target end is replaced with a [MASK] mark to obtain a target language sentence t m with a mask mark, and the source language new sentence s” after the attack and the target language sentence t m with a mask mark are spliced to form an input sequence input t of a TLM {x1,…,x″ n ,…,x l ,[SEP],y1,…,y m-1 ,[MASK],y p ,…,y l}; The mask mark in the target language is predicted by using the cross-language prediction ability of the TLM, and the output of the TLM, i.e., the target language new sentence t’ after the attack, is obtained;
[0068] Step3.2.5, obtaining the final adversarial instance D” through the source language new sentence s” after the attack and the target language new sentence t’ after the attack;
[0069] After the above steps, the source language new sentence s” after the attack predicted by the MLM and the target language new sentence t’ after the attack predicted by the TLM are obtained, and finally the stability of the newly generated sentence pair is evaluated by the iterative back-translation evaluation model. When the stability is poor, the adversarial instance does not meet the similarity principle of the corpus and will be discarded.
[0070] The application also provides a data enhancement system based on a bidirectional back-translation adversarial network, which comprises a module for executing the data enhancement method based on the bidirectional back-translation adversarial network.
[0071] The application further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the data enhancement method based on the bidirectional translation adversarial network when executing the program.
[0072] The application further provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the data enhancement method based on the bidirectional translation adversarial network when executed by a processor.
[0073] The application further provides a computer program product comprising a computer program, wherein the computer program implements the data enhancement method based on the bidirectional translation adversarial network when executed by a processor.
[0074] In order to verify the effectiveness of the application, the following experiments are performed:
[0075] Experiment 1, comparative experiment:
[0076] The comparative experiment shows that an effective data enhancement strategy has a great influence on improving the robustness of a model under limited data resources.
[0077] Table 2 shows the results of the comparative experiment
[0078]
[0079] Although the EDA method aims to increase data diversity by introducing random noise, the actual effect is counterproductive due to excessive noise, resulting in a decrease in model performance. This also shows that the EDA method is not an optimal data enhancement method for a data set that already contains noise. On the contrary, other data enhancement methods have certain improvements compared to the baseline model, and the improvement of the method of the application is the most obvious, which also indicates the effectiveness of the adversarial data enhancement and the specific strategy proposed by the application in enhancing the robustness of the model.
[0080] Experiment 2, influence of the screening threshold hyperparameter when screening adversarial instances on the data enhancement results:
[0081] The purpose of the experiment is to explore the use of different translation similarities for screening, which are defined in the second section. The purpose is to study the degree of quality decline between the enhanced instances D' and the original data D before and after the adversarial attack when the quality of the selected adversarial samples is too low or too high.
[0082] The average drop in performance (MPD) metric was calculated using the same approach as in the literature. The MPD values for the enhanced overall corpus are shown in Table 3. Higher MPD values indicate more effective adversarial data after enhancement. Due to the characteristics of the data set and the translation model, the similarity between D' and D should not be too small or too large. The similarity threshold used was between 11 and 60, with an interval of 10. The reason for choosing a BLEU lower bound of 11 for filtering is that this is the lowest filtering value for the original back-translation model, which is the lowest capability of the model.
[0083] Table 3: Performance of data sets filtered by different similarity thresholds
[0084]
[0085] To explore the effect of the proportion of adversarial data on model performance improvement, the enhanced data with different thresholds mentioned above were fused with the original data and tested on the same ALT data set. The experimental results show that as the similarity threshold increases, the MPD value increases, but the adversarial data after filtering becomes severely insufficient. Even if the adversarial instances become more effective, such data is difficult to support model improvement in robustness.
[0086] It also shows that as the similarity threshold increases, although the selected adversarial samples are of higher quality in theory, the average drop in performance (MPD) between the enhanced instances D' and the original data set D increases, reflecting the more significant positive impact of the enhanced adversarial data on model performance. However, the limitation of this method is that as the threshold increases, the number of adversarial samples that meet the conditions decreases significantly, resulting in a decrease in the proportion of filtered data to the total data. This reduction in data volume may weaken the role of adversarial samples in improving model robustness, especially when the proportion of adversarial samples is low, the positive impact of adversarial samples on overall model performance may not be enough to offset the negative effects caused by the reduction in the number of adversarial samples. At the same time, this also shows that the filtering operation needs to consider the balance between data quality and data size.
[0087] Experiment Three: Effect of Different Data Fusion Strategies on Model Performance
[0088] To verify the effect of different data fusion methods on model robustness, three different fusion methods were designed to test their impact on model performance:
[0089] Directly using augmented data: First, the data set generated by the data augmentation method of the application is used as the only input source for model training. In this method, the original data is not directly included in the training set, but relies entirely on augmented data. The purpose of this strategy is to evaluate the direct impact of relying solely on augmented data on the performance of the translation model.
[0090] Training after fusing augmented data with original data: a method of training a machine translation model using the data generated by the data augmentation technique and the original data set. The goal of this strategy is to combine the additional diversity provided by augmented data with the reliability of the original data to build a more generalizable translation model. By fusing augmented data and original data to create a more comprehensive training set, the performance of the model is tested.
[0091] Fine-tuning a pre-trained model using augmented data: a method of fine-tuning a baseline model trained on original data using data augmented by the application. The fine-tuning stage focuses on adjusting model parameters so that the model better adapts to new samples and contexts shown by augmented data while retaining knowledge obtained on original data. This strategy combines the basics of original data with the characteristics of augmented data to enhance the model's ability to adapt to different text styles and expressions.
[0092] Mixing noise data with original data: To verify the impact of increased data size on evaluation scores when using fused data, this method randomly samples data of the same size as the augmented data from the discarded approximately 320K noise corpus and fuses it with the source language data as a control experiment. The experimental results are shown in Table 4.
[0093] Table 4 shows the performance of different data fusion methods
[0094]
[0095] As can be seen from the data in the table, the method of training after fusing augmented data with original data performs best among these strategies, indicating that when data is augmented, maintaining the original data while introducing augmented data can effectively improve the performance and generalization ability of the model.
[0096] Experiment three, generalization performance experiment of the model:
[0097] To verify the generalization ability of the model improved by the method of the application under different unseen perturbations, the performance on TED, QET, WIKI, Tanzil, etc. data sets (Malay-Chinese as experimental data) is increased, and the results are shown in Table 5.
[0098] Table 5 shows the generalization performance of the model of the application on different data sets
[0099]
[0100] To further verify the effectiveness of the data enhanced by the method of the present application, the IWSLT (English-German, English-Chinese) is combined as experimental data, and the results are shown in Table 6.
[0101] Table 6 Performance of enhanced data of the present application combined with other data sets
[0102]
[0103] From Tables 5 and 6, it can be seen that the method of the present application shows better generalization performance than the baseline model (Base) on different data sets. The method of the present application not only performs well on a single data set, but also shows strong generalization ability on multiple different, unseen perturbed data sets. This proves the effectiveness of the method of the present application, which can maintain the stability and accuracy of the model in a diversified data environment, especially the performance on the Tanzil data set is the most significant, with an improvement rate reaching the highest among all data sets, which may be attributed to the fact that the method of the present application has a more effective adaptation and improvement strategy for the perturbation of this particular data set.
[0104] The specific embodiments of the present application are described in detail above in combination with the accompanying drawings, but the present application is not limited to the above-mentioned embodiments, and various changes can be made within the knowledge of those skilled in the art without departing from the spirit of the present application.
Claims
1. A data augmentation method based on bidirectional back-translation adversarial networks, characterized in that: The method includes: Step 1: Use an adversarial back-translation method to back-translate the pseudo-parallel data in the source and target languages to enhance the instances. Step 2: By introducing adversarial examples that maintain semantic consistency, a wider range of training data is provided for model training; Step 3: Integrate bidirectional translation and MLM to generate augmented data, and use this data to retrain the model.
2. The data augmentation method based on bidirectional back-translation adversarial networks according to claim 1, characterized in that: Step 1 includes: Step 1.1: Construct a module capable of back-translation in two directions, including the BT module, which translates from the source language to the target language and back to the source language. s ={M s→t M t→s }, and the BT back-translation module that translates from the target language to the source language and back to the target language. t ={M t→s M s→t }; where M s→t This refers to a model that translates a source language sentence s into a target language sentence t, while M... t→s This refers to a model that translates a target language sentence t back into a source language sentence s; Step 1.2: Take the set of sentences to be attacked, D = {s, t}, and generate a new parallel sentence pair through the back-translation process BT(·), denoted as the back-translation instance D' = {s′, t′}, where s is the source language sentence, and s' is the sentence generated by s through BT. s The reconstructed translation, where t is the target language sentence and t' is the target language sentence translated via BT. t Translation after process reconstruction.
3. The data augmentation method based on bidirectional back-translation adversarial networks according to claim 1, characterized in that: Step 2 includes: Step 2.1: By constructing example sentences for reverse translation, the reconstructed translation D' is compared with D to calculate the model's translation bias. By analyzing this bias, the accuracy and reliability of the model's sentence-to-sentence translation ability can be effectively evaluated. Step 2.2: When the similarity between D' and D is high, D' is considered a good adversarial instance and is selected as the adversarial target. It will then be used to perform adversarial attacks through the adversarial attack module to enhance the corpus data. The similarity calculation standard is the calculation of the BLEU value.
4. The data augmentation method based on bidirectional back-translation adversarial networks according to claim 1, characterized in that: In Step 3, when the back-translation instance D′ selected from the corpus passes the screening, it will enter the adversarial attack module for attack, and the adversarial instance D” will be generated by using the adversarial attack module; the adversarial attack module in data augmentation includes the source language masking language model MLM and the source language → target language translation direction translation model TLM; In generating adversarial instances, a masked language model (MLM) is used to determine the most vulnerable parts of the example, and then the target words are replaced by a translation language model (TLM) in the source language → target language translation direction.
5. The data augmentation method based on bidirectional back-translation adversarial networks according to claim 4, characterized in that: Step 3 includes: Step 3.1: The Masked Language Model (MLM) attacks the source language sentence s by randomly masking words at random positions in the input sentence and then using contextual information to predict the words at the masked positions. Step 3.2: Utilize the cross-language features of TLM to assist in generating adversarial instances t'′ on the target language side; after obtaining a set of replacement words, determine the attack location and then attack sentence pair D.
6. The data augmentation method based on bidirectional back-translation adversarial networks according to claim 5, characterized in that: In Step 3.2, the attack includes the following steps: Step 3.2.1: Determine the source language sentence s = {x1, ..., x...} n ,…,x l The word attack position n in}; Using a statistical phrase alignment tool, alignment relationships between source and target language sentences in sentence pair D are generated, and the attacked source language word x is then identified through these alignment relationships. n ; Step 3.2.2, Attack source language sentence s; The source language sentence s is used to obtain the attacked source language word x using an alignment tool. n Attack position n, replace the word at attack position n with the new word x″ predicted by MLM. n The resulting new sentence in the source language, 's', is formed after the confrontation. Step 3.2.3: Identify the source language word x being attacked. n Alignment relationships {y1,…,y1,…,y2,…3 ... m ,…,y p ,…,y l }; Determine the attack position mp of the target language, and obtain the attack position mp of the target language corresponding to the word attack position n of the attack vocabulary in the source language through alignment relationship; Step 3.2.4, Attack the target language sentence t; After obtaining the target language attack position as mp, for the target language sentence t={y1,…,y m ,…,y p ,…,y l The masking operation is performed, replacing the target phrase with the [MASK] marker, resulting in the target language sentence t with the mask marker. m By splicing the new source language sentence "s" after adversarial processing and the target language sentence "t" with masked tags... m The input sequence that constitutes a TLM t ={x1,…,x″ n ,…,x l ,[SEP],y1,…,y m-1 [MASK],y p ,…,y l }; By using the cross-lingual prediction capabilities of TLM to predict mask tokens in the target language, the output of TLM is obtained, which is the new sentence t'′ in the target language after adversarial processing; Step 3.2.5: By combining the new source language sentence s” after adversarial communication with the new target language sentence t'′ after adversarial communication, the final adversarial instance D” is obtained. After the above steps, we obtain the new source language sentence s” predicted by MLM and the new target language sentence t’′ predicted by TLM. Finally, the stability of the newly generated sentence pairs is evaluated by the iterative back-translation evaluation model. When the stability is poor, the adversarial instance does not meet the similarity principle of the corpus and will be discarded.
7. A data augmentation system based on bidirectional back-translation adversarial networks, characterized in that, The system includes a module for performing the data augmentation method based on bidirectional back-translation adversarial networks as described in any one of claims 1 to 6.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the data augmentation method based on bidirectional back-translation adversarial networks as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the data augmentation method based on bidirectional back-translation adversarial networks as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the data augmentation method based on bidirectional back-translation adversarial networks as described in any one of claims 1 to 6.