Corpus enhancement method and apparatus, electronic device, and storage medium

By generating positive and negative adversarial samples in neural machine translation, the problem of corpus quality fluctuation caused by word substitution is solved, and the robustness and accuracy of the translation model are improved.

CN115510879BActive Publication Date: 2026-05-22AGRICULTURAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
AGRICULTURAL BANK OF CHINA
Filing Date
2022-10-20
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

In existing technologies, the quality of enhanced corpora generated through word substitution fluctuates greatly, affecting the robustness and accuracy of neural machine translation models.

Method used

By obtaining candidate replacement phrases for the phrases to be replaced in the original corpus, adversarial samples of the source language are generated, as well as positive and negative adversarial samples. Based on these samples, the augmented corpus is determined, thereby improving the stability of the corpus quality and the robustness of the translation model.

Benefits of technology

It improves the quality and stability of the augmented corpus, enhances the robustness and accuracy of the translation model, and solves the problem of decreased accuracy of the translation model caused by word substitution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115510879B_ABST
    Figure CN115510879B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a corpus enhancement method and device, electronic equipment and a storage medium. At least one candidate replacement phrase of a source language to-be-replaced phrase in an original corpus is obtained; an adversarial sample of the source language is generated according to the candidate replacement phrase; a forward adversarial sample and a reverse adversarial sample are generated according to the adversarial sample; and an enhanced corpus is determined according to the original corpus, the forward adversarial sample and the reverse adversarial sample. Embodiments of the present application improve the quality of the enhanced corpus, and further improve the robustness and accuracy of the translation model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to data augmentation technology, and more particularly to a corpus augmentation method, apparatus, electronic device, and storage medium. Background Technology

[0002] In recent years, Neural Machine Translation (NMT) has demonstrated advanced performance in translating many language pairs. As the application of NMT deepens, higher demands are being placed on translation accuracy.

[0003] The accuracy of neural machine translation heavily relies on parallel corpora. Therefore, corpus augmentation is necessary. Typically, word substitution is used for corpus augmentation. However, word substitution results in highly variable quality of the augmented corpus, affecting the robustness of the translation model and consequently reducing its accuracy. Summary of the Invention

[0004] This application provides a corpus enhancement method, apparatus, electronic device, and storage medium to improve the quality of the enhanced corpus and enhance the robustness and accuracy of the translation model.

[0005] In a first aspect, embodiments of this application provide a corpus enhancement method, which includes:

[0006] Obtain at least one candidate replacement phrase for the phrase to be replaced in the source language of the original corpus;

[0007] Based on the candidate replacement phrases, generate adversarial examples of the source language;

[0008] Generate positive and negative adversarial samples based on adversarial samples;

[0009] The augmented corpus is determined based on the original corpus, positive adversarial samples, and negative adversarial samples.

[0010] Secondly, embodiments of this application also provide a corpus enhancement device, which includes:

[0011] The candidate replacement phrase acquisition module is used to acquire at least one candidate replacement phrase for the phrase to be replaced in the source language of the original corpus.

[0012] The adversarial sample generation module is used to generate adversarial samples of the source language based on candidate replacement phrases;

[0013] A bidirectional translation module is used to generate positive and negative adversarial samples based on adversarial samples;

[0014] The augmented corpus determination module is used to determine the augmented corpus based on the original corpus, positive adversarial samples, and negative adversarial samples.

[0015] Thirdly, embodiments of this application also provide an electronic device, which includes:

[0016] One or more processors;

[0017] Storage device for storing one or more programs;

[0018] When one or more programs are executed by one or more processors, the one or more processors implement any of the corpus enhancement methods provided in the embodiments of this application.

[0019] Fourthly, embodiments of this application also provide a storage medium including computer-executable instructions, which, when executed by a computer processor, are used to perform any of the corpus enhancement methods provided in embodiments of this application.

[0020] This application obtains at least one candidate replacement phrase for the source language phrase to be replaced in the original corpus; generates adversarial examples of the source language based on the candidate replacement phrases, generating adversarial examples at the phrase level, which can improve the stability of the candidate augmented corpus quality; generates positive and negative adversarial examples based on the adversarial examples, generating bidirectional adversarial examples, which can improve the robustness of the translation model trained on the augmented corpus; and determines the augmented corpus based on the original corpus, positive adversarial examples, and negative adversarial examples, enriching the corpus samples and improving the accuracy of the translation model trained on the augmented corpus. Therefore, the technical solution of this application solves the problem that the quality of the augmented corpus generated by word replacement fluctuates greatly, which affects the robustness of the translation model and thus reduces the accuracy of the translation model, achieving the effect of improving the stability of the augmented corpus quality and improving the robustness and accuracy of the translation model. Attached Figure Description

[0021] Figure 1 This is a flowchart of a corpus enhancement method according to Embodiment 1 of this application;

[0022] Figure 2 This is a flowchart of a corpus enhancement method according to Embodiment 2 of this application;

[0023] Figure 3 This is a flowchart of a corpus enhancement method according to Embodiment 3 of this application;

[0024] Figure 4 This is a flowchart of a corpus enhancement method according to Embodiment 4 of this application;

[0025] Figure 5 This is a schematic diagram of the structure of a corpus enhancement device according to Embodiment 5 of this application;

[0026] Figure 6This is a schematic diagram of the structure of an electronic device according to Embodiment Six of this application. Detailed Implementation

[0027] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0028] It should be noted that the terms "first" and "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0029] Example 1

[0030] Figure 1 This is a flowchart of a corpus augmentation method provided in Embodiment 1 of this application. This embodiment can be applied to the case of data augmentation of the corpus used to train a language translation model. The method can be executed by a corpus augmentation device, which can be implemented in software and / or hardware and specifically configured in an electronic device, such as a computer.

[0031] See Figure 1 The corpus enhancement method shown includes the following steps:

[0032] S110. Obtain at least one candidate replacement phrase for the phrase to be replaced in the source language of the original corpus.

[0033] The original corpus can be parallel corpora, including two different languages ​​with the same meaning. The original corpus can be used to train a language translation model to improve its accuracy. The source language is the original sentence obtained in another language. That is, the source language is the language to be translated during the translation process, or the language input to the language translation model. The phrase to be replaced can be a phrase in the source language that can be substituted, used to generate adversarial examples in the source language. The phrase to be replaced can be a phrase in a sentence that is susceptible to being replaced. Candidate replacement phrases can be phrases that replace phrases in the source language. For example, candidate replacement phrases can be generated using a trained language translation model to improve generation efficiency. For example, BERT (Bidirectional Encoder Representations from Transformer, a pre-trained language model) can be used. Each phrase to be replaced can have multiple candidate replacement phrases. Each sentence in the source language can include multiple phrases to be replaced.

[0034] S120. Generate adversarial examples of the source language based on the candidate replacement phrases.

[0035] Adversarial examples of the source language can be sentences that replace the original phrases in the source language with candidate replacement phrases. For example, to ensure semantic similarity, candidate replacement phrases can be filtered, replacing the original phrases in the source language with candidate phrases that meet preset conditions. For instance, to ensure translation accuracy, gradients can be used as a preset condition to filter candidate replacement phrases.

[0036] S130. Generate positive adversarial samples and negative adversarial samples based on adversarial samples.

[0037] Positive adversarial examples can be formed by using adversarial examples in the source language as input to a translation model to obtain the output, and using the input as the source language and the output as the target language. Reverse adversarial examples can be formed by using adversarial examples in the source language to translate the target language, and using the output as the source language and the input as the target language. Reverse adversarial examples perturb the source language and can be used to enhance the robustness of the translation.

[0038] S140. Determine the augmented corpus based on the original corpus, positive adversarial samples, and negative adversarial samples.

[0039] Positive and negative adversarial samples enrich the original corpus, but may introduce semantic biases. Therefore, it is necessary to select samples with semantic similarity meeting preset criteria from the positive and negative adversarial samples, based on the original corpus, as augmented samples. For example, the augmented corpus can be determined based on the original corpus, positive and negative adversarial samples, using semantic enhancement methods, to ensure that the meanings of the positive and negative adversarial samples in the augmented corpus maintain a higher semantic similarity to the original corpus.

[0040] The technical solution of this embodiment obtains at least one candidate replacement phrase for the phrase to be replaced in the source language from the original corpus; generates adversarial samples of the source language based on the candidate replacement phrases, generating adversarial samples at the phrase level, which can improve the stability of the candidate augmented corpus quality; generates positive and negative adversarial samples based on the adversarial samples, generating bidirectional adversarial samples, which can improve the robustness of the translation model trained on the augmented corpus; and determines the augmented corpus based on the original corpus, positive adversarial samples, and negative adversarial samples, enriching the corpus samples and improving the accuracy of the translation model trained on the augmented corpus. Therefore, the technical solution of this application solves the problem that the quality of the augmented corpus generated by word replacement fluctuates greatly, which affects the robustness of the translation model and thus reduces the accuracy of the translation model, achieving the effect of improving the stability of the augmented corpus quality and improving the robustness and accuracy of the translation model.

[0041] Example 2

[0042] Figure 2 This is a flowchart of a corpus enhancement method provided in Embodiment 2 of this application. The technical solution of this embodiment is further refined based on the above technical solution.

[0043] Furthermore, "obtaining at least one candidate replacement phrase of the source language in the original corpus" is refined to: "determining at least one source language in the sample corpus; generating at least one candidate replacement phrase of at least one source language in the sample corpus through a phrase generation model" to determine the candidate replacement phrase;

[0044] Furthermore, the step of “generating adversarial samples of the source language based on candidate replacement phrases” is further refined into: “determining the phrase selection parameters of each candidate replacement phrase; wherein, the phrase selection parameters include word features and / or gradient norm; selecting the target replacement phrase from each candidate replacement phrase based on the phrase selection parameters of the candidate replacement phrases; replacing the phrase to be replaced in the original corpus with the target replacement phrase to obtain adversarial samples of the original corpus”, in order to generate adversarial samples of the original corpus.

[0045] See Figure 2 One corpus enhancement method shown includes:

[0046] S210. Identify at least one phrase in the source language of the sample corpus that needs to be replaced.

[0047] The phrases to be replaced are phrases in easily affected positions in the source language of the sample corpus. Randomly selecting positions to replace phrases in certain sentences may lead to worse translation performance of the trained translation model. Therefore, phrases in easily affected positions within the sentence are selected as the phrases to be replaced. For example, the phrases to be replaced can be determined based on their gradients. For instance, a preset gradient threshold of 0.5 is used; if the gradient of a phrase is greater than or equal to 0.5, the phrase is identified as a phrase to be replaced; if the gradient of a phrase is less than 0.5, the phrase is no longer identified as a phrase to be replaced.

[0048] In one optional implementation example, determining at least one phrase to be replaced in the source language of the sample corpus includes: obtaining the gradient norm corresponding to each phrase in the source language of the sample corpus; and determining the phrase to be replaced in the source language of the sample corpus based on a preset gradient norm threshold and the gradient norm.

[0049] Different positions within a sentence have different gradient norms; a larger gradient norm indicates instability at that position. Therefore, the phrase to be replaced can be determined by obtaining the gradient norms of phrases at different positions. Specifically, the gradient norm of each phrase in the source language of the sample corpus can be calculated using a mathematical model. A preset gradient norm threshold can be used to determine the phrases to be replaced in the source language of the sample corpus. For example, if the gradient norm of each phrase in the source language of the sample corpus is greater than or equal to the preset gradient norm threshold, then that phrase is determined to be replaced; if the gradient norm of each phrase in the source language of the sample corpus is less than the preset gradient norm threshold, then that phrase is no longer determined to be replaced.

[0050] Obtain the gradient norm of each phrase in the source language of the sample corpus; determine the phrases to be replaced in the source language of the sample corpus based on the preset gradient norm threshold and gradient norm, which can improve the quality of adversarial samples obtained by replacing the phrases to be replaced.

[0051] S220. Generate at least one candidate replacement phrase for at least one phrase to be replaced using a phrase generation model.

[0052] Phrase generation models can be pre-trained language models used to generate candidate replacement phrases for a phrase to be replaced. A high-quality language model requires billions of monolingual data points for training, and training the model consumes significant time and computational resources. For example, a BERT pre-trained model can be used as a language generation model to generate at least one candidate replacement phrase for at least one phrase to be replaced. A sentence can have at least one phrase to be replaced, and each phrase to be replaced can have at least one candidate replacement phrase.

[0053] S230. Determine the phrase selection parameters for each candidate replacement phrase; wherein, the phrase selection parameters include wording features and / or gradient norm.

[0054] Phrase selection parameters can be used to evaluate candidate replacement phrases, providing a basis for subsequently selecting candidate replacement phrases to replace the phrase to be replaced. For example, phrase selection parameters include word features and / or gradient norm.

[0055] Phrasing features are used to represent the semantic features of phrases and can be obtained through the average word embedding method. Gradient norm is used to represent the degree of influence of a phrase on a sentence and can be calculated through a mathematical model.

[0056] S240. Select the target replacement phrase from each candidate replacement phrase according to the phrase selection parameters of the candidate replacement phrases.

[0057] The target replacement phrase can be a phrase selected from among the candidate replacement phrases, used to replace the phrase to be replaced at the corresponding position in the source language of the sample corpus. Based on phrase selection parameters, the candidate replacement phrases with more similar semantics are selected as the target replacement phrase. For example, candidate replacement phrases with similar semantics can be selected based on wording features, and further, based on a pre-set gradient norm threshold, candidate replacement phrases greater than or equal to the pre-set gradient norm threshold are selected as the target replacement phrase.

[0058] S250. Replace the phrase to be replaced in the original corpus with the target replacement phrase to obtain adversarial samples of the original corpus.

[0059] Adversarial examples of the original corpus can be the corpus after replacing the phrase to be replaced in the original corpus with the target phrase. For example, replacing the phrase to be replaced in the original corpus with at least one target replacement phrase yields at least one adversarial example.

[0060] S260. Generate positive adversarial samples and negative adversarial samples based on adversarial samples.

[0061] S270. Determine the augmented corpus based on the original corpus, positive adversarial samples, and negative adversarial samples.

[0062] The technical solution of this embodiment, by identifying at least one phrase to be replaced in the source language of the sample corpus, and generating at least one candidate replacement phrase for the at least one phrase to be replaced using a phrase generation model, can quickly generate high-quality candidate replacement phrases, improving the efficiency and quality of candidate replacement phrase generation, and thus improving the efficiency and quality of adversarial samples generated from the original corpus. Simultaneously, generating phrase-level candidate replacement phrases can improve the quality of the augmented samples. By determining the phrase selection parameters for each candidate replacement phrase, including word choice features and / or gradient norm, and selecting a target replacement phrase from each candidate replacement phrase according to the phrase selection parameters, the phrase to be replaced in the original corpus is replaced with the target replacement phrase, resulting in adversarial samples of the original corpus. Selecting the target replacement phrase from the dimensions of word choice features and / or gradient norm improves the semantic similarity of the target replacement phrase and the semantic similarity of the obtained adversarial samples of the original corpus, thereby improving the quality of the adversarial samples of the original corpus.

[0063] Example 3

[0064] Figure 3 This is a flowchart of a corpus enhancement method provided in Embodiment 3 of this application. The technical solution of this embodiment is further refined based on the above technical solution.

[0065] Furthermore, the process of "generating positive and negative adversarial samples based on adversarial samples" is further refined into: "translating the source language of the adversarial sample into the target language to obtain the first translation result; generating positive adversarial samples based on the adversarial samples and the first translation result; translating the first translation result back into the source language to obtain the second translation result; and generating negative adversarial samples based on the adversarial samples and the second translation result," thereby generating negative adversarial samples and further enriching the original corpus.

[0066] See Figure 3 One corpus enhancement method shown includes:

[0067] S310. Obtain at least one candidate replacement phrase for the phrase to be replaced in the source language of the original corpus.

[0068] S320. Generate adversarial examples of the source language based on candidate replacement phrases.

[0069] S330. Translate the source language of the adversarial sample into the target language to obtain the first translation result.

[0070] The source language can be the language to be translated in the original corpus. For example, if the translation process is English to Chinese, then the source language is English. The target language can be the target language to be translated from the original corpus. For example, if the translation process is English to Chinese, then the target language is Chinese. The first translation result can be the translation result obtained by translating the source language in the adversarial sample. Specifically, the first translation result can be obtained by translating the source language according to a general encoder and decoder.

[0071] S340. Based on the adversarial examples and the first translation result, generate positive adversarial examples.

[0072] Positive adversarial examples can be adversarial examples obtained from the translation direction from the source language to the target language. The sample formed by the source language and the first translation result in the adversarial example is regarded as a positive adversarial example.

[0073] S350. Translate the first translation result into the source language to obtain the second translation result.

[0074] The second translation result can be the translation obtained by inputting the first translation result into a universal encoder and decoder. The universal encoder and decoder can perform bidirectional translation, that is, it can translate from the source language to the target language, or it can translate from the target language to the source language.

[0075] S360 generates reverse adversarial examples based on adversarial examples and second translation results.

[0076] The second translation result is used as the source language, and the language corresponding to the second translation result is used as the target language. Reverse adversarial samples are generated to ensure that the source and target languages ​​in the reverse adversarial samples, forward adversarial samples and the original corpus are consistent.

[0077] S370. Determine the augmented corpus based on the original corpus, positive adversarial samples, and negative adversarial samples.

[0078] The technical solution of this embodiment involves translating the source language of the adversarial sample into the target language to obtain a first translation result; generating a positive adversarial sample based on the adversarial sample and the first translation result; translating the first translation result back into the source language to obtain a second translation result; and generating a negative adversarial sample based on the adversarial sample and the second translation result. Through a bidirectional generation method, adversarial samples in both the source-to-target and target-to-source directions are generated from the original corpus. The adversarial pair of target-to-source transformation is a slight perturbation to the original data, which can be used to improve the robustness of source-to-target transformation.

[0079] Example 4

[0080] Figure 4This is a flowchart of a corpus enhancement method provided in Embodiment 4 of this application. The technical solution of this embodiment is further refined based on the above technical solution.

[0081] Furthermore, the phrase "determine the augmented corpus based on the original corpus, positive adversarial samples, and negative adversarial samples" is further refined to: "Based on the semantic encoder, establish the adjacency semantic spaces of the original corpus, positive adversarial samples, and negative adversarial samples respectively; based on the Gaussian cyclic chain mixture algorithm, determine the augmented corpus from the original corpus, positive adversarial samples, and negative adversarial samples according to each adjacency semantic space" to determine the augmented corpus.

[0082] See Figure 4 One corpus enhancement method shown includes:

[0083] S410. Obtain at least one candidate replacement phrase for the phrase to be replaced in the source language of the original corpus.

[0084] S420. Generate adversarial examples of the source language based on candidate replacement phrases.

[0085] S430. Generate positive adversarial samples and negative adversarial samples based on adversarial samples.

[0086] S440. Based on the semantic encoder, the adjacency semantic space of the original corpus, positive adversarial samples, and negative adversarial samples are established respectively.

[0087] A semantic encoder can be used to establish a unified semantic space that contains both the source and target languages. The adjacency semantic space can be centered on the observed original corpus, positive adversarial samples, and negative adversarial samples, with a sufficient number of synonym variants to determine the augmented corpus. Inputting the original corpus, positive adversarial samples, and negative adversarial samples into the semantic encoder yields the adjacency semantic space of the original corpus, positive adversarial samples, and negative adversarial samples.

[0088] In one optional embodiment, based on a semantic encoder, adjacency semantic spaces for the original corpus, positive adversarial samples, and negative adversarial samples are established respectively, including: obtaining an optimized semantic encoder based on tangent learning; and establishing adjacency semantic spaces for the original corpus, positive adversarial samples, and negative adversarial samples through the optimized semantic encoder.

[0089] Tangent learning can be used as an optimization algorithm to improve the accuracy of adjacent semantic regions built by the semantic encoder. It treats the tangent points of adjacent semantic regions as critical states of semantic equivalence, making it easy for vectors in adjacent semantic regions to cover multiple synonymous variants with the same semantic meaning. The optimized semantic encoder is obtained through tangent learning. The original corpus, positive adversarial examples, and negative adversarial examples are then input into the optimized semantic encoder to obtain the optimized adjacent semantic space of the original corpus, positive adversarial examples, and negative adversarial examples.

[0090] By learning based on tangents, an optimized semantic encoder is obtained. Using the optimized semantic encoder, an adjacency semantic space is established for the original corpus, positive adversarial samples, and negative adversarial samples. This can improve the semantic similarity in the adjacency semantic space and improve the quality of the subsequently determined augmented corpus.

[0091] S450, based on the Gaussian Mixture Cyclic Chain algorithm, determines the enhanced corpus from the original corpus, positive adversarial samples, and negative adversarial samples according to each adjacent semantic space.

[0092] The Gaussian Mixture Recurrent Chain Algorithm (GSM-LOC) can be used to simulate a recurrent chain that generates a reliable sequence of vectors, where the current vector depends on the previous vector. This dependency relationship allows the Gaussian distribution to reach a steady state. Sampling stops when identical samples are encountered, and the algorithm is used to obtain augmented corpora from adjacent semantic regions. Specifically, the GSM-LOC samples a set of vectors from adjacent semantic regions, merges each sampled vector into the decoder through a broadcast ensemble network, and determines the augmented corpora from the original corpus, positive adversarial examples, and negative adversarial examples. The GSM-LOC transforms a discrete space into a continuous space, thereby improving the model's generalization ability.

[0093] In the context of natural language processing, discrete operations such as adding, deleting, reordering, or replacing words in the original sentence often lead to significant semantic changes. Therefore, by combining phrase-level positive and negative adversarial sample generation methods, continuous semantic augmentation is performed to address the lack of diversity in augmented training samples in discrete space and the difficulty in preserving the original meaning of augmented text in discrete space.

[0094] The technical solution of this embodiment establishes adjacency semantic spaces for the original corpus, positive adversarial samples, and negative adversarial samples based on a semantic encoder. Then, based on a Gaussian mixture cyclic chain algorithm, it determines augmented corpus from the original corpus, positive adversarial samples, and negative adversarial samples according to each adjacency semantic space. By performing continuous semantic augmentation on the original corpus, positive adversarial samples, and negative adversarial samples through the semantic encoder and the Gaussian mixture cyclic chain algorithm, the quality of the augmented corpus is improved, and the generalization ability of the model trained using the augmented corpus is enhanced.

[0095] Example 5

[0096] Figure 5 The diagram shown is a structural schematic of a corpus enhancement device provided in Embodiment 5 of this application. This embodiment is applicable to the case of data augmentation of the corpus used in training a language translation model. The specific structure of the corpus enhancement device is as follows:

[0097] The candidate replacement phrase acquisition module 510 is used to acquire at least one candidate replacement phrase of the phrase to be replaced in the source language of the original corpus.

[0098] The adversarial sample generation module 520 is used to generate adversarial samples of the source language based on candidate replacement phrases;

[0099] The bidirectional translation module 530 is used to generate positive and negative adversarial samples based on adversarial samples;

[0100] The augmented corpus determination module 540 is used to determine the augmented corpus based on the original corpus, positive adversarial samples, and negative adversarial samples.

[0101] The technical solution of this embodiment obtains at least one candidate replacement phrase for the phrase to be replaced in the source language from the original corpus through a candidate replacement phrase acquisition module; generates adversarial samples in the source language based on the candidate replacement phrases through an adversarial sample generation module, generating adversarial samples at the phrase level, which can improve the stability of the candidate augmented corpus quality; generates positive and negative adversarial samples based on the adversarial samples through a bidirectional translation module, which can improve the robustness of the translation model trained on the augmented corpus; and determines the augmented corpus based on the original corpus, positive and negative adversarial samples through an augmented corpus determination module, enriching the corpus samples and improving the accuracy of the translation model trained on the augmented corpus. Therefore, the technical solution of this application solves the problem that the quality of the augmented corpus generated by word replacement fluctuates greatly, which affects the robustness of the translation model and thus reduces the accuracy of the translation model, achieving the effect of improving the stability of the augmented corpus quality and improving the robustness and accuracy of the translation model.

[0102] Optionally, the candidate replacement phrase acquisition module 510 includes:

[0103] The phrase to be replaced unit is used to determine at least one phrase to be replaced in the source language of the sample corpus;

[0104] A candidate replacement phrase generation unit is used to generate at least one candidate replacement phrase for at least one phrase to be replaced by a phrase generation model.

[0105] Optionally, the phrase to be replaced determination unit includes:

[0106] The gradient norm determination subunit is used to obtain the gradient norm of each phrase in the source language of the sample corpus.

[0107] The gradient norm comparison subunit is used to determine the phrases to be replaced in the source language of the sample corpus based on a preset gradient norm threshold and gradient norm.

[0108] Optionally, the adversarial example generation module 520 includes:

[0109] A parameter selection unit is selected to determine the phrase selection parameters for each candidate replacement phrase; wherein, the phrase selection parameters include wording features and / or gradient norm;

[0110] The target replacement phrase selection unit is used to select the target replacement phrase from each candidate replacement phrase according to the phrase selection parameters of the candidate replacement phrases;

[0111] The phrase replacement unit is used to replace the phrase to be replaced in the original corpus with the target replacement phrase, thus obtaining adversarial samples of the original corpus.

[0112] Optional, the bidirectional translation module 530 includes:

[0113] The first translation result acquisition unit is used to translate the source language of the adversarial sample into the target language to obtain the first translation result;

[0114] A positive adversarial example generation unit is used to generate positive adversarial examples based on adversarial examples and the first translation result;

[0115] The second translation result generation unit is used to translate the first translation result into the source language to obtain the second translation result.

[0116] The reverse adversarial sample generation unit is used to generate reverse adversarial samples based on adversarial samples and second translation results.

[0117] Optional, the enhanced corpus determination module 540 includes:

[0118] The adjacency semantic space establishment unit is used to establish the adjacency semantic space of the original corpus, positive adversarial samples, and negative adversarial samples based on the semantic encoder.

[0119] The augmented corpus determination unit is used to determine the augmented corpus from the original corpus, positive adversarial samples, and negative adversarial samples based on the Gaussian mixture cyclic chain algorithm and each adjacent semantic space.

[0120] Optionally, the adjacency semantic space building unit includes:

[0121] The semantic encoder optimization subunit is used to obtain the optimized semantic encoder based on tangent learning;

[0122] The adjacency semantic space establishment subunit is used to establish the adjacency semantic space of the original corpus, positive adversarial samples, and negative adversarial samples through the optimized semantic encoder.

[0123] The corpus enhancement apparatus provided in this application embodiment can execute the corpus enhancement method provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects for executing the corpus enhancement method.

[0124] Example 6

[0125] Figure 6 This is a schematic diagram of the structure of an electronic device provided in Embodiment Six of this application, as shown below. Figure 6 As shown, the electronic device includes a processor 610, a memory 620, an input device 630, and an output device 640; the number of processors 610 in the electronic device can be one or more. Figure 6 Taking a processor 610 as an example; the processor 610, memory 620, input device 630, and output device 640 in the electronic device can be connected via a bus or other means. Figure 6 Taking the example of a connection between China and Israel via a bus.

[0126] The memory 620, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the corpus enhancement method in the embodiments of this application (e.g., candidate replacement phrase acquisition module 510, adversarial sample generation module 520, bidirectional translation module 530, and enhanced corpus determination module 540). The processor 610 executes various functional applications and data processing of the electronic device by running the software programs, instructions, and modules stored in the memory 620, thereby implementing the aforementioned corpus enhancement method.

[0127] The memory 620 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a given function; the data storage area may store data created based on terminal usage. Furthermore, the memory 620 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 620 may further include memory remotely located relative to the processor 610, which can be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0128] Input device 630 can be used to receive input character information and generate key signal inputs related to user settings and function control of the electronic device. Output device 640 may include display devices such as a display screen.

[0129] Example 7

[0130] Embodiment 7 of this application also provides a storage medium containing computer-executable instructions. When executed by a computer processor, the computer-executable instructions are used to perform a corpus enhancement method. The method includes: obtaining at least one candidate replacement phrase of the source language phrase to be replaced in the original corpus; generating adversarial samples of the source language based on the candidate replacement phrases; generating positive adversarial samples and negative adversarial samples based on the adversarial samples; and determining the enhanced corpus based on the original corpus, the positive adversarial samples, and the negative adversarial samples.

[0131] Of course, the computer-executable instructions provided in the embodiments of this application are not limited to the method operations described above, but can also perform related operations in the corpus enhancement method provided in any embodiment of this application.

[0132] Based on the above description of the implementation methods, those skilled in the art can clearly understand that this application can be implemented using software and necessary general-purpose hardware, and of course, it can also be implemented using hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0133] It is worth noting that in the embodiments of the search device described above, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of this application.

[0134] Note that the above are merely preferred embodiments and the technical principles employed in this application. Those skilled in the art will understand that this application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of this application. Therefore, although this application has been described in detail through the above embodiments, this application is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of this application, the scope of which is determined by the scope of the appended claims.

Claims

1. A corpus enhancement method, characterized in that, include: Obtain at least one candidate replacement phrase for the phrase to be replaced in the source language of the original corpus; Based on the candidate replacement phrases, generate adversarial examples of the source language; Based on the adversarial samples, generate positive and negative adversarial samples; The enhanced corpus is determined based on the original corpus, the positive adversarial samples, and the negative adversarial samples; The step of determining the enhanced corpus based on the original corpus, the positive adversarial samples, and the negative adversarial samples includes: Based on the semantic encoder, the adjacency semantic space of the original corpus, the positive adversarial sample, and the negative adversarial sample are established respectively. Based on the Gaussian Mixture Cyclic Chain Algorithm, an enhanced corpus is determined according to the adjacency semantic space, the original corpus, the positive adversarial samples, and the negative adversarial samples.

2. The method according to claim 1, characterized in that, The step of obtaining at least one candidate replacement phrase for the phrase to be replaced in the source language of the original corpus includes: Identify at least one phrase in the source language of the original corpus that needs to be replaced; At least one candidate replacement phrase is generated by the phrase generation model for the at least one phrase to be replaced.

3. The method according to claim 2, characterized in that, The step of determining at least one phrase to be replaced in the source language of the original corpus includes: Obtain the gradient norm of each phrase in the source language of the original corpus; Based on the preset gradient norm threshold and the gradient norm, the phrases to be replaced in the source language of the original corpus are determined.

4. The method according to claim 1, characterized in that, The step of generating adversarial samples of the original corpus based on the candidate replacement phrases includes: Determine the phrase selection parameters for each of the candidate replacement phrases; wherein the phrase selection parameters include wording features and / or gradient norm; Based on the phrase selection parameters of the candidate replacement phrases, a target replacement phrase is selected from each of the candidate replacement phrases; The phrase to be replaced in the original corpus is replaced with the target replacement phrase to obtain adversarial samples of the original corpus.

5. The method according to any one of claims 1-4, characterized in that, The generation of positive and negative adversarial samples based on the adversarial samples includes: The source language of the adversarial sample is translated into the target language to obtain the first translation result; Based on the adversarial sample and the first translation result, a positive adversarial sample is generated; The first translation result is translated into the source language to obtain the second translation result; Based on the adversarial sample and the second translation result, a reverse adversarial sample is generated.

6. The method according to claim 1, characterized in that, The step of establishing adjacency semantic spaces for the original corpus, the positive adversarial samples, and the negative adversarial samples based on a semantic encoder includes: Based on tangent learning, an optimized semantic encoder is obtained; The optimized semantic encoder is used to establish the adjacency semantic space of the original corpus, the positive adversarial sample, and the negative adversarial sample.

7. A corpus enhancement device, characterized in that, include: The candidate replacement phrase acquisition module is used to acquire at least one candidate replacement phrase for the phrase to be replaced in the source language of the original corpus. An adversarial example generation module is used to generate adversarial examples of the source language based on the candidate replacement phrases; A bidirectional translation module is used to generate positive and negative adversarial samples based on the adversarial samples. An enhanced corpus determination module is used to determine enhanced corpus based on the original corpus, the positive adversarial sample, and the negative adversarial sample. The enhanced corpus determination module includes: The adjacency semantic space establishment unit is used to establish the adjacency semantic spaces of the original corpus, the positive adversarial sample, and the negative adversarial sample based on the semantic encoder. The augmented corpus determination unit is used to determine augmented corpus based on the Gaussian Mixture Cyclic Chain algorithm, according to each of the adjacency semantic spaces, the original corpus, the positive adversarial samples, and the negative adversarial samples.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements a corpus enhancement method as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements a corpus enhancement method as described in any one of claims 1-6.