A method, device and electronic device for generating encoded text based on adversarial training
By optimizing the adversarial generation network and performing gradient attacks separately, the robustness problem caused by dropout fixation in pre-trained models is solved, and the text understanding and coding efficiency of the model is improved, especially in bond trading business.
Patent Information
- Application Number
- CN202210188203.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-28
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-02-28
AI Technical Summary
When the existing adversarial training methods are applied to pre-trained models, the robustness of the model is reduced, especially in pre-trained models such as BERT, fixed dropout mask causes dropout to lose meaning, affecting the robustness and learning ability of the model.
FreeLB adversarial training method is adopted to optimize the adversarial generation network, so that the discard mask in each training is not fixed, and the output distribution of the network is consistent through JS divergence constraints, and gradient attacks are carried out on token embedding and position embedding separately to generate target pre-trained models.
The pre-trained model's understanding of text and coding efficiency are improved, and the model's robustness and adversarial training speed are improved. Especially in the secondary trading business of financial bonds, the accuracy of trading factor extraction is increased by 2%-5%.
Smart Images

Figure CN114692570B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a method, apparatus, and electronic device for generating encoded text based on adversarial training. Background Art
[0002] In the field of deep learning, adversarial training is an effective method to improve the robustness of a model. The adversarial training method was first used in the field of computer vision. GAN is a relatively successful example in adversarial training. It generates samples that can deceive the discriminative network through a generative network, while the discriminative network discriminates the authenticity of the samples generated by the generative network, so that the samples generated by the generative network cannot deceive the discriminative network. Under the interaction of the discriminative network and the generative network, the recognition ability of the discriminative network is continuously improved, which is the idea of adversarial training. In the field of natural language processing, adversarial training mainly perturbs the embedding to generate adversarial samples, making the model more robust in semantic understanding of the embedding.
[0003] There are many existing adversarial training methods for the field of natural language processing, such as FGM, PGD, FreeAT, YOPO, FreeLB, etc. FGM takes the same step in each direction to find the direction with the fastest gradient descent. However, the disadvantage of FGM is that it is difficult to find the optimal point within the constraint by taking only one step. Therefore, using a distributed calculation method when calculating the adversarial perturbation is obviously a good improvement method, which is one of the innovations of PGD, and when the adversarial perturbation exceeds the perturbation radius, it is mapped back to the maximum perturbation sphere. FreeLB believes that it is unreasonable for PGD to only use the gradient of the last step to update the perturbation when updating the perturbation, and the gradient of each step should have an impact on the update of the parameters, rather than accumulating to the last step for update. Therefore, FreeLB weights and averages the gradients obtained in each step and uses this gradient to update the parameters of the model. In addition, when learning the perturbation, FreeLB no longer calculates the gradient with respect to the model parameters, but calculates the gradient with respect to the perturbation. The innovation of YOPO lies in proposing two adversarial regularization losses and decoupling between each layer when optimizing the network parameters. These two innovations improve the generalization ability of the model and the speed of adversarial training.
[0004] Existing adversarial training methods have achieved good results in natural language processing. However, there is a problem when applying them to pre-trained models. That is, pre-trained models such as BERT are essentially a two-stage NLP model. The first stage is called: Pre-training, which is similar to WordEmbedding. It uses existing unlabeled corpora to train a language model. The second stage is called: Fine-tuning. It uses the pre-trained language model to complete specific downstream NLP tasks. The training cost of pre-training is very high. Generally, the models pre-trained by Google are directly used. While the cost of fine-tuning is relatively low. During the fine-tuning stage, dropout is adopted. Dropout means that during the training process of a deep learning network, for neural network units, they are temporarily discarded from the network with a certain probability. Note that it is temporary. For stochastic gradient descent, since it is randomly discarded, each mini-batch is training a different network. Different dropouts will make the structure of each layer of the network different, and the input gradients obtained have a lot of noise. Therefore, to solve this problem, existing adversarial training methods will fix the dropout mask when ascending the gradient at each step. This method of fixing the dropout mask does relieve the noise of the input gradient to a certain extent, but it makes dropout lose its original meaning. Fixing the dropout mask is equivalent to no longer using dropout, reducing the robustness of the model.
[0005] Therefore, when adopting dropout and adversarial training, the JS divergence can be used to make the distributions of the ascending gradients at each step as similar as possible. Secondly, in pre-trained models, the adversarial training method perturbs the embeddings. The embeddings in pre-training are the sum of token embeddings, position embeddings, and segment embeddings. Therefore, the adversarial training perturbs all three of them. However, perturbing the three embeddings simultaneously may prevent the model from learning semantic information well. Because when the token embedding is changed, the position embedding and segment embedding are also being changed, resulting in overly difficult learning and a deterioration in the model's learning ability. Therefore, separately perturbing the three embeddings may be a better choice. In addition, perturbing the segment embedding does not lead to a significant improvement in the model's semantic understanding. So it is not necessary to perturb the segment embedding, which can also improve the speed of adversarial training.
[0006] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention
[0007] In view of the above deficiencies of the existing technology, the present invention provides a method, device and electronic device for generating encoded text based on adversarial training, aiming to solve the problem that when the adversarial training method is applied to a pre-trained model in the existing technology, the robustness of the model is reduced.
[0008] The technical solution of the present invention is as follows:
[0009] The first embodiment of the present invention provides a method for generating encoded text based on adversarial training, the method comprising:
[0010] Pre-construct an adversarial generation network;
[0011] Optimize the adversarial generation network to generate an optimized adversarial generation network, wherein the dropout mask in each training in the optimized adversarial generation network is not fixed;
[0012] Perform adversarial training on the pre-trained model according to the adversarial generation network to generate a target pre-trained model;
[0013] Input the bond information to be processed into the target pre-trained model to generate encoded text.
[0014] Further, the pre-constructing an adversarial generation network includes:
[0015] Pre-construct an adversarial generation network based on the FreeLB adversarial training method.
[0016] Further, the optimizing the adversarial generation network to generate an optimized adversarial generation network, wherein the dropout mask in each training in the optimized adversarial generation network is not fixed, includes:
[0017] Optimize the adversarial generation network based on the FreeLB adversarial training method, modify the position of the fixed model dropout to be non-fixed, so as to realize that the dropout mask in each training in the optimized adversarial generation network is not fixed;
[0018] Use the JS divergence to constrain the output of the model to generate an optimized adversarial generation network.
[0019] Further, the performing adversarial training on the pre-trained model according to the adversarial generation network to generate a target pre-trained model includes:
[0020] Perform adversarial training on the BERT pre-trained model according to the adversarial generation network to generate a target pre-trained model.
[0021] Further, the adversarial training of the BERT pre-trained model according to the adversarial generation network to generate a target pre-trained model includes:
[0022] Obtain bond information samples, input the bond information samples into the input layer of the BERT pre-trained model to generate input texts;
[0023] Input the input texts into the embedding layer of the BERT pre-trained model;
[0024] Perform adversarial training on the embedding layer of the BERT pre-trained model according to the adversarial generation network to generate a target pre-trained model.
[0025] Further, the adversarial training of the embedding layer of the BERT pre-trained model according to the adversarial generation network to generate a target pre-trained model includes:
[0026] Perform two gradient attacks on the embedding layer of the BERT pre-trained model according to the adversarial generation network to generate perturbed gradients;
[0027] Update the parameters of the pre-trained model according to the perturbed gradients to generate a target pre-trained model.
[0028] Further, if the embedding of the BERT pre-trained model consists of token embedding, position embedding, and segment embedding, then performing two gradient attacks on the embedding layer of the BERT pre-trained model according to the adversarial generation network to generate perturbed gradients includes:
[0029] Perform the first gradient attack on the token embedding of the BERT pre-trained model according to the adversarial generation network to generate the first gradient after the first perturbation;
[0030] Perform the second gradient attack on the position embedding of the BERT pre-trained model according to the adversarial generation network to generate the second gradient after the second perturbation.
[0031] Another embodiment of the present invention provides an encoded text generation device based on adversarial training. The device includes:
[0032] A network construction module for pre-constructing an adversarial generation network;
[0033] A network optimization module for optimizing the adversarial generation network to generate an optimized adversarial generation network, where the dropout masks are not fixed in each training in the optimized adversarial generation network;
[0034] An adversarial training module, configured to perform adversarial training on a pre-trained model according to an adversarial generation network to generate a target pre-trained model;
[0035] An encoding module, configured to input the bond information to be processed into the target pre-trained model to generate an encoded text.
[0036] Another embodiment of the present invention provides an electronic device, which includes at least one processor; and,
[0037] A memory communicatively connected to the at least one processor; wherein,
[0038] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above-mentioned method for generating an encoded text based on adversarial training.
[0039] Another embodiment of the present invention further provides a non-volatile computer-readable storage medium, which stores computer-executable instructions, and when the computer-executable instructions are executed by one or more processors, the one or more processors can be enabled to execute the above-mentioned method for generating an encoded text based on adversarial training.
[0040] Advantageous effects: In the embodiments of the present invention, instead of using a fixed dropout mask, a probability distribution measurement index or mean square error, etc. is used to measure the difference between the model outputs of two different dropouts, so that dropout can play its own role and the model outputs of two different dropouts can be kept as consistent as possible, improving the text understanding ability of the pre-trained model and the text encoding efficiency. Description of the Drawings
[0041] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:
[0042] Figure 1 is a flowchart of a preferred embodiment of a method for generating an encoded text based on adversarial training according to the present invention;
[0043] Figure 2 is a schematic diagram of the position of the first gradient attack in a specific application embodiment of a method for generating an encoded text based on adversarial training according to the present invention;
[0044] Figure 3 is a schematic diagram of the position of the second gradient attack in a specific application embodiment of a method for generating an encoded text based on adversarial training according to the present invention;
[0045] Figure 4 is a schematic diagram of the functional modules of a preferred embodiment of a device for generating an encoded text based on adversarial training according to the present invention;
[0046] Figure 5 This is a schematic diagram of the hardware structure of a preferred embodiment of an electronic device according to the present invention. Detailed implementation manners
[0047] To make the objectives, technical solutions and effects of the present invention clearer and more definite, the present invention will be further described in detail below. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0048] The embodiments of the present invention will be introduced below with reference to the accompanying drawings.
[0049] In view of the above problems, an embodiment of the present invention provides a method for generating encoded text based on adversarial training. Please refer to Figure 1 , Figure 1 This is a flowchart of a preferred embodiment of a method for generating encoded text based on adversarial training according to the present invention. As Figure 1 shown, it includes:
[0050] Step S100: Pre-construct an adversarial generation network;
[0051] Step S200: Optimize the adversarial generation network to generate an optimized adversarial generation network, where the dropout mask is not fixed in each training of the optimized adversarial generation network;
[0052] Step S300: Perform adversarial training on the pre-trained model according to the adversarial generation network to generate a target pre-trained model;
[0053] Step S400: Input the bond information to be processed into the target pre-trained model to generate encoded text.
[0054] Specifically, in the embodiment of the present invention, it is applied to bond information extraction. The present invention optimizes the adversarial training method and is applicable to any pre-trained network model similar to the BERT model, and has particularly obvious effects on the text data in the field of spot bond transactions.
[0055] The mathematical formula for adversarial training is as follows:
[0056]
[0057] Among them, E7 (x,y) represents the input vector, minE (x,y):D [] represents the minimization of the mathematical expectation of the model, r advrepresents the perturbation of the word vector, θ represents the parameters of the model, y is the true label. In fact, adversarial training is essentially a min-max problem. This formula is mainly divided into two parts, one is the maximization of the internal loss function, and the other is the minimization of the external empirical risk. Describing the idea of adversarial training in one sentence, it is to perform gradient ascent (increase the loss) on the input to make the input as different from the original as possible, and perform gradient descent (decrease the loss) on the parameters to make the model as able to correctly identify as possible.
[0058] This method no longer fixes the dropout mask, but uses a probability distribution measurement index or mean squared error, etc. to constrain the output differences of the models with two different dropouts, so that dropout can play its own role and also make the outputs of the models with two different dropouts as consistent as possible. Among them, the dropout mask masks the positions of the dropout neurons, similar to adding a neuron, so that the model doesn't know that a neuron is missing here.
[0059] In one embodiment, an adversarial generation network is pre-constructed, including:
[0060] An adversarial generation network based on the FreeLB adversarial training method is pre-constructed.
[0061] In specific implementation, FreeLB is an improvement based on PGD. The Projected Gradient Descent (PGD) method uses multiple perturbations, with each perturbation taking only a small step. When the range after multiple perturbations exceeds the specified perturbation range, it is mapped back into the circle of the perturbation range. However, PGD updates the parameters using only the gradient of the last perturbation, while FreeLB updates the parameters by taking the average gradient of multiple iterations.
[0062] The mathematical formula of PGD is as follows:
[0063]
[0064] The mathematical formula of FreeLB is as follows:
[0065]
[0066] Where K represents the number of perturbations.
[0067] In one embodiment, the adversarial generation network is optimized to generate an optimized adversarial generation network. In the optimized adversarial generation network, the dropout mask is not fixed in each training, including:
[0068] Optimize the adversarial generative network based on the FreeLB adversarial training method by modifying the position of the fixed model dropout to be unfixed, so that the dropout mask in each training of the optimized adversarial generative network is not fixed;
[0069] Use the JS divergence to constrain the output of the model to generate an optimized adversarial generative network.
[0070] Specifically, when FreeLB is used in pre-trained models such as BERT, since the BERT model uses dropout during finetuning to make the model performance more superior, but the use of dropout will make the models updated by FreeLB each time when performing gradient ascent to find the maximum perturbation inconsistent, resulting in an increase in input noise. Fine tuning means that after the model is pre-trained, during downstream tasks, the model parameter weights will be adjusted for the corresponding tasks to make the model more accurate on the corresponding tasks.
[0071] The most common practice of FreeLB is to fix the position of the model dropout, so that the model is consistent during finetuning. However, fixing the position of the model dropout makes dropout lose its original meaning. Although the position of dropout during the fine tuning process may be different from that during pre-training, it still cannot play the role of dropout. To solve this problem, this solution introduces the JS divergence. The JS divergence is a method to measure the similarity of two probability distributions and is a variant based on the KL divergence.
[0072]
[0073] where KL is the KL divergence,
[0074]
[0075] Therefore, for N samples, its loss is,
[0076]
[0077] where x represents the input data, X represents the set of data, Y(x) and Z(x) represent the probability distributions corresponding to data Y , Z respectively, and Rs(θ) represents the magnitude of the KL divergence of the model; in this way, by maximizing the JS divergence, the network output distribution of FreeLB during each step of gradient ascent can be constrained to achieve the effect of distribution convergence. The objective function is as follows.
[0078]
[0079] Equation 7 solves the problem of increased input noise caused by dropout during each step of gradient ascent in FreeLB, and is also applicable to adversarial training methods that perform multiple gradient ascents to find the maximum perturbation.
[0080] The improvement to the FreeLB method is not only applicable to pre-trained models such as BERT, but also to models with any network structure.
[0081] In one embodiment, adversarial training is performed on a pre-trained model according to an adversarial generation network to generate a target pre-trained model, including:
[0082] Performing adversarial training on the BERT pre-trained model according to the adversarial generation network to generate a target pre-trained model.
[0083] In specific implementation, the gradient attack scheme of the embodiments of the present invention is applicable to pre-trained models similar to BERT, but not limited to BERT, and is also applicable to Roberta, xlm-roberta, XLNET, etc.
[0084] In one embodiment, adversarial training is performed on the BERT pre-trained model according to the adversarial generation network to generate a target pre-trained model, including:
[0085] Obtaining a bond information sample, inputting the bond information sample into the input layer of the BERT pre-trained model to generate an input text;
[0086] Inputting the input text into the embedding layer of the BERT pre-trained model;
[0087] Performing adversarial training on the embedding layer of the BERT pre-trained model according to the adversarial generation network to generate a target pre-trained model.
[0088] In specific implementation, in computer vision, noise can be added to the original image, but it does not affect the nature of the original image. In the field of NLP (Natural Language Processing), noise cannot be directly added to the word encoding because word embeddings are essentially one-hot encodings. If the above noise is added to the one-hot, it will cause ambiguity in the original sentence. Therefore, a natural idea is to add perturbations to the word embeddings (word vectors). Inputting the input text into the embedding layer of the BERT pre-trained model; performing adversarial training on the embedding layer of the BERT pre-trained model according to the adversarial generation network to generate a target pre-trained model.
[0089] In one embodiment, adversarial training is performed on the embedding layer of the BERT pre-trained model according to the adversarial generation network to generate a target pre-trained model, including:
[0090] Perform two gradient attacks on the embedding layer of the BERT pre-trained model according to the adversarial generative network to generate perturbed gradients.
[0091] Update the parameters of the pre-trained model according to the perturbed gradients to generate a target pre-trained model.
[0092] In specific implementation, during adversarial training, when performing gradient attacks, the attack is on the embedding after adding the three. However, it is not reasonable to perform gradient attacks on the embedding after adding the three, because the original intention of gradient attacks is to semantically change the input words so that the model misinterprets the words, that is, "university" is misinterpreted by the model as other meanings except "university". In this case, performing attacks on the embedding obtained by adding the token embedding, position embedding, and segment embedding seems to make the whole attack more complex, causing the model to also misinterpret the position and segment. Segment position segmentation (or called segment): which sentence the current word belongs to. Therefore, this solution adopts a two-stage attack method to generate perturbed gradients; update the parameters of the pre-trained model according to the perturbed gradients to generate a target pre-trained model.
[0093] In one embodiment, the embedding of the BERT pre-trained model consists of token embedding, position embedding, and segment embedding. Then, performing two gradient attacks on the embedding layer of the BERT pre-trained model according to the adversarial generative network to generate perturbed gradients includes:
[0094] Perform the first gradient attack on the token embedding of the BERT pre-trained model according to the adversarial generative network to generate the first gradient after the first perturbation.
[0095] Perform the second gradient attack on the position embedding of the BERT pre-trained model according to the adversarial generative network to generate the second gradient after the second perturbation.
[0096] In specific implementation, such as Figure 2 and Figure 3As shown, in pre-trained models such as BERT, the embedding input to the model is obtained by adding token embedding, position embedding, and segment embedding. For the first time, a gradient attack is first performed on the token embedding, and the perturbed gradient is used to update the parameters. As mentioned above, the purpose is to semantically change the input words so that the words are misunderstood by the model, further enhancing the model's semantic understanding ability. The second gradient attack is to attack the position embedding, aiming to perturb the position of each character in the corpus. For example, "I like China" may become "I am fond of the country of China", which is similar to data augmentation for the model, enabling the model to resist such perturbations and understand the correct meaning of the sentence.
[0097] The mathematical formula for the first gradient attack is as follows:
[0098]
[0099] Among them, E token represents the input word vector, r adv represents the perturbation of the word vector, θ represents the parameters of the model, (f(E token +r adv ; θ) represents the output result of the model, y is the true label, L(f(E token +r adv ; θ), y) represents the loss between the model and the true label, K represents the number of perturbations, represents the maximization of the loss, ε represents the maximum range of the word vector perturbation, Rs(θ) represents the KL divergence size of the model, minE (x,y):D [] represents the minimization of the mathematical expectation of the model.
[0100] The mathematical formula for the second gradient attack is as follows:
[0101]
[0102] Among them, E token represents the input word vector.
[0103] Combining the above methods, the model of the new technical solution we proposed is shown as follows:
[0104]
[0105] Among them,
[0106]
[0107]
[0108] In the embodiment of the present invention, perturbations are respectively performed on toke embedding and position embedding, rather than perturbing the sum of toke embedding, position embedding, and segment embedding. Perturbing the token embedding is to increase the difficulty of the model's understanding of text semantics, while the effect of perturbing the position embedding is similar to shuffling each character in the text to augment the data and further improve the model's text understanding ability.
[0109] The embodiment of the present invention provides a method for generating encoded text based on adversarial training, which optimizes the FreeLB adversarial training method and is applicable to any network model similar to the BERT model, and is particularly effective for text data in the field of cash bond transactions. This method no longer fixes the dropout mask, but uses a probability distribution measurement index or mean squared error, etc. to constrain the output differences of the models with two different dropouts, enabling dropout to play its own role and making the outputs of the models with two different dropouts as consistent as possible. Moreover, in this solution, perturbations are respectively performed on toke embedding and position embedding, rather than perturbing the sum of toke embedding, position embedding, and segment embedding. Perturbing the token embedding is to increase the difficulty of the model's understanding of text semantics, while the effect of perturbing the position embedding is similar to shuffling each character in the text to augment the data and further improve the model's text understanding ability, and the accuracy of extracting transaction elements in the secondary trading of financial bonds has been increased by more than 2% - 5%.
[0110] It should be noted that there is not necessarily a certain sequence among the above steps. Those of ordinary skill in the art can understand according to the description of the embodiments of the present invention that in different embodiments, the above steps can have different execution sequences, that is, they can be executed in parallel or exchanged, etc.
[0111] Another embodiment of the present invention provides a device for generating encoded text based on adversarial training, as Figure 4 shown. The device 1 includes:
[0112] A network construction module 11, configured to pre-construct an adversarial generation network;
[0113] A network optimization module 12 is used to optimize the adversarial generation network to generate an optimized adversarial generation network, and the dropout masks in each training of the optimized adversarial generation network are not fixed;
[0114] An adversarial training module 13 is used to perform adversarial training on the pre-trained model according to the adversarial generation network to generate a target pre-trained model;
[0115] An encoding module 14 is used to input the bond information to be processed into the target pre-trained model to generate an encoded text.
[0116] For the specific implementation manner, refer to the method embodiments, which will not be elaborated here.
[0117] Another embodiment of the present invention provides an electronic device, as Figure 5 shown, the electronic device 10 includes:
[0118] One or more processors 110 and a memory 120, Figure 5 Taking one processor 110 as an example for introduction, the processor 110 and the memory 120 can be connected through a bus or other means, Figure 5 Taking the connection through the bus as an example here.
[0119] The processor 110 is used to complete various control logics of the electronic device 10. It can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a single-chip microcomputer, an ARM (Acorn RISCMachine), or other programmable logic devices, discrete gate or transistor logic, discrete hardware control, or any combination of these components. Additionally, the processor 110 can also be any conventional processor, microprocessor, or state machine. The processor 110 can also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration.
[0120] The memory 120, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions corresponding to the encoding text generation method based on adversarial training in the embodiments of the present invention. The processor 110 executes various functional applications and data processing of the device 10 by running the non-volatile software programs, instructions, and units stored in the memory 120, that is, implements the encoding text generation method based on adversarial training in the above method embodiments.
[0121] The memory 120 may include a program storage area and a data storage area. The program storage area may store an operating device and application programs required for at least one function. The data storage area may store data created according to the use of the device 10 and the like. In addition, the memory 120 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 120 may optionally include a memory remotely provided with respect to the processor 110, and these remote memories may be connected to the device 10 through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0122] One or more units are stored in the memory 120 and, when executed by one or more processors 110, perform the method for generating encoded text based on adversarial training in any of the above method embodiments. For example, perform the method steps S100 to S400 described above. Figure 1 in the method.
[0123] Embodiments of the present invention provide a non-volatile computer-readable storage medium storing computer-executable instructions that, when executed by one or more processors, perform, for example, the method steps S100 to S400 described above. Figure 1 in the method.
[0124] By way of example, non-volatile storage media can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) as an external cache memory. By way of illustration and not limitation, RAM can be obtained in many forms such as synchronous RAM (SRAM), dynamic RAM, (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and direct Rambus RAM (DRRAM). The memory controls or memories of the disclosed operating environments herein are intended to include one or more of these and / or any other suitable types of memories.
[0125] Another embodiment of the present invention provides a computer program product. The computer program product includes a computer program stored on a non-volatile computer-readable storage medium. The computer program includes program instructions that, when executed by a processor, cause the processor to perform the method for generating encoded text based on adversarial training in the above method embodiments. For example, perform the method steps described above.Figure 1 The method steps S100 to S400 in
[0126] The embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0127] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the relevant technology, can be embodied in the form of a software product. This computer software product can exist in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.
[0128] Among other things, conditional language such as "can", "be able to", "may", or "could" generally aims to convey that a particular embodiment can include (while other embodiments do not include) a particular feature, element, and / or operation, unless specifically stated otherwise or otherwise understood within the context in which it is used. Thus, such conditional language generally also aims to imply that the feature, element, and / or operation are needed for one or more embodiments or that one or more embodiments must include logic for determining whether these features, elements, and / or operations are included or will be performed in any particular embodiment, with or without input or prompting.
[0129] What has been described in this specification and the drawings herein includes examples of a method and apparatus for generating encoded text based on adversarial training. Of course, it is not possible to describe every conceivable combination of elements and / or methods for the purpose of describing the various features of the present disclosure, but it can be recognized that many additional combinations and permutations of the disclosed features are possible. Therefore, it is obvious that various modifications can be made to the present disclosure without departing from the scope or spirit of the present disclosure. In addition, or in the alternative, other embodiments of the present disclosure may be apparent from consideration of the specification and drawings and practice of the present disclosure as presented herein. The intention is that the examples presented in the specification and drawings are considered illustrative in all respects and not restrictive. Although specific terms are employed herein, they are used in a general and descriptive sense and not for purposes of limitation.
Claims
1. A method for generating encoded text based on adversarial training, characterized in that, The method includes: Pre - construct an adversarial generation network; Optimize the adversarial generation network to generate an optimized adversarial generation network, where the dropout mask is not fixed in each training of the optimized adversarial generation network; Perform adversarial training on the pre - trained model according to the adversarial generation network to generate a target pre - trained model; Input the bond information to be processed into the target pre - trained model to generate an encoded text; The pre - constructing the adversarial generation network includes: Pre - construct an adversarial generation network based on the FreeLB adversarial training method; The optimizing the adversarial generation network to generate an optimized adversarial generation network, where the dropout mask is not fixed in each training of the optimized adversarial generation network, includes: Optimize the adversarial generation network based on the FreeLB adversarial training method, and modify the position of the fixed model dropout to be unfixed, so as to realize that the dropout mask is not fixed in each training of the optimized adversarial generation network; Use the JS divergence to constrain the output of the model to generate an optimized adversarial generation network.
2. The method according to claim 1, wherein The performing adversarial training on the pre - trained model according to the adversarial generation network to generate a target pre - trained model includes: Perform adversarial training on the BERT pre - trained model according to the adversarial generation network to generate a target pre - trained model.
3. The method according to claim 2, wherein The performing adversarial training on the BERT pre - trained model according to the adversarial generation network to generate a target pre - trained model includes: Obtain a bond information sample, input the bond information sample into the input layer of the BERT pre - trained model to generate an input text; Input the input text into the embedding layer of the BERT pre - trained model; Perform adversarial training on the embedding layer of the BERT pre - trained model according to the adversarial generation network to generate a target pre - trained model.
4. The method according to claim 3, characterized in that, The performing adversarial training on the embedding layer of the BERT pre - trained model according to the adversarial generation network to generate a target pre - trained model includes: Perform two gradient attacks on the embedding layer of the BERT pre - trained model according to the adversarial generation network to generate a perturbed gradient; Update the parameters of the pre - trained model according to the perturbed gradient to generate a target pre - trained model.
5. The method according to claim 4, wherein If the embedding of the BERT pre - trained model consists of token embedding, position embedding, and segment embedding, then performing two gradient attacks on the embedding layer of the BERT pre - trained model according to the adversarial generation network to generate a perturbed gradient includes: Perform the first gradient attack on the token embedding of the BERT pre - trained model according to the adversarial generation network to generate the first gradient after the first perturbation; Perform the second gradient attack on the position embedding of the BERT pre - trained model according to the adversarial generation network to generate the second gradient after the second perturbation.
6. An encoded text generation device based on adversarial training, characterized in that The device includes: A network construction module for pre - constructing an adversarial generation network; A network optimization module for optimizing the adversarial generation network to generate an optimized adversarial generation network, where the dropout mask is not fixed in each training of the optimized adversarial generation network; An adversarial training module, configured to perform adversarial training on a pre-trained model according to an adversarial generation network to generate a target pre-trained model; An encoding module, configured to input the bond information to be processed into the target pre-trained model to generate an encoded text; The pre-constructed adversarial generation network includes: A pre-constructed adversarial generation network based on the FreeLB adversarial training method; Optimizing the adversarial generation network to generate an optimized adversarial generation network, where the dropout mask is not fixed in each training in the optimized adversarial generation network, including: Optimizing the adversarial generation network based on the FreeLB adversarial training method, modifying the position of the fixed model dropout to be unfixed, so as to realize that the dropout mask is not fixed in each training in the optimized adversarial generation network; Using the JS divergence to constrain the output of the model to generate an optimized adversarial generation network.
7. An electronic device, characterized in that, The electronic device includes at least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor, so that the at least one processor can execute the adversarial training-based encoded text generation method according to any one of claims 1-5.
8. A non-volatile computer-readable storage medium, characterized in that, The non-volatile computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by one or more processors, the one or more processors can execute the adversarial training-based encoded text generation method according to any one of claims 1-5.
Citation Information
Patent Citations
High-resolution image generation method based on generative adversarial network
CN111563841A
Visual mileage calculation method based on generative adversarial network
CN112102399A