A text parsing method, device and electronic device based on adversarial training

By randomly setting gradient attacks at any layer of the neural network and controlling the number of occurrences, the problem of insufficient internal layer perturbation resistance and high cost in the NLP field in adversarial training is solved, and more efficient model training and more accurate text parsing are achieved.

CN114398869BActive Publication Date: 2025-06-10BEIJING KUAQUO INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210039449.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-13
Publication Date
2025-06-10
Estimated Expiration
2042-01-13

AI Technical Summary

Technical Problem

The existing adversarial training methods are insufficient to resist internal layer perturbations of neural networks in the NLP field, and the adversarial sample generation process is high, resulting in increased model training burden and time.

Method used

By randomly setting gradient attacks at any layer of the original neural network and controlling the number of gradient attacks according to the probability, constraints are used to ensure that the target neural network output type is consistent with the original network, thereby achieving adversarial training.

Benefits of technology

Reduces the computational burden and time of adversarial training, improves the training speed of the model, and improves the analytical accuracy in financial text data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114398869B_ABST
    Figure CN114398869B_ABST
Patent Text Reader

Abstract

The present invention discloses a text parsing method, device and electronic device based on adversarial training, including: randomly setting a gradient attack at any layer of the original neural network according to a first probability in advance; obtaining training samples, and controlling the occurrence times of the gradient attack in several trainings according to a second probability; using a constraint condition to constrain the target neural network after the gradient attack, so that the output type of the target neural network is the same as the output type of the original neural network; obtaining the text to be parsed, inputting the text to be parsed into the target neural network, and obtaining the parsed text according to the output. The embodiments of the present invention can adjust the probability of the occurrence of the gradient attack, so as to realize that interference may or may not occur between different batches of data in the same round of training and between different training rounds of the same batch of data; it is not only adapted to financial text data with different noises, but also improves the training speed to a certain extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a text parsing method, apparatus, and electronic device based on adversarial training. Background Art

[0002] In the field of deep learning, adversarial training is an effective method to improve the stability of algorithms, and it was first applied to the field of computer vision. It constructs adversarial samples through gradient attacks, so that the neural network can be more robust in the face of noise or perturbations. In the field of natural language processing, since the input of the neural network is discrete symbols, it is impossible to construct adversarial samples through gradient attacks. However, experiments have shown that if we attack the embedding layer of the neural network, it also has a certain effect. Therefore, currently in the field of natural language processing, adversarial training is achieved by attacking the embedding layer. At present, the core of the adversarial training method in the NLP research field lies in how to find the maximum perturbation.

[0003] Existing adversarial training methods mainly include FGSM, FGM, PGD, FreeAT, YOPO, FreeLb, SMART, etc. The current improvement directions and main innovation points are: how to find the maximum perturbation and how to improve the speed of adversarial training. For example: PGD, FreeLB, and SMART technologies find the maximum perturbation by gradually adjusting the perturbation multiple times in a loop. First, the method of calculating the adversarial perturbation in one step by FGSM and FGM is difficult to obtain the optimal point within the constraint. Second, the improved method of PGD is to calculate step by step and map the perturbation exceeding the perturbation radius back to the maximum perturbation sphere. However, PGD updates the perturbation by only using the gradient of the last time, while FreeLB takes the weighted average of the gradients of each time. Second, the SMART adversarial training method is different from other adversarial training methods. It proposes two adversarial regularization losses and directly adds them to the loss function of the model. Its essence is to reduce the empirical risk of the model and the structural risk through regularization, thereby improving the generalization ability of the model. Second, in terms of improving the speed of adversarial training, YOPO has achieved relatively good results. The advantage of this method is that when optimizing the network parameters, the layers are decoupled, which means that the construction of adversarial samples and the calculation of the loss brought by adversarial samples do not need to pass through all layers of the neural network completely.

[0004] Although existing adversarial training methods have achieved good results, in the field of NLP, first of all, all techniques assume that the attack occurs in the embedding layer, that is, the gradient attack only targets the embedding. In fact, the gradient attack in the NLP field may not only occur in the embedding layer, but also in the middle of the network, which brings a problem that the neural network may not be strong enough to resist perturbations occurring in the middle layer of the network. Secondly, existing methods perform adversarial training on each batch of data, and the generation of adversarial samples is a high-cost process, which leads to an increase in the training burden and training time of the model. Since the perturbations are random, a certain degree of randomness should also be added during adversarial training, and it is not necessary to add perturbations every time, which can reduce the training burden to a certain extent. Finally, the probability distribution of the output of the neural network after adding perturbations should also be as close as possible to the output of the network without perturbations to ensure that the network can remain stable under a certain degree of perturbation. However, the core of current methods lies in how to find the maximum perturbation, without considering how the output of the model after adding perturbations can be consistent with the original output.

[0005] The adversarial training of existing neural networks increases the training burden of the model and the training time.

[0006] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0007] In view of the above deficiencies of the existing technology, the present invention provides a text parsing method, device and electronic device based on adversarial training, aiming to solve the problem of the large training burden and long training time of the neural network under adversarial training in the existing technology.

[0008] The technical solution of the present invention is as follows:

[0009] The first embodiment of the present invention provides a text parsing method based on adversarial training, the method includes:

[0010] Randomly set the gradient attack at any layer of the original neural network according to the first probability in advance;

[0011] Obtain training samples, and control the occurrence times of the gradient attack in several trainings according to the second probability;

[0012] Use constraint conditions to constrain the target neural network after the gradient attack, so that the output type of the target neural network is the same as the output type of the original neural network;

[0013] Obtain the text to be parsed, input the text to be parsed into the target neural network, and obtain the parsed text according to the output.

[0014] Further, randomly setting the gradient attack at any layer of the original neural network according to the first probability includes:

[0015] Randomly setting the gradient attack at any layer of the original neural network according to the binomial distribution probability.

[0016] Further, obtaining the training samples and controlling the occurrence times of the gradient attack in several trainings according to the second probability includes:

[0017] Obtaining the training samples, controlling the occurrence of the gradient attack in the current training according to the second probability, and obtaining the training result of the current time, where the training result includes the occurrence or non-occurrence of the gradient attack;

[0018] Obtaining the training results of several trainings and obtaining the occurrence times of the gradient attack according to the training results.

[0019] Further, obtaining the training samples and controlling the occurrence times of the gradient attack in several trainings according to the second probability includes:

[0020] Obtaining the training samples, controlling the occurrence of the gradient attack in the current training according to the second probability, and obtaining the training result of the current time, where the training result includes the occurrence or non-occurrence of the gradient attack;

[0021] Obtaining the training results of several trainings and obtaining the occurrence times of the gradient attack according to the training results.

[0022] Further, constraining the target neural network after the gradient attack by using the probability distribution index includes:

[0023] Constraining the target neural network after the gradient attack by using the JS divergence or the KL divergence.

[0024] Further, randomly setting the gradient attack at any layer of the original neural network according to the binomial distribution probability includes:

[0025] Setting that the gradient attack occurs only in one network layer during one network backpropagation. For a neural network with a total of N layers, the probability that the gradient attack occurs in the l-th layer is p l , and p l obeys the Bernoulli distribution, then the obtained adversarial training model is:

[0026]

[0027] where θ represents the parameters of the model, represents the average of all samples on the dataset D, ||Δ|| is the norm of the perturbation, ∈ is the maximum norm of the perturbation, p l is the occurrence probability of the l-th layer and its value is 0 or 1, Δl is the perturbation of the l-th layer, y is the true label of the sample x, and f is the overall mapping of the model.

[0028] Further, the obtaining of the training samples and controlling the occurrence times of the gradient attack according to the second probability in several trainings includes:

[0029] Let the adversarial training model under the gradient attack be:

[0030]

[0031] The training model without the gradient attack is:

[0032] L2 = E (x,y):D [L(f(x; θ), y)]

[0033] Then the model of the randomness scheme is as follows:

[0034] L 3 = pL 1 + (1 - p)L 2

[0035] where p follows a Bernoulli distribution; if the occurrence probability is p, the non-occurrence probability is 1 - p.

[0036] Another embodiment of the present invention provides a text parsing device based on adversarial training, and the device includes:

[0037] A gradient attack setting module, configured to randomly set a gradient attack on any layer of the original neural network according to the first probability in advance;

[0038] A training control module, configured to obtain training samples and control the occurrence times of the gradient attack in several trainings according to the second probability;

[0039] A constraint module, configured to constraint the target neural network after the gradient attack by using a constraint condition, so that the output type of the target neural network is the same as the output type of the original neural network;

[0040] A parsing module, configured to obtain the text to be parsed, input the text to be parsed into the target neural network, and obtain the parsed text according to the output.

[0041] Another embodiment of the present invention provides an electronic device, and the electronic device includes at least one processor; and,

[0042] a memory communicatively connected to the at least one processor; wherein,

[0043] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above-mentioned text parsing method based on adversarial training.

[0044] Another embodiment of the present invention further provides a non-volatile computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by one or more processors, the one or more processors can be enabled to execute the above-mentioned text parsing method based on adversarial training.

[0045] Beneficial effects: Embodiments of the present invention can adjust the probability of gradient attacks occurring, and realize that interference may or may not occur between different batches of data in the same round of training and between different training rounds of the same batch of data; it is not only suitable for financial text data with different noises, but also improves the training speed to a certain extent. Description of the Drawings

[0046] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:

[0047] Figure 1 is a flowchart of a preferred embodiment of a text parsing method based on adversarial training of the present invention;

[0048] Figure 2 is a schematic diagram of a model in which gradient attacks occur in different network layers in a specific application embodiment of a text parsing method based on adversarial training of the present invention;

[0049] Figure 3 is a schematic flowchart of controlling the occurrence times of gradient attacks in several trainings according to a second probability in a specific application embodiment of a text parsing method based on adversarial training of the present invention;

[0050] Figure 4a is a schematic diagram of the network output probability distribution without gradient attacks in a specific application embodiment of a text parsing method based on adversarial training of the present invention;

[0051] Figure 4b is a schematic diagram of the network output probability distribution with gradient attacks in a specific application embodiment of a text parsing method based on adversarial training of the present invention;

[0052] Figure 4c is a schematic diagram of the network output probability distribution after using the network without attacks as the teacher network to guide the network with gradient attacks in a specific application embodiment of a text parsing method based on adversarial training of the present invention;

[0053] Figure 5Schematic diagram of functional modules of a preferred embodiment of a text parsing device based on adversarial training according to the present invention;

[0054] Figure 6 Schematic diagram of the hardware structure of a preferred embodiment of an electronic device according to the present invention. Detailed implementation manners

[0055] To make the objectives, technical solutions and effects of the present invention clearer and more definite, the present invention will be further described in detail below. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0056] The embodiments of the present invention will be introduced below with reference to the accompanying drawings.

[0057] The embodiments of the present invention provide a text parsing method based on adversarial training. Please refer to Figure 1 , Figure 1 which is a flowchart of a preferred embodiment of a text parsing method based on adversarial training according to the present invention. As Figure 1 shown, it includes the steps:

[0058] Step S100: Randomly set a gradient attack at any layer of the original neural network according to the first probability in advance;

[0059] Step S200: Obtain training samples and control the occurrence times of the gradient attack in several trainings according to the second probability;

[0060] Step S300: Use a constraint condition to constrain the target neural network after the gradient attack, so that the output type of the target neural network is the same as the output type of the original neural network;

[0061] Step S400: Obtain the text to be parsed, input the text to be parsed into the target neural network, and obtain the parsed text according to the output.

[0062] Specifically, when implemented, the embodiments of the present invention are applied to a text parsing method based on adversarial training, and the adversarial training adopts random gradient attack adversarial training. The adversarial training method of the embodiments of the present invention is applicable to any type of network and has obvious effects on text data with more noise in the financial field. The gradient attack is a method of interfering with a deep learning model. It superimposes gradients on the network layer of the model through backpropagation to generate noise inside the model. The present invention first arranges the random occurrence of the gradient attack at any layer of the network according to a certain probability. Furthermore, the occurrence of the gradient attack in each training is controlled by a certain probability, which alleviates the problem of time-consuming adversarial training to a certain extent. Finally, a probability distribution measurement index or mean square error, etc. is used to constrain the output of the interfered network to be consistent with the output of the original network.

[0063] In one embodiment, randomly setting a gradient attack at any layer of the original neural network according to a first probability in advance includes:

[0064] Randomly setting a gradient attack at any layer of the original neural network according to the binomial distribution probability in advance.

[0065] In specific implementation, the first probability is set in advance to be the probability that follows the binomial distribution. This solves the problem that existing methods only focus on how to find the maximum perturbation while ignoring the location where the gradient attack occurs.

[0066] In one embodiment, randomly setting a gradient attack at any layer of the original neural network according to the binomial distribution probability in advance includes:

[0067] Setting that the gradient attack only occurs in one network layer during one backpropagation of the network. If a neural network has N layers, the probability that the gradient attack occurs in the l-th layer is p l , and p l follows the Bernoulli distribution, then the obtained adversarial training model is:

[0068]

[0069] where θ represents the parameters of the model, represents the average of all samples on the dataset D, ||Δ|| is the norm of the perturbation, ∈ is the maximum norm of the perturbation, p l is the occurrence probability of the l-th layer and its value is 0 or 1, Δ l is the perturbation of the l-th layer, y is the true label of the sample x, and f is the overall mapping of the model.

[0070] In specific implementation, adversarial training is to solve a max-min problem in mathematical modeling. Specifically, it is

[0071]

[0072] where y is the training set label, L is the loss function, and f is the neural network. The meaning of the whole formula is to find an optimal parameter θ on the training dataset D to minimize the empirical risk and at the same time achieve the minimum structural risk for the perturbation △. The structural risk refers to the confidence of the model for data outside the training samples. Geometrically, it means that the decision boundary of the model is relatively smooth, and the distance from the samples of each category reaches a compromise maximum without severely favoring any one category.

[0073] As can be seen from the above formula, the perturbation is directly added to the input layer or the embedding layer, without introducing the concept of randomness, and only simply forces the network output loss after adding the perturbation to be the smallest.

[0074] First, assume that the gradient attack occurs only in one network layer during a single backpropagation of the network. If a neural network has N layers, the probability that the gradient attack occurs in the l-th layer is p l , and p l follows a Bernoulli distribution. Then, our new model can be obtained as follows:

[0075]

[0076] where θ represents the parameters of the model, represents the average over all samples in the dataset D, ||Δ|| is the norm of the perturbation, ∈ is the maximum norm of the perturbation, p l is the occurrence probability of the l-th layer, which takes values of 0 or 1, Δ l is the perturbation of the l-th layer, y is the true label of the sample x, and f is the overall mapping of the model. The schematic diagram of the model is as shown in Figure 2 .

[0077] The gradient attack can occur randomly with a certain probability in any layer of the network, not just limited to the embedding layer. In this way, our network can defend against perturbations from both the lower layers and the internal layers of the network. When we set we then obtain the original scheme of adversarial training.

[0078] In one embodiment, training samples are obtained, and the occurrence times of the gradient attack in several trainings are controlled according to the second probability, including:[[]]

[0079] Training samples are obtained, the occurrence of the gradient attack in the current training is controlled according to the second probability, and the training result of the current time is obtained. The training result includes that the gradient attack occurs or the gradient attack does not occur;

[0080] The training results of several trainings are obtained, and the occurrence times of the gradient attack are obtained according to the training results.

[0081] In specific implementation, the existing adversarial training gradient attack occurs throughout the entire process of neural network training, that is, the gradient attack is used for each batch of data in each round of training. This increases the computational burden of training and is not necessary either, because when specifically using the neural network for inference, perturbations caused by the gradient attack do not occur every time. Therefore, we propose a solution that introduces randomness to the gradient training as a whole. Specifically, we let the model trigger adversarial training with a certain probability during the training process.

[0082] In a further embodiment, training samples are obtained, and the occurrence times of the gradient attack in several trainings are controlled according to the second probability, including:[[]]

[0083] Let the adversarial training model under the gradient attack be:[[]]

[0084]

[0085] The training model without gradient attack is:

[0086] L2 = E (x,y):D [L(f(x; θ), y)]

[0087] Then the model of the randomness scheme is as follows:

[0088] L 3 = pL 1 + (1 - p)L 2

[0089] where p follows a Bernoulli distribution; if the probability of occurrence is p, then the probability of non-occurrence is 1 - p.

[0090] In specific implementation, let the adversarial training model under gradient attack be

[0091]

[0092] The training model without gradient attack is

[0093] L2 = E (x,y):D [L(f(x; θ), y)]

[0094] Then the model of our randomness scheme is as follows

[0095] L 3 = pL 1 + (1 - p)L 2

[0096] where p follows a Bernoulli distribution. The Bernoulli distribution is a discrete random distribution, that is, a random event has only two possibilities of occurrence or non-occurrence. If the probability of occurrence is p, then the probability of non-occurrence is 1 - p. The flowchart is as Figure 3 shown.

[0097] In one embodiment, constraint conditions are used to constrain the target neural network after gradient attack, so that the output type of the target neural network is the same as that of the original neural network, including:

[0098] Using a probability distribution metric or mean square error to constrain the target neural network after gradient attack, so that the output type of the target neural network is the same as that of the original neural network.

[0099] In specific implementation, constraint the output of the network affected by interference such as the probability distribution measurement index or mean square error to be consistent with the output of the original network, so as to ensure that the network not only has internal stability but also has stable output after gradient attack occurs.

[0100] To ensure the consistency of the output, we adopt the teacher network in knowledge distillation, which was initially used in knowledge distillation to guide the compressed network on how to learn similar network outputs. Based on the idea of the teacher-student network, we regard the network under non-gradient attacks as the teacher network, which can also guide the network under gradient attacks on how to learn a similar probability distribution. By definition, the teacher network refers to a model that can well fit the data in a certain domain, but its network itself is very complex and large, making it difficult to be deployed and used in industrial production. Specifically, assuming there is no gradient attack, the network output probability distribution is as Figure 4a shown, and the network output probability distribution under gradient attack is as Figure 4b shown. Obviously, at this time, the probability distribution has seriously deviated from the correct output, resulting in the model making incorrect judgments (judging the text information of category C as category B). We can use the teacher network to guide the network to still maintain a similar probability output under gradient attacks. As Figure 4c shown, under such a probability distribution, the model can still make correct judgments and will not misjudge the text information of category C as category B.

[0101] In one embodiment, a probability distribution metric is used to constrain the target neural network after gradient attack, including:

[0102] Using JS divergence or KL divergence to constrain the target neural network after gradient attack.

[0103] Specifically, when implementing, let the output of the network under non-gradient attack be Y = f(x, h, θ), Z = f(x, θ). Then, use JS divergence to measure these two probability distributions. The full name of JS divergence is Jensen-Shannon divergence, which is a metric used to measure two probability distributions and solves the asymmetry of KL divergence. The specific formula is defined as follows,

[0104]

[0105] where KL is KL divergence. The full name of KL divergence is Kullback-Leibler divergence, also known as relative entropy,

[0106] which is the difference in the information entropy of two probability distributions. The specific formula is as follows,

[0107]

[0108] It is a measure for calculating the similarity of two probability distributions.

[0109] Therefore, for N samples, its loss is,

[0110]

[0111] In this way, by maximizing the JS divergence, the output distribution of the network under the gradient attack can be constrained to achieve the effect of distribution convergence.

[0112] Combining the above methods, the embodiments of the present invention can be described by the following model

[0113]

[0114] In specific implementation, the probability distribution measurement index adopts the JS divergence or the KL divergence. As can be seen from the above method embodiments, the present invention provides a text parsing method based on adversarial training, including an adversarial training method in which gradient attacks occur randomly during training; an adversarial training method in which gradient attacks occur in a random layer of the network; regarding the non-interfered network as a teacher network, and guiding the probability distribution of the network output when gradient attacks occur to still be close to the output of the teacher network. Thus, the accuracy of extracting transaction elements in the secondary trading business of financial bonds is improved by more than 2% - 5%.

[0115] The present invention can adjust the probability of gradient attacks occurring, so that interference may or may not occur between different batches of data in the same round of training and between different training rounds of the same batch of data. It is not only adapted to financial text data with different noises, but also improves the training speed to a certain extent.

[0116] The random gradient attack is not only reflected in when to trigger, but also in the randomness of the interference occurring in the network layer. This method adjusts the possibility of gradient attacks occurring in different layers through another probability. In this way, it can cope with the noise from the system internal, rather than just the perturbation brought by financial text data in the embedding layer.

[0117] In one training step, when the network does not have a gradient attack, as a teacher network, it guides the network output with gradient attacks to be as close as possible to the output of the teacher network. The measurement of the output distribution can be determined according to the specific task. For classification tasks, the JS divergence can be used, and for regression tasks, the mean square error can be used, etc. If the data label is a continuous real value, it is a regression task; if the data label is a discrete specific value, it is a classification task.

[0118] It should be noted that there is not necessarily a certain order among the above steps. Those of ordinary skill in the art can understand according to the description of the embodiments of the present invention that in different embodiments, the above steps can have different execution orders, that is, they can be executed in parallel or exchanged, etc.

[0119] Another embodiment of the present invention provides a text parsing device based on adversarial training, as Figure 5 shown, the device 1 includes:

[0120] The gradient attack setting module 11 is configured to randomly set a gradient attack on any layer of the original neural network in advance according to a first probability;

[0121] The training control module 12 is configured to obtain training samples and control the occurrence times of the gradient attack in several trainings according to a second probability;

[0122] The constraint module 13 is configured to constrain the target neural network after the gradient attack by using constraint conditions, so that the output type of the target neural network is the same as the output type of the original neural network;

[0123] The parsing module 14 is configured to obtain the text to be parsed, input the text to be parsed into the target neural network, and obtain the parsed text according to the output.

[0124] For the specific implementation manner, please refer to the method embodiments, which will not be elaborated here.

[0125] Another embodiment of the present invention provides an electronic device, as Figure 6 shown, the electronic device 10 includes:

[0126] One or more processors 110 and a memory 120, Figure 6 Taking one processor 110 as an example for introduction, the processor 110 and the memory 120 can be connected through a bus or other means, Figure 6 Taking the connection through the bus as an example.

[0127] The processor 110 is configured to complete various control logics of the electronic device 10. It can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a single-chip microcomputer, an ARM (Acorn RISC Machine), or other programmable logic devices, discrete gate or transistor logic, discrete hardware control, or any combination of these components. Moreover, the processor 110 can also be any conventional processor, microprocessor, or state machine. The processor 110 can also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration.

[0128] The memory 120, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as program instructions corresponding to the text parsing method based on adversarial training in the embodiments of the present invention. The processor 110 executes various functional applications and data processing of the device 10 by running the non-volatile software programs, instructions, and units stored in the memory 120, that is, implements the text parsing method based on adversarial training in the above method embodiments.

[0129] The memory 120 may include a program storage area and a data storage area. The program storage area may store an operating device and application programs required for at least one function. The data storage area may store data created according to the use of the device 10, etc. In addition, the memory 120 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 120 may optionally include a memory remotely disposed relative to the processor 110, and these remote memories may be connected to the device 10 through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0130] One or more units are stored in the memory 120 and, when executed by one or more processors 110, execute the adversarial training-based text parsing method in any of the above method embodiments. For example, execute the Figure 1 method steps S100 to S400 described above.

[0131] Embodiments of the present invention provide a non-volatile computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are executed by one or more processors. For example, execute the Figure 1 method steps S100 to S400 described above.

[0132] By way of example, non-volatile storage media can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) as an external cache memory. By way of illustration and not limitation, RAM can be obtained in many forms such as synchronous RAM (SRAM), dynamic RAM, (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), Synchl ink DRAM (SLDRAM), and direct Rambus (Rambus) RAM (DRRAM). The memory controls or memories of the operating environments disclosed herein are intended to include one or more of these and / or any other suitable types of memories.

[0133] Another embodiment of the present invention provides a computer program product. The computer program product includes a computer program stored on a non-volatile computer-readable storage medium. The computer program includes program instructions that, when executed by a processor, cause the processor to execute the adversarial training-based text parsing method of the above method embodiment. For example, execute the method steps S100 to S400 described above. Figure 1 in the method.

[0134] The embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0135] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the related technology can be embodied in the form of a software product. The computer software product can exist in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.

[0136] Among other things, conditional language such as "can", "be able to", "may", or "could", unless specifically stated otherwise or otherwise understood within the context in which it is used, generally is intended to convey that a particular embodiment can include (while other embodiments do not include) a particular feature, element, and / or operation. Thus, such conditional language generally is also intended to imply that the feature, element, and / or operation are in some way required for one or more embodiments or that one or more embodiments must include logic for determining whether these features, elements, and / or operations are included or will be performed in any particular embodiment, with or without input or prompting.

[0137] What has been described herein in this specification and the drawings includes examples of a text parsing method and apparatus capable of providing adversarial training. Of course, it is not possible to describe every conceivable combination of components and / or methods for the purpose of describing the various features of the present disclosure, but it will be recognized that many additional combinations and permutations of the disclosed features are possible. Accordingly, it is evident that various modifications can be made to the present disclosure without departing from the scope or spirit thereof. Additionally, or in the alternative, other embodiments of the present disclosure may be apparent from consideration of the specification and drawings and practice of the disclosure as presented herein. It is intended that the examples presented in this specification and the drawings be considered illustrative in all respects and not restrictive. Although specific terms are employed herein, they are used in a generic and descriptive sense and not for purposes of limitation.

Claims

1. A text parsing method based on adversarial training, characterized in that, the method includes: randomly setting gradient attacks at any layer of the original neural network according to the first probability in advance; obtaining training samples and controlling the occurrence times of gradient attacks in several trainings according to the second probability; constraining the target neural network after gradient attacks by using constraint conditions so that the output type of the target neural network is the same as that of the original neural network; obtaining the text to be parsed, inputting the text to be parsed into the target neural network, and obtaining the parsed text according to the output.

2. The method according to claim 1, characterized in that, the randomly setting gradient attacks at any layer of the original neural network according to the first probability in advance includes: randomly setting gradient attacks at any layer of the original neural network according to the binomial distribution probability in advance.

3. The method according to claim 2, characterized in that, the obtaining training samples and controlling the occurrence times of gradient attacks in several trainings according to the second probability includes: obtaining training samples, controlling the occurrence of gradient attacks in the current training according to the second probability, and obtaining the training result of the current time, where the training result includes the occurrence or non-occurrence of gradient attacks; obtaining the training results of several trainings and obtaining the occurrence times of gradient attacks according to the training results.

4. The method according to claim 3, characterized in that, the constraining the target neural network after gradient attacks by using constraint conditions so that the output type of the target neural network is the same as that of the original neural network includes: constraining the target neural network after gradient attacks by using a probability distribution index or mean square error so that the output type of the target neural network is the same as that of the original neural network.

5. The method according to claim 4, characterized in that, the constraining the target neural network after gradient attacks by using a probability distribution index includes: constraining the target neural network after gradient attacks by using JS divergence or KL divergence.

6. A text parsing device based on adversarial training, characterized in that, the device includes: a gradient attack setting module for randomly setting gradient attacks at any layer of the original neural network according to the first probability in advance; a training control module for obtaining training samples and controlling the occurrence times of gradient attacks in several trainings according to the second probability; a constraint module for constraining the target neural network after gradient attacks by using constraint conditions so that the output type of the target neural network is the same as that of the original neural network; a parsing module for obtaining the text to be parsed, inputting the text to be parsed into the target neural network, and obtaining the parsed text according to the output.

7. An electronic device, characterized in that, the electronic device includes at least one processor; and, a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the text parsing method based on adversarial training according to any one of claims 1-5.

8. A non-volatile computer-readable storage medium, characterized in that, the non-volatile computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by one or more processors, the one or more processors can be caused to execute the adversarial training-based text parsing method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Injection randomization confrontation training method

    CN110222502A

  • Adversarial attack defense method based on adversarial sample training

    CN110334808A