Adversarial Interpolation Sequence Labeling Data Augmentation Method, Device, Equipment and Medium

By co-interpolation, the candidate word vectors that conform to context semantic constraints are generated and interpolated, the problem of overfitting the sequence labeling model under low resource conditions is solved, and the accuracy of the model is improved.

CN113297355BActive Publication Date: 2025-07-22CHINA PING AN LIFE INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202110724373.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-29
Publication Date
2025-07-22
Estimated Expiration
2041-06-29

AI Technical Summary

Technical Problem

The prior art sequence labeling models are easily overfitted under low resource conditions, resulting in performance not meeting expectations, and the existing data enhancement methods fail to effectively consider context and task characteristics, affecting the accuracy of the model.

Method used

Adversarial interpolation-based sequence annotation data enhancement method is used to generate candidate word vectors that conform to context semantic constraints through pre-training language models, and use adversarial interpolation to generate more difficult sample data to improve model training effect.

Benefits of technology

The generated enhanced sample data can better regularize the model, improve the performance of sequence models under low resource conditions, and solve the problem that less labeled data affects the accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113297355B_ABST
    Figure CN113297355B_ABST
Patent Text Reader

Abstract

The present invention discloses a sequence annotation data augmentation method, device, equipment and medium based on adversarial interpolation. The method includes: obtaining first sample data containing sequence annotations; inputting the first sample data into a preset language model to output candidate word vectors that conform to the context semantic constraints, and forming enhanced second sample data according to the candidate word vectors; using the method of adversarial interpolation to interpolate the first sample data and the second sample data to obtain interpolated enhanced sample data. According to the sequence annotation data augmentation method provided by the embodiments of the present disclosure, a language model is used to provide candidate word vectors that conform to the context constraints, and adversarial interpolation is used to consider the task characteristics, so as to generate more difficult samples that are likely to cause misjudgment by machine learning algorithms, improve the effect of the sequence model under low resources, and solve the problem that the lack of labeled data affects the accuracy of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sequence annotation, and in particular to a sequence annotation data augmentation method, device, equipment and medium based on adversarial interpolation. Background Art

[0002] Sequence models have a wide range of application scenarios in Chinese word segmentation, named entity recognition, entity and relationship extraction, etc. When using sequence annotation in an online scenario, there will be a problem of few annotation data (low resources). In the case of low resources, for example, each annotation has only a small number of samples, the model may overfit and its performance may not meet expectations. This overfitting situation is more obvious in the case of scarce data, such as the extreme case where each category has only 5 samples. Facing a low-resource application scenario with scarce annotation data, data augmentation is an effective technical method, which can use a very small amount of annotated corpus to obtain a basic model with certain performance, help break through the low-resource dilemma, reduce the need for annotation, and quickly enter the iterative development of model optimization.

[0003] However, it is difficult for the existing data augmentation methods to augment the data of sequence annotation. When augmenting sequence data, it is necessary to consider the context and task characteristics. The previous method of augmenting classification samples cannot achieve the expected effect because it ignores the task characteristics. And the data augmentation based on interpolation uses two real samples of different classes to generate an interpolated sample, and different "difficulty" levels of samples will be generated due to different interpolation ratios, thus affecting the effect of the sequence annotation model. Summary of the Invention

[0004] Embodiments of the present disclosure provide a sequence annotation data augmentation method, device, equipment and medium based on adversarial interpolation. It solves the technical problem that the sample data of the sequence annotation model is small and affects the model training effect. To have a basic understanding of some aspects of the disclosed embodiments, a simple summary is given below. This summary part is not a general review, nor is it intended to identify key / important constituent elements or delineate the protection scope of these embodiments. Its sole purpose is to present some concepts in a simple form as a prelude to the subsequent detailed description.

[0005] In a first aspect, embodiments of the present disclosure provide a sequence annotation data augmentation method based on adversarial interpolation, including:

[0006] Obtain first sample data including sequence annotation;

[0007] Input the first sample data into a preset language model, and output candidate word vectors that meet the context semantic constraints, and form augmented second sample data according to the candidate word vectors;

[0008] Interpolate the first sample data and the second sample data using the adversarial interpolation method to obtain the enhanced sample data after interpolation.

[0009] In an optional embodiment, the preset language model includes a prediction layer, a sorting and selection layer, a normalization layer, and a replacement layer connected in sequence.

[0010] In an optional embodiment, input the first sample data into the preset language model to output candidate word vectors that conform to the context semantic constraints, including:

[0011] The prediction layer predicts the masked words in the first sample data according to the context semantic constraints, and gives the possible words corresponding to the masked words and their probabilities;

[0012] The sorting and selection layer sorts the probabilities corresponding to each possible word from largest to smallest, and selects a preset number of possible words with larger probabilities;

[0013] The normalization layer normalizes the selected possible words and their corresponding probabilities to obtain a normalized probability distribution;

[0014] The replacement layer forms candidate word vectors from the normalized possible words and their corresponding probabilities, and replaces the masked words with the candidate word vectors.

[0015] In an optional embodiment, the normalization layer is used to normalize the selected possible words and their corresponding probabilities, including:

[0016] The normalization layer performs normalization processing through the Softmax function, and the formula for normalization processing is as follows:

[0017]

[0018] Among them, S i represents the Softmax value of the possible word i, e i represents the exponent of the possible word i, ∑ j e j represents the sum of the exponents of all selected possible words.

[0019] In an optional embodiment, interpolating the first sample data and the second sample data using the adversarial interpolation method includes:

[0020] Perform random interpolation according to the random interpolation ratio in the Beta distribution and the first sample data and the second sample data;

[0021] Adjust the random interpolation ratio through the gradient descent method to obtain the latest interpolation ratio in the adversarial direction;

[0022] Perform interpolation operation again according to the latest interpolation ratio to obtain the enhanced sample data after interpolation.

[0023] In an optional embodiment, performing random interpolation according to the random interpolation ratio in the Beta distribution and the first sample data and the second sample data includes:

[0024] Randomly extracting two samples from the mixed sample data of the first sample data and the second sample data;

[0025] Randomly extracting an interpolation ratio from the Beta distribution to obtain a random interpolation ratio;

[0026] Performing random interpolation according to the extracted sample data, the random interpolation ratio, and the mixup algorithm.

[0027] In an optional embodiment, adjusting the random interpolation ratio by the gradient descent method to obtain the latest interpolation ratio in the adversarial direction includes:

[0028] Calculating the interpolation loss at each position according to a preset loss function;

[0029] Taking the partial derivative of the random interpolation ratio, and calculating the current gradient according to the partial derivative value of the random interpolation ratio and the loss value;

[0030] Updating the random interpolation ratio according to the obtained gradient to obtain the latest interpolation ratio in the adversarial direction.

[0031] In a second aspect, an embodiment of the present disclosure provides a sequence annotation data augmentation device based on adversarial interpolation, including:

[0032] An acquisition module, configured to acquire first sample data including sequence annotations;

[0033] A first data augmentation module, configured to input the first sample data into a preset language model, output candidate word vectors that conform to the context semantic constraints, and form augmented second sample data according to the candidate word vectors;

[0034] A second data augmentation module, configured to perform interpolation on the first sample data and the second sample data by using an adversarial interpolation method to obtain interpolated augmented sample data.

[0035] In a third aspect, an embodiment of the present disclosure provides a computer device, including a memory and a processor. When computer-readable instructions stored in the memory are executed by the processor, the processor is caused to execute the steps of the sequence annotation data augmentation method based on adversarial interpolation provided in the above embodiment.

[0036] Fourthly, embodiments of the present disclosure provide a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the sequence annotation data augmentation method based on adversarial interpolation provided in the above embodiments.

[0037] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:

[0038] The sequence annotation data augmentation method based on adversarial interpolation provided by the embodiments of the present disclosure uses a pre-trained language model to provide candidate word vectors that conform to context constraints, and uses adversarial interpolation to consider task characteristics, thereby generating more difficult samples that are likely to cause misjudgment by machine learning algorithms, so as to better regularize the model and improve the effect of the sequence model under low resources, and solve the problem that the lack of labeled data affects the accuracy of the model.

[0039] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.

[0041] Figure 1 is an implementation environment diagram of a sequence annotation data augmentation method based on adversarial interpolation shown according to an exemplary embodiment;

[0042] Figure 2 is an internal structure diagram of a computer device shown according to an exemplary embodiment;

[0043] Figure 3 is a flowchart of a sequence annotation data augmentation method based on adversarial interpolation shown according to an exemplary embodiment;

[0044] Figure 4 is a schematic diagram of a method for obtaining candidate word vectors according to a language model shown according to an exemplary embodiment;

[0045] Figure 5 is a schematic diagram of an adversarial interpolation method shown according to an exemplary embodiment;

[0046] Figure 6 is a schematic diagram of a sequence annotation sample shown according to an exemplary embodiment;

[0047] Figure 7 is a schematic diagram of a pre-trained language model shown according to an exemplary embodiment;

[0048] Figure 8It is a schematic diagram of random interpolation shown according to an exemplary embodiment;

[0049] Figure 9 It is a schematic structural diagram of a sequence annotation data enhancement device based on adversarial interpolation shown according to an exemplary embodiment. Detailed implementation manners

[0050] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0051] It can be understood that the terms "first", "second", etc. used in the present application may be used herein to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish a first element from another element. For example, without departing from the scope of the present application, the first field and algorithm determination module may be referred to as the second field and algorithm determination module, and similarly, the second field and algorithm determination module may be referred to as the first field and algorithm determination module.

[0052] Figure 1 It is an implementation environment diagram of a sequence annotation data enhancement method based on adversarial interpolation shown according to an exemplary embodiment. As Figure 1 shown, in this implementation environment, it includes a server 110 and a terminal 120.

[0053] The server 110 is a device for sequence annotation data enhancement based on adversarial interpolation, such as a computer device such as a computer used by a technician. A data enhancement tool is installed on the server 110. An application that needs to perform data enhancement is installed on the terminal 120. When data enhancement services need to be provided, a technician can send a request for providing data enhancement services on the computer device 110. The request carries a request identifier. The computer device 110 receives the request and obtains the sequence annotation data enhancement method based on adversarial interpolation stored in the computer device 110. Then, this method is used to drive the dialogue management engine platform to complete the human-computer dialogue.

[0054] It should be noted that the terminal 120 and the computer device 110 may be a smart phone, a tablet computer, a notebook computer, a desktop computer, etc., but are not limited thereto. The computer device 110 and the terminal 120 can be connected through Bluetooth, USB (Universal Serial Bus), or other communication connection methods, and the present invention does not make any limitations here.

[0055] Figure 2 It is an internal structural diagram of a computer device shown according to an exemplary embodiment. As Figure 2As shown in the figure, the computer device includes a processor, a non-volatile storage medium, a memory, and a network interface connected via a system bus. Among them, the non-volatile storage medium of the computer device stores an operating system, a database, and computer-readable instructions. The database can store a sequence of control information. When the computer-readable instructions are executed by the processor, the processor can implement a sequence annotation data enhancement method based on adversarial interpolation. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The memory of the computer device can store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can execute a sequence annotation data enhancement method based on adversarial interpolation. The network interface of the computer device is used to communicate with the terminal. Those skilled in the art can understand that Figure 2 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0056] Next, in conjunction with the attached Figure 3 - attached Figure 8 , a detailed introduction will be given to the sequence annotation data enhancement method based on adversarial interpolation provided by the embodiments of the present application. This method can rely on a computer program to implement and can run on a data transmission device based on the von Neumann architecture. The computer program can be integrated into an application or run as an independent tool class application.

[0057] Please refer to Figure 3 , which is a schematic flowchart of a sequence annotation data enhancement method based on adversarial interpolation provided by the embodiments of the present application. As Figure 3 shown, the method of the embodiments of the present application may include the following steps:

[0058] S301 Obtain first sample data containing sequence annotations.

[0059] In a possible implementation manner, first, obtain first sample data containing sequence annotations, where the sentences in the first sample data already contain labeled word tags.

[0060] In a possible implementation manner, the first sample data can be obtained from a sequence annotation sample database, or those skilled in the art can label the first sample data by themselves. For example, as shown in Figure 6 a piece of labeled data for Chinese word segmentation, each word contains a corresponding word tag. I is a single-character word, B is the start of a word, and E is the end of a word. It can be seen that when enhancing sequence data, it is necessary to consider the context and task characteristics.

[0061] S302 inputs the first sample data into a preset language model, outputs candidate word vectors that conform to the context semantic constraints, and forms enhanced second sample data based on the candidate word vectors.

[0062] In a possible implementation, the first sample data obtained in the above steps is input into a preset language model. Among them, the preset language model can obtain candidate words according to the semantic constraints of the context, and form enhanced second sample data based on the obtained candidate words.

[0063] Optionally, the pre-trained language model can be a BERT language model. The BERT model converts each word in the text into a one-dimensional vector by querying the word vector table as the model input; the model output is the vector representation of each input word after fusing the full-text semantic information. In addition, the model input includes two other parts in addition to the word vectors:

[0064] 1. Text vector: The value of this vector is automatically learned during the model training process, used to depict the global semantic information of the text, and fused with the semantic information of single words / terms.

[0065] 2. Position vector: Since the semantic information carried by words / terms in different positions of the text is different (for example: "I love you" and "You love me"), therefore, the BERT model attaches a different vector to words / terms in different positions for distinction.

[0066] Finally, the BERT model takes the sum of the word vector, text vector, and position vector as the model input. For different NLP tasks, the model input will have fine-tuning, and the utilization of the model output will also be different.

[0067] In addition, the BERT language model is a masked language training model. In the original training text, 15% of the words are randomly selected as the objects to be masked. During the training process, it is not known which words it will predict? Which words are original? Which words are masked, and which words are replaced with other words? It is in such a highly uncertain situation that the model can quickly learn the distributed context semantics of the word, and try its best to learn the appearance of the original language speaking. At the same time, because only 15% of the words in the original text participate in the masking operation, it will not damage the expression ability and language rules of the original language. Therefore, the BERT language model can obtain predicted words according to the context constraints. In a possible implementation, those skilled in the art can also use other masked language models for training.

[0068] Figure 7 It is a schematic diagram of a pre-trained language model shown according to an exemplary embodiment, as Figure 7As shown, the preset language model includes a prediction layer, a sorting and selection layer, a normalization layer, and a replacement layer connected in sequence.

[0069] Figure 4 It is a schematic diagram of a method for obtaining candidate word vectors according to a language model shown in an exemplary embodiment. As Figure 4 shown, the method includes:

[0070] S401 The prediction layer predicts the masked words in the first sample data according to the context semantic constraints, and gives the possible words and probabilities corresponding to the masked words.

[0071] For example, A = [I], B = [like]. Assuming that word A is randomly selected for enhancement, some candidate words for word A are given according to the context. The specific process is as follows:

[0072] First is the Predict prediction stage. Mask the word "I" at position A. Use the pre-trained language model to give the word probability distribution at this position according to the context, that is, give the possible words at position A and the probabilities corresponding to these words. Specifically, in an exemplary scenario, the dimension of the word list is V, and the size of the word list in the figure is V = 3. The possible words and probability distribution obtained at position A are {you: 0.3, I: 0.4, he: 0.3}.

[0073] S402 The sorting and selection layer sorts the probabilities corresponding to each possible word from largest to smallest, and selects a preset number of possible words with larger probabilities.

[0074] The sorting and selection layer mainly performs the Top-k operation, that is, takes the top-k results in the predicted word probability distribution. For example, in the figure, k = 2. Then the possible words obtained after the Top-k operation are {you: 0.3, I: 0.4}.

[0075] S403 The normalization layer normalizes the selected possible words and their corresponding probabilities to obtain a normalized probability distribution.

[0076] Furthermore, the Softmax function is used for normalization. The Softmax function is a generalization of the logistic function and is widely used especially in multi-classification scenarios. It maps some inputs to real numbers between 0 and 1, and the normalization ensures that the sum is 1. Therefore, the sum of the probabilities of multi-classification is also exactly 1.

[0077] Suppose there is an array V, and Vi represents the i-th element in V. Then the Softmax value of this element is:

[0078]

[0079] The Softmax value of this element is the ratio of the exponent of this element to the sum of the exponents of all elements.

[0080] Specifically, the obtained possible words and probability distribution are normalized according to the above formula, and the new probability distribution is {you: 0.4, me: 0.6}.

[0081] The S404 replacement layer forms a candidate word vector from the normalized possible words and their corresponding probabilities, and replaces the masked word with the candidate word vector.

[0082] Finally, in the Replace step, the top-k candidate words and the re-normalized probabilities are combined into a new vector, "you × 0.4 + me × 0.6", and this new vector is used to replace the vector of the original word (me).

[0083] According to the above steps, after the operation of the pre-trained language model, a vector representation of each input sentence is obtained, which is a matrix of D × S, where S is the sentence length, and the sentence lengths are uniformly processed to the same length. The enhanced second sample data is obtained.

[0084] According to this step, the semantic constraints of the context can be considered, and the pre-trained language model is used to provide candidate words to obtain better enhanced data.

[0085] S303 interpolates the first sample data and the second sample data using the method of adversarial interpolation to obtain the interpolated enhanced sample data.

[0086] Furthermore, the enhanced second sample data and the labeled first sample data are mixed to obtain the mixed third sample data, and then the obtained mixed sample data is interpolated based on the adversarial interpolation method to obtain the interpolated enhanced sample data.

[0087] Figure 5 It is a schematic diagram of an adversarial interpolation method shown according to an exemplary embodiment, as Figure 5 shown, and this method includes:

[0088] S501 performs random interpolation according to the random interpolation ratio in the Beta distribution and the first sample data and the second sample data.

[0089] Randomly extract two samples from the mixed sample data of the first sample data and the second sample data; randomly extract an interpolation ratio from the Beta distribution to obtain the random interpolation ratio; perform random interpolation according to the extracted sample data, random interpolation ratio, and the mixup algorithm.

[0090] Figure 8 It is a schematic diagram of a random interpolation shown according to an exemplary embodiment, as Figure 8As shown, the specific steps of the mixup method include first randomly selecting two samples {xi, yi} and {xj, yj} from the obtained mixed third sample, and obtaining g after encoding the two inputs xi and xj through the network k (x i ) and g k (x j )

[0091] Then randomly sample a random interpolation ratio λ from the Beta distribution, and this value belongs to [0, 1], as shown in the following formula:

[0092] λ ∼ Beta(α, α)

[0093] Interpolate the representations g k (x i ) and g k (x j ) according to the interpolation ratio λ to obtain the fused

[0094]

[0095] At the same time, interpolate the corresponding labels yi and yj of xi and xj to obtain

[0096]

[0097] and which is equivalent to the new augmented data.

[0098] S502 adjusts the random interpolation ratio through the gradient descent method to obtain the latest interpolation ratio in the adversarial direction.

[0099] Furthermore, in order to improve the training effect of the model, the embodiments of the present disclosure introduce an adversarial operation, search and adjust the interpolation ratio λ in the adversarial direction to generate samples with higher difficulty. Adversarial samples refer to samples that cause machine learning algorithms to make misjudgments. The model should be trained with samples of higher difficulty.

[0100] Specifically, first calculate the loss of each current position according to the following formula,

[0101]

[0102] where θ is the parameter of the model, i and j are the numbers of the true labeled data, f rand represents the random interpolation operation, λ represents the interpolation ratio, η represents the adversarial noise, λ ∼ Beta(α, α), Beta(α, α) represents the beta distribution, and l mix represents the loss function of the interpolation.

[0103] Then, take the partial derivative with respect to the interpolation ratio according to the following formula to obtain its current gradient;

[0104]

[0105] where η represents the gradient, and Δλ represents the partial derivative of the random interpolation ratio, represents the loss value.

[0106] Finally, update the interpolation ratio according to the gradient calculated by the following formula;

[0107] λ′ = λ + εη

[0108] where λ′ represents the latest interpolation ratio, ε represents the step size, and η represents the gradient.

[0109] According to this step, the interpolation ratio can be updated by the gradient descent method to obtain the latest interpolation ratio in the adversarial direction.

[0110] S503 Re - perform the interpolation operation according to the latest interpolation ratio to obtain the enhanced sample data after interpolation.

[0111] In a possible implementation, according to the latest obtained interpolation ratio, re - perform the interpolation operation according to the following formula:

[0112]

[0113]

[0114] where λ′ represents the latest interpolation ratio, {xi, yi} and {xj, yj} represent the sampled data, Beta(α, α) represents the beta distribution, g k (x i ) and g k (x j ) represent the data after network encoding of xi and xj, represents the enhanced data obtained by interpolating and fusing the representation g k (x i ) and g k (x j ) of the word at position K according to the interpolation ratio λ′, represents the enhanced data obtained by interpolating and fusing the corresponding labels yi and yj of xi and xj.

[0115] The newly obtained enhanced data is generated based on the interpolation ratio in the adversarial direction. Therefore, the latest obtained enhanced data is more difficult, more suitable for training the model, can better regularize the model, and improve the performance of the sequence model under low resources.

[0116] Further, the sequence labeling model is trained using the enhanced sample data after adversarial interpolation, and the loss function is minimized to obtain the trained sequence labeling model. Among them, the training method of the sequence labeling model is the same as that of the prior art, and the embodiments of the present disclosure will not elaborate in detail.

[0117] The sequence labeling data enhancement method based on adversarial interpolation provided by the embodiments of the present disclosure uses a pre-trained model to provide candidate word vectors that conform to context constraints, and uses adversarial interpolation to consider task characteristics, thereby generating more difficult samples that are likely to cause misjudgment by machine learning algorithms, so as to better regularize the model and improve the performance of the sequence model under low resources, and solve the problem that the lack of labeled data affects the accuracy of the model.

[0118] The following is the device embodiment of the present application, which can be used to execute the method embodiment of the present invention. For details not disclosed in the device embodiment of the present invention, please refer to the method embodiment of the present invention.

[0119] Please refer to Figure 9 , which shows a schematic structural diagram of a sequence labeling data enhancement device based on adversarial interpolation provided by an exemplary embodiment of the present invention. As Figure 9 shown, the sequence labeling data enhancement device based on adversarial interpolation can be integrated into the above-mentioned computer device 110, and specifically can include an acquisition module 901, a first data enhancement module 902, and a second data enhancement module 903.

[0120] The acquisition module 901 is configured to acquire first sample data including sequence labeling;

[0121] The first data enhancement module 902 is configured to input the first sample data into a preset language model, output candidate word vectors that conform to context semantic constraints, and form enhanced second sample data according to the candidate word vectors;

[0122] The second data enhancement module 903 is configured to interpolate the first sample data and the second sample data by using the method of adversarial interpolation to obtain the enhanced sample data after interpolation.

[0123] In an optional embodiment, the preset language model includes a prediction layer, a sorting and selection layer, a normalization layer, and a replacement layer connected in sequence.

[0124] In an optional embodiment, the preset language model in the first data enhancement module 902 specifically includes:

[0125] The prediction layer is configured to predict the masked words in the first sample data according to context semantic constraints, and give the possible words and probabilities corresponding to the masked words;

[0126] The sorting and selection layer is used to sort the probabilities corresponding to each possible word from largest to smallest, and select a preset number of possible words with larger probabilities;

[0127] The normalization layer is used to perform normalization processing on the selected possible words and their corresponding probabilities to obtain a normalized probability distribution;

[0128] The replacement layer is used to form candidate word vectors with the normalized possible words and their corresponding probabilities, and replace the masked words with the candidate word vectors.

[0129] In an optional embodiment, the normalization layer is used to perform normalization processing on the selected possible words and their corresponding probabilities, including:

[0130] The normalization layer performs normalization processing through the Softmax function, and the formula for the normalization processing is as follows:

[0131]

[0132] where S i represents the Softmax value of the possible word i, e i represents the exponent of the possible word i, and ∑ j e j represents the sum of the exponents of all the selected possible words.

[0133] In an optional embodiment, the second data augmentation module 903 is specifically used for:

[0134] Performing random interpolation according to the random interpolation ratio in the Beta distribution and the first sample data and the second sample data;

[0135] Adjusting the random interpolation ratio through the gradient descent method to obtain the latest interpolation ratio in the adversarial direction;

[0136] Performing interpolation operation again according to the latest interpolation ratio to obtain the interpolated augmented sample data.

[0137] In an optional embodiment, performing random interpolation according to the random interpolation ratio in the Beta distribution and the first sample data and the second sample data includes:

[0138] Randomly extracting two samples from the mixed sample data of the first sample data and the second sample data;

[0139] Randomly extracting an interpolation ratio from the Beta distribution to obtain a random interpolation ratio;

[0140] Performing random interpolation according to the extracted sample data, the random interpolation ratio, and the mixup algorithm.

[0141] In an optional embodiment, the random interpolation ratio is adjusted by the gradient descent method to obtain the latest interpolation ratio in the adversarial direction, including:

[0142] Calculating the interpolation loss at each position according to a preset loss function;

[0143] Taking the partial derivative of the random interpolation ratio, and calculating the current gradient according to the partial derivative value and the loss value of the random interpolation ratio;

[0144] Updating the random interpolation ratio according to the obtained gradient to obtain the latest interpolation ratio in the adversarial direction.

[0145] The sequence annotation data augmentation device based on adversarial interpolation provided by the embodiments of the present disclosure uses a pre-trained model to provide candidate word vectors that conform to context constraints, and uses adversarial interpolation to consider task characteristics, so as to generate more difficult samples that are likely to cause misjudgment by machine learning algorithms, thereby better regularizing the model and improving the effect of the sequence model under low resources, and solving the problem that the lack of labeled data affects the accuracy of the model.

[0146] It should be noted that when the sequence annotation data augmentation device based on adversarial interpolation provided in the above embodiments executes the sequence annotation data augmentation method based on adversarial interpolation, only the above-mentioned division of each functional module is used as an example. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the sequence annotation data augmentation device based on adversarial interpolation provided in the above embodiments and the embodiments of the sequence annotation data augmentation method based on adversarial interpolation belong to the same concept, and the implementation process thereof is detailed in the method embodiments and will not be repeated here.

[0147] In one embodiment, a computer device is proposed. The computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented: obtaining first sample data including sequence annotation; inputting the first sample data into a preset language model, outputting candidate word vectors that conform to context semantic constraints, and forming enhanced second sample data according to the candidate word vectors; using the method of adversarial interpolation to interpolate the first sample data and the second sample data to obtain the interpolated enhanced sample data.

[0148] In an optional embodiment, the preset language model includes a prediction layer, a sorting and selection layer, a normalization layer, and a replacement layer connected in sequence.

[0149] In an optional embodiment, inputting the first sample data into a preset language model and outputting candidate word vectors that conform to context semantic constraints includes:

[0150] The prediction layer predicts the masked words in the first sample data based on the context semantic constraints, and gives the possible words corresponding to the masked words and their probabilities;

[0151] The sorting and selection layer sorts the probabilities corresponding to each possible word from largest to smallest, and selects a preset number of possible words with larger probabilities;

[0152] The normalization layer normalizes the selected possible words and their corresponding probabilities to obtain a normalized probability distribution;

[0153] The replacement layer forms a candidate word vector with the normalized possible words and their corresponding probabilities, and replaces the masked words with the candidate word vector.

[0154] In an optional embodiment, the normalization layer is used to normalize the selected possible words and their corresponding probabilities, including:

[0155] The normalization layer performs normalization through the Softmax function, and the formula for normalization is as follows:

[0156]

[0157] where S i represents the Softmax value of the possible word i, e i represents the exponent of the possible word i, and ∑ j e j represents the sum of the exponents of all selected possible words.

[0158] In an optional embodiment, the method of adversarial interpolation is used to interpolate the first sample data and the second sample data, including:

[0159] Perform random interpolation according to the random interpolation ratio in the Beta distribution and the first sample data and the second sample data;

[0160] Adjust the random interpolation ratio through the gradient descent method to obtain the latest interpolation ratio in the adversarial direction;

[0161] Perform interpolation operation again according to the latest interpolation ratio to obtain the interpolated enhanced sample data.

[0162] In an optional embodiment, performing random interpolation according to the random interpolation ratio in the Beta distribution and the first sample data and the second sample data includes:

[0163] Randomly extract two samples from the mixed sample data of the first sample data and the second sample data;

[0164] Randomly extract an interpolation ratio from the Beta distribution to obtain a random interpolation ratio;

[0165] Perform random interpolation based on the extracted sample data, the random interpolation ratio, and the mixup algorithm.

[0166] In an optional embodiment, the random interpolation ratio is adjusted by the gradient descent method to obtain the latest interpolation ratio in the adversarial direction, including:

[0167] Calculate the interpolation loss at each position according to a preset loss function;

[0168] Take the partial derivative of the random interpolation ratio, and calculate the current gradient according to the partial derivative value of the random interpolation ratio and the loss value;

[0169] Update the random interpolation ratio according to the obtained gradient to obtain the latest interpolation ratio in the adversarial direction.

[0170] In an embodiment, a storage medium storing computer-readable instructions is proposed. When the computer-readable instructions are executed by one or more processors, the one or more processors perform the following steps: obtain first sample data including sequence annotation; input the first sample data into a preset language model, output candidate word vectors that conform to the context semantic constraints, and form enhanced second sample data according to the candidate word vectors; perform interpolation on the first sample data and the second sample data by using the adversarial interpolation method to obtain the interpolated enhanced sample data.

[0171] In an optional embodiment, the preset language model includes a prediction layer, a sorting and selection layer, a normalization layer, and a replacement layer connected in sequence.

[0172] In an optional embodiment, inputting the first sample data into the preset language model and outputting candidate word vectors that conform to the context semantic constraints includes:

[0173] The prediction layer predicts the masked words in the first sample data according to the context semantic constraints, and gives the possible words corresponding to the masked words and the probabilities;

[0174] The sorting and selection layer sorts the probabilities corresponding to the respective possible words from largest to smallest, and selects a preset number of possible words with larger probabilities;

[0175] The normalization layer normalizes the selected possible words and their corresponding probabilities to obtain a normalized probability distribution;

[0176] The replacement layer forms candidate word vectors with the normalized possible words and their corresponding probabilities, and replaces the masked words with the candidate word vectors.

[0177] In an optional embodiment, the normalization layer is used to normalize the selected possible words and their corresponding probabilities, including:

[0178] The normalization layer performs normalization through the Softmax function, and the formula for normalization is as follows:

[0179]

[0180] Among them, S i represents the Softmax value of the possible word i, e i represents the exponent of the possible word i, and ∑ j e j represents the sum of the exponents of all selected possible words.

[0181] In an optional embodiment, the first sample data and the second sample data are interpolated by using the method of adversarial interpolation, including:

[0182] Performing random interpolation according to the random interpolation ratio in the Beta distribution and the first sample data and the second sample data;

[0183] Adjusting the random interpolation ratio by the gradient descent method to obtain the latest interpolation ratio in the adversarial direction;

[0184] Performing an interpolation operation again according to the latest interpolation ratio to obtain the enhanced sample data after interpolation.

[0185] In an optional embodiment, performing random interpolation according to the random interpolation ratio in the Beta distribution and the first sample data and the second sample data includes:

[0186] Randomly extracting two samples from the mixed sample data of the first sample data and the second sample data;

[0187] Randomly extracting an interpolation ratio from the Beta distribution to obtain the random interpolation ratio;

[0188] Performing random interpolation according to the extracted sample data, the random interpolation ratio, and the mixup algorithm.

[0189] In an optional embodiment, adjusting the random interpolation ratio by the gradient descent method to obtain the latest interpolation ratio in the adversarial direction includes:

[0190] Calculating the interpolation loss at each position according to the preset loss function;

[0191] Taking the partial derivative of the random interpolation ratio, and calculating the current gradient according to the partial derivative value of the random interpolation ratio and the loss value;

[0192] Updating the random interpolation ratio according to the obtained gradient to obtain the latest interpolation ratio in the adversarial direction.

[0193] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disc, a read-only memory (ROM), etc., or a random access memory (RAM), etc.

[0194] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0195] The above embodiments only represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the invention patent of the present invention should be subject to the appended claims.

Claims

1. A sequence annotation data augmentation method based on adversarial interpolation, characterized in that , including: Obtain first sample data including sequence annotation; wherein, each word includes a corresponding word label, I is a single-character word, B is the start of a word, and E is the end of a word; Input the first sample data into a preset language model, output candidate word vectors that conform to the context semantic constraints, and form enhanced second sample data according to the candidate word vectors; Interpolate the first sample data and the second sample data by using the method of adversarial interpolation to obtain interpolated enhanced sample data; including: performing random interpolation according to the random interpolation ratio in the Beta distribution and the first sample data and the second sample data; calculating the interpolation loss at each position according to a preset loss function; taking the partial derivative of the random interpolation ratio, and calculating the current gradient according to the partial derivative value of the random interpolation ratio and the loss value; updating the random interpolation ratio according to the obtained gradient to obtain the latest interpolation ratio in the adversarial direction; performing interpolation operation again according to the latest interpolation ratio to obtain interpolated enhanced sample data.

2. The method according to claim 1, wherein The preset language model includes a prediction layer, a sorting and selection layer, a normalization layer, and a replacement layer connected in sequence.

3. The method according to claim 2, characterized in that, The step of inputting the first sample data into a preset language model and outputting candidate word vectors that conform to the context semantic constraints includes: The prediction layer predicts the masked words in the first sample data according to the context semantic constraints, and gives the possible words corresponding to the masked words and their probabilities; The sorting and selection layer sorts the probabilities corresponding to each possible word from large to small, and selects a preset number of possible words with larger probabilities; The normalization layer performs normalization processing on the selected possible words and their corresponding probabilities to obtain a normalized probability distribution; The replacement layer forms candidate word vectors by using the normalized possible words and their corresponding probabilities, and replaces the masked words with the candidate word vectors.

4. The method according to claim 3, wherein The normalization layer is used to perform normalization processing on the selected possible words and their corresponding probabilities, including: The normalization layer performs normalization processing through the Softmax function, and the formula for normalization processing is as follows: Among them, S i represents the Softmax value of the possible word i, and e i represents the exponent of the possible word i, and ∑ j e j represents the sum of the exponents of all selected possible words.

5. The method according to claim 1, wherein Performing random interpolation according to the random interpolation ratio in the Beta distribution and the first sample data and the second sample data includes: Randomly extract two samples from the mixed sample data of the first sample data and the second sample data; Randomly extract an interpolation ratio from the Beta distribution to obtain a random interpolation ratio; Perform random interpolation according to the extracted sample data, the random interpolation ratio, and the mixup algorithm.

6. A sequence annotation data augmentation device based on adversarial interpolation, characterized in that , including: An acquisition module, configured to obtain first sample data including sequence annotation; wherein, each word includes a corresponding word label, I is a single-character word, B is the start of a word, and E is the end of a word; A first data enhancement module, configured to input the first sample data into a preset language model, output candidate word vectors that conform to the context semantic constraints, and form enhanced second sample data according to the candidate word vectors; The second data augmentation module is used to interpolate the first sample data and the second sample data by using the method of adversarial interpolation to obtain the augmented sample data after interpolation, and includes: performing random interpolation according to the random interpolation ratio in the Beta distribution and the first sample data and the second sample data; calculating the interpolation loss at each position according to a preset loss function; taking the partial derivative of the random interpolation ratio, and calculating the current gradient according to the partial derivative value of the random interpolation ratio and the loss value; updating the random interpolation ratio according to the obtained gradient to obtain the latest interpolation ratio in the adversarial direction; and re-performing the interpolation operation according to the latest interpolation ratio to obtain the augmented sample data after interpolation.

7. A computer device, comprising a memory and a processor, wherein computer-readable instructions are stored in the memory, and when the computer-readable instructions are executed by the processor, the processor is caused to execute the steps of the sequence annotation data augmentation method based on adversarial interpolation according to any one of claims 1 to 5.

8. A storage medium storing computer-readable instructions, wherein when the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the sequence annotation data augmentation method based on adversarial interpolation according to any one of claims 1 to 5.