Continuous Relation Extraction Method and System Based on Multi-Head Self-Attention Mechanism Adapter
By using a combination of multi-head self-attention mechanism adapter and pre-trained model in entity relationship extraction, the problems of catastrophic forgetting and data distribution imbalance are solved, and efficient parameter fine-tuning and accuracy improvement are achieved.
Patent Information
- Application Number
- CN202211632267.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-19
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-12-19
AI Technical Summary
The existing technology has catastrophic forgetting problems in entity relationship extraction, resulting in a sharp decline in the performance of the model for historical tasks after training on new tasks, and the data distribution is unbalanced and computational cost is high.
The continuous relationship extraction method based on the multi-head self-attention mechanism adapter is adopted, and efficient parameter fine-tuning is performed through the combination of pre-trained models and multiple A-Adapters. Each task is trained using independent A-Adapters to avoid catastrophic forgetting.
It effectively avoids catastrophic forgetting problems, improves the performance and accuracy of the model on historical tasks, and reduces the file size and parameter volume of the model, reducing the calculation cost and storage cost.
Smart Images

Figure CN116186171B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of information extraction, and particularly relates to a continuous relation extraction method and system based on a multi-head self-attention mechanism adapter. Background Art
[0002] Entity relation extraction is one of the core tasks of information extraction, aiming to extract entities and relations from unstructured text to form entity relation triples, and thus become the basic work for constructing a large-scale knowledge base. Current research mainly includes pipeline methods and joint learning methods. By using various encoding methods such as sequence labeling methods, table filling methods, fragment classification, pointer networks, etc., combined with specific decoding strategies, many problems such as nested entities, exposure bias, redundant calculations, and overlapping relations can be effectively solved. Although these methods can achieve good performance in a limited domain, they need to construct datasets with predefined entity types and relation types, which cannot meet the dynamic real-world needs. In response to this situation, a simple approach is to mix the data of historical tasks and new tasks and retrain the model. However, since the data volume of each task may not be the same, it will bring serious data distribution imbalance problems. At the same time, since the entity types and relation types labeled for each task are different, there will be a large amount of "unlabeled" data after data mixing, that is, the same text is labeled with different entities and relations in different tasks, seriously affecting the learning of each relation by the model and resulting in poor performance. Moreover, as the number of tasks increases, the computational cost and storage cost will also increase continuously. In addition, this method also requires sharing data when sharing the model, thus bringing the risk of data leakage.
[0003] To solve the above problems, the lifelong entity relation extraction task is utilized, that is, learning step by step from an infinite data stream, enabling the model to retain the knowledge of historical tasks while learning new tasks. A major challenge faced by this task is catastrophic forgetting, that is, retraining the model of historical tasks on the data of new tasks will lead to a sharp decline in the performance of the new model on historical tasks, and in more serious cases, it may even be unable to extract any results of historical tasks. This is an inherent defect of deep neural network models because the network structure and parameters determine the capacity of the model to remember knowledge, and even a slight change may affect the model's performance. Current research mainly alleviates the problem of catastrophic forgetting through three methods: the memory replay method, the regularization method, and the parameter isolation method. The memory replay method means that when the model learns a new task, it replays the typical samples of the remembered historical tasks to alleviate forgetting. The problem is that as the number of tasks increases, the number of samples to be remembered will become larger and larger. The regularization method means that by introducing additional regularization terms, the update of model parameters is restricted, ensuring data security and reducing memory consumption. The parameter isolation method means that a part of the model parameters is allocated to each task, and they do not affect each other, but this method limits the number of tasks. Summary of the Invention
[0004] Therefore, the present invention provides a continuous relation extraction method and system based on a multi-head self-attention mechanism adapter. In the training stage, multiple A-Adapters are used for each task to perform efficient parameter fine-tuning, and the catastrophic forgetting problem in model training for entity relation extraction is effectively avoided through the independent training of each task model.
[0005] According to the design scheme provided by the present invention, a continuous relation extraction method based on a multi-head self-attention mechanism adapter is provided, including the following contents:
[0006] Use a pre-trained model as the backbone network to construct a continuous relation extraction model. The pre-trained model includes multiple Transformer layers for encoding and decoding input feature vectors, and an adapter is used to learn task knowledge. Among them, the adapter includes: a dimensionality reduction projection layer for reducing the dimensionality of the input feature vector, a multi-head self-attention layer for mining the context semantic information representation of the feature vector, a dimensionality increase projection layer for restoring the mined context semantic representation to the dimension of the input feature vector, and a residual structure for optimizing the residual connection in model training;
[0007] Regard each task of continuous relation extraction as a multi-classification problem, and use the combined training of multiple adapters to learn the preset task knowledge in the target text.
[0008] As the continuous relation extraction method based on the multi-head self-attention mechanism adapter in the present invention, further, in the combined training to learn the model of the target task, each adapter is provided with a Transformer layer corresponding to the parallel combination relationship in the pre-trained model, and the input of the current adapter is the superposition of the output vector of the previous adapter in the combined training and the output of the Transformer layer parallel to the previous adapter.
[0009] As the continuous relation extraction method based on the multi-head self-attention mechanism adapter in the present invention, further, for the target text, first freeze the parameters of the pre-trained model, then add position markers of the head and tail entities around the entity of the input text sentence, and use the BERT model to extract the feature vector of the text sentence.
[0010] As the continuous relation extraction method based on the multi-head self-attention mechanism adapter in the present invention, further, the pre-trained model is denoted as: h = [h1, h2, …, h N , where represents the output vector of the l-th Transformer layer of the pre-trained model, and represent the trainable parameters of the model, LayerNorm(·) represents the normalization operation, h and d respectively represent the dimension of the hidden layer vector and the sentence representation vector of the pre-trained model, N represents the total number of Transformer layers in the pre-trained model, h 11 , h 21 represent the vector representations corresponding to the positions of the head and tail entities.
[0011] As the continuous relation extraction method based on the multi-head self-attention mechanism adapter of the present invention, further, for the input text, the input vector representation of the adapter corresponding to the l-th layer Transformer layer in the continuous relation extraction model is: I l = h l-1 + O l-1 , where O l = LayerNorm(I l + ((I l W down )W MultiHead )W up ), represents the output vector of the (l - 1)-th Transformer layer of the pre-trained model, represents the output vector of the adapter corresponding to the (l - 1)-th layer, represents the trainable parameter of the dimensionality reduction projection layer, represents the trainable parameter of the multi-head attention layer, represents the trainable parameter of the dimensionality increase projection layer, O ldenotes the output vector of the adapter corresponding to the l-th layer, and md denotes the dimension after the input vector is dimensionally reduced.
[0012] As the continuous relation extraction method based on the multi-head self-attention mechanism adapter of the present invention, further, cross-entropy is used as the target loss function in the training of the continuous relation extraction model, and the target loss function is expressed as: p k =[p0, p1, … p R-1 represents the classification result of the k-th task, and p i ∈p k represents the probability that the input sentence belongs to the i -th class, R is the number of predefined relations, and O last denotes the output vector of the adapter corresponding to the last layer, and y = [y0, y1, …, y R-1 is the one-hot representation of the true label of the input sentence.
[0013] Further, the present invention also provides a continuous relation extraction system based on the multi-head self-attention mechanism adapter, including: a model construction module and a relation extraction module, where
[0014] The model construction module is used to construct a continuous relation extraction model by using a pre-trained model as the backbone network, and the pre-trained model includes multiple Transformer layers for encoding and decoding input feature vectors, and an adapter is used to learn task knowledge, where the adapter includes: a dimensionality reduction projection layer for dimensionality reduction of the input feature vector, a multi-head self-attention layer for mining the context semantic information representation of the feature vector, a dimensionality increase projection layer for restoring the mined context semantic representation to the dimension of the input feature vector, and a residual structure for optimizing the residual connection in model training;
[0015] The relation extraction module is used to regard each task of continuous relation extraction as a multi-classification problem, and use the combined training of multiple adapters to learn the preset task knowledge in the target text.
[0016] Further, the present invention also provides a multi-objective task recognition method, including:
[0017] First, collect sample data containing preset task objects, and use the above continuous relation extraction method to construct a multi-objective task recognition model;
[0018] Next, use the sample data to train the multi-objective task recognition model;
[0019] Then, input the data to be processed into the trained multi-objective task recognition model, and use the trained multi-objective task recognition model to identify and obtain the data containing the expected objects.
[0020] Advantages of the present invention:
[0021] The pre-trained model of the present invention is used as the backbone network to keep all parameters frozen. Each task is independently trained using a model composed of multiple A-Adapters to learn specific task knowledge. This can not only ensure a high accuracy of the entity relationship extraction model, but also reduce the file size and the number of parameters of the model. During actual use, the number of layers of A-Adapters and the dimension of A-Adapters can be appropriately increased within an acceptable range. Further, experimental results prove that adding a multi-head attention mechanism layer in the bottleneck structure can not only enhance the fitting ability of A-Adapters well, but also reduce the risk of overfitting of A-Adapters, and has good application prospects. Description of the drawings
[0022] Figure 1 Schematic diagram of the continuous relationship extraction process based on the multi-head self-attention mechanism adapter in the embodiment;
[0023] Figure 2 Schematic diagram of the adapter structure in the embodiment;
[0024] Figure 3 Schematic diagram of the continuous relationship extraction model framework in the embodiment;
[0025] Figure 4 Schematic diagram of the average training time curve of the A-Adapter model and the baseline model on the historical tasks of the FewRel dataset in the embodiment;
[0026] Figure 5 Schematic diagram of the average training time curve of the A-Adapter model and the baseline model on the historical tasks of the TACRED dataset in the embodiment;
[0027] Figure 6 Schematic diagram of the video memory usage when the A-Adapter model and the baseline model finish training on the FewRel dataset and the TACRED dataset in the embodiment;
[0028] Figure 7 Schematic diagram of the comparison of the file sizes of the A-Adapter model and the BERT model in the embodiment. Detailed implementation manners
[0029] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the drawings and technical solutions.
[0030] To address the catastrophic forgetting problem in continuous relation extraction, the catastrophic forgetting problem in continuous entity relation extraction can be alleviated by combining prototype representation and memory replay. In the embodiments of this case, the A-Adapter is applied to the continuous entity relation extraction task. Refer to Figure 1 As shown, a continuous relation extraction method based on a multi-head self-attention mechanism adapter is provided, including:
[0031] S101. Use a pre-trained model as the backbone network to construct a continuous relation extraction model. The pre-trained model includes multiple Transformer layers for encoding and decoding input feature vectors, and use an adapter to learn task knowledge. Among them, the adapter includes: a dimensionality reduction projection layer for reducing the dimension of the input feature vector, a multi-head self-attention layer for mining the context semantic information representation of the feature vector, an upsampling projection layer for restoring the mined context semantic representation to the dimension of the input feature vector, and a residual structure for model training optimization and residual connection;
[0032] S102. Treat each task of continuous relation extraction as a multi-classification problem, and use the combined training of multiple adapters to learn the preset task knowledge in the target text.
[0033] The existing continuous entity relation extraction task is to use a relation classification model to identify the relationship between two entities. Formally, given a series of K tasks {T1, T2, …, T K}, the kth task T k includes a training set a validation set and a predefined relation set where k ∈ [1, K], represents the ith input sentence of task T k , represents the ith input entity pair of task T k , represents the true relation label of the ith input data of task T k , represents the number of the training set of task T k , represents the number of the validation set of task T k , represents the number of the relation set of task T k . The goal of continuous entity relation extraction is to learn a relation multi-classification model f. After training on the training set Train k of task T k , given a sentence and an entity pair x p ∈Valid p≤k , the relationship between the entity pairs can be predicted as Rather than just y p ∈R k In the embodiments of this case, A-Adapter is a general, flexible, and lightweight module. Different from the bottleneck structures used in previously proposed adapters, A-Adapter adds a multi-head attention layer to enhance feature vector learning in the bottleneck structure, rather than adding a non-linear activation function or a Transformer layer. It can satisfy the independent training and learning of each task and effectively avoid the problem of catastrophic forgetting.
[0034] As a preferred embodiment, further, in the model for learning the target task by combined training, each adapter is provided with a Transformer layer corresponding to the parallel combination relationship in the pre-trained model, and the input of the current adapter is the superposition of the vector output of the previous adapter and the output of the Transformer layer parallel to the previous adapter in the combined training.
[0035] A-Adapter can be set in parallel with any layer of the pre-trained model through a configuration file without affecting the structure of the pre-trained model. Therefore, A-Adapter can be combined with any pre-trained model as a pluggable module. A-Adapter is a module with a bottleneck structure. Refer to Figure 2 As shown, it mainly includes four components, namely a dimensionality reduction projection layer, a multi-head self-attention layer, a dimensionality increase projection layer, and a residual structure. To limit the number of parameters of A-Adapter, the input feature vector first passes through the dimensionality reduction projection layer for dimensionality reduction, then obtains a feature vector containing richer context semantic information through the multi-head self-attention layer, and then restores the feature vector to the same dimension as the input feature vector through the dimensionality increase projection layer, ensuring the consistency of the dimensions of the input vector and the output vector. Finally, the residual structure reduces network degradation, making the model have better performance.
[0036] Specifically, A-Adapter first uses the to project the d-dimensional feature vector of the input sample onto where d m << d, d m represents the hidden layer dimension of A-Adapter, and the calculation formula is shown in Equation (1).
[0037] X'←XW down (1)
[0038] Then, s self-attention heads are used to perform self-attention calculation on to obtain with a richer context semantic representation, where The calculation formulas are shown in Equations (2) and (3).
[0039]
[0040]
[0041] Finally, using the project the onto Then add it to the input X of the residual connection and input it into the LayerNorm layer to obtain the final output vector The calculation formula is shown in Equation (4).
[0042] Y = LayerNorm(X + X”W up ) (4)
[0043] Taking the BERT model as an example, d = 768, s = 12, the present invention sets d m to 72, so d k = d v = d m / s = 6, the number of parameters of A-Adapter is approximately (d × d m ) + (3 × s × d m × d m / s + d m × d m ) + (d m × d) = 131,328 ≈ 0.125M, and a Transformer layer is mainly composed of a multi-head attention layer and two feed-forward layers, and its number of parameters is approximately (3 × s × d × d / s + d × d) + (d × 4 × d + 4 × d × d) = 7,077,888 = 6.75M. The number of parameters per layer is reduced by approximately 53.9 times. And the present invention only uses 4 layers of A-Adapter in each task and combines it with the pre-trained model, so the introduced number of parameters is approximately 0.5M, accounting for 0.46% of the total number of parameters of the BERT model, and the introduced number of parameters is very small.
[0044] See Figure 3As shown in the figure, in the embodiments of this case, the A-Adapter is used in the model for continuous relation extraction tasks. The pre-trained model with all parameters frozen is on the far left. Each task is trained using a separate combined A-Adapter model, sharing the lower layers of the Transformer in the pre-trained model. Each A-Adapter is combined in parallel with the upper layers of the Transformer in the pre-trained model. The input is the superposition of the output of the previous pre-trained model's hidden layer and the output of the previous A-Adapter. The pre-trained model, as the backbone network, keeps all parameters frozen. Each task uses a model composed of multiple A-Adapters for independent training to learn specific task knowledge. Each A-Adapter is combined in parallel with the Transformer layers in any pre-trained model. During the training phase, each task uses multiple combined A-Adapters for efficient parameter fine-tuning.
[0045] Furthermore, in the embodiments of this case, for the target text, first freeze the parameters of the pre-trained model, then add position markers of the head and tail entities around the entity in the input text sentence, and use the BERT model to extract the feature vectors of the text sentence.
[0046] The BERT model can be used to encode the input sentence to extract feature information. BERT is a bidirectional encoding representation model composed of multiple Transformer layers, with excellent language representation ability, and has refreshed the historical best results in 11 NLP tasks. To improve the performance of the relation classification task, special markers are added around the entities in the input sentence to enhance the context information contained in the encoded vector. [E11], [E12], [E21], [E22] can be used to represent the start and end positions of the head entity and the tail entity respectively, and the vector representations of [E11] and [E21] are used to replace the vector representation of the entity. Therefore, the vector representation of an input sentence is as shown in Equation (5), and the output vectors of all layers of the pre-trained model are as shown in Equation (6).
[0047]
[0048] h = [h1, h2, …, h N (6)
[0049] where represents the output vector of the l-th Transformer layer of the pre-trained model, and represent trainable parameters. LayerNorm(·) represents the layer normalization operation. h and d represent the dimensions of the hidden layer vector of the pre-trained model and the sentence representation vector respectively. To ensure that this method is consistent with the vector dimension directly encoded by the pre-trained model, generally, the sizes of h and d can be set to be equal.
[0050] Each task is trained using multiple A - Adapters in combination. The input vector of each A - Adapter is the superposition of the output of the hidden layer of the previous pre - trained model and the output of the previous A - Adapter. The calculation formula is shown in Equation (7). The output vector is the hidden layer vector obtained after the input vector passes through the dimensionality reduction projection layer, the multi - head self - attention layer, the dimensionality increase projection layer and the residual structure, and then the result of layer normalization after being superimposed with the input vector. The calculation formula is shown in Equation (8).
[0051] I l = h l-1 + O l-1 (7)
[0052] O l = LayerNorm(I l + ((I l W down )W MultiHead )W up ) (8)
[0053] Where represents the input vector of the l - th layer A - Adapter, represents the output vector of the (l - 1) - th Transformer layer of the pre - trained model, represents the output vector of the (l - 1) - th layer A - Adapter, represents the trainable parameter of the dimensionality reduction projection layer, represents the trainable parameter of the multi - head attention layer, represents the trainable parameter of the dimensionality increase projection layer, O l represents the output vector of the l - th layer A - Adapter, md represents the dimension after the input vector is reduced in dimension. By setting md << d, the size of the A - Adapter parameter quantity can be restricted.
[0054] Each task of continuous relation extraction can be regarded as a multi - classification problem. Therefore, the model of each task uses softmax to calculate the probability that the input sentence belongs to a certain category. The calculation formula is shown in Equation (9). The cross - entropy loss function is used as the objective function. The calculation formula is shown in Equation (10).
[0055]
[0056]
[0057] Where p k = [p0, p1,... p R-1 represents the classification result of the k - th task, p i ∈ p kThe probability that the input sentence belongs to the i-th class, where R is the predefined number of relationships, O last represents the output vector of the last layer of the A-Adapter, and y = [y0, y1, …, y R-1 is the one-hot representation of the true label of the input sentence. When the input sentence belongs to the i-th class, y i = 1, otherwise y i = 0.
[0058] Furthermore, based on the above method, an embodiment of the present invention further provides a continuous relationship extraction system based on a multi-head self-attention mechanism adapter, including: a model construction module and a relationship extraction module, where,
[0059] The model construction module is used to construct a continuous relationship extraction model by using a pre-trained model as the backbone network. The pre-trained model includes multiple Transformer layers for encoding and decoding input feature vectors, and an adapter is used to learn task knowledge, where the adapter includes: a dimensionality reduction projection layer for reducing the dimensionality of the input feature vector, a multi-head self-attention layer for mining the context semantic information representation of the feature vector, an upsampling projection layer for restoring the mined context semantic representation to the dimension of the input feature vector, and a residual structure for optimizing the residual connection during model training;
[0060] The relationship extraction module is used to treat each task of continuous relationship extraction as a multi-classification problem, and learn the preset task knowledge in the target text through the combined training of multiple adapters.
[0061] Furthermore, an embodiment of the present invention further provides a multi-objective task recognition method, including:
[0062] First, collect sample data containing preset task objects, and use the above continuous relationship extraction method to construct a multi-objective task recognition model;
[0063] Next, use the sample data to train the multi-objective task recognition model;
[0064] Then, input the data to be processed into the trained multi-objective task recognition model, and use the trained multi-objective task recognition model to identify and obtain the data containing the expected objects.
[0065] Continuous relation extraction enables the model to have the ability to recognize all historical task data after sequential task training. For example, assume that Task 1 recognizes picture data containing target objects of chickens and ducks, Task 2 recognizes picture data containing target objects of cats and dogs, and Task 3 recognizes picture data containing target objects of tigers and lions. Through continuous relation extraction, it is expected that after the model finishes training on Task 1, it can recognize chickens and ducks; after continuing to train on Task 2, it can recognize chickens, ducks, cats, and dogs, and so on. However, training with a single model on long-term tasks will be affected by the catastrophic forgetting problem. With this solution, only an Adapter adapter needs to be trained separately for each task to learn specific task knowledge. For example, using Adapter_1 to train Task 1 can recognize chickens and ducks; using Adapter_2 to train Task 2 can recognize cats and dogs, and so on. The models between different tasks are not affected, so the catastrophic forgetting problem can be effectively avoided, which is convenient for applications in target location recognition scenarios such as image classification.
[0066] To verify the effectiveness of the solution in this case, the following further explanation is made in combination with experimental data:
[0067] In the experiment, two publicly available datasets, FewRel and TACRED, which are widely used for continuous relation extraction, are used. Among them, FewRel is a large-scale dataset released by Han et al. when applying few-shot learning to the relation extraction task for the first time. This dataset uses Wikipedia as the corpus and Wikidata as the knowledge graph, and is constructed by means of distant supervision and multiple rounds of manual annotation. It contains 100 relations, and each relation contains 700 instances. And the data of 80 relations are selected, and the data of each relation are added to the training set, validation set, and test set according to the ratio of 3:1:1. The TACRED dataset is created by sampling sentences that mention pairs from the TACKBP newswire and web forum corpus, and is one of the largest and most widely used sentence-level relation extraction datasets, containing a total of 42 relations and 106,264 examples. Since the distribution of this dataset is unbalanced, 40 relations are selected during the data processing, and the training set and test set are divided according to the ratio of 4:1, and the training samples of each relation are limited to within 320, and the test samples are limited to within 40. The specific division of the dataset is listed in Table 1.
[0068] Table 1 Dataset Division
[0069]
[0070]
[0071] The experiment uses the average accuracy to measure the impact of catastrophic forgetting. After training the model on the training set of the new task, the performance of the model on the current task is evaluated by the average accuracy on the test set of the current task, and the performance of the model on the historical tasks is evaluated by the average accuracy on the merged dataset of the historical task test sets.
[0072] The model uses Pytoych to build a deep learning neural network. The bert_base_uncased pre-trained model is used as the backbone network. All parameters are pre-frozen before training the model. The dimension of the A-Adapter hidden layer is 72. Dai et al. found through experiments that most of the fact-related knowledge is distributed in the top layer of the Transformer. Therefore, 4 layers of A-Adapter are used for each task to be combined in parallel with the last 4 Transformer layers of bert-base-uncased respectively. The maximum length of the input sentence is 256. The batch size used to load the dataset is 64, and the number of gradient accumulations is 4. The AdamW optimizer is used to optimize the loss function, and the learning rate is dynamically adjusted using the cosine decay method. The learning rate and weight decay rate of the A-Adapter layer are set to 1e-4, and the learning rate and weight decay rate of the classification layer are set to 1e-3. The hardware environment uses the Nvidia A100 graphics card provided by the OpenI platform, with 80GB of video memory, a 24-core CPU, and 64GB of memory.
[0073] The model is fully trained for 5 epochs. When loading data in each epoch, it is divided into 10 tasks in total. Each task in the FewRel dataset has 8 relations, and each task in the TACRED dataset has 4 relations. The relations divided for each task adopt a completely random sampling strategy, and the random seed is set to 2021 to ensure that the random sequence of relations is exactly the same as that of Cui et al. and Zhao et al. Each task is trained for 10 loops, and the trained model will be inferentially verified on the test set of the current task and the training set merged from all the test sets of the historical tasks.
[0074] Table 2 Accuracy (%) of the A-Adapter model and the baseline model on the inclusion relationship of the current task
[0075]
[0076] Table 2 shows the comparison results of the A-Adapter model and the baseline model on the FewRel dataset and the TACRED dataset for the current task, that is, after the training of each task is completed, the performance of the verification model on the test set of the current task is verified. Table 2 shows the comparison results of the A-Adapter model and the baseline model on the FewRel dataset and the TACRED dataset for historical tasks, that is, after the training of each task is completed, the performance of the verification model on the test sets of all historical tasks is verified. Since Cui et al. and Zhao et al. only listed the comparison results of the model's performance on historical tasks, the present invention reproduced the official open-source code to obtain the results of the model on the current task and historical tasks respectively. and The average value of the training results of 5 rounds is taken for each result.
[0077] The following conclusions can be drawn from Table 2:
[0078] (1) The results of the first task T1 are obtained based on the full fine-tuning of the pre-trained language model. The accuracy rates of the RP-CRE and CRL models on the FewREL and TACRED test sets both exceed 97%. Although the A-Adapter model differs from the RP-CRE and CRL models by 1%-2.1%, considering that the former's number of parameters only accounts for about 1% of the BERT model's parameters, this gap is acceptable.
[0079] (2) After task T1, the RP-CRE and CRL models continue to be fine-tuned on the new task training set. It can be found that the performance of the model on each task generally gets worse and worse. Compared with task T1, the accuracy rate of the CRL model on task T8 has decreased by 16.3%. This shows that continuously fine-tuning the pre-trained language model will seriously damage the knowledge of the pre-trained language model, which is also a catastrophic forgetting problem. Since the MA-Adapter-CRE model uses independent A-Adapters for training on each task, the accuracy rate of the A-Adapter model remains at about 95% on each task.
[0080] Table 3 shows the comparison results of the A-Adapter model on historical tasks, that is, after each task is completed, the extraction effect of the verification model is verified after mixing all the historical task test sets. As listed in Table 3, where T1-2 means the historical tasks are T1 and T2, T1-3 means the historical tasks are T1, T2 and T3, and so on. The following conclusions can be drawn from Table 3:
[0081] (1) As the number of tasks increases, the accuracy of the baseline model on historical tasks gradually decreases. On the one hand, it is because the pre-trained language model is affected by the catastrophic forgetting problem, that is, after the baseline model fine-tunes the pre-trained language model, it destroys the knowledge of the pre-trained language model, resulting in a decrease in the accuracy of the model on each task. On the other hand, it is because the relation classifier is affected by the catastrophic forgetting problem. After the model learns new tasks, it destroys the understanding of the knowledge of historical tasks. The accuracy of the A-Adapter model on the historical task mixed test set has always remained above 90%, far exceeding the baseline model, proving the superiority of this method. This is because the relation classifier for each task only has a good classification effect on the dataset under that task. Therefore, when a piece of text and a pair of entities are input into the relation classifiers of all historical tasks, the prediction result with the highest probability can be selected to obtain the prediction result of the input data.
[0082] (2) It can be found from Table 3 that the accuracy of the A-Adapter model on the FewRel dataset task T1-4 and the accuracy on the TACRED dataset task T1-9 have increased by 0.5% and 0.3% respectively compared to the previous task. By analyzing the data in Table 2, it is found that the reason is that the accuracy of the A-Adapter model on the FewRel dataset task T4 and the accuracy on the TACRED dataset task T9 have increased by 4.7% and 3.7% respectively compared to the previous task. This shows that the overall extraction effect of the A-Adapter model can be improved by adjusting the A-Adapter for specific tasks.
[0083] Table 3 Accuracy (%) of the A-Adapter model and the baseline model on the inclusion relationship of historical tasks
[0084]
[0085] To explore the computational efficiency of the A-Adapter model, the experiment also counted the time used by the A-Adapter model in each round of training. As Figure 4 and Figure 5As shown, the average training time for each result on the curve is taken over 5 rounds. It can be seen from the figure that the training time of the A-Adapter model is significantly less than that of the baseline model. The RP-CRE model takes a long time because after the model finishes training on each task, it also needs to screen typical historical task data as memory samples, calculate the prototype vector representation of the relationship, perform replay operations on the memory samples, and optimize the embedding vectors of new data. These operations are all time-consuming. The CRL model takes even longer because when training the memory network, it uses contrastive learning to calculate the similarity between each memory sample, and in addition, it needs to use the knowledge distillation method to make the model replay the memory samples. The A-Adapter model does not need to screen, store, and replay typical samples, so the training time is greatly reduced. Intuitively, in the validation stage, the A-Adapter for each task needs to be used to predict each input sample. As the number of tasks increases, the validation time will become longer and longer. However, since the pre-trained model is shared among all tasks, in the validation stage, only the pre-trained model needs to be used to encode the input sample once to obtain the output vector of each Transformer layer, which can then be shared with the A-Adapter model for each task. Note that the A-Adapter model designed for each task in the present invention has only 4 layers and the hidden layer dimension is only 72, so the encoding time using A-Adapter is very little.
[0086] To explore the computational efficiency of the A-Adapter model, the experiment further counted the video memory usage and model file size of the A-Adapter model after each round. As Figure 6 shown, compared with the RP-CRE model and the CRL model, the video memory used by the A-Adapter model is reduced by 2060MB - 2368MB. This is because the model of the present invention does not need to store historical data, and the number of parameters of the A-Adapter model used for each task is only about 0.5M, and the increase in the occupied video memory can be basically ignored. As Figure 7 shown, the size of the bert_base_uncased model file used in the solution of this case is 421MB, while the size of the model file using 4-layer A-Adapter is only 4.4MB, greatly reducing the space occupancy.
[0087] There are many factors affecting the performance and size of the A-Adapter model, such as the position of parallel combination with the pre-trained model, the size of the hidden layer dimension, and the components in the bottleneck structure. By designing multiple groups of ablation experiments, the influence degree of these factors on the A-Adapter model is verified. The experimental results are listed in Table 4.
[0088] (1) The pre-trained bert_base_uncased model used in the experiment has a total of 12 layers. A-Adapter is combined in parallel with the last four layers of the pre-trained model. Therefore, the first experiment is to combine A-Adapter with the first four layers, the middle four layers, the odd layers, and the even layers of the pre-trained model in parallel to study the impact of different combination positions on the model.
[0089] (2) The dimension size of A-Adapter is set to 72, which is 6 times the number of self-attention heads, which is 12. Therefore, the second experiment is to set the dimension size of A-Adapter to 1 times, 3 times, 12 times, and 24 times the number of self-attention heads, that is, 12, 36, 144, 288, to study the impact of different A-Adapter dimensions on the model.
[0090] (3) Add a multi-head attention mechanism layer to the bottleneck structure to enhance the fitting ability of A-Adapter. Therefore, the third experiment is to replace the multi-head self-attention mechanism layer with a Relu non-linear activation function layer and a Transformer layer to study the impact of different bottleneck structures on the model.
[0091] Table 4 Results of ablation experiments
[0092]
[0093]
[0094] The following conclusions can be drawn from Table 4:
[0095] (1) Combining A-Adapter with the bottom four layers of the pre-trained model has the worst effect, followed by the middle four layers, and then the last four layers, indicating that using A-Adapter to fit the top layers of the pre-trained model can better adapt to downstream tasks; the odd layers and the even layers have basically the same effect and are a little better than the model of the present invention, indicating that increasing the number of layers of A-Adapter can improve the effect, but it will increase the encoding time and the number of model parameters;
[0096] (2) Both too low or too high dimensions of A-Adapter will have a negative impact on the model. The former will cause the model to underfit, and the latter will cause the model to overfit. According to the experimental results of the present invention, setting the dimension of A-Adapter to 3 to 12 times the number of attention heads can achieve better results;
[0097] (3) After replacing the multi-head attention mechanism layer with the Relu non-linear activation function layer, the performance of the model on each task decreased by basically 15%, indicating that the simple bottleneck structure model has a poor fitting effect, and introducing the attention mechanism can significantly improve the effect of the model; replacing the multi-head attention mechanism layer with the Transformer layer has a worse effect. On the one hand, it shows that the Transformer layer has a high dimension and overfitting occurs. On the other hand, it shows that the forward feedback layer of the Transformer layer duplicates the function of the dimensionality increase projection layer in the A-Adapter, and at the same time increases the computational complexity and the number of model parameters.
[0098] In summary, the solution of this case to combine the A-Adapter and the last four Transformer layers of the pre-trained model in parallel and set the A-Adapter to 72 is an empirical result, which can not only ensure that the model has a high accuracy rate, but also reduce the file size and the number of parameters of the model. In the actual use process, the number of layers of the A-Adapter and the dimension of the A-Adapter can be appropriately increased within an acceptable range. The experimental results prove that adding a multi-head attention mechanism layer in the bottleneck structure can not only enhance the fitting ability of the A-Adapter well, but also reduce the risk of overfitting of the A-Adapter.
[0099] Unless otherwise specifically stated, the relative steps, numerical expressions and values of the components and steps set forth in these embodiments do not limit the scope of the present invention.
[0100] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0101] The units and method steps of each embodiment described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art can use different methods to implement the described functions for each specific application, but such implementation is not considered to exceed the scope of the present invention.
[0102] Those of ordinary skill in the art can understand that all or part of the steps in the above method can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a magnetic disk, or an optical disc, etc. Optionally, all or part of the steps of the above embodiments can also be implemented using one or more integrated circuits. Correspondingly, each module / unit in the above embodiments can be implemented in the form of hardware or in the form of a software function module. The present invention is not limited to any specific form of combination of hardware and software.
[0103] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A continuous relation extraction method based on a multi-head self-attention mechanism adapter, characterized in that, It includes the following content: Use a pre-trained model as the backbone network to construct a continuous relation extraction model. The pre-trained model contains multiple Transformer layers for encoding and decoding input feature vectors, and an adapter is used to learn task knowledge. The adapter includes: a dimensionality reduction projection layer for reducing the dimensionality of the input feature vectors, a multi-head self-attention layer for mining the context semantic information representation of the feature vectors, an upsampling projection layer for restoring the mined context semantic representation to the dimension of the input feature vectors, and a residual structure for optimizing the residuals in model training; the pre-trained model is denoted as: h = [h1, h2, …, h N , represents the output vector of the l-th Transformer layer of the pre-trained model, and represent the trainable parameters of the model. LayerNorm(·) represents the normalization operation. h and d represent the hidden layer vector dimension and the sentence representation vector dimension of the pre-trained model respectively. N represents the total number of Transformer layers in the pre-trained model. h 11 , h 21 represent the vector representations corresponding to the head and tail entity positions; the input vector representation of the l-th layer Transformer layer corresponding to the adapter in the continuous relation extraction model is: I l = h l-1 + O l-1 , O l = LayerNorm(I l + ((I l W down )W MultiHead )W up ), represents the output vector of the (l - 1)-th Transformer layer of the pre-trained model, represents the output vector of the adapter corresponding to the (l - 1)-th layer, represents the trainable parameter of the dimensionality reduction projection layer, represents the trainable parameter of the multi-head attention layer, represents the trainable parameter of the upsampling projection layer. O l represents the output vector of the adapter corresponding to the l-th layer. dm represents the dimension after reducing the dimensionality of the input vector; cross-entropy is used as the objective loss function in the training of the continuous relation extraction model. The objective loss function is expressed as: p k = [p0, p1,... p R-1 represents the classification result of the k-th task. p i ∈ p k represents the probability that the input sentence belongs to the i-th class. R is the predefined number of relations. O last Denote the output vector of the corresponding adapter for the last layer, y = [y0, y1, …, y R-1 is the one-hot representation of the true label of the input sentence; For the target text, first freeze the parameters of the pre-trained model, then add position markers of the head and tail entities around the entity of the input text sentence, and use the BERT model to extract the feature vectors of the text sentence. Each task of continuous relation extraction is regarded as a multi-classification problem, and the combination training of multiple adapters is used to learn the preset task knowledge in the target text.
2. The continuous relation extraction method based on a multi-head self-attention mechanism adapter according to claim 1, characterized in that, In the model trained by the combination training to learn the target task, each adapter is provided with a Transformer layer corresponding to the parallel combination relationship in the pre-trained model, and the input of the current adapter is the superposition of the vector output of the previous adapter and the output of the Transformer layer parallel to the previous adapter in the combination training.
3. A continuous relation extraction system based on a multi-head self-attention mechanism adapter, characterized in that, It includes: a model construction module and a relation extraction module, where A model construction module is used to construct a continuous relation extraction model by using a pre-trained model as the backbone network. The pre-trained model includes multiple Transformer layers for encoding and decoding input feature vectors, and an adapter is used to learn task knowledge. The adapter includes: a dimensionality reduction projection layer for reducing the dimension of the input feature vector, a multi-head self-attention layer for mining the context semantic information representation of the feature vector, an upsampling projection layer for restoring the mined context semantic representation to the dimension of the input feature vector, and a residual structure for optimizing the residual connection during model training. The pre-trained model is denoted as: h = [h1, h2, …, h N , denotes the output vector of the l-th Transformer layer of the pre-trained model, and denote the trainable parameters of the model. LayerNorm(·) represents the normalization operation. h and d represent the hidden layer vector dimension and the sentence representation vector dimension of the pre-trained model respectively. N represents the total number of Transformer layers in the pre-trained model. h 11 , h 21 denote the vector representations corresponding to the head and tail entity positions. The input vector representation of the adapter corresponding to the l-th Transformer layer in the continuous relation extraction model is: I l = h l-1 + O l-1 , O l = LayerNorm(I l + ((I l W down )W MultiHead )W up ), denotes the output vector of the (l - 1)-th Transformer layer of the pre-trained model, denotes the output vector of the adapter corresponding to the (l - 1)-th layer, denotes the trainable parameter of the dimensionality reduction projection layer, denotes the trainable parameter of the multi-head attention layer, denotes the trainable parameter of the upsampling projection layer. O l denotes the output vector of the adapter corresponding to the l-th layer. dm represents the dimension after reducing the dimension of the input vector. During the training of the continuous relation extraction model, cross-entropy is used as the target loss function, and this target loss function is expressed as: p k = [p0, p1,... p R-1 represents the classification result of the k-th task. p i ∈ p k represents the probability that the input sentence belongs to the i-th class. R is the predefined number of relations. O last Denote the output vector of the corresponding adapter in the last layer, y = [y0, y1, …, y R-1 is the one-hot representation of the true label of the input sentence; The relation extraction module is used for the target text. First, freeze the parameters of the pre-trained model, then add position markers of the head and tail entities around the entity of the input text sentence, and use the BERT model to extract the feature vectors of the text sentence. Each task of continuous relation extraction is regarded as a multi-classification problem, and the combination training of multiple adapters is used to learn the preset task knowledge in the target text.
4. A multi-objective task recognition method, characterized in that, It includes: First, collect sample data containing preset task objects, and use the continuous relation extraction method described in claim 1 to construct a multi-target task recognition model; Next, use the sample data to train the multi-target task recognition model; Then, input the data to be processed into the trained multi-target task recognition model, and use the trained multi-target task recognition model to identify and obtain the data containing the expected object.
5. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; The memory is used to store computer programs; The processor is used to execute the program stored on the memory and implement the method steps described in any one of claims 1 to 2 when the program is executed.
6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the method steps described in any one of claims 1 to 2 are implemented.
Citation Information
Patent Citations
Class label identification method and device based on pre-training language model
CN114995903A
Drug discovery method and apparatus based on relation extraction and machine learning, and device
WO2022047972A1