Relationship extraction method, system, medium and device based on cascaded second-order screening
Through the cascading second-order screening and word-level layer attention mechanism, the problems of poor generalization ability and great noise influence of existing relationship extraction methods are solved, and higher relationship extraction accuracy and model generalization ability are achieved.
Patent Information
- Application Number
- CN202411707362.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-27
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-11-27
AI Technical Summary
The existing relationship extraction methods mainly rely on traditional training strategies, limit the generalization ability of the algorithm, and are difficult to adapt to different supervised learning methods, and the noise influence in remote supervision methods is great.
A relational extraction method based on cascading second-order screening is adopted, through two cascading screening, the first screening result is used as a hint, the second screening result is used as the final predicted value, and combined with the word-level layer attention mechanism, the global and local context semantics in the sentence are captured.
It improves the accuracy of relationship extraction, enhances the generalization ability of the model, can better adapt to the relationship extraction tasks of full supervision and remote supervision, and reduces the impact of noise on the model.
Smart Images

Figure CN119204019B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of natural language processing, and in particular relates to a relationship extraction method, system, medium and equipment based on cascaded second-order screening. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] Relation extraction is a very important task in natural language processing, especially in the context of the explosive growth of more and more information today.
[0004] Fully supervised relation extraction is to determine the relationship category between each entity pair given correctly labeled entities. Currently, there are two main methods for fully supervised relation extraction: pipeline and joint.
[0005] The remote supervision method uses a large amount of data labeled with heuristic rules for training, so it does not require high-quality labeled data like full supervision. However, the data labeled with heuristic rules often contains a lot of noise, so the contribution of the remote supervision relationship extraction method is mainly to reduce the impact of noise on the model.
[0006] Currently, relation extraction mainly relies on traditional training strategies, that is, obtaining entity and relation semantic representations through feature extraction, and then predicting relation categories. The model is continuously optimized through gradient descent and back propagation until convergence. This method belongs to first-order classification, and the model only predicts once to get the result. Whether it is fully supervised or remotely supervised relation extraction, feature engineering is performed on the basis of first-order classification to improve the training effect. This model limits the generalization ability of the algorithm, because they can usually only optimize feature engineering under a specific supervised learning framework and are difficult to adapt to other supervised learning methods. Summary of the invention
[0007] In order to solve the technical problems existing in the above-mentioned background technology, the present invention provides a relationship extraction method, system, medium and equipment based on cascaded second-order screening. Through two cascade screenings, the result of the first screening (filter type) is used as a prompt, and the result of the second screening is used as the final prediction value to fully utilize the equivalent or inclusion connections between different relationships and improve the accuracy of relationship extraction.
[0008] In order to achieve the above object, the present invention adopts the following technical solution:
[0009] A first aspect of the present invention provides a relationship extraction method based on cascaded second-order screening, which comprises:
[0010] Get a sentence and insert entity tags before and after each entity in the sentence to get a sentence sequence;
[0011] Based on the sentence sequence, the relationship type of the sentence is predicted through the relationship extraction model;
[0012] Among them, the relationship extraction model extracts entity features and full sentence representation features from the sentence sequence, and then performs vector splicing to obtain first-order classification features; based on the first-order classification features, the filter type is obtained through activation functions and independent variable maximum functions; through several layers of encoders, several feature matrices are extracted from the sentence sequence and stacked to obtain feature blocks; based on the feature blocks, fine-grained features are obtained through the word-level attention mechanism, and the fine-grained features and the entity features are spliced to obtain second-order classification features; based on the second-order classification features and the filter type, the relationship type of the sentence is predicted.
[0013] Furthermore, the entity tags inserted before and after each entity are and ,in, represents the kth entity, , Represents a set of predefined entity types.
[0014] Furthermore, the word-level attention mechanism is expressed as: ;in, represents the feature block, represent tanh Activation function, represent sigmoid Activation function, represents average pooling, M represents maximum pooling, and represents the weight, and represents the convolution operation, Represents fine-grained features.
[0015] Furthermore, the prediction of the relationship type of the sentence adopts an activation function and a function for maximizing independent variables.
[0016] Furthermore, the loss function used in the training of the relationship extraction model is: ,in, , = - - - ,in, and represents the noise parameter, n represents the number of filter types, is the true label of the filter type of sample x, represents the softmax activation function, represent sigmoid function, and Represents the number of relationship types in the filter. is the score of each relationship type, is the score of each filter type, Representation sample The true label of the relationship type, Representation sample The true label of the relationship type, Representation sample The true label of the relationship type.
[0017] A second aspect of the present invention provides a relation extraction system based on cascaded second-order screening, comprising:
[0018] A data acquisition module is configured to: acquire sentences and insert entity tags before and after each entity in the sentences to obtain a sentence sequence;
[0019] A relation extraction module is configured to: predict the relation type of a sentence based on a sentence sequence through a relation extraction model;
[0020] Among them, the relationship extraction model extracts entity features and full sentence representation features from the sentence sequence, and then performs vector splicing to obtain first-order classification features; based on the first-order classification features, the filter type is obtained through activation functions and independent variable maximum functions; through several layers of encoders, several feature matrices are extracted from the sentence sequence and stacked to obtain feature blocks; based on the feature blocks, fine-grained features are obtained through the word-level attention mechanism, and the fine-grained features and the entity features are spliced to obtain second-order classification features; based on the second-order classification features and the filter type, the relationship type of the sentence is predicted.
[0021] Furthermore, the entity tags inserted before and after each entity are and ,in, represents the kth entity, , Represents a set of predefined entity types.
[0022] Furthermore, the word-level attention mechanism is expressed as: ;in, represents the feature block, represent tanh Activation function, represent sigmoid Activation function, represents average pooling, M represents maximum pooling, and represents the weight, and represents the convolution operation, Represents fine-grained features.
[0023] A third aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the relationship extraction method based on cascaded second-order screening as described above.
[0024] The fourth aspect of the present invention provides a computer device, including a computer-readable storage medium, a processor, and a computer program stored on the computer-readable storage medium and executable on the processor, wherein when the processor executes the program, the steps in the relationship extraction method based on cascaded second-order screening as described above are implemented.
[0025] Compared with the prior art, the present invention has the following beneficial effects:
[0026] The present invention uses two cascade screenings, takes the result of the first screening as a hint, and uses the result of the second screening as the final prediction value, so as to make full use of the equivalent or inclusion relationship between different relations and improve the accuracy of relation extraction.
[0027] The present invention jointly trains the screening tasks of two granularities, hoping that the loss offset generated by the fine-grained screening task during the training process can promote the learning of the coarse-grained screening task and improve the accuracy of the relationship extraction model.
[0028] The present invention proposes a word-level layer attention mechanism to capture the rich global and local contextual semantics in a sentence and improve the accuracy of relation extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0030] Figure 1 is a flow chart of a relationship extraction method based on cascaded second-order screening according to the first embodiment of the present invention;
[0031] Figure 2 is a flowchart of the relationship extraction model training of the first embodiment of the present invention;
[0032] Figure 3 is an overview diagram of a relationship extraction method based on cascaded second-order screening according to the first embodiment of the present invention;
[0033] Figure 4is an overview diagram of the word-level attention mechanism of the first embodiment of the present invention;
[0034] Figure 5 It is a structural diagram of a computer device according to a fourth embodiment of the present invention. DETAILED DESCRIPTION
[0035] To make the purpose, technical solution and advantages of the embodiments of the present invention more clear, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0036] It should be noted that the following detailed descriptions are all illustrative and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.
[0037] Embodiment 1
[0038] This embodiment provides a relationship extraction method based on cascaded second-order screening.
[0039] Existing relationship extraction learning methods are mainly divided into full supervision and semi-supervision. However, there are differences in the networks under different supervised learning methods, and they cannot be applied to both supervised learning modes at the same time. Therefore, this embodiment changes the traditional training strategy and proposes a cascaded second-order screening network model suitable for full supervision and remote supervision.
[0040] Since each row of the feature graph formed by text data represents a word, and the semantics of each word requires all columns to be represented, the entire feature graph is not only arranged in order but also interconnected. In view of this feature, this embodiment proposes a word-level layer attention mechanism to capture the rich global and local contextual semantics in the sentence.
[0041] The relationship extraction method based on cascaded second-order screening provided in this embodiment explores a new problem-solving paradigm - weakening the role of feature engineering, solving the differences between models under different supervised learning, and thus proposing a universal relationship extraction method under different supervised learning. The traditional first-order classification method ignores the intrinsic connection between different relationships (for example, there are connections such as equivalence or inclusion between different relationships), and this connection is an extremely important hidden information for the classification of relationships. Therefore, this embodiment proposes a cascaded second-order screening method to make full use of this hidden information. In general, this method uses two cascade screenings, uses the results of the first screening as a hint, and uses the results of the second screening as the final prediction value.
[0042] In this embodiment, the SemEval-2010 Task 8 and NYT10 datasets, which are widely used in fully supervised and remotely supervised relationship extraction tasks, are selected respectively.
[0043] Among them, SemEval-2010 Task 8 has 10 relationship types, including 9 annotated types and Other types. It is a small dataset with accurate annotations. The relationship type refers to the relationship between entity pairs in a sentence. Taking the SemEval-2010 Task 8 dataset as an example, it includes 10 relationships such as Cause-Effect, Component-Whole, Message-Topic, and Product-Producer. The task goal of relationship extraction is to predict which relationship type r the entity pairs in a sentence have.
[0044] Among them, NYT10 aligns the unlabeled New York Times Corpus with the remote knowledge base, and both the training set and the test set are remotely supervised and annotated.
[0045] Suppose, one of the sentence instances S=(In 2009 Speaker Morgan once again won re-election to lead the Navajo Nation Council.).
[0046] This embodiment provides a relationship extraction method based on cascaded second-order screening, such as Figure 2 As shown, the following steps are included:
[0047] Step 1: Perform coarse-grained screening tasks.
[0048] For a given input sequence S, the goal of the coarse-grained screening task is to filter the input sequence S into the corresponding "filter". Different filters Contains different sets of relationships That is, the output of the coarse-grained screening task is .
[0049] Furthermore, the operation in step 1 is specifically implemented as follows:
[0050] Step 101: Pre-processing before input.
[0051] This embodiment specifies a set of special entity marking symbols, namely , , , where ek represents the kth entity, , Represents a set of predefined entity types.
[0052] Inserting special entity tags before and after two entities in a sentence helps the model better identify entity information. The process can be expressed as: = .
[0053] Therefore, the sentence instance S is formalized as: ( In, 2009, <e1:person> , Speaker, Morgan,< / e1:person> , once, again, won, re-election, to, lead, the, <e2:org> , Navajo,Nation, Council,< / e2:org> , .).
[0054] Step 102: Obtain the first-order classification features of the sentence instance. This feature is composed of the entity feature and the whole sentence representation feature. The specific implementation process is as follows: .in is a sequence of sentences The first-order classification feature representation obtained after encoding by the pre-trained model BERT (Bidirectional Encoder Representation from Transformers, a natural language processing model based on deep learning proposed by Google in 2018, represents vector concatenation, is the representational feature of the whole sentence. It is the characteristic representation of the entity.
[0055] Step 103: Based on the first-order classification features Predict the "filter" type corresponding to the input sequence S.
[0056] First go through softmax The activation function obtains the probability distribution of each "filter" type ; then based on argmax Function (finding the maximum function of independent variables) to get the prediction result . The process can be expressed as: ,in, express softmax Activation function.
[0057] Step 2: Perform fine-grained screening tasks.
[0058] set up Stands for "Filter" A set of predefined relation types in the task is obtained by receiving the output of the coarse-grained filtering task. , for different "filters" , for the sequence in the "filter" Perform more detailed screening and finally determine the relationship type r to which the entity pairs in the input sequence belong. That is, the output of this task is .
[0059] Furthermore, the operation in step (2) is specifically implemented as follows:
[0060] Step 201: Obtain feature block X. For a pre-trained model BERT that includes L layers of encoders connected in series, each layer of encoder can output a feature matrix , input sequence After BERT encoding, L feature matrices can be obtained, namely , stack the L feature matrices to obtain the feature block X. The process can be expressed as: ,in, Represents a stacking operation of features.
[0061] Step 202: Input the feature block X into the word-level attention mechanism, make full use of the feature representations of different granularities at each layer of the pre-trained model BERT, and then obtain fine-grained features Y with rich semantics.
[0062] Among them, the word-level attention mechanism, such as Figure 4 As shown. In particular, for the feature block X (i.e., the L-layer feature matrix, whose size is the sentence length × hidden dimension), the attention mechanism uses local average pooling and local maximum pooling instead of global average pooling and global maximum pooling, with the purpose of compressing the feature map to a column vector of size length×1 (length refers to the maximum length of a sentence) instead of a 1×1 vector. This not only collects global token (character) level information, but also retains the order information of the original text features. Most importantly, by training, the value of the column vector can be regarded as the weight of each token. In a sentence, meaningful words will be given a larger weight, while relatively meaningless words will be assigned a smaller weight. After the pooling operation, in order to obtain local information at the token level, a convolution kernel is introduced, and a convolution kernel of size (k, 1) is used to collect local connections between k words. Because the relationship judgment of a sentence is often expressed in a few local words. The process can be expressed as: ; .in, LA represents the word-level attention mechanism, represent tanh Activation function, represent sigmoid Activation function, is a set of trainable parameters, represents average pooling, M stands for max pooling; and represents the weight, , ; C refers to the number of hidden layers of feature blocks input to the layer attention mechanism, where C=L, L-layer feature matrix; represents MLP (Multi-layer Perceptron); r is a reduction ratio used to form a bottleneck-shaped gating mechanism; and Represents a convolution operation; + represents an Add operation. X is processed by the word-level attention layer to obtain the attention vector s. Then, X is multiplied by the corresponding element in the attention vector to scale each channel of the input feature to obtain the fine-grained feature Y.
[0063] Step 203: Obtain the second-order classification features of the sentence instance.
[0064] The second-order classification features of a sentence instance are composed of entity features and fine-grained features Y with rich semantics. The specific implementation process is as follows: ;in, It is the second-order classification feature of the sentence instance.
[0065] Step 204: Based on the second-order classification features and filter type, predict the relationship type r corresponding to the input sequence S. First, according to The value of In "Filter" The corresponding relationship set The probability distribution of , , then based on argmax Function to get the predicted relationship . The process can be expressed as: , For example, if i=1, then the input sequence S will be in the filter The set of relations in In the filter, if i=2, the input sequence S will be classified. The set of relations in and so on.
[0066] Step 3: Calculate the loss of the coarse-grained screening task.
[0067] In this task, the goal is to filter the given input sequence S into the corresponding "filter" , which is a regular classification task. Therefore, the loss of the coarse-grained screening task can be expressed as: , ;in, represent softmax Activation function, is the true label of the filter type of sample x, It is the score of each "filter" obtained after the coarse-grained screening task. Representatives passed softmax The predicted value after activation function After the effect of , the sum of the probability values of each category is 1.
[0068] Step 4: Loss calculation for fine-grained screening tasks.
[0069] This task plays a connecting role and aims to receive the output of the coarse-grained screening task. , then determine the sequence Belongs to "Filter" Therefore, the fine-grained screening task has The branch loss, Represents the number of types of "filters". In this embodiment, The value of is 3. It is worth noting that in the "Filter" There is only one relationship type "Other", so in the "Filter" In , the branch loss is a binary cross entropy loss. The loss of the fine-grained screening task can be expressed as: = - = ; ; ; .
[0070] in, represent softmax Activation function, represent sigmoid Activation function, and represents the number of relationship categories in the corresponding "filter", Indicates that it belongs to the filter Sample The true label of the relationship type, Indicates that it belongs to the filter Sample The true label of the relationship type, Indicates that it belongs to the filter Sample The true label of the relationship type, is the score of each relationship type obtained after the fine-grained screening task. Since there is a binary cross entropy loss in the total loss, the branch loss is used As the activation function, It is a common S-shaped function that can map variables to intervals It is used for binary classification tasks.
[0071] Step 5: Joint training of coarse-grained screening tasks and fine-grained screening tasks.
[0072] This embodiment jointly trains the two-granularity screening tasks, hoping that the loss offset generated by the fine-grained screening task during the training process can promote the learning of the coarse-grained screening task. Since the incorrect classification of the coarse-grained screening task will cause error propagation, resulting in an increase in the error of the fine-grained screening task, the strategy of jointly training the two tasks is adopted, and the error of the fine-grained screening task is used to prompt the coarse-grained screening task to make corrections in time. Therefore, the total loss in the training phase is the sum of the losses of the two screening tasks. This embodiment uses a multi-task likelihood method to adaptively balance the two losses. That is: ≈ . Where W represents the network weight, Representation vector The cth element of represents the noise parameter.
[0073] Step 6: Use the trained relation extraction model to predict the relation type of the sentence, such as Figure 1 and Figure 3 As shown, the method includes: the relation extraction model extracts entity features and whole sentence representation features from the sentence sequence, and then performs vector concatenation to obtain first-order classification features; based on the first-order classification features, the filter type is obtained through activation functions and independent variable maximum functions; for the sentence sequence, several feature matrices are obtained through several layers of encoders, and they are stacked to obtain feature blocks; based on the feature blocks, fine-grained features are obtained through a word-level attention mechanism, and the fine-grained features and the entity features are concatenated to obtain second-order classification features; based on the second-order classification features and the filter type, the relationship type of the sentence is predicted.
[0074] It should be noted that the formula derivation in step 4 is:
[0075] set up , , , then the original formula is: = .
[0076] If you will Transformed to ,but: =-[ …+ +…+ +…+ +…+ ]- = .
[0077] It can be seen that formula I Formula II, if formula I Formula II only needs a simple scaling to be established. Zoom out times, will Zoom out times. = = = .
[0078] Table 1 shows the experimental results of the model of this embodiment and the existing models on the SemEval-2010 Task 8 dataset. Compared with previous work, the model of this embodiment has achieved similar or more advanced results in three evaluation parameters. It can be observed that the model of this embodiment has reached more than 90 in both precision and recall indicators. Therefore, the ability of this model to find positive samples and predict positive samples is relatively good, and it is a more balanced model compared with other models. Specifically, in terms of precision, the model of this embodiment has reached 91.71, which is not as good as the REDN model (94.2), but compared with other models, it is also an excellent result, which has increased by +3.71%, +4.68%, +2.27%, +2.07%, and +1.81% respectively. In terms of recall, the model of this embodiment has achieved similar recall values (90.97, 90.59) as some existing models. It is also noted that REDN, which performs very well in precision, has some shortcomings in recall. Similarly, the model of this embodiment is slightly lower than SPOT in recall, but far exceeds SPOT in precision. Since the model of this embodiment achieves good results in both precision and recall without obvious shortcomings, the model of this embodiment achieves an F1 value of 91.18, surpassing the existing models.
[0079] Table 1. Comparison with existing methods on the SemEval-2010 Task 8 dataset
[0080]
[0081] Among them, Att-RCNN is a relation extraction model that combines RNN (recurrent neural network) and CNN (convolutional neural network); MALNet is a new multi-head attention long short-term memory relation extraction model with filtering mechanism; R-BERT is a relation extraction model based on BERT fine-tuning; BERTEM+MTB is a relation extraction model based on blank matching training; BERT-ECM is a relation extraction model based on piecewise convolution; REDN is a relation extraction model based on asymmetric kernel inner product function; RELA is a relation extraction model with label enhancement; KLG is a relation extraction model based on label graph; SPOT is an information extraction model based on knowledge enhancement.
[0082] Table 2 shows the experimental results on NYT10. The model of this embodiment has an improvement of 1.8ptAUC compared with the HFMRE model, surpassing the current most advanced model. In terms of Micro f1 index, the model of this embodiment achieved the best result on the NYT10 dataset, greatly surpassing the existing methods.
[0083] Table 2. Comparison with existing methods on the NYT10 dataset
[0084]
[0085] Among them, AUC is the area under the ROC (Receiver Operating Characteristic curve), which is used to measure the overall performance of the model; RL-DSRE is a remote supervision relationship extraction model based on reinforcement learning; RESIDE is a remote supervision relationship extraction model using knowledge base information; Bert+Att is a remote supervision relationship extraction model based on attention mechanism; Bert+Avg is a remote supervision relationship extraction model based on average method; Bert+One is a remote supervision relationship extraction model based on at least one assumption; CIL is a remote supervision relationship extraction model based on contrastive learning; PARE is a remote supervision relationship extraction model based on paragraph summary; HFMRE is a remote supervision relationship extraction model based on Huffman tree.
[0086] Embodiment 2
[0087] This embodiment provides a relationship extraction system based on cascaded second-order screening, which specifically includes:
[0088] A data acquisition module is configured to: acquire sentences and insert entity tags before and after each entity in the sentences to obtain a sentence sequence;
[0089] A relation extraction module is configured to: predict the relation type of a sentence based on a sentence sequence through a relation extraction model;
[0090] Among them, the relationship extraction model extracts entity features and full sentence representation features from the sentence sequence, and then performs vector splicing to obtain first-order classification features; based on the first-order classification features, the filter type is obtained through activation functions and independent variable maximum functions; through several layers of encoders, several feature matrices are extracted from the sentence sequence and stacked to obtain feature blocks; based on the feature blocks, fine-grained features are obtained through the word-level attention mechanism, and the fine-grained features and the entity features are spliced to obtain second-order classification features; based on the second-order classification features and the filter type, the relationship type of the sentence is predicted.
[0091] It should be noted here that each module in this embodiment corresponds to each step in Example 1 one by one, and the specific implementation process is the same, which will not be repeated here.
[0092] Embodiment 3
[0093] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps in the relationship extraction method based on cascaded second-order screening as described in the first embodiment above are implemented.
[0094] Embodiment 4
[0095] This embodiment provides a computer device, such as Figure 5 As shown, it includes a computer-readable storage medium 1003, a processor 1001, a communication interface 1002, and a computer program stored on the computer-readable storage medium 1003 and executable on the processor 1001, wherein the processor 1001, the communication interface 1002, and the computer-readable storage medium 1003 can be connected via a bus or other means. The communication interface 1002 is used to receive and send data, and when the processor 1001 executes the program, the steps in the relationship extraction method based on cascaded second-order screening as described in the first embodiment above are implemented.
[0096] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A relation extraction method based on cascaded second-order screening, characterized in that: include: Get a sentence and insert entity tags before and after each entity in the sentence to get a sentence sequence; Based on the sentence sequence, the relationship type of the sentence is predicted through the relationship extraction model; Among them, after the relation extraction model extracts entity features and full sentence representation features from the sentence sequence, it performs vector concatenation to obtain first-order classification features; based on the first-order classification features, the filter type is obtained through activation functions and independent variable maximum functions; through several layers of encoders, several feature matrices are extracted from the sentence sequence and stacked to obtain feature blocks; based on the feature blocks, fine-grained features are obtained through the word-level attention mechanism, and the fine-grained features and the entity features are concatenated to obtain second-order classification features; based on the second-order classification features and the filter type, the relationship type of the sentence is predicted; Different filters contain different sets of relationships; The relation extraction model extracts entity features and full sentence representation features from the sentence sequence, and then performs vector concatenation to obtain first-order classification features. The specific implementation is as follows: in, is a sequence of sentences The first-order classification feature representation obtained after encoding by the pre-trained model BERT, represents vector concatenation, is the representational feature of the whole sentence. It is the characteristic representation of the entity; Based on the first-order classification features, the filter type is obtained through the activation function and the maximum function of the independent variable; the specific implementation is: First go through softmax The activation function obtains the probability distribution of each filter type ; then based on argmax Function gets the prediction result ; The process is expressed as: ,in, express softmax Activation function; The fine-grained features and the entity features are concatenated to obtain second-order classification features, specifically: ;in, It is the second-order classification feature of the sentence instance; according to The value of In the filter The corresponding relationship set The probability distribution of , , then based on argmax Function to get the predicted relationship ; The process is expressed as: , ; If i=1, then the input sequence S will be in the filter The set of relations in In the filter, if i=2, the input sequence S will be classified. The set of relations in and so on; The word-level attention mechanism is expressed as: ;in, represents the feature block, represent tanh Activation function, represent sigmoid Activation function, represents average pooling, M represents maximum pooling, and represents the weight, and represents the convolution operation, Represents fine-grained features; For the joint training of coarse-grained screening tasks and fine-grained screening tasks, a strategy of joint training of the two tasks is adopted, and the error of the fine-grained screening task is used to prompt the coarse-grained screening task to make timely corrections.
2. The relationship extraction method based on cascaded second-order screening according to claim 1, characterized in that: The entity tags inserted before and after each entity are and ,in, represents the kth entity, , Represents a set of predefined entity types.
3. The relationship extraction method based on cascaded second-order screening according to claim 1, characterized in that: The prediction of the relationship type of the sentence adopts an activation function and a function for maximizing independent variables.
4. The relationship extraction method based on cascaded second-order screening according to claim 1, characterized in that: The loss function used in the training of the relationship extraction model is: ,in, , = - - - ,in, and represents the noise parameter, n represents the number of filter types, is the true label of the filter type of sample x, represent softmax Activation function, represent sigmoid Type function, and Represents the number of relationship types in the filter. is the score of each relationship type, is the score of each filter type, Representation sample The true label of the relationship type, Representation sample The true label of the relationship type, Representation sample The true label of the relationship type.
5. A relation extraction system based on cascaded second-order screening, characterized in that: include: A data acquisition module is configured to: acquire sentences and insert entity tags before and after each entity in the sentences to obtain a sentence sequence; A relation extraction module is configured to: predict the relation type of a sentence based on a sentence sequence through a relation extraction model; Among them, after the relation extraction model extracts entity features and full sentence representation features from the sentence sequence, it performs vector concatenation to obtain first-order classification features; based on the first-order classification features, the filter type is obtained through activation functions and independent variable maximum functions; through several layers of encoders, several feature matrices are extracted from the sentence sequence and stacked to obtain feature blocks; based on the feature blocks, fine-grained features are obtained through the word-level attention mechanism, and the fine-grained features and the entity features are concatenated to obtain second-order classification features; based on the second-order classification features and the filter type, the relationship type of the sentence is predicted; Different filters contain different sets of relationships; The relation extraction model extracts entity features and full sentence representation features from the sentence sequence, and then performs vector concatenation to obtain first-order classification features. The specific implementation is as follows: in, is a sequence of sentences The first-order classification feature representation obtained after encoding by the pre-trained model BERT, represents vector concatenation, is the representational feature of the whole sentence. It is the characteristic representation of the entity; Based on the first-order classification features, the filter type is obtained through the activation function and the maximum function of the independent variable; the specific implementation is: First go through softmax The activation function obtains the probability distribution of each filter type ; then based on argmax Function gets the prediction result ; The process is expressed as: ,in, express softmax Activation function; The fine-grained features and the entity features are concatenated to obtain second-order classification features, specifically: ;in, It is the second-order classification feature of the sentence instance; according to The value of In the filter The corresponding relationship set The probability distribution of , , then based on argmax Function to get the predicted relationship ; The process is expressed as: , ; If i=1, then the input sequence S will be in the filter The set of relations in In the filter, if i=2, the input sequence S will be classified. The set of relations in and so on; The word-level attention mechanism is expressed as: ;in, represents the feature block, represent tanh Activation function, represent sigmoid Activation function, represents average pooling, M represents maximum pooling, and represents the weight, and represents the convolution operation, Represents fine-grained features; For the joint training of coarse-grained screening tasks and fine-grained screening tasks, a strategy of joint training of the two tasks is adopted, and the error of the fine-grained screening task is used to prompt the coarse-grained screening task to make timely corrections.
6. The relationship extraction system based on cascaded second-order screening as claimed in claim 5, characterized in that: The entity tags inserted before and after each entity are and ,in, represents the kth entity, , Represents a set of predefined entity types.
7. The relationship extraction system based on cascaded second-order screening as claimed in claim 5, characterized in that: The word-level attention mechanism is expressed as: ;in, represents the feature block, represent tanh Activation function, represent sigmoid Activation function, represents average pooling, M represents maximum pooling, and represents the weight, and represents the convolution operation, Represents fine-grained features.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in the relationship extraction method based on cascaded second-order screening as described in any one of claims 1 to 4 are implemented.
9. A computer device comprising a computer-readable storage medium, a processor, and a computer program stored on the computer-readable storage medium and executable on the processor, characterized in that: When the processor executes the program, the steps in the relationship extraction method based on cascaded second-order screening as described in any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Entity relationship extraction method based on fine-grained semantic information enhancement
CN113051929A