Few-Shot Relation Extraction Method and Device Based on Multi-Knowledge Enhanced Prototype Network
By building a multi-knowledge enhancement prototype network, using a priori knowledge of multi-grained entity types and relationship descriptions, combined with a dual-comparative learning strategy, the accuracy problem of relationship extraction under the condition of few samples is solved, and more efficient text entity recognition and relationship classification are achieved.
Patent Information
- Application Number
- CN202510523318.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-24
AI Technical Summary
The prior art is difficult to effectively capture and understand the characteristics of all relationships in text under the condition of few samples, resulting in poor performance of models in the face of unseen relationship identification and classification.
Build a multi-knowledge enhancement prototype network, adopting semantic encoder, multi-knowledge enhancement learning module, dual-contrast learning module and relation prediction module. By introducing multi-grained entity types and relationship descriptions as prior knowledge, combined with prompt learning design, knowledge enhancement learning of examples and prototypes is carried out to achieve double-contrast learning to improve the accuracy of relationship extraction.
It improves the accuracy of text entity recognition and relationship classification, can fully capture and understand the characteristics of all relationships in the text under a small number of samples, and enhances the generalization ability of the model.
Smart Images

Figure CN120068874B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of natural language processing, and particularly to a few-shot relation extraction method and device based on a multi-knowledge enhanced prototype network. Background Art
[0002] The relation extraction task aims to identify and classify the relationships between different entities in text. Most existing relation extraction methods, especially those based on pre-trained language models (such as the Bidirectional Encoder Representations from Transformers model BERT), have achieved significant performance improvements. However, these methods usually require a large amount of high-quality labeled data to train the model to obtain good performance. Humans can effectively learn new knowledge through a small number of instance samples, mainly due to their rich prior knowledge and background knowledge, enabling them to reason and generalize. For example, after knowing that a person is a doctor, humans can naturally infer that this person may work in a hospital and have medical knowledge. This reasoning ability based on prior knowledge enables humans to be more flexible and efficient when facing new relation extraction tasks. In order to enable machines to learn new relationships through a small number of labeled instances like humans and apply them to real life, the few-shot relation extraction task has become a current research hotspot.
[0003] Few-shot relation extraction aims to train a model by using a small amount of labeled data so that it can quickly learn and make effective inferences under new relation types. Although previous work has improved the performance of few-shot relation extraction by introducing external prior knowledge and semantic representations of pre-trained language models, it still faces the problem of insufficient semantic learning due to the lack of sample quantity. Especially in the expression of complex relation categories, it is difficult for the model to fully capture and understand the characteristics of all relationships through a small number of samples, resulting in poor performance of the model when identifying and classifying unseen relationships. Summary of the Invention
[0004] Based on this, it is necessary to provide a few-shot relation extraction method and device based on a multi-knowledge enhanced prototype network for the above technical problems.
[0005] A few-shot relation extraction method based on a multi-knowledge enhanced prototype network, the method is applied to the text entity recognition and classification scenario, and includes the following steps:
[0006] Divide a data set containing multiple relation categories and corresponding instances into a training set, a validation set, and a test set according to the categories;
[0007] Randomly extract multiple meta-tasks from each of the partitioned datasets, and introduce multi-granularity entity types and relationship descriptions as prior knowledge; among them, each meta-task consists of a support set and a query set, the support set and the query set contain multiple relationship categories and multiple instances corresponding to each relationship category, and each instance consists of a sentence and the relationship category of the entity pair in the sentence;
[0008] Construct a multi-knowledge enhanced prototype network composed of a semantic encoder, a multi-knowledge enhanced learning module, a dual-contrast learning module, and a relationship prediction module; among them, the semantic encoder is used to encode and generate the feature representations of instances and prior knowledge; the multi-knowledge enhanced learning module is used to perform knowledge enhanced learning on the support set instances and query set instances according to the prior knowledge, and generate a prototype enhanced representation according to the obtained enhanced representation of the support set instances; the dual-contrast learning module is used to perform relationship category discrimination learning from two different levels of instances and prototypes respectively by using instance-based contrast learning and prototype-based contrast learning; the relationship prediction module is used to predict the relationship category to which the query set instance belongs according to the similarity between the prototype enhanced representation and the enhanced representation of the query set instance;
[0009] Input the meta-tasks and prior knowledge randomly extracted from the training set and the validation set into the multi-knowledge enhanced prototype network, and construct a comprehensive loss function including the dual-contrast learning loss and the classification loss of relationship prediction for network training and evaluation until a trained network is obtained to perform few-shot relationship extraction tasks.
[0010] In one embodiment, randomly extract multiple meta-tasks from each of the partitioned datasets, and introduce multi-granularity entity types and relationship descriptions as prior knowledge, including:
[0011] Randomly extract multiple meta-tasks from the training set, validation set, and test set respectively for training, validation, and testing. Each meta-task for few-shot relationship extraction consists of a support set and a query set ; among them, the support set contains relationship categories, and the relationship category set is denoted as , and the support set of each relationship category is composed of support set instances of the current relationship category ; the query set also contains the same relationship categories, but the instances in the query set are randomly extracted from the remaining instances of each relationship category query set instances ; among them, represents the i th relationship category, Indicating the relationship category of the k th support set instance respectively indicates the relationship categories of the sentences and entity pairs in Indicating the j th query set instance respectively indicates the relationship categories of the sentences and entity pairs in
[0012] Support set and query set each instance in is composed of a sentence containing the head entity and the tail entity as well as the relationship category indicating the relationship between the head entity and the tail entity ; the goal of the meta-task is to use a very small number of labeled data in the support set to predict the relationship category between the head entity and the tail entity in the sentence of the query set instance ; among them, the relationship categories extracted by each meta-task are different ;
[0013] Furthermore, multi-granularity entity types and relationship descriptions are introduced as prior knowledge; among them, multi-granularity entity types represent all entity type information contained in an entity represents the number of entity types is the i th entity type and ; the relationship description consists of the relationship name of the relationship categoryand the corresponding detailed description
[0014] In one of the embodiments, the steps of the semantic encoder encoding and generating the feature representations of the instance and the prior knowledge include:
[0015] Using the pre-trained language model BERT as the semantic encoder
[0016] For the instances in the support set and the query set, first splice the sentence in the instance with the prompt text with the entity pair to form a prompt template with entity information , and the expression is:
[0017] ;
[0018] Prompt template here Adopt the relational triple structure of entity pairs as the prompt text to activate the relational representation of entity pairs in the instance; among them, [CLS] is called the classification token, which is used to represent the start of the input sequence; [SEP] is called the separator token, which is used to separate two sentences or represent the end of a single sentence; is called the mask token, which is used to represent some chunks of words that require the model to predict the relationship between two entities; then, the prompt template is input into BERT for encoding, and the contextual representation of is used as the instance representation that describes the semantic relationship between entity pairs , and the expression is:
[0019] ;
[0020] Among them, represents BERT; the instance representation includes the support set instance representation and the query set instance representation;
[0021] For the relationship descriptions in the prior knowledge, first, the relationship name of each relationship category and the detailed description are concatenated together to form the input sequence ; then is input into BERT for encoding, and the contextual representation of is used as the relationship representation that depicts the relationship description , and the expression is:
[0022] ;
[0023] For the multi-granularity entity types in the prior knowledge , first, each entity type is transformed into format, then input into BERT for encoding, and the contextual representation of is used as the entity type representation that describes the entity type , and the expression is:
[0024] .
[0025] In one of the embodiments, the steps of the multi-knowledge enhanced learning module for performing knowledge enhanced learning on the support set instance and the query set instance according to the prior knowledge include:
[0026] For the relationship category thek One support set instance , using known relationship categories to guide the selection of multi-granularity entity types; that is, in the multi-knowledge reinforcement learning module, first calculate the entity type representation and the relationship representation corresponding to its belonging relationship category The correlation coefficient between them , and the expression is:
[0027] ;
[0028] Among them, represents the similarity measurement function, that is, the dot product operation; , represents the number of entity types; the entity , and represent the head entity and the tail entity respectively;
[0029] Then, based on the correlation coefficient , selectively fuse all entity types by weighted summation to obtain the entity type representation related to the relationship category in the support set instance, and the expression is:
[0030] ;
[0031] Therefore, the head entity type representation and the tail entity type representation in the support set instance are respectively and ;
[0032] Furthermore, in the multi-knowledge reinforcement learning module, the knowledge graph embedding learning technology is adopted to model the relationship category into the distance transformation from the head entity type to the tail entity type, that is, using the tail entity type representation subtract the head entity type representation to obtain the relationship representation of the belonging relationship category, and the expression is:
[0033] ;
[0034] Finally, introduce into the support set instance representation to obtain the enhanced support set instance representation , and the expression is:
[0035] ;
[0036] Among them, represents an instance belonging to the relationship category in the support set;
[0037] For the query set instance , since the relationship category and its corresponding relationship description are unknown, the query set instance is used as the key value to learn the correlation between the entity type and the instance itself; that is, in the multi-knowledge enhanced learning module, first, based on the attention mechanism, calculate the entity type representation of the head entity or the tail entity in the query set instance and the query set instance representation The correlation coefficient between , the expression is:
[0038] ;
[0039] Then, through weighted summation, obtain the entity type representation of the query set instance , the expression is:
[0040] ;
[0041] Therefore, the head entity type representation and the tail entity type representation in the query set instance are respectively and ;
[0042] Similarly, further adopt the knowledge graph embedding learning technology in the multi-knowledge enhanced learning module to obtain the relationship representation related to the entity type as ;
[0043] Finally, the enhanced representation of the query set instance is obtained through the following calculation , the expression is:
[0044] .
[0045] In one of the embodiments, the multi-knowledge enhanced learning module generates a prototype enhanced representation according to the obtained enhanced representation of the support set instance, including:
[0046] Obtain the enhanced representation of the support set instance , and average all the enhanced representations of the support set instances belonging to the relationship category to obtain the prototype representation of this relationship category ;
[0047] Then integrate the relationship representation corresponding to the relationship category into the corresponding prototype representation to obtain the prototype enhanced representation , the expression is:
[0048] ;
[0049] Among them, , is the support set for the relationship category , and represents the total number of support set instances belonging to the relationship category .
[0050] In one embodiment, the steps of the dual contrast learning module for differentiating relationship categories from two different levels of instances and prototypes by using instance-based contrast learning and prototype-based contrast learning respectively include:
[0051] Instance-based contrast learning: First, select a support set instance of the relationship category as the origin instance, and use the support set instances belonging to this relationship category as positive examples. These support set instances belong to the same category as the origin instance. At the same time, regard the support set instances belonging to other relationship categories as negative examples . These support set instances do not belong to the same category as the origin instance;
[0052] Next, use the dot product operation to measure the similarity between the augmented representation of the origin instance and the augmented representations of all positive and negative examples to obtain the similarity between positive and negative sample pairs of the relationship category , which is used to calculate the instance-based contrast loss . The expression is:
[0053] ;
[0054] where the symbol represents the dot product operation;
[0055] Prototype-based contrast learning: For the relationship category and its corresponding prototype augmented representation , select as the positive example, and use the prototype augmented representations of the remaining relationship categories as negative examples; where and , is the set of relationship categories;
[0056] Then use the dot product operation to measure the prototype augmented representation of the relationship category The similarity between the prototype-enhanced representations of all relationship categories is obtained for the relationship categories to obtain positive and negative sample pairs, which are respectively expressed as:
[0057] ;
[0058] ;
[0059] Among them, is the positive sample pair, is the negative sample pair;
[0060] Finally, calculate the prototype-based contrastive loss , which is used to help the network better distinguish the feature differences between different category prototypes. The expression is:
[0061] .
[0062] In one embodiment, the steps for the relationship prediction module to predict the relationship category to which the query set instance belongs according to the similarity between the prototype-enhanced representation and the query set instance-enhanced representation include:
[0063] When performing relationship classification based on the prototype-enhanced representation of the relationship category , the probability distribution that the query set instance belongs to the relationship category is mainly calculated by computing the similarity between the query set instance-enhanced representation and the prototype-enhanced representation of the relationship category , and converting it into a probability representation through the Softmax function. The expression is:
[0064] ;
[0065] Among them, represents the similarity metric function, i.e., the Euclidean distance;
[0066] Therefore, for the relationship prediction task, the cross-entropy loss function is used to calculate the classification loss of the relationship prediction to evaluate the relationship type to which the query set instance belongs. The expression is:
[0067] ;
[0068] Among them, represents the indicator function; if the relationship type is the correct label, then ; otherwise, ; is a query set, is a set of relationship categories.
[0069] In one embodiment, the comprehensive loss function of the multi-knowledge enhanced prototype network is expressed as:
[0070] ;
[0071] wherein, and are the instance-based contrast loss and the prototype-based contrast loss respectively, and represent weight coefficients.
[0072] A few-shot relation extraction device based on a multi-knowledge enhanced prototype network, the device is applied to the text entity recognition and classification scenario, and the device includes:
[0073] A dataset division module, configured to divide a dataset containing multiple relationship categories and corresponding instances into a training set, a validation set, and a test set according to categories;
[0074] A preprocessing module, configured to randomly extract multiple meta-tasks from each of the divided datasets, and introduce multi-granularity entity types and relationship descriptions as prior knowledge; wherein, each meta-task consists of a support set and a query set, and the support set and the query set contain multiple relationship categories and multiple instances corresponding to each relationship category, and each instance consists of a sentence and the relationship category of the entity pair in the sentence;
[0075] A network construction module, configured to construct a multi-knowledge enhanced prototype network composed of a semantic encoder, a multi-knowledge enhanced learning module, a dual contrast learning module, and a relationship prediction module; wherein, the semantic encoder is configured to encode and generate feature representations of instances and prior knowledge; the multi-knowledge enhanced learning module is configured to perform knowledge enhanced learning on the support set instances and the query set instances according to the prior knowledge, and generate a prototype enhanced representation according to the enhanced representation of the support set instances obtained; the dual contrast learning module is configured to perform relationship category discrimination learning from two different levels of instances and prototypes respectively by using instance-based contrast learning and prototype-based contrast learning; the relationship prediction module is configured to predict the relationship category to which the query set instance belongs according to the similarity between the prototype enhanced representation and the enhanced representation of the query set instance;
[0076] A task execution module, configured to input the meta-tasks and prior knowledge randomly extracted from the training set and the validation set into the multi-knowledge enhanced prototype network, and construct a comprehensive loss function including a dual contrast learning loss and a classification loss of relationship prediction for network training and evaluation until a trained network is obtained to perform the few-shot relation extraction task.
[0077] In the above-mentioned few-shot relation extraction method and device based on the multi-knowledge enhanced prototype network, a multi-knowledge enhanced prototype network composed of a semantic encoder, a multi-knowledge enhanced learning module, a dual contrastive learning module, and a relation prediction module is constructed. This network adopts prompt learning to design a prompt template with entity information to activate the knowledge in the pre-trained language model, and can obtain a more accurate instance semantic representation. At the same time, two kinds of prior knowledge, multi-granularity entity types and relation descriptions, are introduced to enhance the semantic representations of instances and prototypes. In addition, a dual contrastive learning module based on instances and prototypes is designed to learn the class distinctiveness and distinguishability of instance representations and prototype representations from two different levels of instances and prototypes, so as to fully capture and understand the characteristics of all relations in the text based on a small number of samples, and improve the accuracy of text entity recognition and relation classification prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] Figure 1 It is a schematic flowchart of a few-shot relation extraction method based on a multi-knowledge enhanced prototype network in an embodiment;
[0079] Figure 2 It is a schematic diagram of the architecture of a multi-knowledge enhanced prototype network in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0080] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0081] In one embodiment, as Figure 1 shown, a few-shot relation extraction method based on a multi-knowledge enhanced prototype network is provided, including the following steps:
[0082] Step S1, divide a data set containing multiple relation categories and corresponding instances into a training set, a validation set, and a test set according to categories.
[0083] Among them, dividing the data set according to categories can ensure that the relation categories in the test set and the training set are non-overlapping. Such a setting can ensure that the network model makes accurate predictions and evaluations when facing unseen relation categories.
[0084] Step S2, randomly extract multiple meta-tasks from each of the divided data sets, and introduce multi-granularity entity types and relation descriptions as prior knowledge.
[0085] Specifically, step S2 includes: randomly extracting multiple meta-tasks from the training set, validation set, and test set for training, validation, and testing. In a typical N-way-K-shot (N classes and K samples) setting, each meta-task for few-shot relation extraction consists of a support set and a query set ; among them, the support set contains relation classes, and the set of relation classes is denoted as . The support set of each relation class is composed of of the current relation class support set instances ; the query set also contains the same relation classes, but the instances in the query set are randomly extracted from the remaining instances of each relation class query set instances ; among them, represents the i nd relation class, represents the th support set instance of the relation class k , respectively represent the sentence in and the relation class of the entity pair in the sentence; j represents the th query set instance, respectively represent
[0086] The support set and each instance in the query set is composed of a sentence containing a head entity and a tail entity and the relation class representing the relationship between the head entity and the tail entity ; the goal of the meta-task is to use the extremely small amount of labeled data in the support set to predict the relation class of the head entity and the tail entity in the sentence of the query set instance ; among them, the relation classes extracted by each meta-task are different. This random extraction method of meta-tasks helps to improve the generalization ability of the model, enabling it to adapt to different combinations of relation classes.
[0087] Furthermore, to obtain a better prototype representation, multi-granularity entity types and relationship descriptions are introduced as prior knowledge. Among them, multi-granularity entity types include two different levels of entity types: coarse-grained and fine-grained, providing more fine-grained entity types for the relationship representation of instances. For each entity it may be assigned a coarse-grained entity type or a fine-grained entity type. To unify all possible entity types of the entity the multi-granularity entity type represents all entity type information contained in an entity , represents the number of entity types, is the i th entity type and . The relationship description consists of the relationship name of the relationship category and the corresponding detailed description . These information summarize the specific characteristics and attributes of the relationship category, helping to deeply understand the relationship category.
[0088] Step S3: Construct a multi-knowledge enhanced prototype network composed of a semantic encoder, a multi-knowledge enhanced learning module, a dual contrast learning module, and a relationship prediction module.
[0089] As Figure 2 shown, the semantic encoder is used to encode and generate the feature representations of instances and prior knowledge; the multi-knowledge enhanced learning module is used to perform knowledge enhanced learning on the support set instances and query set instances according to the prior knowledge, and generate a prototype enhanced representation according to the obtained enhanced representation of the support set instances; the dual contrast learning module is used to perform relationship category discrimination learning from two different levels of instances and prototypes respectively by using instance-based contrast learning and prototype-based contrast learning; the relationship prediction module is used to predict the relationship category to which the query set instance belongs according to the similarity between the prototype enhanced representation and the enhanced representation of the query set instance. The specific implementation processes of the components in the multi-knowledge enhanced prototype network are as follows:
[0090] (1) Semantic encoder: The pre-trained language model has learned rich language knowledge from large-scale unlabeled text data through self-supervised learning, so it can understand the relationships between words in a sentence and obtain a more accurate semantic representation. With the emergence of GPT-3 (Generative Pretrained Transformer), it has promoted a new learning method - prompt learning, which guides the model to generate specific types of outputs by providing prompt texts for the input of the model. Therefore, the multi-knowledge enhanced prototype network proposed in this application uses BERT as the semantic encoder to learn the relevant context semantic representations of instances and prior knowledge, and designs a prompt template with entity information as the instance input through prompt learning to guide the model to output feature representations related to the relationship category.
[0091] Specifically, the steps for the semantic encoder to encode and generate the feature representation of an instance include:
[0092] For the instances in the support set and the query set, first, the sentences in the instance are concatenated with the hint text with the entity pair to form a hint template with entity information , and the expression is:
[0093] ;
[0094] Here, the hint template adopts the relational triple structure of the entity pair as the hint text to activate the relational representation of the entity pair in the instance; among them, [CLS] is called the classification token, which is used to represent the start of the input sequence; [SEP] is called the separator token, which is used to separate two sentences or represent the end of a single sentence; is called the mask token, which is used to represent some chunks of words that require the model to predict the relationship between two entities; then, the hint template is input into BERT for encoding, and the contextual representation is used as the instance representation describing the semantic relationship between entity pairs , and the expression is:
[0095] ;
[0096] Among them, represents BERT; the instance representation includes the support set instance representation and the query set instance representation.
[0097] Specifically, the steps for the semantic encoder to encode and generate the feature representation of prior knowledge include:
[0098] For the relationship descriptions in the prior knowledge, first, the relationship name of each relationship category and the detailed description are concatenated together to form the input sequence ; then is input into BERT for encoding, and the contextual representation is used as the relationship representation depicting the relationship description , and the expression is:
[0099] .
[0100] Furthermore, considering that an entity may have multiple fine-grained entity types in addition to the coarse-grained entity type, and the degree of correlation between different entity types and relationship categories also varies. Therefore, a semantic encoder is used to encode each entity type in the multi-granularity entity types separately.
[0101] For the multi-granularity entity types in the prior knowledge , each entity type is first transformed into the format and then input into BERT for encoding, and the context representation is used as the entity type representation describing the entity type
[0102] .
[0103] (2) Multi-knowledge enhanced learning module: To learn more representative instance representations and prototype representations, the multi-knowledge enhanced prototype network proposes a multi-knowledge enhanced learning module that performs knowledge-enhanced learning on instances and prototypes by introducing two types of prior knowledge: multi-granularity entity types and relationship descriptions. For instances, the attention mechanism is first used to select more important entity types from the multi-granularity entity types to obtain new entity type representations; then, knowledge graph embedding technology is used to enhance the instance representation by calculating the relationship category representation corresponding to the head entity type representation and the tail entity type representation. For prototypes, the relationship representation learned from the relationship description is integrated into the prototype representation to enhance the prototype representation by highlighting the unique features of the relationship category.
[0104] Specifically, the steps for the multi-knowledge enhanced learning module to perform knowledge-enhanced learning on the support set instances according to the prior knowledge include:
[0105] For the th k support set instance of the relationship category , the known relationship category is used to guide the selection of multi-granularity entity types; that is, in the multi-knowledge enhanced learning module, first, the correlation coefficient between the th entity type representation and the relationship representation corresponding to its belonging relationship category is calculated based on the attention mechanism, and the expression is:
[0106] ;
[0107] where represents the similarity measurement function, that is, the dot product operation; , represents the number of entity types; the entity , and represent the head entity and the tail entity, respectively;
[0108] Then, based on the correlation coefficient , all entity types are selectively fused by weighted summation to obtain the entity type representation related to the relationship category in the support set instance , and the expression is:
[0109] ;
[0110] Therefore, the head entity type representation and the tail entity type representation in the support set instance are and ;
[0111] Furthermore, in the multi-knowledge enhanced learning module, the knowledge graph embedding learning technology is adopted to model the relationship category as a distance transformation from the head entity type to the tail entity type, that is, the tail entity type representation is subtracted from the head entity type representation to obtain the relationship representation of the belonging relationship category, and the expression is:
[0112] ;
[0113] Finally, is introduced into the support set instance representation to obtain the enhanced support set instance representation , and the expression is:
[0114] ;
[0115] where represents an instance in the support set belonging to the relationship category .
[0116] Specifically, the steps of the multi-knowledge enhanced learning module for knowledge enhanced learning of the query set instance according to prior knowledge include:
[0117] For the query set instance , since the relationship category and its corresponding relationship description are unknown, the query set instance representation is used as the key value to learn the correlation between the entity type and the instance itself; that is, in the multi-knowledge enhanced learning module, first, the entity type representation of the head entity or the tail entity in the query set instance and the query set instance representation are used to calculate the correlation coefficient therebetween, and the expression is:
[0118] ;
[0119] Then, through weighted summation, the entity type representation of the query set instance is obtained , and the expression is:
[0120] ;
[0121] Therefore, the head entity type representation and the tail entity type representation in the query set instance are respectively and ;
[0122] Similarly, the knowledge graph embedding learning technology is further adopted in the multi-knowledge enhanced learning module to obtain the relationship representation related to the entity type as ;
[0123] Finally, the enhanced representation of the query set instance is obtained through the following calculation , and the expression is:
[0124] .
[0125] Specifically, the multi-knowledge enhanced learning module generates a prototype enhanced representation based on the obtained enhanced representation of the support set instance, including:
[0126] Since the relationship description generalizes the specific features of the relationship category, it brings considerable benefits to correcting the prototype representation of the relationship category. First, obtain the enhanced representation of the support set instance , and average all the enhanced representations of the support set instances belonging to the relationship category to obtain the prototype representation of this relationship category ; then integrate the relationship representation of the relationship category into the corresponding prototype representation to obtain the prototype enhanced representation , and the expression is:
[0127] ;
[0128] Among them, , is the support set of the relationship category , represents the total number of support set instances belonging to the relationship category .
[0129] (3) Dual contrast learning module: Although prior knowledge can highlight the essential features of relationship categories and is very helpful for learning more representative instance and prototype representations. However, when the semantic representations of support set instances of the same category are quite different, it is still necessary to highlight the common features of the same-class instance representations and prototype representations, as well as the distinguishable features of different-class instance representations and prototype representations. Therefore, the multi-knowledge enhanced prototype network designs a dual contrast learning strategy based on instances and prototypes to learn unique and distinguishable feature representations of instances and prototypes, thereby improving the accuracy of few-shot learning.
[0130] Specifically, starting from the instance representation level, instance-based contrast learning is designed to learn the differences between within-class and between-class instances in the support set instances, thereby highlighting the category uniqueness and distinguishability in the instance representation. Instance-based contrast learning includes the following steps:
[0131] First, select a certain support set instance of the relationship category as the origin instance, and take the support set instances belonging to this relationship category as positive examples, and these support set instances belong to the same category as the origin instance; at the same time, regard the support set instances belonging to other relationship categories as negative examples ; these support set instances do not belong to the same category as the origin instance; is the total number of relationship categories in the meta-task;
[0132] Next, use the dot product operation to measure the similarity between the enhanced representation of the origin instance and the enhanced representations of all positive and negative examples to obtain the similarity between positive and negative sample pairs of the relationship category for calculating the instance-based contrast loss , and the expression is:
[0133] ;
[0134] where the symbol represents the dot product operation.
[0135] Specifically, an ideal prototype should contain an immutable category representation that can be distinguished from other categories. Therefore, starting from the prototype representation level, prototype-based contrast learning is designed to highlight the uniqueness and distinguishability of category prototypes. Prototype-based contrast learning includes the following steps:
[0136] For the relationship category and its corresponding prototype enhanced representation , select As positive examples, and the remaining relation categories of the prototype-enhanced representation as negative examples; among them, and , is the set of relation categories;
[0137] Then, the dot product operation is used to measure the similarity between the prototype-enhanced representation of the relation category and the prototype-enhanced representations of all relation categories, obtaining positive and negative sample pairs of the relation category , which are respectively expressed as:
[0138] ;
[0139] ;
[0140] Among them, is the positive sample pair, is the negative sample pair;
[0141] Finally, the prototype-based contrastive loss is calculated to help the network better distinguish the feature differences between different category prototypes. The expression is:
[0142] .
[0143] (4) Relation prediction module: When performing relation classification based on the prototype-enhanced representation of the relation category , the probability distribution that the query set instance belongs to the relation category is mainly calculated by measuring the similarity between the enhanced representation of the query set instance and the prototype-enhanced representation of the relation category , and converting it into a probability representation through the Softmax function. The expression is:
[0144] ;
[0145] Among them, represents the similarity measurement function, i.e., the Euclidean distance;
[0146] Therefore, for the relation prediction task, the cross-entropy loss function is used to calculate the classification loss of the relation prediction to evaluate the relation type to which the query set instance belongs. The expression is:
[0147] ;
[0148] in, represents the indicator function; if the relationship type is the correct label, then ;otherwise, ; For the query set, A collection of relationship categories.
[0149] In step S4, the meta-tasks and prior knowledge randomly extracted from the training set and the validation set are input into the multi-knowledge enhanced prototype network, and a comprehensive loss function including the double contrast learning loss and the classification loss of relationship prediction is constructed to perform network training and evaluation until a trained network is obtained to perform the few-sample relationship extraction task.
[0150] Specifically, the comprehensive loss function of the multi-knowledge enhanced prototype network is It is expressed as:
[0151] ;
[0152] in, and They are instance-based contrast loss and prototype-based contrast loss, and Represents the weight coefficient.
[0153] In summary, the present application provides a method for extracting few-sample relations based on a multi-knowledge enhanced prototype network. In the constructed multi-knowledge enhanced prototype network, the pre-trained language model is activated by prompt learning to obtain a more accurate semantic representation, and multi-granularity entity types and relationship information are introduced as external knowledge to enhance the representation effect of instances and prototypes. At the same time, a dual contrast learning strategy is designed, which combines class-independent contrast learning and class-specific contrast learning to effectively solve the problem of intra-class instances with large syntactic differences and inter-class instances with similar semantics, thereby enhancing the compactness of intra-class instance features and widening the distance between inter-class prototypes, improving the accuracy of text entity recognition and relationship classification prediction.
[0154] Furthermore, experiments are conducted on two public datasets, FewRel 1.0 and FewRel 2.0, to evaluate the performance of the multi-knowledge enhanced prototype network constructed in this application.
[0155] The FewRel 1.0 dataset contains 100 relation categories, with 700 instances for each relation category. These instances are all extracted from Wikipedia articles. According to the official evaluation setting, Fewrel 1.0 is divided into a training set, a validation set, and a test set based on 100 relation categories. Among them, the training set contains 64 relation categories, the validation set contains 16 relation categories, and the test set contains 20 relation categories, with 700 instances for each relation category.
[0156] The FewRel 2.0 dataset mainly focuses on the problem of domain adaptability. Its training set is the same as that of FewRel 1.0; the validation set is the SemEval-2010 task 8 dataset, which contains 17 relation categories, with 520 instances for each relation category; the test set is the PubMed dataset, which comes from biomedical literature and contains 25 relation categories, with 100 instances for each relation category.
[0157] It is worth noting that all the data of FewRel 1.0 belongs to the Wikipedia domain, while the validation set and test set of FewRel 2.0 come from the biomedical domain. The domain difference brings more challenges to the model's rapid learning and generalization, so FewRel 2.0 is more challenging than FewRel 1.0.
[0158] To obtain the fine-grained entity types of entities, WikiData (a general domain knowledge base) and UMLS (Unified Medical Language System, a specific domain knowledge base) are used as external knowledge bases during the experiment. WikiData is a free knowledge base containing structured data of various Wikimedia projects, which contains a large amount of information on various topics, including people, places, concepts, etc. WikiData is used to obtain the fine-grained entity types of general domain entities. UMLS is a knowledge base in the biomedical and health fields collected and constructed by experts, which contains a large number of detailed definitions of medical entities, related concepts, and semantic relationships between different entities. UMLS is used to obtain the fine-grained entity types of specific domain entities.
[0159] In the experiment, four N-way-K-shot few-shot learning settings, namely 5-way-1-shot (hereinafter referred to as 5-w-1-s, and the same applies to others), 5-way-5-shot, 10-way-1-shot, and 10-way-5-shot, were set to evaluate the performance of the multi-knowledge enhanced prototype network. During training, 20,000 meta-tasks were randomly selected from the training set to train the network model, and 1,000 meta-tasks were randomly selected from the validation set to evaluate the network model; during testing, 10,000 meta-tasks were randomly selected from the test set for testing. All experiments were trained and evaluated on an NVIDIA GeForce RTX 2080 Ti graphics card.
[0160] In the comparative experiment, the performance was compared and evaluated with three most representative methods, namely some basic models, external information enhanced models, and specific pre-trained models, such as Proto-BERT (prototype BERT model), BERT-PAIR (BERT paired classification model), MAML (model-agnostic meta-learning), GNN (graph neural network), REGRAB (relational graph attention benchmark model), CTEG (context temporal event graph model), ConceptFERE (concept-enhanced entity relation extraction model), HCRP (hierarchical concept relation propagation model), MTB (matching template BERT model), CP (contrast prototype network), and LPD (latent prototype discovery model). The accuracy (Accuracy) was selected as the evaluation metric to evaluate the performance of the model. Tables 1 and 2 list the model performance comparisons of the few-shot relation extraction task on the FewRel 1.0 and FewRel 2.0 test sets respectively. KnowProto and KnowProto+CP in the tables represent the multi-knowledge enhanced prototype network constructed in this application. The main experimental results of various models on the FewRel 1.0 and FewRel 2.0 test sets show that the multi-knowledge enhanced prototype networks KnowProto and KnowProto+CP proposed in this application achieved the best performance under all N-way-K-shot few-shot learning settings.
[0161] Table 1 Comparison results of the accuracy of different network models on the FewRel 1.0 test set
[0162]
[0163] Table 2 Comparison results of the accuracy of different network models on the FewRel 2.0 test set
[0164]
[0165] In one embodiment, a few-shot relation extraction device based on a multi-knowledge enhanced prototype network is provided, including:
[0166] A dataset partitioning module for partitioning a dataset containing multiple relation categories and corresponding instances into a training set, a validation set, and a test set according to categories;
[0167] A preprocessing module for randomly extracting multiple meta-tasks from each of the partitioned datasets and introducing multi-granularity entity type and relation descriptions as prior knowledge; wherein each meta-task consists of a support set and a query set, the support set and the query set contain multiple relation categories and multiple instances corresponding to each relation category, and each instance consists of a sentence and the relation category of the entity pair in the sentence;
[0168] A network construction module for constructing a multi-knowledge enhanced prototype network composed of a semantic encoder, a multi-knowledge enhanced learning module, a dual contrast learning module, and a relation prediction module; wherein the semantic encoder is used to encode and generate feature representations of instances and prior knowledge; the multi-knowledge enhanced learning module is used to perform knowledge enhanced learning on the support set instances and the query set instances according to the prior knowledge, and generate a prototype enhanced representation according to the enhanced representation of the support set instances obtained; the dual contrast learning module is used to perform relation category discrimination learning from two different levels of instances and prototypes respectively by using instance-based contrast learning and prototype-based contrast learning; the relation prediction module is used to predict the relation category to which the query set instance belongs according to the similarity between the prototype enhanced representation and the enhanced representation of the query set instance;
[0169] A task execution module for inputting the meta-tasks and prior knowledge randomly extracted from the training set and the validation set into the multi-knowledge enhanced prototype network, and constructing a comprehensive loss function including a dual contrast learning loss and a classification loss of relation prediction for network training and evaluation until a trained network is obtained to perform few-shot relation extraction tasks.
[0170] For the specific limitations of the few-shot relation extraction device based on the multi-knowledge enhanced prototype network, reference can be made to the limitations of the few-shot relation extraction method based on the multi-knowledge enhanced prototype network in the above text, which will not be elaborated here. Each module in the above few-shot relation extraction device based on the multi-knowledge enhanced prototype network can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form so that the processor can call and execute the operations corresponding to the above modules.
[0171] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered that the scope recorded in this specification.
[0172] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A few-shot relation extraction method based on a multi-knowledge enhanced prototype network, characterized in that The method is applied to the scenario of text entity recognition and classification, and the method includes: Dividing a data set containing multiple relationship categories and corresponding instances into a training set, a validation set, and a test set according to categories; Randomly extracting multiple meta-tasks from each of the divided data sets, and introducing multi-granularity entity types and relationship descriptions as prior knowledge; wherein, each meta-task consists of a support set and a query set, and the support set and the query set contain multiple relationship categories and multiple instances corresponding to each relationship category, and each instance consists of a sentence and the relationship category of the entity pair in the sentence; Constructing a multi-knowledge enhanced prototype network composed of a semantic encoder, a multi-knowledge enhanced learning module, a dual contrast learning module, and a relationship prediction module; wherein, the semantic encoder is used to encode and generate feature representations of instances and prior knowledge; the multi-knowledge enhanced learning module is used to perform knowledge enhanced learning on the support set instances and the query set instances according to the prior knowledge, and generate a prototype enhanced representation according to the enhanced representation of the support set instances obtained; the dual contrast learning module is used to perform relationship category discrimination learning from two different levels of instances and prototypes respectively by using instance-based contrast learning and prototype-based contrast learning; the relationship prediction module is used to predict the relationship category to which the query set instance belongs according to the similarity between the prototype enhanced representation and the enhanced representation of the query set instance; Inputting the meta-tasks and prior knowledge randomly extracted from the training set and the validation set into the multi-knowledge enhanced prototype network, and constructing a comprehensive loss function including a dual contrast learning loss and a classification loss for relationship prediction to perform network training and evaluation until a trained network is obtained to perform few-shot relationship extraction tasks.
2. The method according to claim 1, wherein Randomly extracting multiple meta-tasks from each of the divided data sets, and introducing multi-granularity entity types and relationship descriptions as prior knowledge, including: Randomly extract multiple meta-tasks from the training set, validation set, and test set for training, validation, and testing. Each meta-task of few-shot relation extraction consists of a support set and a query set ; among them, the support set contains relation categories, and the set of relation categories is denoted as , and the support set of each relation category is composed of support set instances of the current relation category ; the query set also contains the same relation categories, but the instances in the query set are randomly selected from the remaining instances of each relation category query set instances ; among them, represents the i th relation category, represents the th support set instance of the relation category k , respectively represent the sentence in and the relation category of the entity pair in the sentence; represents the j th query set instance, respectively represent the sentence in and the relation category of the entity pair in the sentence; Support set and query set Each instance in is composed of a sentence containing a head entity and a tail entity and the relationship category indicating between the head entity and the tail entity The goal of the meta - task is to use a very small number of labeled data in the support set to predict the relationship category of the sentence of the query set instance between the head entity and the tail entity in; among them, the relationship categories extracted by each meta - task are different; Furthermore, multi-granularity entity types and relationship descriptions are introduced as prior knowledge; among them, multi-granularity entity types represent all entity type information contained in an entity , represents the number of entity types is the i th entity type and ; the relationship description consists of the relationship name of the relationship category and the corresponding detailed description .
3. The method according to claim 2, wherein The steps of the semantic encoder encoding and generating feature representations of instances and prior knowledge include: Using the pre-trained language model BERT as the semantic encoder; For the instances in the support set and the query set, first, the sentences in the instances are concatenated with the hint text with the entity pair to form a hint template with entity information , and the expression is: ; Prompt template here Adopt the relational triple structure of entity pairs as the prompt text to activate the relational representation of entity pairs in the instance; among them, [CLS] is called the classification token, which is used to represent the start of the input sequence; [SEP] is called the separator token, which is used to separate two sentences or represent the end of a single sentence; is called the mask token, which is used to represent some chunks of words that require the model to predict the relationship between two entities; then, the prompt template is input into BERT for encoding, and the contextual representation is used as the instance representation describing the semantic relationship between entity pairs , and the expression is: ; Among them, represents BERT; the instance represents including the support set instance representation and the query set instance representation; For the relationship descriptions in prior knowledge, first, the relationship name and the detailed description of each relationship category are concatenated together to form an input sequence ; then, the input sequence is fed into BERT for encoding, and the context representation is used as the relationship representation to characterize the relationship description, with the expression being: ; For the multi-granularity entity types in prior knowledge First, each entity type is converted into format, then input into BERT for encoding, and the context representation is used as the entity type representation describing the entity type , and the expression is: 。 4. The method according to claim 3, wherein The steps of the multi-knowledge enhanced learning module performing knowledge enhanced learning on the support set instances and the query set instances according to the prior knowledge include: For the relationship category the k th support set instance , use the known relationship category to guide the selection of multi-granularity entity types; that is, in the multi-knowledge reinforcement learning module, first calculate the th entity type representation and the relationship representation corresponding to its belonging relationship category The correlation coefficient between them is expressed as: ; Among them, represents the similarity metric function, i.e., the dot product operation; , represents the number of entity types; the entity , and respectively represent the head entity and the tail entity; Then, based on the correlation coefficient , all entity types are selectively fused in a weighted summation manner to obtain the entity type representation related to the relationship category in the support set instances , and the expression is as follows: ; Therefore, the head entity type representation and the tail entity type representation in the support set instance are respectively and ; Furthermore, the knowledge graph embedding learning technology is adopted in the multi-knowledge enhanced learning module, and the relationship category is modeled as a distance transformation from the head entity type to the tail entity type, that is, the tail entity type representation is used subtract the head entity type representation to obtain the relationship representation of the belonging relationship category , and the expression is: ; Finally, is introduced into the support set instance representation to obtain the enhanced representation of the support set instance , and the expression is: ; Among them, represents an instance belonging to the relationship category in the support set; For the query set instance , since the relationship category and its corresponding relationship description are unknown, the query set instance is used as the key value to learn the correlation between the entity type and the instance itself; that is, in the multi-knowledge reinforcement learning module, first calculate the entity type representation of the head entity or the tail entity in the query set instance based on the attention mechanism and the query set instance representation The correlation coefficient between , the expression is: ; Then, through weighted summation, the entity type representation of the query set instance is obtained , and the expression is: ; Therefore, the head entity type representation and the tail entity type representation in the query set instance are respectively and ; Similarly, the knowledge graph embedding learning technology is further adopted in the multi-knowledge reinforcement learning module to obtain the relationship related to the entity type as ; Finally, the enhanced representation of the query set instance is obtained through the following calculations , and the expression is: 。 5. The method according to claim 4, wherein The steps of the multi-knowledge enhanced learning module generating a prototype enhanced representation according to the enhanced representation of the support set instances obtained include: Obtain the enhanced representation of the support set instances , and average all the enhanced representations of the support set instances belonging to the relationship category to obtain the prototype representation of this relationship category ; Then integrate the relationship category corresponding relationship representation into the corresponding prototype representation to obtain the prototype enhanced representation , and the expression is: ; Among them, , is the support set of the relationship category , and represents the total number of support set instances belonging to the relationship category .
6. The method according to claim 5, wherein The steps of the dual contrast learning module performing relationship category discrimination learning from two different levels of instances and prototypes respectively by using instance-based contrast learning and prototype-based contrast learning include: Instance-based contrastive learning: First, select a support set instance of a certain relationship category as the origin instance. Consider the number of support set instances belonging to this relationship category as positive examples. These support set instances belong to the same category as the origin instance. At the same time, consider the support set instances belonging to other relationship categories as negative examples . These support set instances do not belong to the same category as the origin instance; is the total number of relationship categories in the meta-task; Next, use the dot product operation to measure the augmented representation of the origin instance and the augmented representations of all positive and negative examples to obtain the similarity between them, and get the relationship category The similarity between positive and negative sample pairs is used to calculate the instance-based contrastive loss , and the expression is: ; Among them, the symbol represents the dot product operation; Prototype-based contrastive learning: for relation categories and its corresponding prototype-enhanced representation , select as the positive example, and use the prototype-enhanced representations of the remaining relation categories as negative examples; among them, and , is the set of relation categories; Then, the dot product operation is used to measure the relationship category of the prototype-enhanced representation and the similarity between the prototype-enhanced representations of all relationship categories, obtaining the positive and negative sample pairs of the relationship category which are respectively represented as: ; ; Among them, is a positive sample pair, is a negative sample pair; Finally, calculate the prototype-based contrastive loss , which helps the network better distinguish the feature differences between different class prototypes. The expression is as follows: 。 7. The method according to claim 6, wherein The steps of the relationship prediction module predicting the relationship category to which the query set instance belongs according to the similarity between the prototype enhanced representation and the enhanced representation of the query set instance include: When performing relation classification based on the prototype-enhanced representation of the relation category the probability distribution of the query set instance belonging to the relation category is mainly to calculate the similarity between the enhanced representation of the query set instance and the prototype-enhanced representation of the relation category, and convert it into a probability representation through the Softmax function. The expression is: ; Among them, represents a similarity metric function, i.e., the Euclidean distance; Therefore, for the relationship prediction task, the cross-entropy loss function is used to calculate the classification loss of relationship prediction to evaluate the relationship type to which the query set instance belongs. The expression is as follows: ; Among them, represents an indicator function; if the relationship type is the correct label, then ; otherwise, ; is a query set, is a set of relationship categories.
8. The method according to claim 7, wherein Comprehensive Loss Function of Multi-Knowledge Enhanced Prototype Network It is expressed as: ; Among them, and are the instance-based contrastive loss and the prototype-based contrastive loss respectively, and represent the weight coefficients.
9. A few-shot relation extraction device based on a multi-knowledge enhanced prototype network, characterized in that, The device is applied to the scenario of text entity recognition and classification, and the device includes: A data set division module, configured to divide a data set containing multiple relationship categories and corresponding instances into a training set, a validation set, and a test set according to categories; A preprocessing module, configured to randomly extract multiple meta-tasks from each of the divided data sets, and introduce multi-granularity entity types and relationship descriptions as prior knowledge; wherein, each meta-task consists of a support set and a query set, and the support set and the query set contain multiple relationship categories and multiple instances corresponding to each relationship category, and each instance consists of a sentence and the relationship category of the entity pair in the sentence; The network construction module is used to construct a multi-knowledge enhanced prototype network composed of a semantic encoder, a multi-knowledge enhanced learning module, a dual contrast learning module, and a relationship prediction module. Among them, the semantic encoder is used to encode and generate the feature representations of instances and prior knowledge. The multi-knowledge enhanced learning module is used to perform knowledge enhanced learning on the support set instances and query set instances according to the prior knowledge, and generate a prototype enhanced representation based on the enhanced representation of the support set instances obtained. The dual contrast learning module is used to perform relationship category discrimination learning from two different levels of instances and prototypes respectively by using instance-based contrast learning and prototype-based contrast learning. The relationship prediction module is used to predict the relationship category to which the query set instance belongs according to the similarity between the prototype enhanced representation and the enhanced representation of the query set instance. The task execution module is used to input the meta-tasks and prior knowledge randomly extracted from the training set and the validation set into the multi-knowledge enhanced prototype network, and construct a comprehensive loss function including the dual contrast learning loss and the classification loss of relationship prediction for network training and evaluation until a trained network is obtained to perform the few-shot relationship extraction task.
Citation Information
Patent Citations
Cross-domain small sample relation extraction method and device for learning fine-grained general knowledge
CN118674036A
Small sample remote sensing image scene classification method based on embedding smoothing graph neural network
WO2023087558A1