Method and device for extracting few-sample relation based on multi-knowledge enhanced prototype network
By using a multi-knowledge enhancement prototype network in the small-sample relationship extraction task, combining the prior knowledge of multi-grained entity types and relationship descriptions, knowledge enhancement and comparison learning is carried out, the problem of insufficient semantic learning caused by insufficient sample number in the existing technology is solved, and the accuracy of relationship extraction is improved.
Patent Information
- Application Number
- CN202510523318.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-24
AI Technical Summary
The prior art is difficult to fully capture and understand the characteristics of complex relationships through a small number of samples in the small sample relationship extraction task, resulting in poor performance of models when identifying and classifying unseen relationships.
Using a method based on multi-knowledge enhancement prototype network, a multi-knowledge enhancement prototype network consisting of a semantic encoder, a multi-knowledge enhancement learning module, a dual-contrast learning module and a relation prediction module is constructed to carry out knowledge enhancement learning and contrast learning to improve the relationship extraction ability of the model.
Through multi-knowledge enhancement of prototype networks, it is possible to capture and understand the characteristics of relationships in text more effectively with few samples, and improve the accuracy of text entity recognition and relationship classification prediction.
Smart Images

Figure CN120068874A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of natural language processing, and particularly to a few-shot relation extraction method and device based on a multi-knowledge enhanced prototype network. Background Art
[0002] The relation extraction task aims to identify and classify the relationships between different entities in text. Most existing relation extraction methods, especially those based on pre-trained language models (such as the Bidirectional Encoder Representations from Transformers BERT), have achieved significant performance improvements. However, these methods usually require a large amount of high-quality labeled data to train the model to obtain good performance. Humans can effectively learn new knowledge through a small number of instance samples, which mainly benefits from humans' rich prior knowledge and background knowledge, enabling them to reason and generalize. For example, after knowing that a person is a doctor, humans can naturally infer that this person may work in a hospital and have medical knowledge. This reasoning ability based on prior knowledge enables humans to be more flexible and efficient when facing new relation extraction tasks. In order to enable machines to learn new relationships through a small number of labeled instances like humans and apply them to real life, the few-shot relation extraction task has become a current research hotspot.
[0003] Few-shot relation extraction aims to train a model by using a small amount of labeled data so that it can quickly learn and make effective inferences under new relation types. Although previous work has improved the performance of few-shot relation extraction by introducing external prior knowledge and the semantic representations of pre-trained language models, it still faces the problem of insufficient semantic learning caused by insufficient sample quantity. Especially in the expression of complex relation categories, it is difficult for the model to fully capture and understand the characteristics of all relationships through a small number of samples, resulting in poor performance of the model when identifying and classifying unseen relationships. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide a few-shot relation extraction method and device based on a multi-knowledge enhanced prototype network.
[0005] A few-shot relation extraction method based on a multi-knowledge enhanced prototype network, the method is applied to the text entity recognition and classification scenario, and includes the following steps: Divide a data set containing multiple relation categories and corresponding instances into a training set, a validation set, and a test set according to categories; Randomly extract multiple meta-tasks from each of the divided data sets, and introduce multi-granularity entity types and relation descriptions as prior knowledge; where each meta-task consists of a support set and a query set, and the support set and the query set contain multiple relation categories and multiple instances corresponding to each relation category, and each instance consists of a sentence and the relation category of the entity pair in the sentence; A multi-knowledge enhanced prototype network consisting of a semantic encoder, a multi-knowledge enhanced learning module, a dual contrast learning module and a relationship prediction module is constructed; wherein the semantic encoder is used to encode the feature representation of the generated instance and the prior knowledge; the multi-knowledge enhanced learning module is used to perform knowledge enhancement learning on the support set instance and the query set instance according to the prior knowledge, and generate the prototype enhanced representation according to the obtained support set instance enhanced representation; the dual contrast learning module is used to use instance-based contrast learning and prototype-based contrast learning to perform relationship category distinction learning from two different levels, instance and prototype; the relationship prediction module is used to predict the relationship category to which the query set instance belongs according to the similarity between the prototype enhanced representation and the query set instance enhanced representation; The meta-tasks and prior knowledge randomly extracted from the training set and the validation set are input into the multi-knowledge enhanced prototype network, and a comprehensive loss function including the double contrastive learning loss and the classification loss of relationship prediction is constructed for network training and evaluation until a well-trained network is obtained to perform the few-sample relationship extraction task.
[0006] In one embodiment, multiple meta-tasks are randomly extracted from each divided data set, and multi-granularity entity types and relationship descriptions are introduced as prior knowledge, including: Multiple meta-tasks are randomly extracted from the training set, validation set, and test set for training, validation, and testing. Each meta-task of the few-shot relation extraction consists of a support set and a queryset Composition; among them, the support set Include relation categories, and the relation category set is represented as , each relationship category The support set By current relationship category of Support set examples Composition; query set Also includes the same relation categories, but the instances in the query set are randomly drawn from the remaining instances of each relation category. QuerySet instances constitute; among them, Indicates i relationship categories, Represents relationship category No. k Support set examples, Respectively The sentences in and the relationship categories of entity pairs in the sentences; Indicates j QuerySet instances, Respectively The sentences in and the relationship categories of entity pairs in the sentences; Support set and queryset Each instance in is composed of a header entity and tail entity Sentences and the relationship between the head entity and the tail entity composition; the goal of the meta-task is to use the support set Very few labeled data are used to predict the sentences of query set instances The relationship type between the head entity and the tail entity ; Among them, each meta-task extracts The relationship categories are different; Furthermore, multi-granularity entity types and relationship descriptions are introduced as prior knowledge; Represents an entity All entity type information contained, Indicates the number of entity types, For the i entity types and ; The relationship description consists of the relationship name of the relationship category and the corresponding detailed description composition.
[0007] In one embodiment, the step of encoding the semantic encoder to generate feature representations of instances and prior knowledge includes: Use the pre-trained language model BERT as the semantic encoder; For the instances in the support set and query set, firstly, the sentences in the instance With entity pairs The prompt texts are stitched together to form a prompt template with entity information , the expression is: ; Tip template here The relation triple structure of entity pairs is used as the prompt text to activate the relation representation of entity pairs in the instance; [CLS] is called the classification tag, which is used to indicate the beginning of the input sequence; [SEP] is called the separator tag, which is used to separate two sentences or indicate the end of a single sentence; It is called a mask tag, which is used to indicate some chunks of words that require the model to predict the relationship between two entities; then, the prompt template Input into BERT for encoding and adopt Contextual representation of As an instance representation describing the semantic relationship between entity pairs , the expression is: ; Among them, represents BERT; the instance representation includes the support set instance representation and the query set instance representation; For the relationship description in the prior knowledge, first, the relationship name of each relationship category and the detailed description are concatenated together to form the input sequence ; then the input is fed into BERT for encoding, and the context representation is used as the relationship representation characterizing the relationship description , the expression is: ; For the multi-granularity entity types in the prior knowledge , each entity type is first transformed into format, then fed into BERT for encoding, and the context representation is used as the entity type representation describing the entity type , the expression is: .
[0008] In one embodiment, the steps of the multi-knowledge reinforcement learning module for performing knowledge reinforcement learning on the support set instances and query set instances according to the prior knowledge include: For the th support set instance k of the relationship category , the known relationship category is used to guide the selection of multi-granularity entity types; that is, in the multi-knowledge reinforcement learning module, first, the correlation coefficient between the th entity type representation and the relationship representation corresponding to its relationship category is calculated based on the attention mechanism, and the expression is: ; Among them, represents the similarity measurement function, that is, the dot product operation; , represents the number of entity types; the entity , and respectively represent the head entity and the tail entity; Then, based on the correlation coefficient , all entity types are selectively fused by weighted summation to obtain the entity type representation related to the relationship category in the support set instances , and the expression is: ; Therefore, the head entity type representation and the tail entity type representation in the support set instances are respectively and ; Furthermore, in the multi-knowledge enhanced learning module, the knowledge graph embedding learning technology is adopted to model the relationship category as the distance transformation from the head entity type to the tail entity type, that is, the tail entity type representation is subtracted from the head entity type representation to obtain the relationship representation belonging to the relationship category , and the expression is: ; Finally, is introduced into the support set instance representation to obtain the enhanced support set instance representation , and the expression is: ; Among them, represents an instance belonging to the relationship category in the support set; For the query set instance , since the relationship category and its corresponding relationship description are unknown, the query set instance representation is used as the key value to learn the correlation between the entity type and the instance itself; that is, in the multi-knowledge enhanced learning module, first, based on the attention mechanism, the correlation coefficient between the entity type representation of the head entity or the tail entity in the query set instance and the query set instance representation is calculated, and the expression is: ; Then, through weighted summation, the entity type representation of the query set instance is obtained, and the expression is: Therefore, the head entity type representation and the tail entity type representation in the query set instance are respectively and; Similarly, further in the multi-knowledge enhanced learning module, the knowledge graph embedding learning technology is adopted to obtain the relationship representation related to the entity type as ; Finally, the enhanced representation of the query set instance is obtained through the following calculations , and the expression is: .
[0009] In one embodiment, the multi-knowledge enhanced learning module generates a prototype enhanced representation based on the obtained enhanced representation of the support set instance, including: Obtain the enhanced representation of the support set instance , and average all the enhanced representations of the support set instances belonging to the relationship category to obtain the prototype representation of this relationship category ; Then integrate the relationship representation corresponding to the relationship category into the corresponding prototype representation to obtain the prototype enhanced representation , and the expression is: ; Among them, , is the support set of the relationship category , represents the total number of support set instances belonging to the relationship category .
[0010] In one embodiment, the steps of the double contrast learning module for differentiating relationship categories from two different levels of instances and prototypes by using instance-based contrast learning and prototype-based contrast learning respectively include: Instance-based contrast learning: First, select a certain support set instance of the relationship category as the origin instance, and use the support set instances belonging to this relationship category as positive examples, and these support set instances belong to the same category as the origin instance; at the same time, regard the support set instances belonging to other relationship categories as negative examples , and these support set instances do not belong to the same category as the origin instance; is the total number of relationship categories in the meta-task; Next, use the dot product operation to measure the similarity between the enhanced representation of the origin instance and the enhanced representations of all positive and negative examples to obtain the similarity between positive and negative sample pairs of the relationship category , which is used to calculate the instance-based contrast loss , and the expression is: ; Among them, the symbol represents the dot product operation; Prototype-based contrastive learning: For the relationship category and its corresponding prototype enhanced representation , select as the positive example, and use the remaining relationship categories 's prototype enhanced representations as negative examples; among them, and , is the set of relationship categories; Then use the dot product operation to measure the similarity between the prototype enhanced representation of the relationship category and the prototype enhanced representations of all relationship categories, and obtain the positive and negative sample pairs of the relationship category , which are respectively expressed as: ; ; ; Among them, is the positive sample pair, is the negative sample pair; Finally, calculate the prototype-based contrastive loss , which is used to help the network better distinguish the feature differences between different category prototypes. The expression is: .
[0011] In one embodiment, the steps of the relationship prediction module predicting the relationship category to which the query set instance belongs according to the similarity between the prototype enhanced representation and the query set instance enhanced representation include: When performing relationship classification based on the prototype enhanced representation of the relationship category , the probability distribution that the query set instance belongs to the relationship category is mainly to calculate the similarity between the query set instance enhanced representation and the prototype enhanced representation of the relationship category , and convert it into a probability representation through the Softmax function. The expression is: ; ; Among them, represents the similarity measurement function, that is, the Euclidean distance; Therefore, for the relationship prediction task, use the cross-entropy loss function to calculate the classification loss of the relationship prediction for the query set instance Evaluate the ownership relationship type, and the expression is: ; Among them, represents the indicator function; if the relationship type is the correct label, then ; otherwise, ; is the query set, is the set of relationship categories.
[0012] In one embodiment, the comprehensive loss function of the multi-knowledge enhanced prototype network is expressed as: ; Among them, and are the instance-based contrast loss and the prototype-based contrast loss respectively, and represent the weight coefficients.
[0013] A few-shot relation extraction device based on a multi-knowledge enhanced prototype network, the device is applied to the text entity recognition and classification scenario, and the device includes: A dataset division module, which is used to divide the dataset containing multiple relationship categories and corresponding instances into a training set, a validation set, and a test set according to categories; A preprocessing module, which is used to randomly extract multiple meta-tasks from each of the divided datasets, and introduce multi-granularity entity types and relationship descriptions as prior knowledge; among them, each meta-task consists of a support set and a query set, and the support set and the query set contain multiple relationship categories and multiple instances corresponding to each relationship category, and each instance consists of a sentence and the relationship category of the entity pair in the sentence; A network construction module, which is used to construct a multi-knowledge enhanced prototype network composed of a semantic encoder, a multi-knowledge enhanced learning module, a double contrast learning module, and a relationship prediction module; among them, the semantic encoder is used to encode and generate the feature representations of instances and prior knowledge; the multi-knowledge enhanced learning module is used to perform knowledge enhanced learning on the support set instances and query set instances according to the prior knowledge, and generate prototype enhanced representations according to the obtained support set instance enhanced representations; the double contrast learning module is used to perform relationship category discrimination learning from two different levels of instances and prototypes respectively by using instance-based contrast learning and prototype-based contrast learning; the relationship prediction module is used to predict the relationship category to which the query set instance belongs according to the similarity between the prototype enhanced representation and the query set instance enhanced representation; A task execution module, which is used to input the meta-tasks and prior knowledge randomly selected from the training set and the validation set into a multi-knowledge enhanced prototype network, and construct a comprehensive loss function including a dual contrastive learning loss and a classification loss for relation prediction to train and evaluate the network until a trained network is obtained to perform few-shot relation extraction tasks.
[0014] In the above-mentioned few-shot relation extraction method and device based on a multi-knowledge enhanced prototype network, a multi-knowledge enhanced prototype network composed of a semantic encoder, a multi-knowledge enhanced learning module, a dual contrastive learning module, and a relation prediction module is constructed. This network adopts prompt learning to design a prompt template with entity information to activate the knowledge in the pre-trained language model, and can obtain a more accurate instance semantic representation; at the same time, two prior knowledges of multi-granularity entity types and relation descriptions are introduced to enhance the semantic representation of instances and prototypes; and a dual contrastive learning module based on instances and prototypes is designed to learn the class distinctiveness and distinguishability of instance representations and prototype representations from two different levels of instances and prototypes, so that the characteristics of all relations in the text can be fully captured and understood based on a small number of samples, and the accuracy of text entity recognition and relation classification prediction is improved. Description of the Drawings
[0015] Figure 1 It is a schematic flowchart of a few-shot relation extraction method based on a multi-knowledge enhanced prototype network in an embodiment; Figure 2 It is a schematic architecture diagram of a multi-knowledge enhanced prototype network in an embodiment. Detailed Embodiments
[0016] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0017] In one embodiment, as Figure 1 shown, a few-shot relation extraction method based on a multi-knowledge enhanced prototype network is provided, including the following steps: Step S1, divide the data set containing multiple relation categories and corresponding instances into a training set, a validation set, and a test set according to categories.
[0018] Among them, dividing the data set according to categories can ensure that the relation categories in the test set and the training set are non-overlapping. Such a setting can ensure that the network model makes accurate predictions and evaluations when facing unseen relation categories.
[0019] Step S2, randomly extract multiple meta-tasks from each of the divided data sets, and introduce multi-granularity entity types and relation descriptions as prior knowledge.
[0020] Specifically, step S2 includes: randomly extracting multiple meta-tasks from the training set, validation set, and test set for training, validation, and testing. In a typical N-way-K-shot (N categories and K samples) setting, each meta-task for few-shot relation extraction consists of a support set and a query set ; among them, the support set contains relation categories, and the set of relation categories is denoted as , and the support set of each relation category is composed of support set instances of the current relation category ; the query set also contains the same relation categories, but the instances in the query set are randomly extracted from the remaining instances of each relation category query set instances ; among them, represents the i th relation category, represents the relation category 's k th support set instance, respectively represent the sentence in and the relation category of the entity pair in the sentence; j represents the th query set instance, respectively represent
[0021] the sentence in and the relation category of the entity pair in the sentence. Each instance in the support set and the query set is composed of a sentence containing a head entity and a tail entity and the relation category indicating the relationship between the head entity and the tail entity ; the goal of the meta-task is to use a very small amount of labeled data in the support set to predict the relation category between the head entity and the tail entity in the sentence of the query set instance ; among them, the
[0022] Furthermore, to obtain a better prototype representation, multi-granularity entity types and relationship descriptions are introduced as prior knowledge. Among them, multi-granularity entity types include two different levels of entity types, namely coarse-grained and fine-grained, providing more fine-grained entity types for the relationship representation of instances. For each entity it may be assigned a coarse-grained entity type or a fine-grained entity type. To unify all possible entity types of an entity the multi-granularity entity type represents all entity type information contained in an entity , represents the number of entity types, is the i th entity type and . The relationship description consists of the relationship name of the relationship category and the corresponding detailed description . These information summarize the specific characteristics and attributes of the relationship category, which helps to deeply understand the relationship category.
[0023] Step S3, construct a multi-knowledge enhanced prototype network composed of a semantic encoder, a multi-knowledge enhanced learning module, a dual contrast learning module, and a relationship prediction module.
[0024] As Figure 2 shown, the semantic encoder is used to encode and generate the feature representations of instances and prior knowledge; the multi-knowledge enhanced learning module is used to perform knowledge enhanced learning on the support set instances and query set instances according to the prior knowledge, and generate a prototype enhanced representation according to the obtained enhanced representation of the support set instances; the dual contrast learning module is used to perform relationship category discrimination learning from two different levels of instances and prototypes respectively by using instance-based contrast learning and prototype-based contrast learning; the relationship prediction module is used to predict the relationship category to which the query set instance belongs according to the similarity between the prototype enhanced representation and the enhanced representation of the query set instance. The specific implementation processes of the components in the multi-knowledge enhanced prototype network are as follows: (1) Semantic encoder: The pre-trained language model has learned rich language knowledge from large-scale unlabeled text data through self-supervised learning, so as to be able to understand the relationship between words in a sentence and obtain a more accurate semantic representation. With the emergence of GPT-3 (Generative Pretrained Transformer), it has promoted a new learning method - prompt learning, which guides the model to generate specific types of outputs by providing prompt texts for the input of the model. Therefore, the multi-knowledge enhanced prototype network proposed in this application uses BERT as the semantic encoder to learn the relevant context semantic representations of instances and prior knowledge, and designs a prompt template with entity information as the instance input through prompt learning to guide the model to output feature representations related to the relationship category.
[0025] Specifically, the steps for the semantic encoder to encode and generate the feature representation of an instance include: For the instances in the support set and the query set, first, the sentences in the instance are concatenated with the prompt text with the entity pair to form a prompt template with entity information, ; Here, the prompt template adopts the relational triple structure of the entity pair as the prompt text to activate the relational representation of the entity pair in the instance; among them, [CLS] is called the classification token, which is used to represent the start of the input sequence; [SEP] is called the separator token, which is used to separate two sentences or represent the end of a single sentence; is called the mask token, which is used to represent some chunks that require the model to predict the relationship between two entities; then, the prompt template is input into BERT for encoding, and the contextual representation is used as the instance representation that describes the semantic relationship between entity pairs , and the expression is: ; Among them, represents BERT; the instance representation includes the support set instance representation and the query set instance representation.
[0026] Specifically, the steps for the semantic encoder to encode and generate the feature representation of prior knowledge include: For the relationship descriptions in the prior knowledge, first, the relationship name of each relationship category and the detailed description are concatenated together to form the input sequence ; then is input into BERT for encoding, and the contextual representation is used as the relationship representation that depicts the relationship description , and the expression is: .
[0027] Furthermore, considering that an entity may have multiple fine-grained entity types in addition to the coarse-grained entity type, and the degree of correlation between different entity types and relationship categories is also different. Therefore, the semantic encoder is used to encode each entity type in the multi-grained entity type separately.
[0028] For the multi-grained entity types in the prior knowledge , each entity type is first converted into format, then input into BERT for encoding, and contextual representation is used as the entity type representation describing the entity type , and the expression is: .
[0029] (2) Multi-knowledge enhanced learning module: In order to learn more representative instance representations and prototype representations, the multi-knowledge enhanced prototype network proposes a multi-knowledge enhanced learning module, which performs knowledge-enhanced learning on instances and prototypes by introducing two types of prior knowledge: multi-granularity entity types and relationship descriptions. For instances, the attention mechanism is first used to select more important entity types from the multi-granularity entity types to obtain new entity type representations; then, knowledge graph embedding technology is used to enhance the instance representation by calculating the relationship category representation corresponding to the head entity type representation and the tail entity type representation. For prototypes, the relationship representations learned from the relationship descriptions are integrated into the prototype representations to enhance the prototype representations by highlighting the unique features of the relationship categories.
[0030] Specifically, the steps for the multi-knowledge enhanced learning module to perform knowledge-enhanced learning on the support set instances according to the prior knowledge include: For the th k support set instance of the relationship category , the known relationship category is used to guide the selection of multi-granularity entity types; that is, in the multi-knowledge enhanced learning module, first, the correlation coefficient between the relationship representation corresponding to the th entity type representation is calculated based on the attention mechanism; where represents the similarity measurement function, i.e., the dot product operation; , represents the number of entity types; entity , and represent the head entity and the tail entity respectively; Then, based on the correlation coefficient , all entity types are selectively fused in a weighted summation manner to obtain the entity type representation related to the relationship category in the support set instance, and the expression is: ; Therefore, the head entity type representation and the tail entity type representation in the support set instance are respectively and ; Furthermore, in the multi-knowledge enhanced learning module, the knowledge graph embedding learning technology is adopted to model the relationship category as the distance transformation from the head entity type to the tail entity type, that is, the tail entity type representation is subtracted from the head entity type representation to obtain the relationship representation of the belonging relationship category, and the expression is: ; Finally, is introduced into the support set instance representation to obtain the enhanced representation of the support set instance, and the expression is: ; Among them, represents an instance in the support set that belongs to the relationship category.
[0031] Specifically, the steps of the multi-knowledge enhanced learning module for performing knowledge enhanced learning on the query set instance according to prior knowledge include: For the query set instance , since the relationship category and its corresponding relationship description are unknown, the query set instance representation is used as the key value to learn the correlation between the entity type and the instance itself; that is, in the multi-knowledge enhanced learning module, first, based on the attention mechanism, the entity type representation of the head entity or the tail entity in the query set instance and the query set instance representation are used to calculate the correlation coefficient therebetween, and the expression is: ; Then, through weighted summation, the entity type representation of the query set instance is obtained, and the expression is: ; Therefore, the head entity type representation and the tail entity type representation in the query set instance are respectively and ; Similarly, in the multi-knowledge enhanced learning module, the knowledge graph embedding learning technology is further adopted to obtain the relationship representation related to the entity type as ; Finally, the enhanced representation of the query set instance is obtained through the following calculation, and the expression is: .
[0032] Specifically, the multi-knowledge enhanced learning module generates a prototype enhanced representation based on the obtained support set instance enhanced representation, including: Since the relationship description generalizes the specific features of the relationship category, it brings considerable benefits to correcting the prototype representation of the relationship category. First, obtain the support set instance enhanced representation , and average all the support set instance enhanced representations belonging to the relationship category to obtain the prototype representation of this relationship category ; then integrate the relationship representation of the relationship category into the corresponding prototype representation to obtain the prototype enhanced representation , and the expression is: ; where , is the support set of the relationship category , represents the total number of support set instances belonging to the relationship category .
[0033] (3) Dual contrast learning module: Although prior knowledge can highlight the essential features of the relationship category and is very helpful for learning more representative instance and prototype representations. However, when the semantic representations of the same-class support set instances vary greatly, it is still necessary to highlight the common features of the same-class instance representations and prototype representations and the distinguishable features of different-class instance representations and prototype representations. Therefore, the multi-knowledge enhanced prototype network designs a dual contrast learning strategy based on instances and prototypes to learn unique and distinguishable feature representations of instances and prototypes, thereby improving the accuracy of few-shot learning.
[0034] Specifically, starting from the instance representation level, instance-based contrast learning is designed to learn the differences between in-class and inter-class instances in the support set instances, thereby highlighting the category uniqueness and distinguishability in the instance representation. Instance-based contrast learning includes the following steps: First, select a certain support set instance of the relationship category as the origin instance, and regard the support set instances belonging to this relationship category as positive examples, and these support set instances belong to the same category as the origin instance; at the same time, regard the support set instances belonging to other relationship categories as negative examples , and these support set instances do not belong to the same category as the origin instance; is the total number of relationship categories in the meta-task; Next, use the dot product operation to measure the origin instance 's enhanced representation and the enhanced representations of all positive and negative examples to obtain the relationship category The similarity between positive and negative sample pairs, which is used to calculate the instance-based contrast loss , and the expression is: ; where the symbol represents the dot product operation.
[0035] Specifically, an ideal prototype should contain an immutable category representation that can be distinguished from other categories. Therefore, from the perspective of prototype representation, prototype-based contrast learning is designed to highlight the uniqueness and distinguishability of category prototypes. Prototype-based contrast learning includes the following steps: For the relationship category and its corresponding prototype enhanced representation , select as the positive example, and use the remaining relationship categories 's prototype enhanced representations as negative examples; where and , is the set of relationship categories; Then use the dot product operation to measure the similarity between the prototype enhanced representation of the relationship category and the prototype enhanced representations of all relationship categories, and obtain the positive and negative sample pairs of the relationship category , which are respectively expressed as: ; ; where is the positive sample pair, is the negative sample pair; Finally, calculate the prototype-based contrast loss , which is used to help the network better distinguish the feature differences between different category prototypes, and the expression is: .
[0036] (4) Relationship prediction module: When performing relationship classification based on the prototype enhanced representation of the relationship category , the probability distribution that the query set instance belongs to the relationship category is mainly calculated by the enhanced representation The prototype enhancement representation of the relationship category and calculate the similarity, and convert it into a probability representation through the Softmax function. The expression is as follows: ; wherein, represents the similarity metric function, that is, the Euclidean distance; Therefore, for the relationship prediction task, the cross-entropy loss function is used to calculate the classification loss of the relationship prediction to evaluate the relationship type to which the query set instance belongs. The expression is as follows: ; wherein, represents the indicator function; if the relationship type is the correct label, then ; otherwise, ; is the query set, is the set of relationship categories.
[0037] Step S4: Input the meta-tasks and prior knowledge randomly selected from the training set and the validation set into the multi-knowledge enhanced prototype network, and construct a comprehensive loss function including the dual contrast learning loss and the classification loss of the relationship prediction for network training and evaluation until a trained network is obtained to perform the few-shot relationship extraction task.
[0038] Specifically, the comprehensive loss function of the multi-knowledge enhanced prototype network is expressed as: ; wherein, and are the instance-based contrast loss and the prototype-based contrast loss respectively, and represent the weight coefficients.
[0039] To sum up, a few-shot relationship extraction method based on a multi-knowledge enhanced prototype network provided by the present application activates the pre-trained language model through prompt learning in the constructed multi-knowledge enhanced prototype network to obtain a more accurate semantic representation, and introduces multi-granularity entity type and relationship information as external knowledge to enhance the representation effects of instances and prototypes. At the same time, a dual contrast learning strategy is designed, combining class-agnostic contrast learning and class-specific contrast learning, which can effectively solve the problems of intra-class instances with large syntactic differences and inter-class instances with similar semantics, thereby enhancing the compactness of intra-class instance features and widening the distance between inter-class prototypes, and improving the accuracy of text entity recognition and relationship classification prediction.
[0040] Furthermore, experiments are conducted on two public datasets, FewRel 1.0 and FewRel 2.0, to evaluate the performance of the multi-knowledge enhanced prototype network constructed in this application.
[0041] The FewRel 1.0 dataset contains 100 relation categories, with 700 instances for each relation category. These instances are all extracted from Wikipedia articles. According to the official evaluation settings, Fewrel 1.0 is divided based on 100 relation categories, and the entire dataset is divided into a training set, a validation set, and a test set. Among them, the training set contains 64 relation categories, the validation set contains 16 relation categories, and the test set contains 20 relation categories, with 700 instances for each relation category.
[0042] The FewRel 2.0 dataset mainly studies the problem of domain adaptability. Its training set is the same as that of FewRel 1.0; the validation set is the SemEval-2010 task 8 dataset, which contains 17 relation categories, with 520 instances for each relation category; the test set is the PubMed dataset, which comes from biomedical literature and contains 25 relation categories, with 100 instances for each relation category.
[0043] It should be noted that all the data in FewRel 1.0 belongs to the Wikipedia domain, while the validation set and test set of FewRel 2.0 come from the biomedical domain. The domain difference poses more challenges to the model's rapid learning and generalization. Therefore, FewRel 2.0 is more challenging than FewRel 1.0.
[0044] To obtain the fine-grained entity types of entities, WikiData (Wikidata) in the general domain and UMLS (Unified Medical Language System) in the specific domain are used as external knowledge bases during the experiment. WikiData is a free knowledge base containing structured data of various Wikimedia projects, which contains a large amount of information on various topics, including people, places, concepts, etc. WikiData is used to obtain the fine-grained entity types of entities in the general domain. UMLS is a knowledge base in the biomedical and health fields collected and constructed by experts, which contains a large amount of detailed definitions of medical entities, related concepts, and semantic relationships between different entities. UMLS is used to obtain the fine-grained entity types of entities in the specific domain.
[0045] In the experiment, four N-way-K-shot few-shot learning settings, namely 5-way-1-shot (hereinafter referred to as 5-w-1-s, and the same applies to others), 5-way-5-shot, 10-way-1-shot, and 10-way-5-shot, were set up to evaluate the performance of the multi-knowledge enhanced prototype network. During training, 20,000 meta-tasks were randomly selected from the training set to train the network model, and 1,000 meta-tasks were randomly selected from the validation set to evaluate the network model; during testing, 10,000 meta-tasks were randomly selected from the test set for testing. All experiments were trained and evaluated on an NVIDIA GeForce RTX 2080 Ti graphics card.
[0046] In the comparative experiment, the performance was compared and evaluated with three groups of the most representative methods, namely some basic models, external information enhanced models, and specific pre-trained models, such as Proto-BERT (prototype BERT model), BERT-PAIR (BERT paired classification model), MAML (model-agnostic meta-learning), GNN (graph neural network), REGRAB (relational graph attention benchmark model), CTEG (context temporal event graph model), ConceptFERE (concept-enhanced entity relation extraction model), HCRP (hierarchical concept relation propagation model), MTB (matching template BERT model), CP (contrastive prototype network), and LPD (latent prototype discovery model). The evaluation metric of accuracy was selected to evaluate the performance of the model. Tables 1 and 2 respectively list the comparison of the model performance on the FewRel 1.0 and FewRel 2.0 test sets for the few-shot relation extraction task. KnowProto and KnowProto+CP in the tables represent the multi-knowledge enhanced prototype network constructed in this application. The main experimental results of various models on the FewRel 1.0 and FewRel 2.0 test sets show that the multi-knowledge enhanced prototype networks KnowProto and KnowProto+CP proposed in this application achieved the best performance under all N-way-K-shot few-shot learning settings.
[0047] Table 1 Comparison results of the accuracy of different network models on the FewRel 1.0 test set
[0048] Table 2 Comparison results of the accuracy of different network models on the FewRel 2.0 test set
[0049] In one embodiment, a few-shot relation extraction device based on a multi-knowledge enhanced prototype network is provided, including: A dataset partitioning module, configured to partition a dataset containing multiple relation categories and corresponding instances into a training set, a validation set, and a test set according to the categories; A preprocessing module, configured to randomly extract multiple meta-tasks from each of the partitioned datasets, and introduce multi-granularity entity types and relation descriptions as prior knowledge; wherein, each meta-task consists of a support set and a query set, and the support set and the query set contain multiple relation categories and multiple instances corresponding to each relation category, and each instance consists of a sentence and the relation category of the entity pair in the sentence; A network construction module, configured to construct a multi-knowledge enhanced prototype network composed of a semantic encoder, a multi-knowledge enhanced learning module, a dual contrast learning module, and a relation prediction module; wherein, the semantic encoder is configured to encode and generate feature representations of instances and prior knowledge; the multi-knowledge enhanced learning module is configured to perform knowledge enhanced learning on the support set instances and the query set instances according to the prior knowledge, and generate a prototype enhanced representation according to the enhanced representation of the support set instances obtained; the dual contrast learning module is configured to perform relation category discrimination learning from two different levels of instances and prototypes respectively by using instance-based contrast learning and prototype-based contrast learning; the relation prediction module is configured to predict the relation category to which the query set instance belongs according to the similarity between the prototype enhanced representation and the enhanced representation of the query set instance; A task execution module, configured to input the meta-tasks and prior knowledge randomly extracted from the training set and the validation set into the multi-knowledge enhanced prototype network, and construct a comprehensive loss function including a dual contrast learning loss and a classification loss of relation prediction for network training and evaluation until a trained network is obtained to perform few-shot relation extraction tasks.
[0050] For the specific limitations of the few-shot relation extraction device based on the multi-knowledge enhanced prototype network, reference may be made to the limitations of the few-shot relation extraction method based on the multi-knowledge enhanced prototype network in the foregoing text, which will not be elaborated here. Each module in the foregoing few-shot relation extraction device based on the multi-knowledge enhanced prototype network may be implemented in whole or in part by software, hardware, and their combination. The foregoing modules may be embedded in or independent of a processor in a computer device in the form of hardware, or may be stored in a memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the foregoing modules.
[0051] The technical features of the above embodiments may be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered that the scope described in this specification.
[0052] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A few-sample relation extraction method based on multi-knowledge enhanced prototype network, characterized in that: The method is applied to text entity recognition and classification scenarios, and the method includes: Divide the data set containing multiple relationship categories and corresponding instances into training set, validation set and test set according to the categories; Multiple meta-tasks are randomly extracted from each divided data set, and multi-granular entity types and relationship descriptions are introduced as prior knowledge; each meta-task consists of a support set and a query set, and the support set and query set contain multiple relationship categories and multiple instances corresponding to each relationship category, and each instance consists of a sentence and a relationship category of an entity pair in the sentence; A multi-knowledge enhanced prototype network consisting of a semantic encoder, a multi-knowledge enhanced learning module, a dual contrast learning module and a relationship prediction module is constructed; wherein the semantic encoder is used to encode the feature representation of the generated instance and the prior knowledge; the multi-knowledge enhanced learning module is used to perform knowledge enhancement learning on the support set instance and the query set instance according to the prior knowledge, and generate the prototype enhanced representation according to the obtained support set instance enhanced representation; the dual contrast learning module is used to use instance-based contrast learning and prototype-based contrast learning to perform relationship category distinction learning from two different levels, instance and prototype; the relationship prediction module is used to predict the relationship category to which the query set instance belongs according to the similarity between the prototype enhanced representation and the query set instance enhanced representation; The meta-tasks and prior knowledge randomly extracted from the training set and the validation set are input into the multi-knowledge enhanced prototype network, and a comprehensive loss function including the double contrastive learning loss and the classification loss of relationship prediction is constructed for network training and evaluation until a well-trained network is obtained to perform the few-sample relationship extraction task.
2. The method according to claim 1, characterized in that: Multiple meta-tasks are randomly extracted from each divided data set, and multi-granular entity types and relationship descriptions are introduced as prior knowledge, including: Multiple meta-tasks are randomly extracted from the training set, validation set, and test set for training, validation, and testing. Each meta-task of the few-shot relation extraction consists of a support set and a queryset Composition; among them, the support set Include relation categories, and the relation category set is represented as , each relationship category The support set By current relationship category of Support set examples Composition; query set Also includes the same relation categories, but the instances in the query set are randomly drawn from the remaining instances of each relation category. QuerySet instances constitute; among them, Indicates i relationship categories, Represents relationship category No. k Support set examples, Respectively The sentences in and the relationship categories of entity pairs in the sentences; Indicates j QuerySet instances, Respectively The sentences in and the relationship categories of entity pairs in the sentences; Support set and queryset Each instance in is a header entity and tail entity Sentences and the relationship between the head entity and the tail entity composition; the goal of the meta-task is to use the support set Very few labeled data are used to predict the sentences of query set instances The relationship type between the head entity and the tail entity ; Among them, each meta-task extracts The relationship categories are different; Furthermore, multi-granularity entity types and relationship descriptions are introduced as prior knowledge; Represents an entity All entity type information contained, Indicates the number of entity types, For the i entity types and ; The relationship description consists of the relationship name of the relationship category and the corresponding detailed description composition.
3. The method according to claim 2, characterized in that The steps of encoding the feature representation of the generated instance and prior knowledge by the semantic encoder include: Use the pre-trained language model BERT as the semantic encoder; For the instances in the support set and query set, firstly, the sentences in the instance With entity pairs The prompt texts are stitched together to form a prompt template with entity information , the expression is: ; Tip template here The relation triple structure of entity pairs is used as the prompt text to activate the relation representation of entity pairs in the instance; [CLS] is called the classification tag, which is used to indicate the beginning of the input sequence; [SEP] is called the separator tag, which is used to separate two sentences or indicate the end of a single sentence; It is called a mask tag, which is used to represent some chunks of words that require the model to predict the relationship between two entities; then, the prompt template Input into BERT for encoding and adopt Contextual representation of As an instance representation describing the semantic relationship between entity pairs , the expression is: ; in, represents BERT; the example represents Including support set instance representation and query set instance representation; For the relationship description in prior knowledge, firstly, each relationship category Relationship name and detailed description Spliced together to form the input sequence ; then Input into BERT for encoding and adopt Contextual representation of Relational Representation as a Characterization of Relational Descriptions , the expression is: ; For multi-granular entity types in prior knowledge , first separate each entity type Convert to format, and then input into BERT for encoding and using Contextual representation of As an entity type representation describing an entity type , the expression is: 。 4. The method according to claim 3, characterized in that The steps of the multi-knowledge enhanced learning module to perform knowledge enhanced learning on the support set instances and the query set instances according to prior knowledge include: For relationship categories No. k Support set examples , using known relationship categories to guide the selection of multi-granular entity types; that is, in the multi-knowledge reinforcement learning module, firstly, the first Entity type representation The relationship representation corresponding to the relationship category to which it belongs The correlation coefficient between , the expression is: ; in, represents the similarity metric function, i.e., the dot product operation; , Indicates the number of entity types; entity , and Represent the head entity and the tail entity respectively; Then, based on the correlation coefficient , selectively fuse all entity types by weighted summation to obtain the entity type representation related to the relationship category in the support set instance , the expression is: ; Therefore, the head entity type representation and the tail entity type representation in the support set instance are and ; Furthermore, the knowledge graph embedding learning technology is used in the multi-knowledge reinforcement learning module to model the relationship category as a distance transformation from the head entity type to the tail entity type, that is, the tail entity type is used to represent Entity type representation minus the header Get the relationship representation of the belonging relationship category , the expression is: ; Finally, Introducing support set instance representation In the example, we get the support set enhanced representation , the expression is: ; in, Indicates that the support set belongs to the relationship category An example of For queryset instances Since the relationship category and its corresponding relationship description are unknown, the query set instance is used to represent As a key value, we learn the correlation between the entity type and the instance itself; that is, in the multi-knowledge reinforcement learning module, we first calculate the entity type representation of the head entity or tail entity in the query set instance based on the attention mechanism. and queryset instance representation The correlation coefficient between , the expression is: ; Then, through weighted summation, the entity type representation of the query set instance is obtained , the expression is: ; Therefore, the head entity type representation and the tail entity type representation in the query set instance are and ; Similarly, the knowledge graph embedding learning technology is further used in the multi-knowledge reinforcement learning module to obtain the relationship representation related to the entity type: ; Finally, the query set instance enhanced representation is obtained by the following calculation , the expression is: 。 5. The method according to claim 4, characterized in that The multi-knowledge reinforcement learning module generates a prototype reinforcement representation based on the obtained support set instance reinforcement representation, including: Get the enhanced representation of the support set instance , and will belong to the relationship category The enhanced representations of all support set instances are averaged to obtain the relationship category The prototype representation of Then the relationship category The corresponding relationship is expressed Integrate into the corresponding prototype representation to obtain the prototype enhanced representation , the expression is: ; in, , For relationship category The support set of Indicates that it belongs to the relationship category The total number of support instances.
6. The method according to claim 5, characterized in that The dual contrastive learning module uses instance-based contrastive learning and prototype-based contrastive learning to perform relationship category distinction learning from two different levels: instance and prototype. The steps include: Instance-based contrastive learning: first select the relationship category An instance of the support set of As the origin instance, it will belong to the relationship category of Support set examples As positive examples, these support set instances belong to the same category as the origin instance; at the same time, The support set instances of the relation categories are regarded as negative examples. , these support set instances do not belong to the same category as the origin instances; is the total number of relation categories in the meta-task; Next, use the dot product operation to measure the origin instance Enhanced representation of and the enhanced representation of all positive and negative examples The similarity between them is used to obtain the relationship category The similarity between positive and negative sample pairs is used to calculate the instance-based contrast loss , the expression is: ; Among them, the symbol Represents the dot product operation; Prototype-based contrastive learning: for relational categories And its corresponding prototype enhanced representation , select As a positive example, and the rest Relationship categories Prototype Enhanced Representation As a negative example; and , is a set of relation categories; Then use the dot product operation to measure the relationship category Prototype Enhanced Representation The similarity between the prototype enhanced representations of all relation categories is obtained. The positive and negative sample pairs are expressed as: ; ; in, is a positive sample pair, is a negative sample pair; Finally, calculate the prototype-based contrast loss , which is used to help the network better distinguish the feature differences between prototypes of different categories. The expression is: 。 7. The method according to claim 6, characterized in that The step of predicting the relationship category to which the query set instance belongs according to the similarity between the prototype enhanced representation and the query set instance enhanced representation by the relationship prediction module includes: Based on the relationship category Prototype Enhanced Representation When classifying relations, query set instances Belongs to the relationship category The probability distribution of Mainly calculate the query set instance enhanced representation Relationship Category Prototype Enhanced Representation The similarity between them is converted into a probability representation through the Softmax function, and the expression is: ; in, represents the similarity metric function, namely the Euclidean distance; Therefore, for the relationship prediction task, the cross entropy loss function is used to calculate the classification loss of relationship prediction To query set instance The relationship type is evaluated, and the expression is: ; in, represents the indicator function; if the relationship type is the correct label, then ;otherwise, ; For the query set, A collection of relationship categories.
8. The method according to claim 7, characterized in that Comprehensive loss function of multi-knowledge enhanced prototype network It is expressed as: ; in, and They are instance-based contrast loss and prototype-based contrast loss, and Represents the weight coefficient.
9. A device for extracting relations from a small number of samples based on a multi-knowledge enhanced prototype network, characterized in that: The device is applied to text entity recognition and classification scenarios, and the device comprises: A data set partitioning module is used to partition a data set containing multiple relationship categories and corresponding instances into a training set, a validation set, and a test set according to the categories; A preprocessing module is used to randomly extract multiple meta-tasks from each divided data set and introduce multi-granularity entity types and relationship descriptions as prior knowledge; each meta-task consists of a support set and a query set, and the support set and query set contain multiple relationship categories and multiple instances corresponding to each relationship category, and each instance consists of a sentence and a relationship category of an entity pair in the sentence; A network construction module is used to construct a multi-knowledge enhanced prototype network consisting of a semantic encoder, a multi-knowledge enhanced learning module, a dual contrast learning module and a relationship prediction module; wherein the semantic encoder is used to encode and generate feature representations of instances and prior knowledge; the multi-knowledge enhanced learning module is used to perform knowledge enhancement learning on support set instances and query set instances based on prior knowledge, and generate prototype enhanced representations based on the obtained support set instance enhanced representations; the dual contrast learning module is used to perform relationship category distinction learning from two different levels, instance and prototype, using instance-based contrast learning and prototype-based contrast learning respectively; the relationship prediction module is used to predict the relationship category to which the query set instance belongs based on the similarity between the prototype enhanced representation and the query set instance enhanced representation; The task execution module is used to input meta-tasks and prior knowledge randomly extracted from the training set and the validation set into the multi-knowledge enhanced prototype network, and construct a comprehensive loss function including double contrastive learning loss and classification loss of relationship prediction to perform network training and evaluation until a trained network is obtained to perform the few-sample relationship extraction task.
Citation Information
Patent Citations
Zero sample relation extraction method and model based on dual contrast learning framework and cross attention module
CN118484537A
Cross-domain small sample relation extraction method and device for learning fine-grained general knowledge
CN118674036A
Small sample remote sensing image scene classification method based on embedding smoothing graph neural network
WO2023087558A1
Cited By
Active element learning-based few-sample entity relationship extraction method and related equipment
CN120873200A