Few-sample triple extraction method and system, computer equipment and storage medium
By combining the prototype network and the model-independent meta-learning algorithm, the small sample triple extraction method is solved, and the entity difference and sequence labeling errors of relation triple extraction under the small sample setting are achieved, and higher quality prototype and triple extraction are achieved.
Patent Information
- Application Number
- CN202510275716.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
Under the few-sample setting, the relationship triple extraction task faces the problems of entity differences, sequence labeling errors, and incomplete relationship prototypes, resulting in poor performance of the model when identifying and extracting the correct entities and relationships.
Prototype network algorithm and model-independent element learning algorithm are adopted to improve prototype quality and triple quality by embedding vector representation, relationship extraction and classification, and entity position recognition.
By explicitly fusion of relational semantic information and using meta-learning algorithms, the representativeness of the prototype and the accuracy of triplets are improved, and the entity differences and sequence labeling errors in few-sample learning are solved.
Smart Images

Figure CN120218068A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular, to a few-shot triple extraction method, system, computer device, and storage medium. Background Art
[0002] Relation triple extraction is to identify structured information in the form of (head entity - relation - tail entity) triples from unstructured text, and has wide applications in knowledge graph construction, question answering systems, etc. in the field of natural language processing. Traditional relation triple extraction methods usually require a large amount of labeled data to effectively train models, presenting significant challenges in scenarios where labeled data is scarce or the acquisition cost is high.
[0003] Therefore, to solve the dilemma of relation triple extraction in few-shot settings, few-shot relation triple extraction emerged and has become a research hotspot in recent years. The core idea of few-shot relation triple extraction is to enable the model to learn from a limited number of labeled samples and thus adapt to previously unseen data. In the relation triple extraction task, this means that the model needs to learn how to identify and extract correct entities and relations from instances with only a small number of labeled samples.
[0004] Previous few-shot relation triple extraction studies usually solely based on metric learning, by comparing the query set between the instance and the relation prototype, the distance between the query label and the entity prototype, and classifying the label according to the distance between them. In addition, some recent research works have tried to simulate the traditional triple extraction task, that is, first identify entities and then classify relations based on the extracted entities. However, this simple method of combining the traditional triple extraction task with few-shot learning will undoubtedly cause some problems.
[0005] (1) There are entity difference problems in few-shot triple extraction, which means that the entities in newly emerging relations may contain entity types completely different from those previously identified entities, because each relation imposes specific constraints on the types of the head entity and the tail entity. Therefore, the entity detector trained based on known relations is likely to be confused when facing the entity types corresponding to new relations.
[0006] (2) To solve problem (1), some studies have adopted the relation-entity extraction mode, or adopted the mutual guidance strategy of relation and entity extraction. These studies still choose to use prototypes to classify entities in the query instance during the entity extraction process, but ignore the consideration of label dependency relationships during the entity sequence annotation process. This indicates that there is no clear constraint on the dependency relationships between sequence annotations, which may lead to obvious sequence annotation errors.
[0007] (3) Implicitly using relational information to constrain relational prototypes will lead to incomplete construction of relational prototypes. The current methods mainly utilize the specified instances in each relation to obtain sentence embedding representations, usually by averaging the embedding representations of these instances to obtain prototypes, which means that implicitly incorporating relational information through contrastive learning, graph learning, or attention to constrain the prototype representation will result in prototypes that are not sufficiently representative. Summary of the Invention
[0008] To solve the above problems, the present invention proposes a few-shot triple extraction method, system, computer device, and storage medium. Based on the characteristics of few-shot learning, combining the prototype network algorithm and the model-agnostic meta-learning algorithm, the prototype network is reasonably utilized in the few-shot relation triple extraction task. The relational semantic information can assist the prototype network, improve the quality of the generated prototypes to make the prototypes more representative, and improve the quality of the finally generated triples.
[0009] The technical solution adopted by the present invention is as follows:
[0010] A few-shot triple extraction method, comprising:
[0011] Embedding vector representation: Obtain support set and query set samples through the base dataset, and utilize a pre-trained language model to generate the embedding vector representation of each sentence in the support set and query set samples;
[0012] Relation extraction and classification: Integrate the relational semantic information between entities through the attention mechanism to form enhanced relational semantic information; perform an average process on the embedding vector representations of the support set samples to obtain initial prototypes, and construct enhanced prototypes in combination with the relational semantic information; compare the similarity between the sentences in the query set and the enhanced prototypes to achieve relation extraction and classification;
[0013] Entity position recognition: Based on the sequence annotation strategy and the model-agnostic meta-learning algorithm, determine the position of each term in the sentence and assign corresponding type labels, and form complete triples according to the identified head entity, relation, and tail entity combinations.
[0014] Further, the obtaining of the support set and query set samples through the base dataset, and the utilization of the pre-trained language model to generate the embedding vector representation of each sentence in the support set and query set samples includes:
[0015] Randomly obtain samples from the base dataset for constructing the support set and query set, and obtain the relation names and relation description information according to all relation categories in the base dataset;
[0016] Input the support set, query set, and relationship description information into the language model for encoding to obtain the embedded vector representations of each sentence. The language model includes the Bert and Roberta models.
[0017] Further, integrating the relationship semantic information between entities through the attention mechanism to form enhanced relationship semantic information; averaging the embedded vector representations of the support set samples to obtain an initial prototype, and constructing an enhanced prototype in combination with the relationship semantic information, including:
[0018] Integrate the semantic information corresponding to the entity relationship through the attention mechanism to detect the classification of the relationship, and perform an averaging operation on the embedded vector representations in the query set to obtain an initial prototype. One initial prototype corresponds to one relationship category, representing a representative vector representation under this relationship category;
[0019] Explicitly fuse the relationship semantic information into the initial prototype to obtain a more representative enhanced prototype.
[0020] Further, based on the sequence annotation strategy and model-agnostic meta-learning algorithm, determine the position of each term in the sentence and assign corresponding type labels, and form complete triples according to the identified head entity, relationship, and tail entity, including:
[0021] Convert the entity position detection problem into a sequence annotation problem. Based on the improved BIO sequence tagging strategy, cross-entropy loss calculation, and Viterbi decoding, locate the position of each term in the sentence and distinguish the type labels, thereby locating the entity position;
[0022] Based on the model-agnostic meta-learning algorithm, find the initialization parameters that can adapt to the new domain according to few samples, so as to complete entity position recognition without knowing the entity type and help generate the final triples.
[0023] A few-shot triple extraction system, including:
[0024] An encoder, configured to obtain support set and query set samples through a basic dataset, and use a pre-trained language model to generate the embedded vector representations of each sentence in the support set and query set samples;
[0025] A relationship extractor, configured to integrate the relationship semantic information between entities through the attention mechanism to form enhanced relationship semantic information; average the embedded vector representations of the support set samples to obtain an initial prototype, and construct an enhanced prototype in combination with the relationship semantic information; compare the similarity between the sentences in the query set and the enhanced prototype to realize the extraction and classification of relationships;
[0026] An entity position recognizer, configured to determine the position of each term in a sentence and assign corresponding type labels based on a sequence annotation strategy and a model-agnostic meta-learning algorithm, and form a complete triple according to the identified head entity, relationship, and tail entity combinations.
[0027] Further, in the encoder, support set and query set samples are obtained from a base dataset, and using a pre-trained language model, an embedding vector representation of each sentence in the support set and query set samples is generated, including:
[0028] Randomly obtain samples from the base dataset to construct the support set and query set, and obtain relationship names and relationship description information according to all relationship categories in the base dataset;
[0029] Input the support set, query set, and relationship description information into the language model for encoding to obtain the embedding vector representation of each sentence, and the language model includes Bert and Roberta models.
[0030] Further, the relationship extractor includes:
[0031] An attention mechanism prototype generation module, configured to integrate semantic information corresponding to entity relationships through an attention mechanism to detect the classification of relationships, and perform an averaging operation on the embedding vector representations in the query set to obtain an initial prototype, where one initial prototype corresponds to one relationship category and represents a representative vector representation under this relationship category;
[0032] A relationship semantic information fusion module, configured to explicitly fuse relationship semantic information into the initial prototype to obtain a more representative enhanced prototype;
[0033] A similarity comparison module, configured to compare the similarity between the sentence to be subjected to relationship extraction and the enhanced prototype to complete relationship extraction and classification.
[0034] Further, the entity position recognizer includes:
[0035] A basic sequence prediction module, configured to convert the entity position detection problem into a sequence annotation problem, and based on an improved BIO sequence tagging strategy, cross-entropy loss calculation, and Viterbi decoding, locate the position of each term in the sentence and distinguish type labels, thereby locating the entity position;
[0036] A meta-learning module, configured to based on a model-agnostic meta-learning algorithm, find initialization parameters that can adapt to a new domain according to few samples, so as to be able to complete entity position recognition without knowing entity types and help generate the final triple.
[0037] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned few-shot triple extraction method is implemented.
[0038] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned few-shot triple extraction method is implemented.
[0039] The beneficial effects of the present invention are as follows:
[0040] 1. A pipeline-style few-shot triple extraction framework is proposed, including an encoder, a relation extractor, and an entity position recognizer. Only the parameters in the encoder are shared among the modules, and the mutual influence is very small.
[0041] 2. In terms of relation classification, a relation extractor that uses an attention module and relation semantic information is proposed to explicitly correct prototypes, and an effective parameter-free method is proposed, which only requires training the parameters of the encoder.
[0042] 3. The entity position recognizer is designed using the model-agnostic meta-learning algorithm, which can promote sequence labeling tasks beneficial to label dependencies.
[0043] In summary, based on the prototype network algorithm and model-agnostic meta-learning, the present invention reasonably connects the encoder, the relation extractor, and the entity position recognizer. The relation classification is completed by comparing the similarity with the generated prototypes, and the entity position is located by decoding the sequence of entity recognition through the Viterbi decoder, thereby forming triples by combining relations and entities. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is the training and testing flowchart of few-shot triple extraction in Embodiment 2 of the present invention.
[0045] Figure 2 It is the implementation flowchart of the encoder in Embodiment 2 of the present invention.
[0046] Figure 3 It is the implementation flowchart of the relation extractor in Embodiment 2 of the present invention.
[0047] Figure 4 It is the implementation flowchart of the entity position recognizer in Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0048] For a clearer understanding of the technical features, objectives, and effects of the present invention, the specific implementation manners of the present invention will now be described. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention, that is, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0049] Embodiment 1
[0050] This embodiment provides a few-shot triple extraction method, including:
[0051] Embedding vector representation: Obtain support set and query set samples through the basic dataset, and use a pre-trained language model to generate the embedding vector representation of each sentence in the support set and query set samples;
[0052] Relationship extraction and classification: Integrate the relational semantic information between entities through the attention mechanism to form enhanced relational semantic information; perform an average process on the embedding vector representations of the support set samples to obtain an initial prototype, and construct an enhanced prototype in combination with the relational semantic information; compare the sentences in the query set with the enhanced prototype to achieve the extraction and classification of relationships;
[0053] Entity position recognition: Based on the sequence labeling strategy and the model-agnostic meta-learning algorithm, determine the position of each term in the sentence and assign corresponding type labels, and form a complete triple according to the identified head entity, relationship, and tail entity combinations.
[0054] Based on the characteristics of few-shot learning, this embodiment combines the prototype network algorithm and the model-agnostic meta-learning algorithm (MAML), reasonably utilizes the prototype network in the few-shot relation triple extraction task, and the relational semantic information can assist the prototype network, improve the quality of the generated prototype to make the prototype more representative, and improve the quality of the finally generated triple.
[0055] It should be noted that the prototype network algorithm is a machine learning method for few-shot learning. Its core idea is to use a set of pre-trained prototypes to represent different categories or concepts, and predict their categories by comparing the similarity between the input instances and these prototypes. The prototype network algorithm performs well in dealing with situations with only a small amount of labeled data and is widely used in tasks such as image classification and text classification.
[0056] In addition, the model-agnostic meta-learning algorithm is a meta-learning algorithm aimed at improving the performance of a model on new tasks through learning from a small number of task samples. Its core idea is to utilize the shared knowledge of a pre-trained model among multiple tasks, thereby reducing the time and computational cost of model training.
[0057] Preferably, support set and query set samples are obtained from a base dataset, and using a pre-trained language model, an embedded vector representation of each sentence in the support set and query set samples is generated, including:
[0058] Samples are randomly obtained from the base dataset to construct the support set and query set, and relation names and relation description information are obtained according to all relation categories in the base dataset;
[0059] The support set, query set, and relation description information are input into the language model for encoding to obtain the embedded vector representation of each sentence. The language model includes Bert and Roberta models.
[0060] Preferably, the relation semantic information between entities is integrated through an attention mechanism to form enhanced relation semantic information; an average process is performed on the embedded vector representation of the support set samples to obtain an initial prototype, and an enhanced prototype is constructed in combination with the relation semantic information, including:
[0061] The semantic information corresponding to the entity relation is integrated through an attention mechanism to detect the classification of the relation, and an average operation is performed on the embedded vector representation in the query set to obtain an initial prototype. One initial prototype corresponds to one relation category, representing a representative vector representation under this relation category;
[0062] The relation semantic information is explicitly fused into the initial prototype to obtain a more representative enhanced prototype.
[0063] Preferably, based on a sequence annotation strategy and a model-agnostic meta-learning algorithm, the position of each term in the sentence is determined and the corresponding type label is assigned, and a complete triple is formed according to the identified head entity, relation, and tail entity, including:
[0064] The entity position detection problem is transformed into a sequence annotation problem. Based on an improved BIO sequence tagging strategy, cross-entropy loss calculation, and Viterbi decoding, the position of each term in the sentence is located and the type label is distinguished, thereby locating the entity position;
[0065] Based on the model-agnostic meta-learning algorithm, the initialization parameters that can adapt to the new domain are found according to few samples, so that the entity position recognition can be completed without knowing the entity type and it helps to generate the final triple.
[0066] Embodiment 2
[0067] This embodiment provides a few-shot triple extraction system, including:
[0068] An encoder, configured to obtain support set and query set samples through a base dataset, and use a pre-trained language model to generate embedding vector representations for each sentence in the support set and query set samples;
[0069] A relation extractor, configured to integrate the relational semantic information between entities through an attention mechanism to form enhanced relational semantic information; perform an average process on the embedding vector representations of the support set samples to obtain an initial prototype, and construct an enhanced prototype in combination with the relational semantic information; compare the similarity between the sentences in the query set and the enhanced prototype to achieve the extraction and classification of relations;
[0070] An entity position recognizer, configured to determine the position of each term in the sentence and assign corresponding type labels based on a sequence annotation strategy and a model-agnostic meta-learning algorithm, and form a complete triple according to the identified head entity, relation, and tail entity combinations.
[0071] In this embodiment, the encoder completes the encoding of the sentence to generate embedding vector representations (Embedding representations) with sentence semantics. In the relation extractor, this embodiment explicitly introduces relational semantic information and combines it with the prototype network algorithm to incorporate the relational semantic information into the prototype to be generated, making the generated prototype more representative and assisting the prototype network algorithm to perform more accurate relation classification. In the entity position recognizer, this embodiment uses a model-agnostic meta-learning algorithm to identify the entity positions in the entity recognition stage, which enables our triple extraction model to identify the positions of the head and tail entities without knowing the entity types and helps generate the final triple results.
[0072] Based on the above technologies, this embodiment proposes a few-shot triple extraction system using the prototype network algorithm and the model-agnostic meta-learning algorithm, reasonably connects the encoder, the relation extractor, and the entity position recognizer, completes relation classification by comparing the similarity with the generated prototype, and locates the entity positions by decoding the sequence of entity recognition through the Viterbi decoder, combines the relation and the entity to form a triple, and forms a few-shot triple extraction system based on the prototype network and the model-agnostic algorithm.
[0073] In this embodiment, by decomposing the triple extraction task, a relation extractor is implemented in sequence to address the challenges related to few-shot relation extraction. First, instance representations of implicit relation semantics are obtained through a prompt encoder and an attention module, and then explicit relation semantic information is fused to enhance the relation prototype. Finally, for few-shot entity span detection, in the case of entity category independence, it is modeled as a sequence labeling problem to avoid dealing with the problem of span overlap. Their call flow is as Figure 1 shown.
[0074] Preferably, in the encoder, through a language model such as Bert or Roberta, combined with the support set and query set samples obtained from the basic dataset, the embedding vector representations corresponding to each sentence in the support set and query set samples are output.
[0075] Preferably, the relation extractor includes:
[0076] An attention mechanism prototype generation module, configured to integrate the semantic information corresponding to the entity relation through the attention mechanism to detect the classification of the relation, and perform an averaging operation on the embedding vector representations in the query set to obtain an initial prototype. One initial prototype corresponds to one relation category, representing a representative vector representation under this relation category;
[0077] A relation semantic information fusion module, configured to explicitly fuse the relation semantic information into the initial prototype to obtain a more representative enhanced prototype;
[0078] A similarity comparison module, configured to compare the similarity between the sentence to be subjected to relation extraction and the enhanced prototype to complete relation extraction and classification.
[0079] Preferably, the entity position recognizer includes:
[0080] A basic sequence prediction module: Since the actually input sentence may not contain the entity type information, a sequence labeling strategy is applied to the entity position recognition task, and this model is entity type independent, so it can adapt to the entity position recognition tasks in most domains. The basic sequence prediction module locates the position of each term in the sentence and distinguishes the type label based on the improved BIO sequence labeling strategy, cross-entropy loss calculation, and Viterbi decoding, so as to locate the entity position.
[0081] A meta-learning module, configured to find the initialization parameters that can adapt to the new domain based on the model-agnostic meta-learning algorithm according to few samples, endowing the model with strong generalization ability and sampling efficiency, so as to complete entity position recognition without knowing the entity type and help generate the final triple.
[0082] Embodiment 3
[0083] Based on Embodiment 2, this embodiment is as follows:
[0084] This embodiment provides a few-shot triple extraction system, including an encoder, a relation extractor, and an entity position recognizer, which are specifically described as follows.
[0085] I. Encoder
[0086] The construction of the dataset is to randomly obtain samples from the basic dataset to construct the support set and the query set, and obtain information such as relation names and relation descriptions according to all relation categories in the dataset; on this basis, a language model such as Bert or Roberta is selected as the encoder. After inputting the support set, the query set, and the relation description information into the encoder, the embedded vector representation of each sentence is obtained. The encoding method of the encoder mainly includes three steps: support set encoding, query set encoding, and relation description information encoding, as Figure 2 shown:
[0087] 1) Adopt the few-shot data acquisition mode of N-way K-shot, select N relation categories from the basic dataset, select K sample instances in each relation type to construct the support set, and input the instance sentences into the encoder to obtain the embedded vector representation of the support set: where d represents the size of the encoder hidden layer.
[0088] 2) Select Q sample instances from the relation range of the same relation category as in the support set in the basic dataset to construct the query set, and input the instance sentences in the query set into the encoder to obtain the embedded vector representation of the query set: {Q j =(q0,q1,…,q Q )∈R d ; j = 1,…,Q}, where d represents the size of the encoder hidden layer.
[0089] 3) For each relation, use a template similar to "name:description" to connect the name and description of the relation, and then input the sequence into the encoder to obtain the representation of the relation. For example, combine the relation name "is located in" and its description "something can be found in a specific thing or location" into the sequence "is located in:something can be found in a specific thing or location". Input this sequence into the encoder to get the representation form of this relation {R=(r 0 ,r 1 ,…,r N )∈R d ; i = 1,…,N}.
[0090] II. Relation Extractor
[0091] After constructing the basic dataset and inputting it into the encoder, the embedded vector representations of the support set, query set, and relationship description information are obtained. Then, since the semantics of the relationship are implicitly contained in the context semantics, an attention module is adopted to integrate the semantic information corresponding to the relationship to detect the classification of the relationship. Subsequently, an averaging operation is performed on the embedded vector representations of K instances in the query set to obtain the initial prototype. One prototype corresponds to one relationship category and represents a representative vector representation under this relationship category. Then, the relationship semantic information is explicitly fused into the initial prototype to obtain a more representative enhanced prototype.
[0092] Therefore, based on the above processing flow, the implementation method of the relationship extractor can be divided into the following steps, as Figure 3 shown:
[0093] (1) Instance prototype acquisition: Since the embedding of the sentence contains the context semantics of the relationship, and the embeddings of the head and tail entities contain the entity type constraints of the relationship, the three can be fused to obtain a representation containing the context semantics of the relationship, called the instance representation.
[0094] (2) Given K instances in the support set , we use the embedded vector at the "[CLS]" position as the embedded vector representing the sentence semantics, and use the average of the embedded vectors at the head entity and tail entity positions to obtain the embedded vectors and
[0095] (3) Then, we can use and to fuse the three to obtain the representation of the instance
[0096] (4) Next, we use the attention mechanism module between the initial embedding of the sentence and the instance representation to obtain the instance representation
[0097]
[0098] (5) Based on all the enhanced instances in the support set S, we calculate the original prototype i of the relationship category r by averaging the K instance representations i belonging to the relationship category r
[0099]
[0100] (6)Next, explicitly introduce relational information to guide the generation of enhanced prototypes. After the above processing, the embedded representation form r of the relational semantic information R is obtained. i and the initial prototypes that contain implicit relational semantics . By combining the two, enhanced relational prototypes can be obtained.
[0101]
[0102] (7)Based on the finally generated prototypes, we adopt an attention mechanism to obtain the enhanced representation of the query set. Given the representation Q j of the instances in the query set, an enhanced query representation is derived using each relational prototype , which focuses on relation-specific semantics.
[0103]
[0104] (8)Similarity comparison: Relational similarity comparison divides the instances in the query set into corresponding relations. In the metric space of the prototype network, the magnitude of similarity is usually used to calculate the similarity between an instance and a prototype. The greater the similarity, the more similar the instance is to the prototype. Similarity measurement methods include dot distance, cosine distance, Euclidean distance, etc. In this paper, Euclidean distance measurement is used to calculate the similarity between prototypes and instances:
[0105]
[0106] where represents the similarity matrix, d represents the dimension of the Bert hidden layer, that is, the embedding dimension of the representation. After obtaining the similarity matrix , the corresponding instance is judged to belong to which relation according to the maximum similarity, and the predicted relation pred is obtained.
[0107] III. Entity Location Recognizer
[0108] The task of the entity location recognizer is to locate the position of each term in the sentence and distinguish its label. Since the entities in the actual input sentence samples do not have entity type labels, a sequence labeling strategy is applied to the entity location detection task. And this model is type-independent, so it can adapt to entity location detection tasks in most domains. Based on the above considerations, this embodiment applies the MAML algorithm to the entity location detection task. In this way, the meta-learning algorithm endows the model with strong generalization ability and sampling efficiency by quickly finding meta-initialization parameters that can adapt to the new domain. This part is mainly divided into two modules: the basic sequence prediction module and the meta-learning module, as Figure 4 shown.
[0109] (1) Basic Sequence Prediction Module
[0110] As described above, in this embodiment, the entity position detection problem is transformed into a sequence labeling problem. Considering that the entities in the triple may contain multiple characters, based on the traditional BIO sequence labeling strategy, a sequence labeling pattern of B1, I1, O, B2, I2 is adopted, where B1 and I1 respectively represent the starting position and internal position of the head entity; B2 and I2 respectively represent the starting position and internal position of the tail entity; O represents other types of characters. For the model in the basic sequence prediction module, this embodiment selects the BertForTokenClassification model as the basic model for the sequence prediction task. The specific sequence prediction steps are as follows:
[0111] 1) Given the input sequence x = {w1, w2,..., w n}, we uniformly refer to the encoded context representations of the sentences in the support set S and the query set Q as h = h i ; i = 1, 2,..., L, where L represents the length of the sentence. Then we use a linear classification layer to determine the distribution probability of w i :
[0112] p(w i ) = softmax(Wh i + b) (6)
[0113] where p(w i ) ∈ R C , C = B1, I1, O, B2, I2 represents the probability that the character w i is the corresponding sequence label.
[0114] 2) For the calculation of the loss value, this part mainly uses the cross-entropy loss. Since the average loss of all tokens weakens the learning process of the token with the highest loss, more attention is placed on learning from the high-loss tokens, which may be corrected during meta-training.
[0115]
[0116] where Θ = {θ, W, b}, θ represents the parameters in the encoder, W and b represent the trainable parameters, and λ is the weight value.
[0117] 3) Viterbi decoding: To maintain the dependency between labels, the Viterbi algorithm is used for decoding. It should be noted that the transition matrix used here has nothing to do with training. Its existence is only to add constraints between labels so that the decoded sequence labels do not violate the BIO sequence labeling pattern.
[0118] (2) Meta-learning module
[0119] The purpose of meta - learning is to obtain a better set of model initialization parameters (i.e., let the model learn the initialization itself). Through N - way, K - shot tasks (training tasks) for meta - learning training, the model learns "prior knowledge", that is, the initialization parameters. The initialization parameters enable our subsequent training on the basic model to be better. Next, this embodiment will briefly introduce the content of the meta - learning part, which is divided into two steps: meta - training and meta - testing.
[0120] (1) Meta - training: In the meta - training stage of meta - learning, the model learns a good initial parameter by training on multiple tasks, enabling it to quickly adapt and achieve good performance when encountering new tasks with only a small amount of data and a few gradient update steps.
[0121] In the specific implementation process, random sampling is performed on the initial dataset to obtain and Then we first perform an inner loop on to update the parameters:
[0122]
[0123] where α represents the learning rate on the support set, and n represents the number of times the parameter is updated n times on . Only the parameters on are updated during this process, while the parameters Θ of the model are evaluated with the goal of minimizing the performance L on . This process is called the outer loop, which is expressed as:
[0124]
[0125] where β is the learning rate on the query set.
[0126] (2) Meta - testing: In this stage, we will first resample the initial dataset to obtain and Then use the inner loop to fine - tune , and then evaluate and predict based on the fine - tuned model
[0127] IV. Experimental Results and Analysis
[0128] (1) Dataset Selection
[0129] According to the settings of the previous work, we adopt the FewRel dataset used in the previous work. This dataset contains 70,000 instances in 100 relations, which is more than any other carefully labeled dataset of the same kind before. In addition, the name and description of each relation are provided, giving additional explanations for each relation. According to the settings of other relation triple extraction experiments, 50 relations are selected from this for training in this embodiment, with 15 of them as the validation set and the remaining 15 as the test set.
[0130] (2) Experimental settings
[0131] Based on the previous experiments, this embodiment sets the experiments to 5-way 5-shot and 10-way 10-shot for evaluation. During the evaluation process, the evaluation will be carried out on three dimensions: entities, relations, and triples, and the main evaluation metric is the F1 score. This embodiment selects the case-insensitive Bert as the encoder, with a hidden layer size of 768; the Adam optimizer is used to train the relation extractor, where the learning rate of the Bert model is 1e-5, and the learning rate of other network layers is 1e-4, with a weight decay of 0.01. The AdamW optimizer is used to train the entity detector, where the learning rates of both the inner loop and the outer loop are 3e-5, with a weight decay of 0.01. We use linear warm-up to change the learning rate, where the first 10% of the steps are warm-up steps. The batch size under all settings is 4.
[0132] (3) Baseline models
[0133] To verify the performance and efficiency of the model, the model is compared with several standard baselines and several few-shot baselines. First, the triple extraction is evaluated by dividing it into two steps: relation extraction and entity span detection. The following models are used as baseline models to compare with the triple extraction effect of this model:
[0134] MPE: A multi-prototype embedding network model for jointly extracting relation triples.
[0135] NNM: Fusing the prototype network and the nearest neighbor matching to extract relation triples.
[0136] Proto: A standard few-shot prototype network model.
[0137] PTN: A novel perspective transfer network (PTN) model.
[0138] CasRel: A cascade binary tagging framework (CasRel) model.
[0139] MLMAN: A multi-level matching and aggregation network for few-shot relation classification.
[0140] StructShot: A simple few-shot named entity recognition system based on nearest neighbor learning and structured reasoning.
[0141] RelATE: A new few-shot relation triple task decomposition strategy - relation first then entity.
[0142] MG-FTE: A mutual guidance-based few-shot learning framework for relation triple extraction.
[0143] PM-FTE: The few-shot meta-learning triple extraction model proposed in this embodiment.
[0144] (4) Experimental results
[0145] The experimental results on the Fewrel dataset are shown in Table 1 and Table 2. It can be observed that the model of this embodiment is always superior to the baseline and shows sufficient stability in all evaluation settings.
[0146] Table 1 - Experimental results of relation extraction and entity location recognition
[0147]
[0148] Comparison of experimental results of relation extraction and entity location recognition: Table 1 shows the experimental results of the model of this embodiment on the relation extraction and entity location detection tasks. It can be observed that the model is superior to these baseline models in few-shot relation extraction and entity location detection. In particular, compared with MPE, PM-FTE has made some improvements (performance improvements of 0.65% and 2.3% respectively in two cases of relation extraction), which shows the superiority of explicitly introducing relation semantic information in the prototype representation. The performance improvement in entity location recognition is more obvious, with improvements of 4.6% and 6.0% respectively compared with PTN in the experiment, further verifying the effectiveness of using MAML to enable the model to quickly adapt to new domains.
[0149] Table 2 - Experimental results of relation extraction and triple extraction
[0150]
[0151]
[0152] Comparison of experimental results of relation extraction and triple extraction: Table 2 shows the experimental results on the relation extraction and triple extraction tasks.
[0153] 1) Comparison with the standard supervised model: Obviously, all few-shot learning-based models outperform Bert and CasRel in two different scenarios. This effectively proves that traditional supervised learning methods cannot solve the few-shot relation triple extraction problem.
[0154] 2) Comparison with few-shot models: The PM-FTE proposed in this embodiment outperforms most existing methods in two scenarios respectively. In all N-way K-shot settings, compared with the relevant ones, the accuracy of PM-FTE is improved by at least 6.93%. However, compared with the state-of-the-art MG-FTE model in the few-shot relation triple extraction task, although PM-FTE is not as good as MG-FTE in the triple extraction task (F1 value is -5.9%), it is better than MG-FTE in the few-shot relation extraction task (F1 value is +5.28%).
[0155] Example 4
[0156] This embodiment is based on Embodiment 1:
[0157] This embodiment provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the few-shot triple extraction method of Embodiment 1. Among them, the computer program can be in the form of source code, object code, executable file or some intermediate form, etc.
[0158] Example 5
[0159] This embodiment is based on Embodiment 1:
[0160] This embodiment provides a computer-readable storage medium, storing a computer program, and when the computer program is executed by a processor, it implements the few-shot triple extraction method of Embodiment 1. Among them, the computer program can be in the form of source code, object code, executable file or some intermediate form, etc. The storage medium includes: any entity or device capable of carrying computer program code, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the storage medium does not include electrical carrier signals and telecommunication signals.
[0161] It should be noted that, for the foregoing method embodiments, for the sake of simplicity of description, they are expressed as a series of action combinations. However, those skilled in the art should be aware that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
Claims
1. A method for extracting triples from a small number of samples, characterized in that: include: Embedded vector representation: obtain support set and query set samples through the basic data set, and use the pre-trained language model to generate embedded vector representations for each sentence in the support set and query set samples; Relation extraction and classification: Integrate the relational semantic information between entities through the attention mechanism to form enhanced relational semantic information; The initial prototype is obtained by averaging the embedding vector representations of the support set samples, and the enhanced prototype is constructed by combining the relational semantic information; Compare the similarity between the sentences in the query set and the enhanced prototypes to extract and classify relations; Entity position recognition: Based on the sequence labeling strategy and model-independent meta-learning algorithm, the position of each word in the sentence is determined and the corresponding type label is assigned, and a complete triple is formed based on the identified head entity, relation, and tail entity combination.
2. The method for extracting triples of a small number of samples according to claim 1, characterized in that: The method of obtaining support set and query set samples through the basic data set and generating an embedded vector representation of each sentence in the support set and query set samples using a pre-trained language model includes: Randomly obtain samples from the basic data set to build the support set and query set, and obtain the relationship name and relationship description information based on all the relationship categories in the basic data set; The support set, query set and relationship description information are input into a language model for encoding to obtain an embedded vector representation of each sentence. The language model includes a Bert model and a Roberta model.
3. The method for extracting triples from a small number of samples according to claim 1, characterized in that: The relational semantic information between entities is integrated through the attention mechanism to form enhanced relational semantic information; The initial prototype is obtained by averaging the embedding vector representations of the support set samples, and the enhanced prototype is constructed by combining the relational semantic information, including: The semantic information corresponding to the entity relationship is integrated through the attention mechanism to detect the classification of the relationship, and the embedding vector representation in the query set is averaged to obtain the initial prototype. An initial prototype corresponds to a relationship category and represents a representative vector representation under this relationship category. The relational semantic information is explicitly fused into the initial prototype to obtain a more representative enhanced prototype.
4. The method for extracting triples from a small number of samples according to claim 1, characterized in that: The sequence labeling strategy and model-independent meta-learning algorithm determine the position of each word in the sentence and assign a corresponding type label, and form a complete triple based on the identified head entity, relationship and tail entity combination, including: The entity position detection problem is transformed into a sequence labeling problem. Based on the improved BIO sequence labeling strategy, cross entropy loss calculation and Viterbi decoding, the position of each word in the sentence is located and the type label is distinguished, thereby locating the entity position. Based on the model-independent meta-learning algorithm, the initialization parameters that can adapt to the new domain are found according to a few samples, so that the entity location recognition can be completed without knowing the entity type, and the final triples can be generated.
5. A few-sample triple extraction system, characterized in that: include: The encoder is configured to obtain support set and query set samples through the basic data set, and use the pre-trained language model to generate an embedding vector representation of each sentence in the support set and query set samples; The relation extractor is configured to integrate the relational semantic information between entities through an attention mechanism to form enhanced relational semantic information; The initial prototype is obtained by averaging the embedding vector representations of the support set samples, and the enhanced prototype is constructed by combining the relational semantic information; Compare the similarity between the sentences in the query set and the enhanced prototypes to extract and classify relations; The entity position identifier, configured as a meta-learning algorithm based on a sequence labeling strategy and a model-independent method, determines the position of each word in a sentence and assigns a corresponding type label, and forms a complete triple based on the identified head entity, relation, and tail entity combination.
6. The few-sample triple extraction system according to claim 5, characterized in that: In the encoder, support set and query set samples are obtained through the basic data set, and the embedded vector representation of each sentence in the support set and query set samples is generated using the pre-trained language model, including: Randomly obtain samples from the basic data set to build the support set and query set, and obtain the relationship name and relationship description information based on all the relationship categories in the basic data set; The support set, query set and relationship description information are input into a language model for encoding to obtain an embedded vector representation of each sentence. The language model includes a Bert model and a Roberta model.
7. The few-sample triple extraction system according to claim 5, characterized in that: The relation extractor comprises: The attention mechanism prototype generation module is configured to integrate the semantic information corresponding to the entity relationship through the attention mechanism to detect the classification of the relationship, and average the embedding vector representations in the query set to obtain the initial prototype. An initial prototype corresponds to a relationship category and represents a representative vector representation under this relationship category. a relational semantic information fusion module, configured to explicitly fuse relational semantic information into the initial prototype to obtain a more representative enhanced prototype; The similarity comparison module is configured to compare the similarity between the sentence to be extracted and the enhanced prototype to complete the relationship extraction and classification.
8. The few-sample triple extraction system according to claim 5, characterized in that: The entity location identifier comprises: The basic sequence prediction module is configured to transform the entity position detection problem into a sequence labeling problem. Based on the improved BIO sequence labeling strategy, cross entropy loss calculation and Viterbi decoding, it locates the position of each word in the sentence and distinguishes the type label, thereby locating the entity position; The meta-learning module is configured to find initialization parameters that can adapt to the new domain based on a few samples based on a model-independent meta-learning algorithm, so that it can complete entity location recognition without knowing the entity type and help generate the final triples.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the method for extracting a few-sample triples according to any one of claims 1 to 4 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for extracting a few-sample triples according to any one of claims 1 to 4 is implemented.
Citation Information
Cited By
Active element learning-based few-sample entity relationship extraction method and related equipment
CN120873200A
Data generation method based on small sample seeds and multi-round reinforcement and electronic equipment
CN121765062A