Entity relation joint extraction method based on multi-head collaborative matrix labeling and multi-feature fusion

By using multi-headed collaborative matrix annotation and multi-feature fusion methods in the Chinese text relationship extraction task, problems such as nested entities, overlapping entity relationships and Chinese entity ambiguity are solved, and more efficient and accurate entity relationship extraction is achieved.

CN119940355AInactive Publication Date: 2025-05-06KUNMING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510098984.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When dealing with the task of extracting Chinese text relationships, the prior art faces problems such as nested entities, overlapping entity relationships, Chinese entity ambiguity and error propagation, resulting in poor extraction performance.

Method used

A joint extraction method of entity relationships based on multi-headed collaborative matrix annotation and multi-feature fusion is adopted. A three-dimensional matrix is ​​constructed for annotation through predefined relationship types and entity type information, and combined with the multi-feature fusion attention mechanism, the entity relationship five-tuple in the text are extracted.

Benefits of technology

It effectively solves the problems of nested entities, overlapping entity relationships and Chinese entity ambiguity, improves the model's ability to extract complex semantic relationships, reduces error propagation, and significantly improves the accuracy and efficiency of entity relationship extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940355A_ABST
    Figure CN119940355A_ABST
Patent Text Reader

Abstract

The invention discloses an entity relationship joint extraction method based on multi-head collaborative matrix labeling and multi-feature fusion. The method comprises the following steps of: extracting quintuple of a head entity, a head entity type, a relationship, a tail entity and a tail entity type from a text; a three-dimensional matrix and an entity-entity type matrix EN * N are constructed and used for clearly marking entity type information corresponding to head and tail entities of the relation; extracting a structured entity relationship quintuple from the unstructured Chinese text, and processing entity ambiguity and entity relationship overlap in the Chinese text; and obtaining and analyzing an experiment result. When executing a Chinese entity relationship extraction task, the model can synchronously identify and extract the type labels corresponding to the entities in the process of extracting the relationship triples, so that the relationship between overlapped entities is effectively limited and distinguished, and the problem of vocabulary ambiguity existing in the Chinese entity relationship extraction task is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of entity relationship extraction tasks, and in particular relates to an entity relationship joint extraction method based on multi-head collaborative matrix annotation and multi-feature fusion. Background Art

[0002] The goal of entity relationship extraction is to identify and extract entity information from unstructured text and reveal the relationship between entities. By extracting entities and their interrelationships, originally isolated entities can be connected through various relationships to build a more complete semantic network. This not only helps to structure the information in the text, but also provides basic support for tasks such as automatic generation of knowledge graphs, information retrieval, and natural language understanding, further promoting the implementation of intelligent applications. In previous research work, entity relationship extraction tasks are mostly carried out in a pipeline manner, that is, the entity relationship extraction task is divided into named entity recognition and relationship classification. Usually, the entity is recognized first, and then the relationship is classified based on the extracted entities. The two are independent of each other. This not only ignores the correlation between entity recognition and relationship prediction, but is also easily affected by the error propagation problem. The entity relationship extraction method based on joint learning regards entity recognition and relationship extraction as a whole task. Through joint modeling, shared feature representation and model structure are used, so that entity recognition and relationship extraction can influence and optimize each other, which can better capture the contextual dependency between entities and relationships and the relationship interaction between multiple entities, thereby improving the extraction performance of the model. In recent years, many unique and innovative joint extraction methods have been proposed, bringing rich solutions to entity relationship extraction tasks. These methods integrate cutting-edge technologies and algorithms, significantly improve the accuracy and efficiency of extraction, and provide more powerful support for processing complex text data. However, most of the current joint extraction methods still face the following major challenges when dealing with Chinese text relationship extraction tasks: (1) Nested entity problem: refers to the situation in which one entity is contained within another entity in the text or multiple entities appear in a hierarchical or nested form. For example, "Shanghai Institute of Materia Medica, Chinese Academy of Sciences" is an entity of type "institution", which contains another entity of type "institution", "Chinese Academy of Sciences", and an entity of type "city", "Shanghai".

[0003] (2) Overlapping entity relationship problem: refers to the situation where an entity in a text may have different relationships with multiple other entities, or multiple relationships share the same entity. According to the degree of entity overlap in the relationship, entity relationship triples can be divided into three types: normal, single entity overlap (SEO), and entity pair overlap (EPO). If there are no overlapping entities in all entity relationship triples in a sentence, the sentence belongs to the normal type; if an entity in a sentence appears in multiple entity relationship triples, the sentence belongs to the single entity overlap type; if there are multiple different relationships between the same entity pair in a sentence, the sentence belongs to the entity pair overlap type.

[0004] (3) Chinese entity ambiguity: In Chinese texts, the semantic association between words is often not intuitive. Because many Chinese characters have multiple meanings, the same word is likely to carry completely different semantic connotations in different contexts. When the same entity expresses different meanings due to contextual differences, the types of relationships it builds with other entities will also change significantly accordingly. Take the entity "apple" as an example. When it semantically refers to the concept of fruit, the types of relationships it may participate in are mainly limited to categories related to commodity transactions and attributes, such as "purchase" and "price"; when "apple" refers to the business entity Apple, the types of relationships it involves are transformed into categories related to corporate creation and management structure, such as "establishment" and "chairman". This semantic ambiguity caused by entity polysemy greatly increases the difficulty of accurately identifying entity relationships in Chinese texts, and poses a severe challenge to the task of entity relationship extraction in Chinese natural language processing.

[0005] (4) Error propagation problem: This means that in a multi-step processing process, if an error occurs in one step, this error will affect the results of the subsequent steps, and this impact may be amplified, resulting in a decrease in the overall accuracy of the final result. For example, the CasRel joint extraction model mainly processes the entity-relationship joint extraction task through two stages: subject identification and relationship classification and object identification. The model first identifies all possible subjects from the sentence, then classifies the relationship based on the identified subject, and identifies the related objects. Any error in subject identification will directly lead to errors in relationship extraction. Therefore, the more errors in subject identification, the more serious the decrease in the accuracy of relationship extraction.

[0006] Therefore, it is urgent to design an entity relationship joint extraction method based on multi-head collaborative matrix labeling and multi-feature fusion to solve the above-mentioned problems. Summary of the invention

[0007] The purpose of the present invention is to provide a method for joint extraction of entity relationships based on multi-head collaborative matrix annotation and multi-feature fusion, which has the advantage of integrating predefined relationship type information and entity type information into the feature vector of the text sequence, thereby mining deeper semantic information of the text and improving the ability to extract complex semantic relationships, thereby solving the problems mentioned in the background technology.

[0008] To achieve the above objectives, the specific technical solution of the entity relationship joint extraction method based on multi-head collaborative matrix annotation and multi-feature fusion of the present invention is as follows: The entity relationship joint extraction method based on multi-head collaborative matrix annotation and multi-feature fusion includes the following steps: Extract the five-tuple of head entity, head entity type, relationship, tail entity and tail entity type from the text; Construct a three-dimensional matrix, the entity-entity type matrix E N×N , used to clearly mark the entity type information corresponding to the head and tail entities of the relationship; Extract structured entity relationship quintuples from unstructured Chinese text and handle entity polysemy and entity relationship overlap in Chinese text; Obtain experimental results and analyze them.

[0009] Furthermore, the five-tuple extraction of the head entity, the head entity type, the relationship, the tail entity and the tail entity type is implemented for the text, including the following steps: Facing a line of length Text sequence input, for each predefined relationship type Both generate two two-dimensional tables: entity-relationship matrix and the entity-entity type matrix ; In these two-dimensional matrices, each element represents a certain type of relationship The lower coordinates are Character label information of Refers to the first The word and The coordinates of the word pairs composed of words; The two predefined labels correspond to the matrix and matrix The purpose of the entity-relationship matrix is ​​to accurately extract all possible entity-relationship triplets in the Chinese text sequence, including nested entities and overlapping entity relationships in the text; Four types of labels are predetermined: -, HB-TB, HB-TE, and HE-TE, which are used to fully represent the position information of the head and tail entities in the relation triple of each word pair in the matrix.

[0010] Further, - indicates that the word is not involved in any entity of the relation triple; HB-TB indicates that the word pair is Represents the first Chinese character position of the head entity The first Chinese character position of the and tail entity ; HB-TE means word pair It is the first Chinese character position of the head entity and the last Chinese character position of the tail entity in a certain relation triple; HE-TE indicates word pair It is the last Chinese character position of the head entity and the last Chinese character position of the tail entity in a certain relation triple.

[0011] Furthermore, structured entity relationship quintuples are extracted from unstructured Chinese text, and entity polysemy and entity relationship overlap in Chinese text are processed, including the following steps: Perform deep encoding operations on the input Chinese text sequence through a specific pre-trained model; At the same time, we innovatively construct a multi-feature fusion attention mechanism; The single-stage scorer is used to simultaneously complete the two key tasks of relationship classification and entity type recognition in a parallel computing manner, which greatly improves the model processing efficiency; A specifically designed loss function is introduced to carry out systematic training and optimization of the entire model.

[0012] Furthermore, a deep encoding operation is performed on the input Chinese text sequence through a specific pre-trained model, including the following steps: For a given Chinese text sequence X=[x1,x2,x3,⋯,x n ], where x i Represents the i-th word in the text, encodes it through the pre-trained model, and embeds the vector of the last hidden layer as the feature representation of each word in the Chinese text sentence: H = ERNIE3.0[x1,x2,x3,⋯,x n ]=[h1,h2,h3,⋯,h n ] where h i is the feature representation of the i-th character, h i ∈R d , d represents the hidden feature dimension of the i-th word.

[0013] Furthermore, a multi-feature fusion attention mechanism is innovatively constructed, including the following steps: One-hot encode the predefined entity type labels and relationship type labels; Each entity type or relationship type is converted into a corresponding high-dimensional embedding vector through the corresponding embedding layer; Through the appropriate linear function layer, these embedding vectors are mapped to the corresponding entity type feature vectors or relationship type feature vectors, and the summary formula is as follows: in, a feature vector representing the entity type, represents the relation type feature vector, and is the unified latent dimension of these feature vectors; represents the corresponding embedding matrix, represents the corresponding linear transformation matrix, |C| represents the set corresponding to the entity type or relationship type, is the dimension of the embedding vector; and is the one-hot encoding representation of the corresponding label; represents the corresponding bias vector.

[0014] Furthermore, the innovative construction of the multi-feature fusion attention mechanism also includes the following steps: By introducing a gating mechanism, the model can selectively adjust the strength of fusion; And in the feature fusion process, the model automatically decides how to use the text features Fusion entity type features and relationship type characteristics information; First of all and Calculate a gating weight and , to indicate that they are in the text features The importance of and Indicates that the text features that have been mapped to the same vector space and (or ) is concatenated to form a joint vector containing joint feature information; and are two learnable weight matrices used to and Perform a linear transformation to remap the concatenated high-dimensional vector into a smaller dimensional space; and Represents the relevant bias term, providing an additional degree of flexibility for linear transformation, helping the model to better fit the data. At the same time, in order to make the gating weight a selectively controlled coefficient, use The function will and Limited to between 0 and 1; Secondly, the multi-head attention mechanism is combined to learn different attention modes from multiple subspaces and calculate the corresponding attention weights. The specific formula is as follows: in, and Indicates The result of the attention heads calculating the entity type features and the relationship type features respectively; Indicates The computation of an attention head; After calculating the results of each attention head, we next concatenate the outputs of all attention heads to obtain the final attention weights: in, Represents a splicing operation, Represents the number of attention heads; Finally, the concatenated attention weights are mapped to the corresponding vector space through a linear layer for final feature fusion. in, and Representing text features After the gated weights are adjusted and combined with the attention weights, and The feature representation of Represents the corresponding attention weight after mapping through the linear layer; Indicates fusion and The final feature representation of The basic information and the information adjusted by the gating mechanism and The feature information is combined together.

[0015] Furthermore, a single-stage scorer is used to simultaneously complete the two key tasks of relation classification and entity type recognition in a parallel computing manner, including the following steps: Through the corresponding linear transformation layer, the entity pairs in the relation triples are mapped to the space of relation scores and entity scores and projected to higher dimensions to enhance the expressiveness of the model and capture more complex relations. Among them, it is assumed that there is an entity pair vector and , respectively representing the The head entity at position The representation of the tail entity at the position; is the entity pair representation vector after projection; It is the projection matrix that is responsible for the corresponding dimensional transformation; represents the relationship score matrix, is the number of types of relations, is the size of the label (tag-size) corresponding to the relationship classification matrix; is the entity score matrix, is the number of entity types, is the number of label types corresponding to the entity recognition matrix; and is a linear mapping matrix that represents the high-dimensional representation of entity pairs Mapped into the corresponding score matrix; Represents the bias vector corresponding to the linear transformation; Then, a specific classifier is used to assign corresponding high-confidence labels to the two score matrices, indicating the existence of a specific relationship and the entity type corresponding to the head and tail entities under the relationship. The classifier uses all relationship representations and all entity type information to calculate each character pair in the two score matrices. The final scoring function is defined as: in, is the corresponding score vector; Indicates the use of dropout to prevent overfitting; After scoring, the relationship score vector and entity type score vector Input the softmax function to predict the corresponding label of the labeled entity pair and the given relationship or entity type. Through the softmax function, the model can determine whether the triple has a certain relationship, as well as the entity type information corresponding to the head and tail entities in the relationship triple; Softmax of the relationship score vector: Through the softmax function, the model converts the score into a normalized probability, indicating the label pair Have a relationship Probability ; Softmax of entity type score vector: Indicates the tag pair The entity type corresponding to the entity represented is Probability .

[0016] Furthermore, a specially designed loss function is introduced to carry out systematic training and optimization of the entire model, including the following steps: For the relationship classification part, a three-dimensional matrix is ​​used to represent the relationship classification results, and a multi-label classification cross entropy loss function is adopted: Where N represents the length of the text sequence; R represents the number of predefined relationship categories; Represents the first and The characters in The one-hot encoding representation corresponding to the real label under the class relationship type; is the probability representation of the 𝑟th type of relationship represented by the 𝑖th and 𝑗th characters predicted by the model; For the entity recognition part, a three-dimensional matrix is ​​also used to represent the recognition results of the entity type, and a multi-label classification cross entropy loss function is used: The loss calculation of the entity recognition part is generally the same as that of the relationship classification part. Represents the first and The entity consisting of characters is The real label one-hot encoding representation corresponding to the entity type under the class relationship type; is the probability representation of the entity composed of the 𝑖th and 𝑗th characters predicted by the model corresponding to the entity type of the 𝑟th relationship; Finally, the two losses are added together using specific weights to obtain the joint loss function: in, and are the weight parameters of entity recognition and relation classification losses, respectively, by minimizing the joint loss to optimize the model parameters.

[0017] Further, obtaining and analyzing the experimental results includes the following steps: Datasets and evaluation metrics; Experimental environment and parameter settings; Model comparison and result analysis; Comparison and analysis of experimental results on complex overlapping entity relationships; Experimental comparison and result analysis of Chinese models; Comparison and analysis of experimental results of various pre-training models.

[0018] The present invention has the following advantages: (1) The present invention designs a multi-head collaborative matrix annotation scheme, which enables the model to simultaneously identify and extract the type labels corresponding to the entities in the process of extracting relationship triples when performing the Chinese entity relationship extraction task, thereby effectively defining and distinguishing the relationship between overlapping entities, and further solving the lexical ambiguity problem in the Chinese entity relationship extraction task. In addition, unlike the traditional BIO annotation method, this three-dimensional matrix annotation method can accurately mark the start and end positions of complex nested entities in the text, and allows the entity boundaries of different relationship triples to be independently annotated, avoiding conflicts in the annotation process.

[0019] (2) The joint entity relationship extraction task is converted into a sequence labeling task in matrix form, and the whole process does not contain any interdependent steps. The Mengzi pre-trained large model is used to encode Chinese text, which achieves efficient representation and processing of text semantic information. Through a unified single-stage model framework, dual multi-label classification is performed on each character in the Chinese text sequence, thereby achieving a close interaction between entities, entity types, and entity pairs. This method naturally solves the problems of error accumulation and overlapping Chinese entity relationships.

[0020] (3) The present invention proposes a multi-feature fusion attention mechanism. This mechanism combines the gated weight unit and the multi-head attention mechanism to adaptively filter and fuse predefined entity type and relationship type feature information in the text feature vector, thereby more comprehensively capturing the semantic features of the text and improving the model's ability to model complex entity relationships. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 It is a schematic diagram of the multi-head cooperative matrix of the present invention; Figure 2 It is a schematic diagram of the overall framework of the model of the present invention; Figure 3 This is a schematic diagram of the collaborative task of the multi-head attention unit and the gated recurrent unit of the present invention; Figure 4 Schematic diagram of the task of selecting entity types and relationship types for the attention mechanism of the present invention; Figure 5 This is a schematic diagram of the comparison of experimental results of the model of the present invention; Figure 6 This is a schematic diagram of the comparison of F1 values ​​of the model of the present invention under different overlap types; Figure 7 A comparison diagram of model effects under different numbers of triples contained in the text of the present invention; DETAILED DESCRIPTION

[0022] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0023] Those skilled in the art will appreciate that, although some embodiments herein include certain features included in other embodiments but not other features, the combination of features of different embodiments is meant to be within the scope of the present invention and form different embodiments. For example, in the claims, any one of the claimed embodiments may be used in any combination.

[0024] Please refer to the attached Figure 1 To Attachment Figure 7 The present invention describes the entity relationship joint extraction method based on multi-head collaborative matrix annotation and multi-feature fusion.

[0025] Extract the five-tuple of head entity, head entity type, relationship, tail entity and tail entity type from the text; specifically, when faced with a length of Text sequence input, for each predefined relationship type Both generate two two-dimensional tables: entity-relationship matrix and the entity-entity type matrix ; In these two-dimensional matrices, each element represents a certain type of relationship The lower coordinates are Character label information of Refers to the first The word and The coordinates of the word pairs composed of words; The two predefined labels correspond to the matrix and matrix The purpose of the entity-relationship matrix is ​​to accurately extract all possible entity-relationship triplets in the Chinese text sequence, including nested entities and overlapping entity relationships in the text; Four types of labels are predetermined: -, HB-TB, HB-TE, and HE-TE, which are used to fully represent the position information of the head and tail entities in the relation triple of each word pair in the matrix.

[0026] Further, - indicates that the word is not involved in any entity of the relation triple; HB-TB indicates that the word pair is Represents the first Chinese character position of the head entity The first Chinese character position of the and tail entity ; HB-TE means word pair It is the first Chinese character position of the head entity and the last Chinese character position of the tail entity in a certain relation triple; HE-TE indicates word pair It is the last Chinese character position of the head entity and the last Chinese character position of the tail entity in a certain relation triple.

[0027] The quintuple pattern can provide additional contextual information, creating favorable conditions for the model to more accurately infer the relationship between entities. The type information of the head entity and the tail entity, as a priori constraint, effectively avoids unreasonable relationship combinations.

[0028] Taking the text "Li Lei is studying for a master's degree at Peking University" as an example, when the model recognizes that the type of "Peking University" is "organization" rather than "place", it can help it correctly predict the "studying" relationship between "Peking University" and "Li Lei" (name). In downstream tasks such as knowledge graph construction, quintuples can embed entity types directly into knowledge graphs, thereby achieving a more nuanced description of knowledge points. For example, in the field of medicine, it is crucial to clearly distinguish the types of "drugs" and "diseases" in the construction of the "drug-treatment-disease" relationship, which helps avoid erroneous inferences and provides strong support for more accurate question and answer. When fusing data from different knowledge graphs, the rich contextual information carried by the quintuple can effectively avoid data inconsistencies caused by type conflicts.

[0029] In summary, quintuple extraction has greatly expanded the representation capabilities based on traditional triples. By combining the auxiliary information of entity types, the accuracy and information completeness of relationship extraction are significantly improved. In downstream tasks such as knowledge graphs, it can not only significantly enhance the accuracy and completeness of data, but also provide support for more complex task scenarios, and effectively promote the performance optimization and model innovation of downstream tasks. In practical application scenarios, quintuple extraction has laid a solid foundation for achieving more refined and semantically rich knowledge representation and reasoning. Therefore, in-depth research on the quintuple joint extraction model is of great significance to promoting the design of complex task solutions, and can contribute important theoretical and practical value to the field of natural language processing.

[0030] In order to mark the entity types of the relation triples and their head and tail entities in the matrix at the same time, corresponding labels need to be designed. These predefined labels should not only indicate the position of the entity to which the character belongs in the relation triple, but also contain the entity type information corresponding to the character. In the currently selected dataset, there are 27 entity type labels. These 27 entity types are combined in no order, and there are 351 combinations of the relationship types between each entity type and other entity types. If the order information of the head and tail entities is taken into account, the number of combinations will reach 702, and it is obviously impossible to set so many label types.

[0031] In order to solve the above problem, a three-dimensional matrix, the entity-entity type matrix E N×N , used to clearly mark the entity type information corresponding to the head and tail entities of the relationship; like Figure 1 As shown, Indicates that the first to third characters in the text sequence constitute an entity, and the entity type is "song", and its corresponding one-hot encoding label is "11". Finally, the two matrices are connected through the position information of the head and tail entities, and predictions are performed in parallel, avoiding the problem of redundant entities and ensuring that the entity type order of the head and tail entities is correct. For example, if in the entity pair-relation annotation matrix, a certain relation triple = HB-TB, = HB-TE, = HE-TE. Then the entity position information in the text sequence can be transferred to the entity-entity type annotation matrix and the corresponding label of the entity can be annotated = 11, = 1. 11 is the label corresponding to the entity type "song", and 1 is the label corresponding to "person".

[0032] Extract structured entity relationship quintuples from unstructured Chinese text and handle entity polysemy and entity relationship overlap in Chinese text. The specific steps include: Perform deep encoding operations on the input Chinese text sequence through a specific pre-trained model; The pre-trained model is the Mengzi large model, a pre-trained model focused on Chinese natural language processing, which is specially trained for Chinese grammar, vocabulary structure, and the differences between written and spoken expressions. In addition, the Mengzi large model also introduces a hierarchical training architecture, which models different levels of information (such as vocabulary, syntax, and semantic layers) in stages, so that the model can better capture language features at different levels. When facing problems such as long sentences or complex sentences that are common in Chinese government texts, Mengzi adopts a multi-task learning framework, and simultaneously learns multiple language tasks (such as sentence sorting, sentence completion, logical relationships between sentences, etc.) during the pre-training process, so that the model can better understand syntactic structure and semantic associations. After fine-tuning on the Chinese government data set, the final pre-trained model selected uses a 12-layer Transformer encoder, each layer containing 12 self-attention heads. This structure can help the model effectively capture semantic and structural information of different granularities in large-scale Chinese corpora, thereby learning more accurate word vectors and semantic relationships.

[0033] The deep encoding operation of the input Chinese text sequence is performed through a specific pre-trained model, including the following steps: For a given Chinese text sequence X=[x1,x2,x3,⋯,x n ], where x i Represents the i-th word in the text, encodes it through the pre-trained model, and embeds the vector of the last hidden layer as the feature representation of each word in the Chinese text sentence: H = ERNIE3.0[x1,x2,x3,⋯,x n ]=[h1,h2,h3,⋯,h n ] where h i is the feature representation of the i-th character, h i ∈R d , d represents the hidden feature dimension of the i-th word.

[0034] At the same time, we innovatively construct a multi-feature fusion attention mechanism; The model proposed in this application involves multiple multi-label classification tasks. In order to ensure that the model can flexibly adjust its attention to specific entity types and relationship type information, the present invention introduces a multi-feature fusion attention mechanism. Before the text feature vector passes through the encoder and enters the scoring stage, the model adaptively enhances the entity type and relationship type features in the Chinese text. This enhancement mechanism can improve the model's ability to capture rich semantic features in Chinese text, and then more accurately identify and extract the relationship between entities, especially when processing complex or ambiguous texts, significantly improving the accuracy of extraction. At the same time, this method avoids repeated calculations for independently obtaining similar features for each task, reducing redundant computing overhead.

[0035] Innovatively construct a multi-feature fusion attention mechanism, including the following steps: One-hot encode the predefined entity type labels and relationship type labels; Each entity type or relationship type is converted into a corresponding high-dimensional embedding vector through the corresponding embedding layer; Through the appropriate linear function layer, these embedding vectors are mapped to the corresponding entity type feature vectors or relationship type feature vectors, and the summary formula is as follows: in, a feature vector representing the entity type, represents the relation type feature vector, and is the unified latent dimension of these feature vectors; represents the corresponding embedding matrix, represents the corresponding linear transformation matrix, |C| represents the set corresponding to the entity type or relationship type, is the dimension of the embedding vector; and is the one-hot encoding representation of the corresponding label; represents the corresponding bias vector.

[0036] In order to integrate the feature information of entity type and relationship type into the text feature vector and adaptively determine their fusion ratio, the present invention designs an attention mechanism. This mechanism introduces a gating mechanism to enable the model to selectively adjust the fusion strength and automatically determine how to fuse entity type features into text features h during feature fusion. and relationship type characteristics In addition, the multi-head attention mechanism allows the model to learn different attention modes from multiple subspaces, thereby enhancing the fusion effect of feature information. During the training phase, the attention mechanism selects entity types and relationship types for the corresponding text based on the principle shown in the following figure: First of all and Calculate a gating weight and , to indicate that they are in the text features The importance of and Indicates that the text features that have been mapped to the same vector space and (or ) is concatenated to form a joint vector containing joint feature information; and are two learnable weight matrices used to and Perform a linear transformation to remap the concatenated high-dimensional vector into a smaller dimensional space; and Represents the relevant bias term, providing an additional degree of flexibility for linear transformation, helping the model to better fit the data. At the same time, in order to make the gating weight a selectively controlled coefficient, use The function will and Limited to between 0 and 1; Secondly, the multi-head attention mechanism is combined to learn different attention modes from multiple subspaces and calculate the corresponding attention weights. The specific formula is as follows: in, and Indicates The result of the attention heads calculating the entity type features and the relationship type features respectively; Indicates The computation of an attention head; After calculating the results of each attention head, we next concatenate the outputs of all attention heads to obtain the final attention weights: in, Represents a splicing operation, Represents the number of attention heads; Finally, the concatenated attention weights are mapped to the corresponding vector space through a linear layer for final feature fusion. in, and Representing text features After the gated weights are adjusted and combined with the attention weights, and The feature representation of Represents the corresponding attention weight after mapping through the linear layer; Indicates fusion and The final feature representation of The basic information and the information adjusted by the gating mechanism and The feature information is combined together.

[0037] The single-stage scorer is used to simultaneously complete the two key tasks of relationship classification and entity type recognition in a parallel computing manner, which greatly improves the model processing efficiency; Traditional entity relationship extraction models usually rely on a single classifier or rule, and usually need to process different tasks in stages, such as identifying entity pairs first and then classifying relationships, or identifying head entities first and then matching corresponding relationships and tail entities. These methods often lead to information loss and error accumulation, and it is difficult to effectively deal with entity overlap problems. The model framework proposed in the present invention uniformly transforms entity recognition and relationship classification tasks into multi-label classification tasks in matrix form, and introduces a single-stage classifier based on scoring. The classifier characterizes the relationship between entity pairs and their corresponding entity type labels in the form of scoring, allowing the model to simultaneously process multiple pairs of entities and their relationships in a unified framework and optimize them, thereby solving the problems of information loss and error accumulation in traditional methods.

[0038] Instead of directly giving a definite predicted label, the scoring mechanism calculates a score representing the possibility of different relationship types for each entity pair, and also calculates the scores of the corresponding entity type labels for the head and tail entities under each relationship type. The final relationship category and the entity type labels of the head and tail entities under the relationship type are all decided based on these scores.

[0039] Using a single-stage scorer, the two key tasks of relation classification and entity type identification are completed simultaneously in a parallel computing manner, including the following steps: Through the corresponding linear transformation layer, the entity pairs in the relation triples are mapped to the space of relation scores and entity scores and projected to higher dimensions to enhance the expressiveness of the model and capture more complex relations. Among them, it is assumed that there is an entity pair vector and , respectively representing the The head entity at position The representation of the tail entity at the position; is the entity pair representation vector after projection; It is the projection matrix that is responsible for the corresponding dimensional transformation; represents the relationship score matrix, is the number of types of relations, is the size of the label (tag-size) corresponding to the relationship classification matrix; is the entity score matrix, is the number of entity types, is the number of label types corresponding to the entity recognition matrix; and is a linear mapping matrix that represents the high-dimensional representation of entity pairs Mapped into the corresponding score matrix; Represents the bias vector corresponding to the linear transformation; Then, a specific classifier is used to assign corresponding high-confidence labels to the two score matrices, indicating the existence of a specific relationship and the entity type corresponding to the head and tail entities under the relationship. The classifier uses all relationship representations and all entity type information to calculate each character pair in the two score matrices. The final scoring function is defined as: in, is the corresponding score vector; Indicates the use of dropout to prevent overfitting; After scoring, the relationship score vector and entity type score vector Input the softmax function to predict the corresponding label of the labeled entity pair and the given relationship or entity type. Through the softmax function, the model can determine whether the triple has a certain relationship, as well as the entity type information corresponding to the head and tail entities in the relationship triple; Softmax of the relationship score vector: Through the softmax function, the model converts the score into a normalized probability, indicating the label pair Have a relationship Probability ; Softmax of entity type score vector: Indicates the tag pair The entity type corresponding to the entity represented is Probability .

[0040] Introduce a specially designed loss function to systematically train and optimize the entire model; The main goal of the entity relationship joint extraction model proposed in the present invention is to identify all possible entity relationship triplets in Chinese text sentences, as well as the corresponding entity types of the head and tail entities in the triples, and the format is (head entity, head entity type, relationship, tail entity type, tail entity). In order to improve the training effect of the model, the present invention uses the cross entropy loss function of multi-label classification to train the relationship classification three-dimensional matrix and the entity recognition three-dimensional matrix respectively. Subsequently, the losses of the two three-dimensional matrices are added according to specific weights to obtain the final joint loss. By minimizing the joint loss, the model parameters are optimized and the model performance is improved.

[0041] Furthermore, a specially designed loss function is introduced to carry out systematic training and optimization of the entire model, including the following steps: For the relationship classification part, a three-dimensional matrix is ​​used to represent the relationship classification results, and a multi-label classification cross entropy loss function is adopted: Where N represents the length of the text sequence; R represents the number of predefined relationship categories; Represents the first and The characters in The one-hot encoding representation corresponding to the real label under the class relationship type; is the probability representation of the 𝑟th type of relationship represented by the 𝑖th and 𝑗th characters predicted by the model; For the entity recognition part, a three-dimensional matrix is ​​also used to represent the recognition results of the entity type, and a multi-label classification cross entropy loss function is used: The loss calculation of the entity recognition part is generally the same as that of the relationship classification part. Represents the first and The entity consisting of characters is The real label one-hot encoding representation corresponding to the entity type under the class relationship type; is the probability representation of the entity composed of the 𝑖th and 𝑗th characters predicted by the model corresponding to the entity type of the 𝑟th relationship; Finally, the two losses are added together using specific weights to obtain the joint loss function: in, and are the weight parameters of entity recognition and relation classification losses, respectively, by minimizing the joint loss to optimize the model parameters.

[0042] Obtaining experimental results and analyzing them includes the following steps: Datasets and evaluation metrics; This paper uses DuIE2.0, a relatively authoritative Chinese public dataset in the field of entity relationship extraction, as an experimental dataset to verify the performance of the model. DuIE2.0 is a Chinese public dataset for relationship extraction launched by Baidu, and its texts mainly come from Baidu Encyclopedia, Baidu Tieba, and Baidu Information Stream. Compared with the earlier version DuIE1.0, DuIE2.0 has expanded on the relationship types, adding more fine-grained and complex relationship types, covering multiple fields such as people, organizations, places, events, etc., making it suitable for a wider range of practical application scenarios.

[0043] Among them, TP (True Positives) indicates the number of samples correctly predicted as positive by the model; FP (False Positives) indicates the number of samples incorrectly predicted as positive by the model; FN (False Negatives) indicates the number of samples incorrectly predicted as negative by the model. Precision measures the proportion of samples predicted correctly among those predicted as positive by the model; Recall reflects the proportion of samples correctly predicted by the model among all samples that are actually positive.

[0044] DuIE2.0 has further improved the annotation quality. The annotation process is more rigorous, ensuring the accuracy of entities and relations, and more manual review and correction are performed for complex relations to reduce noise and annotation errors in the data. The DuIE2.0 dataset provides about 244,000 Chinese text samples, including about 420,000 triples, and predefines 27 entity types and 49 relationship types. Among them, the training set contains 171,135 samples, the validation set has 20,652 samples, and the test set contains 30,126 samples. Considering that the entity-relationship triples in the dataset have different entity-relationship overlaps, and in order to verify that this model has better performance in dealing with relationship overlap problems, this paper divides the text in the dataset into three categories: normal, single entity overlap (SEO), and entity pair overlap (EPO). Since the relation triples in some text sentences in the dataset belong to both EPO type and SEO type, the total number of Normal type, EPO type and SEO type sentences will be slightly larger than the total amount of the corresponding dataset. The relevant statistics of the above dataset division are shown in the following table: When evaluating model performance, the entity relationship quintuple predicted by the model is considered accurate only when the order of entity pairs, relations, head and tail entities, and their corresponding entity types are correct. To measure the prediction effect of the model, this paper uses F1 score (F1 Value, F1), Recall Rate (Rec.) and Precision (Prec.) as evaluation indicators.

[0045] Experimental environment and parameter settings; The experiment in this paper is based on the Windows 11 Professional 64-bit operating system, and the hardware environment includes the 12th GenIntel(R) Core(TM) i7-12700 2.10 GHz processor and the NVIDIA GeForce RTX 3090 graphics card. The experiment uses the PyTorch deep learning framework and uses the Mengzi large model as the Chinese pre-trained language model. In order to avoid overfitting of the model, the training process automatically stops when the performance of the validation set does not improve significantly after more than 15 epochs. In terms of training parameter settings, the training batch size is set to 16, the test batch is 4, the built-in hidden dimension is 312, the random number seed is set to 2024, and the learning rate is 2e-5. The multi-head attention mechanism in the model contains 12 layers and 6 heads, and the total training epochs is set to 100. This paper uses the adaptive moment estimation (Adam) optimizer for model optimization. The main parameter settings in this paper are the results of multiple preview experiments, as shown in the following table: Model comparison and result analysis; In order to evaluate the performance of the model proposed in this paper in the task of extracting entity relations from Chinese text, this paper selected a relatively advanced model framework in recent years for comparative experiments. Due to the lack of authoritative open source models for Chinese joint extraction, this experiment selected five baseline models for English datasets, and made them suitable for Chinese datasets through detailed adjustments such as data format conversion, while retaining the structure of the main part of the model unchanged. Finally, a comparative experiment was conducted on the Chinese public dataset DuIE2.0 to verify the effectiveness of the model. The results of the comparative experiments of each model are shown in the following table: Model Precision Recall F1 Score T / Epoch(min) Casrel 73.75 71.64 72.68 59.1 RIFRE 74.86 71.59 73.19 60.2 OneRel 75.62 73.17 74.37 54.6 PNDec 72.24 69.05 70.64 64.9 RSAN 76.92 70.13 73.37 65.1 Ours 84.15 83.47 83.81 55.8 Experimental results show that the proposed model performs significantly better than all baseline models in the Chinese entity relationship extraction task. Specifically, on the DuIE2.0 dataset, the model's precision, recall, and F1 value in extracting entity relationship quintuples are 7.23% to 11.91%, 10.3% to 14.42%, and 9.44% to 13.17% higher than those of the baseline models in extracting entity relationship triples. This fully proves that the proposed model can not only extract more entity relationship information, but also has a significantly higher accuracy than the baseline model, demonstrating its excellent performance in the Chinese entity relationship extraction task.

[0046] Comparison and analysis of experimental results on complex overlapping entity relationships; To further verify the effectiveness of our model in extracting overlapping triples from Chinese text, we divide the DuIE2.0 dataset into three categories according to entity relationship overlap: Normal, EPO, and SEO, and test the performance of each baseline model in different categories. Figure 5 As shown in the figure, it is obvious that the performance of our model is better than the baseline model under various Chinese overlapping relationship types. Therefore, it can be considered that our model can effectively solve the overlapping triples problem in entity relationship extraction.

[0047] In addition, this paper further evaluates the ability of the proposed model to extract multiple triplets from a single Chinese text sentence. The specific operation is to divide the text in the DuIE2.0 dataset into five categories according to the number of triplets contained, namely, the number of triplets in a single text sentence is N=1, N=2, N=3, N=4 and N≥5. The experimental results are shown in Table 4. The proposed model shows good results under all quantity thresholds. This shows that the proposed model can still maintain high performance when processing different numbers of triplets and has strong generalization ability.

[0048] Experimental comparison and result analysis of Chinese models; Since the model proposed in this paper is specially designed for the characteristics of Chinese text, and considering the significant differences between Chinese and English texts, it is difficult to fully and effectively prove the superiority of the proposed model in the task of extracting entity relations from Chinese text by simply using an open source entity relation extraction model for English on a Chinese public dataset. However, although many entity relation extraction models for Chinese text have emerged in recent years, it is difficult to conduct model experimental comparisons because most of the models are not open source and difficult to fully reproduce.

[0049] It is worth noting that in the existing research on entity relationship extraction of Chinese text, in addition to some difficult-to-obtain self-built domain datasets, many researchers usually choose DuIE1.0, a Chinese public dataset, for experiments. Therefore, this paper chooses to conduct experiments on the DuIE1.0 dataset and compares the experimental results with other entity relationship extraction models also based on the DuIE1.0 dataset to verify the effectiveness and superiority of this model in Chinese text tasks. Model Precision Recall F1 Score BSCRE 81.6 79.5 80.5 ECRE 78.54 83.48 80.94 NPCTS 79.6 82.7 81.1 RWG-LSA 83.98 80.96 82.44 RBEA 82.3 83.2 82.7 Ours 83.62 84.83 84.22 As can be seen from the table above, the model proposed in this paper performs best in Recall (84.83%) and F1 Score (84.22%) on the DuIE1.0 Chinese public dataset, which is 1.6% and 1.5% higher than the RBEA model proposed by Yao Feiyang et al., and 1.3% higher in Precision. Although the accuracy is slightly lower than the RWG-LSA model proposed by Zhang Li et al., with a difference of 0.36%, overall, the model proposed in this paper performs more balanced in the three indicators and can extract more and more comprehensive Chinese entity relationship information, showing its comprehensive advantages in entity relationship extraction tasks, and fully verifying the superior performance of the model in Chinese text tasks.

[0050] Comparison and analysis of experimental results of various pre-training models; This paper also conducted a comparative experiment on the effects of multiple pre-trained models, using ERNIE3.0, RoBERTa and other pre-trained models to conduct comparative evaluation on the DuIE2.0 Chinese public dataset. The experimental results are shown in the following table. As can be seen from the table, the pre-trained model we selected is significantly better than other pre-trained models in terms of entity relationship extraction, and the F1 value under other pre-trained models is improved by 6.27%~11.6%. This result proves the effectiveness and advantages of the Mengzi pre-trained large model we adopted in processing large Chinese datasets, further supports the performance of the model in complex tasks, and ensures that the relationship between entities can be more accurately captured and extracted in a variety of contexts. Pre-Model F1 Score Mengzi(Ours) 83.81 ERNIE3.0 77.54 RoBERTa 75.38 R-BERT 74.73 BERT-based 72.21 Ablation experiment In order to verify the effectiveness of each module of the model proposed in this paper, an ablation experiment was conducted on the DuIE2.0 Chinese public dataset. By gradually removing or modifying certain parts of the model and observing the changes in model performance, the importance of each module can be understood. The experimental results are shown in the following table.

[0051] In the ablation experiment, w / o Ent-Gate and w / o Ent-Att respectively represent the comparative experimental results after removing the gated weight unit and multi-head attention unit of the entity type feature in the multi-feature fusion module. The results show that after removing these two modules, the F1 value of the model on the two tasks decreased by 2.02% and 0.82% respectively. In addition, w / o Rel-Gate and w / oRel-Att represent the comparative experimental results of removing the gated weight unit and multi-head attention unit of the relation type feature, and the F1 value decreased by 1.9% and 0.67% respectively. The ablation experiment results show that in the task of joint extraction of entity relations, it is very necessary to appropriately enhance the specific features of the extraction task, and the collaboration of the gated weight unit and the multi-head attention mechanism can effectively enhance the entity type and relation type feature information in the Chinese text features, thereby improving the model extraction effect. In addition, it can be seen from the experimental results that the gated weight unit plays a greater role in the feature interaction module. Since the feature expressions of entity types and relationship types in Chinese texts are usually more obscure, the gated weight unit plays the role of an "information filter" in the model. By introducing control signals to screen or suppress some features, the model can dynamically generate feature expressions suitable for the current task according to the characteristics of different inputs, retain and enhance important information, and suppress unimportant or irrelevant features, thereby adaptively adjusting the degree of attention to special features, making the final feature expression more targeted and improving the accuracy of information transmission. method Precision Recall F1 Score Ours 84.15 83.47 83.81 w / o Ent- Gate 82.23 81.35 81.79 w / o Ent- Att 83.5 82.49 82.99 w / o Rel-Gate 82.37 81.46 81.91 w / o Rel-Att 83.68 82.61 83.14 w / o Multi-head 81.74 80.61 81.17 w / o ALL 78.64 76.47 77.54 The present invention proposes a joint extraction model based on a multi-head collaborative matrix annotation method for the Chinese entity relationship extraction task. The model makes full use of entity type features by identifying the relationship between entity pairs and the type information of the head and tail entities in their triples in parallel, so as to effectively deal with the ambiguity and entity overlap problems in Chinese text when extracting relationship triples, and finally generates a five-tuple prediction result containing the head entity, the head entity type, the relationship, the tail entity type, and the tail entity. In addition, the model introduces a multi-feature fusion attention mechanism. By combining the multi-head attention mechanism and the gated recurrent unit collaborative task, the predefined entity type and relationship type features are adaptively filtered and fused into the text vector features encoded by the Mengzi pre-trained large model to mine the deep semantic information of Chinese text while reducing redundant calculations. Experimental results show that the performance of the model on Chinese public datasets is significantly better than that of existing methods, which fully verifies the effectiveness of its joint extraction framework and related designs. Future research will further explore more efficient feature fusion methods and applicability on larger-scale and multi-domain datasets to improve the robustness and versatility of the model in practical applications.

[0052] The present invention proposes a joint extraction model based on a multi-head collaborative matrix annotation method for the Chinese entity relationship extraction task. The model makes full use of entity type features by identifying the relationship between entity pairs and the type information of the head and tail entities in their triples in parallel, so as to effectively deal with the ambiguity and entity overlap problems in Chinese text when extracting relationship triples, and finally generates a five-tuple prediction result containing the head entity, the head entity type, the relationship, the tail entity type, and the tail entity. In addition, the model introduces a multi-feature fusion attention mechanism. By combining the multi-head attention mechanism and the gated recurrent unit collaborative task, the predefined entity type and relationship type features are adaptively filtered and fused into the text vector features encoded by the Mengzi pre-trained large model to mine the deep semantic information of Chinese text while reducing redundant calculations. Experimental results show that the performance of the model on Chinese public datasets is significantly better than that of existing methods, which fully verifies the effectiveness of its joint extraction framework and related designs. Future research will further explore more efficient feature fusion methods and applicability on larger-scale and multi-domain datasets to improve the robustness and versatility of the model in practical applications.

[0053] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the embodiments here. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the claims of the present invention.

Claims

1. An entity relationship joint extraction method based on multi-head collaborative matrix annotation and multi-feature fusion, characterized by: The following steps are involved: Extract the five-tuple of head entity, head entity type, relationship, tail entity and tail entity type from the text; Construct a three-dimensional matrix, the entity-entity type matrix E N×N , used to clearly mark the entity type information corresponding to the head and tail entities of the relationship; Extract structured entity relationship quintuples from unstructured Chinese text and handle entity polysemy and entity relationship overlap in Chinese text; Obtain experimental results and analyze them.

2. The entity relationship joint extraction method based on multi-head collaborative matrix annotation and multi-feature fusion according to claim 1 is characterized in that: Extracting the five-tuple of head entity, head entity type, relation, tail entity and tail entity type from the text includes the following steps: Facing a line of length Text sequence input, for each predefined relationship type Both generate two two-dimensional tables: entity-relationship matrix and the entity-entity type matrix ; In these two-dimensional matrices, each element represents a certain type of relationship The lower coordinates are Character label information of Refers to the first The word and The coordinates of the word pairs composed of words; The two predefined labels correspond to the matrix and matrix The purpose of the entity-relationship matrix is ​​to accurately extract all possible entity-relationship triplets in the Chinese text sequence, including nested entities and overlapping entity relationships in the text; Four types of labels are predetermined: -, HB-TB, HB-TE, and HE-TE, which are used to fully represent the position information of the head and tail entities in the relation triple of each word pair in the matrix.

3. The entity relationship joint extraction method based on multi-head collaborative matrix annotation and multi-feature fusion according to claim 2 is characterized in that: - indicates that the word is composed of any entity that does not participate in the relation triple; HB-TB indicates that the word pair is Represents the first Chinese character position of the head entity The first Chinese character position of the and tail entity ; HB-TE means word pair It is the first Chinese character position of the head entity and the last Chinese character position of the tail entity in a certain relation triple; HE-TE indicates word pair It is the last Chinese character position of the head entity and the last Chinese character position of the tail entity in a certain relation triple.

4. The entity relationship joint extraction method based on multi-head collaborative matrix annotation and multi-feature fusion according to claim 1 is characterized in that: Extracting structured entity relationship quintuples from unstructured Chinese text and handling entity polysemy and entity relationship overlap in Chinese text includes the following steps: Perform deep encoding operations on the input Chinese text sequence through a specific pre-trained model; At the same time, we innovatively construct a multi-feature fusion attention mechanism; The single-stage scorer is used to simultaneously complete the two key tasks of relationship classification and entity type recognition in a parallel computing manner, which greatly improves the model processing efficiency; A specifically designed loss function is introduced to carry out systematic training and optimization of the entire model.

5. The entity relationship joint extraction method based on multi-head collaborative matrix annotation and multi-feature fusion according to claim 4 is characterized in that: The input Chinese text sequence is deeply encoded through a specific pre-trained model, including the following steps: For a given Chinese text sequence X=[x1,x2,x3,⋯,x n ], where x i Represents the i-th word in the text, encodes it through the pre-trained model, and embeds the vector of the last hidden layer as the feature representation of each word in the Chinese text sentence: H=ERNIE3.0[x1,x2,x3,⋯,x n ]=[h1,h2,h3,⋯,h n ] where h i is the feature representation of the i-th character, h i ∈R d , d represents the hidden feature dimension of the i-th word.

6. The entity relationship joint extraction method based on multi-head collaborative matrix annotation and multi-feature fusion according to claim 4 is characterized in that: Innovatively construct a multi-feature fusion attention mechanism, including the following steps: One-hot encode the predefined entity type labels and relationship type labels; Each entity type or relationship type is converted into a corresponding high-dimensional embedding vector through the corresponding embedding layer; Through the appropriate linear function layer, these embedding vectors are mapped to the corresponding entity type feature vectors or relationship type feature vectors, and the summary formula is as follows: in, a feature vector representing the entity type, represents the relation type feature vector, and is the unified latent dimension of these feature vectors; represents the corresponding embedding matrix, represents the corresponding linear transformation matrix, |C| represents the set corresponding to the entity type or relationship type, is the dimension of the embedding vector; and is the one-hot encoding representation of the corresponding label; represents the corresponding bias vector.

7. The entity relationship joint extraction method based on multi-head collaborative matrix annotation and multi-feature fusion according to claim 6 is characterized in that: The innovative construction of multi-feature fusion attention mechanism also includes the following steps: By introducing a gating mechanism, the model can selectively adjust the strength of fusion; And in the feature fusion process, the model automatically decides how to use the text features Fusion entity type features and relationship type characteristics information; First of all and Calculate a gating weight and , to indicate that they are in the text features The importance of and Indicates that the text features that have been mapped to the same vector space and (or ) is concatenated to form a joint vector containing joint feature information; and are two learnable weight matrices used to and Perform a linear transformation to remap the concatenated high-dimensional vector into a smaller dimensional space; and Represents the relevant bias term, providing an additional degree of flexibility for linear transformation, helping the model to better fit the data. At the same time, in order to make the gating weight a selectively controlled coefficient, use The function will and Limited to between 0 and 1; Secondly, the multi-head attention mechanism is combined to learn different attention modes from multiple subspaces and calculate the corresponding attention weights. The specific formula is as follows: in, and Indicates The result of the attention heads calculating the entity type features and the relationship type features respectively; Indicates The computation of an attention head; After calculating the results of each attention head, we then concatenate the outputs of all attention heads to get the final attention weights: in, Represents a splicing operation, Represents the number of attention heads; Finally, the concatenated attention weights are mapped to the corresponding vector space through a linear layer for final feature fusion. in, and Representing text features After the gated weights are adjusted and combined with the attention weights, and The feature representation of Represents the corresponding attention weight after mapping through the linear layer; Indicates fusion and The final feature representation of The basic information and the information adjusted by the gating mechanism and The feature information is combined together.

8. The entity relationship joint extraction method based on multi-head collaborative matrix annotation and multi-feature fusion according to claim 4 is characterized in that: Using a single-stage scorer, the two key tasks of relation classification and entity type identification are completed simultaneously in a parallel computing manner, including the following steps: Through the corresponding linear transformation layer, the entity pairs in the relation triples are mapped to the space of relation scores and entity scores and projected to higher dimensions to enhance the expressiveness of the model and capture more complex relations. Among them, it is assumed that there is an entity pair vector and , respectively representing the The head entity at position The representation of the tail entity at the position; is the entity pair representation vector after projection; It is the projection matrix that is responsible for the corresponding dimensional transformation; represents the relationship score matrix, is the number of types of relations, is the size of the label (tag-size) corresponding to the relationship classification matrix; is the entity score matrix, is the number of entity types, is the number of label types corresponding to the entity recognition matrix; and is a linear mapping matrix that represents the high-dimensional representation of entity pairs Mapped into the corresponding score matrix; Represents the bias vector corresponding to the linear transformation; Then, a specific classifier is used to assign corresponding high-confidence labels to the two score matrices, indicating the existence of a specific relationship and the entity type corresponding to the head and tail entities under the relationship. The classifier uses all relationship representations and all entity type information to calculate each character pair in the two score matrices. The final scoring function is defined as: in, is the corresponding score vector; Indicates the use of dropout to prevent overfitting; After scoring, the relationship score vector and entity type score vector Input the softmax function to predict the corresponding label of the labeled entity pair and the given relationship or entity type. Through the softmax function, the model can determine whether the triple has a certain relationship, as well as the entity type information corresponding to the head and tail entities in the relationship triple; Softmax of the relationship score vector: Through the softmax function, the model converts the score into a normalized probability, indicating the label pair Have a relationship Probability ; Softmax of entity type score vector: Indicates the tag pair The entity type corresponding to the entity represented is Probability .

9. The entity relationship joint extraction method based on multi-head collaborative matrix annotation and multi-feature fusion according to claim 4 is characterized in that: Introducing a specially designed loss function, systematically training and optimizing the entire model includes the following steps: For the relationship classification part, a three-dimensional matrix is ​​used to represent the relationship classification results, and a multi-label classification cross entropy loss function is adopted: Where N represents the length of the text sequence; R represents the number of predefined relationship categories; Represents the first and The characters in The one-hot encoding representation corresponding to the real label under the class relationship type; is the probability representation of the 𝑟th type of relationship represented by the 𝑖th and 𝑗th characters predicted by the model; For the entity recognition part, a three-dimensional matrix is ​​also used to represent the recognition results of the entity type, and a multi-label classification cross entropy loss function is used: The loss calculation of the entity recognition part is generally the same as that of the relationship classification part. Represents the first and The entity consisting of characters is The real label one-hot encoding representation corresponding to the entity type under the class relationship type; is the probability representation of the entity composed of the 𝑖th and 𝑗th characters predicted by the model corresponding to the entity type of the 𝑟th relationship; Finally, the two losses are added together using specific weights to obtain the joint loss function: in, and are the weight parameters of entity recognition and relation classification losses, respectively, by minimizing the joint loss to optimize the model parameters.

10. The entity relationship joint extraction method based on multi-head collaborative matrix annotation and multi-feature fusion according to claim 1 is characterized in that: Obtaining experimental results and analyzing them includes the following steps: Datasets and evaluation metrics; Experimental environment and parameter settings; Model comparison and result analysis; Comparison and analysis of experimental results on complex overlapping entity relationships; Experimental comparison and result analysis of Chinese models; Comparison and analysis of experimental results of various pre-training models.

Citation Information

Patent Citations

  • Entity relationship recognition method and device

    CN111476023A

  • Knowledge extraction method and system of entity relationship combined triad

    CN116127921A