A Chinese Medical Entity Relationship Joint Extraction Method and System

By using the Transformer-XL encoder and TPLinker joint decoding framework in the Chinese medical entity relationship joint extraction method, combining the vocabulary enhancement and relational attention mechanism, the accuracy of entity recognition and relationship extraction in Chinese medical text is solved, and more efficient model convergence and professional vocabulary recognition are achieved.

CN114036934BActive Publication Date: 2025-05-27ZHEJIANG UNIV OF TECH

Patent Information

Application Number
CN202111203313.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-15
Publication Date
2025-05-27
Estimated Expiration
2041-10-15

AI Technical Summary

Technical Problem

The existing Chinese medical entity relationship joint extraction method has problems such as inaccurate entity recognition, poor relationship extraction performance, sparse decoding matrix, and redundancy in relationships when dealing with Chinese medical texts.

Method used

A Chinese medical entity relationship joint extraction method based on the Transformer-XL encoder and TPLinker joint decoding framework is adopted, and a vocabulary enhancement and relational attention mechanism is added to improve the recognition accuracy of entity types and boundaries, and solve the problems of sparse decoding matrix and redundancy in relation.

Benefits of technology

It improves the accuracy of entity recognition and relationship extraction in Chinese medical texts, solves the problems of entity nesting and relationship overlap, improves the convergence speed of the model, and alleviates the difficulties of professional vocabulary recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114036934B_ABST
    Figure CN114036934B_ABST
Patent Text Reader

Abstract

A method for jointly extracting Chinese medical entity relationships includes: a medical relationship embedding representation module, a head and tail position acquisition module for head and tail entities in medical texts, a word vector and relative distance calculation module for medical text words, a word vector output module after vocabulary enhancement, a relationship prediction module for medical texts, a character pair vector generation module for medical texts, a subject-predicate-object triple output module, a joint extraction model training module, an F1 score calculation module for the joint extraction model, a module for cyclically training the joint extraction model, and a medical text entity relationship acquisition module. The present invention also includes a system for jointly extracting Chinese medical entity relationships. The present invention solves the problems of entity nesting and relationship overlap in complex sentences in Chinese medical texts, alleviates the sparsity of the TPLinker decoding matrix, improves the convergence speed of the joint extraction model, and alleviates the problem that many professional terms in Chinese medical texts cannot be accurately recognized even in combination with the context through a vocabulary enhancement coding unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This patent relates to the field of natural language processing, and particularly to a method for jointly extracting Chinese medical entity relationships. Background Art

[0002] To construct a knowledge graph in the medical field, it is first necessary to obtain useful information such as entities, relationships, and attributes from a large amount of unstructured data (such as text), that is, information extraction. Entity and relationship extraction are two important subtasks in the information extraction task. According to the different orders of completion of the two subtasks, entity relationship extraction methods can be divided into two methods: pipeline extraction and joint extraction.

[0003] Pipeline extraction, that is, first extracting entities and then extracting relationships, is a relatively traditional extraction method. This extraction method will cause the following three problems: 1) Error accumulation: Errors in entity extraction will affect the accuracy of relationship extraction; 2) Entity redundancy: Pairing the extracted entities two by two and then performing relationship classification, if there is no relationship between entity pairs, redundant information will appear; 3) Interaction loss: The internal connection and dependency relationship between entity extraction and relationship extraction are not considered.

[0004] The joint extraction method compensates for the above three shortcomings to a certain extent. Joint extraction, that is, relational triple extraction (RTE), represents triples in the form of (head entity, relationship, tail entity). Joint extraction can be further divided into joint extraction based on parameter sharing and joint extraction based on joint decoding. The joint extraction model based on shared parameters only shares the parameters of the two models for entity and relationship extraction, such as hidden layer states, etc., and the interaction between the entity model and the relationship model is not strong. In 2017, Zheng et al. first proposed unifying the annotation of entities and relationships, and using the same decoder for both the entity model and the relationship model, that is, joint decoding. However, Zheng et al. directly used relationships as labels, resulting in an entity or a pair of entities not being able to have multiple relationships, that is, the problem of relationship overlap cannot be solved.

[0005] In 2020, the TPLinker joint extraction framework proposed by Yu et al. achieved SOTA in entity relation extraction. It not only solved the problem of relation overlap, but also solved problems such as entity nesting and exposure bias. However, the TPLinker framework still has some drawbacks. TPLinker is more suitable for English texts, and its extraction performance is poor for Chinese texts, especially Chinese medical texts. The Chinese BERT preprocessing model provided by Google can achieve context awareness and improve the effect of Chinese entity recognition to a certain extent. However, there are still many professional terms in Chinese medical texts that cannot be accurately recognized even in combination with the context. In addition, the decoder of the TPLinker framework is relatively complex, and there are problems such as sparse decoding matrix, slow convergence speed, and relation redundancy. Summary of the Invention

[0006] The present invention aims to overcome the above-mentioned drawbacks of the prior art and provides a method for jointly extracting Chinese medical entity relations.

[0007] For Chinese medical texts, based on the Transformer-XL encoder and the TPLinker joint decoding framework, the present invention adds a vocabulary enhancement and a relation attention mechanism. The medical professional vocabulary is introduced through vocabulary enhancement to facilitate the recognition of entity types and entity boundaries. At the same time, relation prediction is performed through the relation attention mechanism to solve the problems of sparse decoding matrix and relation redundancy, and improve the accuracy of entity recognition and relation extraction in Chinese medical texts.

[0008] The technical solution adopted by the present invention to solve its technical problems is as follows:

[0009] A method for jointly extracting Chinese medical entity relations includes the following steps:

[0010] Step 1: Prepare the Chinese medical text Text to be extracted for entity relations. According to the given ontology constraint set (including relation names, head entity types, and tail entity types), use the Chinese BERT model to represent each relation name as an embedding vector to obtain the semantic information of the relation, that is, the relation embedding C = {c 1 , c 2 ,..., c l}, where l is the total number of relations;

[0011] Step 2: Obtain the labeled Chinese medical information extraction data set Data (including the relation names, head entities, and tail entity names and types of each medical text), preprocess Data to obtain the head and tail positions of the head entities and tail entities in each medical text;

[0012] Step 3: Based on the Flat_Lattice structure, perform lexical enhancement on Text and Data. Calculate the 4 relative distances between any two character (or word) vectors of each medical text to represent the possible intersection, inclusion, or separation relationships between the character (or word) vectors, and obtain the character (or word) vectors and their relative distance matrices of each medical text. The specific process is as follows:

[0013] 3.1 For each medical text in Text and Data, use the Chinese BERT model to obtain their respective character vectors;

[0014] 3.2 Obtain the pre-trained Chinese biomedical word vectors. Match each medical text in Text and Data with the word list of the Chinese biomedical word vectors, identify the words with intersections with the word list for lexical enhancement, and obtain the word vectors of each medical text in Text and Data;

[0015] 3.3 Perform head and tail position encoding on the character vectors and word vectors of each medical text in Text and Data to obtain the start and end positions of the characters and words. Use the relative position encoding technology in Flat_Lattice to obtain 4 relative distances i between any two character (or word) vectors x j and x and and put them into the relative distance matrix:

[0016]

[0017] where head[i] and tail[i] represent the head and tail positions of the i-th character (or word) vector x i and head[j] and tail[j] represent the head and tail positions of the j-th character (or word) vector x j . represents the distance from the start position of x i to the start position of x j , represents the distance from the start position of x i to the end position of x j , represents the distance from the end position of x i to the start position of x j , represents the distance from the end position of x i to the end position of x j ;

[0018] Step 4: Take a batch of training datasets from Data, and input the word (or term) vectors Z and position encoding vectors R of its medical text into the Transformer-XL encoder, and output the word vectors H = {h 1 , h 2 , …, h n} after enhancing the medical text vocabulary, where n is the length of the medical text. The Transformer-XL encoder consists of two sub-layers: a self-attention layer and a feed-forward layer. After each sub-layer, a residual connection and layer normalization are connected. For any two word (or term) vectors x i and x j , the position encoding R ij between them is obtained by concatenating four relative distances and in the form of absolute position encoding and then passing through a fully connected layer with ReLU as the activation function:

[0019]

[0020] where W r is the parameter to be trained, and P d adopts absolute position encoding:

[0021]

[0022]

[0023] where d refers to and k is the dimension index inside the position encoding vector (k ∈ [0, (d model - 1) / 2]), and d model = H × d head (d head is the dimension of each head of the multi-head attention mechanism, with a total of H heads);

[0024] The self-attention mechanism based on the position encoding vector R is as follows:

[0025] Attention(A * , V) = Softmax(A * )V,

[0026]

[0027] [Q, K, V] = E x [W q , W k , W v ,

[0028] where W q , W k,Z , Wk,R , u, v, W k , W v are all parameters to be trained. The first two terms of A * are the semantic interaction and position interaction between two words (or terms) respectively, and the last two terms are the global content bias and global position bias;

[0029] Step 5: Predict the relationship based on the relationship embedding C and the medical text word vector H output by the Transformer-XL encoder to obtain a list of predicted relationships. The specific process includes self-attention mechanism, relationship attention mechanism, attention fusion mechanism and relationship prediction:

[0030] 5.1 Input the medical text word vector H into two fully connected layers to obtain the self-attention value A (s) , where the first fully connected layer uses the tanh activation function and the second fully connected layer uses the softmax activation function. Calculate the medical text representation M (s) according to A (s) :

[0031] A (s) = softmax(W 2 tanh(W 1 H)),

[0032] M (s) = A (s) H T ,

[0033] where W 1 and W 2 are parameters to be trained;

[0034] 5.2 Calculate the relationship attention value A (l) and the medical text representation M (l) based on the relationship attention mechanism according to the relationship embedding C and the medical text word vector H:

[0035] A (l) = CH,

[0036] M (l) = A (l) H T ;

[0037] 5.3 Through the attention fusion mechanism, input M (s) and M (l) into a fully connected layer using the sigmoid activation function respectively to obtain α and β, and constrain α and β by α + β = 1 to fuse and obtain M:

[0038] α = sigmoid(M (s) W 3 ),

[0039] β = sigmoid(M (l) W 4 ),

[0040] M = αM (s) + βM (l) ,

[0041] where W 3 and W 4 are parameters to be trained;

[0042] 5.4 Input M into two fully connected layers to obtain the predicted probability of the relationship label The first fully connected layer uses the ReLU activation function, and the second fully connected layer uses the sigmoid activation function:

[0043]

[0044] where, W 5 and W 6 are parameters to be trained. If is greater than the threshold 0.5, it is added to the predicted relationship list;

[0045] Step Six: Concatenate every two word vectors h i and h j output by the Transformer-XL encoder and perform a fully connected layer to obtain the character pair vector h ij :

[0046]

[0047] where the activation function used is tanh, and W h and b h are parameters to be trained;

[0048] Step Seven: Decode the subject-predicate-object triples through the TPLinker decoder that fuses specific relationship embeddings. Mark the head and tail characters of the entity with EH-to-ET, mark the head characters of the head and tail entities of the relationship with SH-to-OH, and mark the tail characters of the head and tail entities of the relationship with ST-to-OT. Among them, the EH-to-ET, SH-to-OH, and ST-to-OT decoders are implemented by an identical fully connected layer:

[0049]

[0050] where, represents the predicted value of the marked character pair h ij , k q represents the embedding of the q-th relationship, and W t , b tis the parameter to be trained, and the activation function used is softmax. The specific decoding process is as follows:

[0051] 7.1) Decode EH-to-ET to obtain all entities and their starting characters in the medical text;

[0052] 7.2) For each relationship in the predicted relationship list, decode ST-to-OT to obtain the ending character pairs of the head and tail entities, store the ending character pairs and the relationship in set O. At the same time, decode SH-to-OH to obtain the starting character pairs of the head and tail entities, match the starting character pairs with the starting characters of all entities, and find the head and tail entities corresponding to the starting character pairs and store them in set S;

[0053] 7.3) Determine whether the ending character pairs of each pair of head and tail entities in S are in O. If so, then determine the triple as (head entity, relationship, tail entity);

[0054] Step Eight: Calculate the total loss function L and perform joint training through the backpropagation algorithm to obtain the joint extraction model:

[0055] L = L rel + L tp ,

[0056]

[0057]

[0058] where L rel is the loss function for relationship prediction, the true value of the q-th relationship the predicted value of the q-th relationship L tp is the loss function after adding relationship prediction. E, H, and T represent EH-to-ET, SH-to-OH, and ST-to-OT respectively, represents the predicted value of the character pair h ij being marked, y ijq represents the true value of the character pair h ij being marked, represents the probability that the character pair h ij is marked as y ijq when decoding the q-th relationship, represents the number of predicted relationships, is the number of head and tail entity types corresponding to the predicted relationship found according to the given ontology constraint set, that is, the number of predicted entity types;

[0059] Step Nine: Take a batch of validation datasets from Data, input the word (or term) vectors of its medical text and their relative distance matrices into the joint extraction model, and calculate the F 1 score of the joint extraction model:

[0060]

[0061] where precision is the precision rate and recall is the recall rate;

[0062] Step ten: Repeat steps four to nine until the predefined F 1 score is exceeded, and save the joint extraction model;

[0063] Step eleven: Input the word (or term) vectors of each medical text word in Text after enhancement and their relative distance matrix into the joint extraction model to obtain entity-relationship triples.

[0064] The technical concept of the present invention is: to complete the joint extraction of Chinese medical entity relationships through vocabulary enhancement encoding, relationship prediction based on relationship attention mechanism, and TPLinker joint decoding framework that fuses specific relationship embeddings. The vocabulary enhancement encoding uses the Flat_Lattice structure and the self-attention mechanism based on relative position encoding proposed in Transformer-XL, integrating character and vocabulary information. The relationship prediction mainly adopts the relationship attention mechanism, combining the semantic information of medical texts and relationships to predict medical relationships. The TPLinker joint decoding represents the character vectors output by Transformer-XL as character pair vectors, fuses specific relationship embeddings, and obtains the head and tail characters of entities, that is, all entities, through EH-to-ET decoding. For each relationship in the predicted relationship list, the tail characters of all head and tail entities are obtained through ST-to-OT decoding, and the head characters of all head and tail entities are obtained through SH-to-OH decoding, thereby extracting the (head entity, relationship, tail entity) triples.

[0065] A method for jointly extracting Chinese medical entity relationships consists of three parts: a lexical enhancement encoding unit, a relationship prediction unit based on a relationship attention mechanism, and a TPLinker joint decoding unit. The lexical enhancement encoding unit uses the Flat_Lattice structure and the self-attention mechanism based on relative position encoding proposed in Transformer-XL, integrating character and professional vocabulary information, which is beneficial to the recognition of Chinese medical entities. The relationship prediction unit mainly adopts the relationship attention mechanism, combining the semantic information of medical texts and relationship labels to predict medical relationships. The TPLinker joint decoding unit represents the word vectors output by Transformer-XL as character pair vectors, integrates specific relationship embeddings, obtains the head and tail characters of entities through EH-to-ET decoding, and for each relationship in the relationship list obtained by the relationship prediction unit, obtains all the tail characters of the head and tail entities through ST-to-OT decoding, and obtains all the head characters of the head and tail entities through SH-to-OH decoding, thereby extracting (head entity, relationship, tail entity) triples. The present invention uses the TPLinker joint decoding unit to solve the problems of entity nesting and relationship overlap in complex Chinese medical texts, introduces relationship prediction based on the relationship attention mechanism and specific relationship embeddings to alleviate the sparsity of the TPLinker decoding matrix, improves the convergence speed of the joint extraction model, and alleviates the problem that many professional vocabulary in Chinese medical texts cannot be accurately recognized even in combination with context through the lexical enhancement encoding unit.

[0066] The present invention also includes a system for implementing the method for jointly extracting Chinese medical entity relationships of the present invention, including: a medical relationship embedding representation module, a head and tail position acquisition module for head and tail entities in medical texts, a medical text word vector and its relative distance calculation module, a word vector output module after lexical enhancement, a relationship prediction module for medical texts, a character pair vector generation module for medical texts, a subject-predicate-object triple output module, a joint extraction model training module, an F 1 score calculation module, a module for cyclically training the joint extraction model, and a medical text entity relationship acquisition module. The above modules respectively correspond to the contents of steps one to eleven of the method of the present invention in sequence.

[0067] The beneficial effects of the present invention are as follows: The present invention uses TPLinker joint decoding to solve the problems of entity nesting and relationship overlap in complex Chinese medical texts, that is, entity pair overlap and single entity overlap, adds relationship prediction based on the relationship attention mechanism, and only decodes the relationships in the predicted relationship list, alleviating the sparsity of the TPLinker decoding matrix and increasing the convergence speed of the model. Adding lexical enhancement in the encoding part is more conducive to the recognition of Chinese medical entities and alleviates the problem that many professional vocabulary in Chinese medical texts cannot be accurately recognized even in combination with context. Brief Description of the Drawings

[0068] Figure 1 is the algorithm block diagram of the present invention.

[0069] Figure 2 is the flow chart of the present invention. Detailed Embodiments

[0070] The present invention will be further described below with reference to the accompanying drawings.

[0071] Refer to Figure 1 and Figure 2 , taking the Chinese medical information consultation system and the Chinese medical information extraction dataset CMeIE as an example, applying the Chinese medical entity relationship joint extraction method based on lexical enhancement and relational attention mechanism of the present invention, a method for constructing a Chinese medical information consultation system is formed, including the following steps:

[0072] Step 1: Prepare the Chinese medical text Text to be extracted for entity relationships. According to the given ontology constraint set (including relationship names, head entity types, and tail entity types), such as the ontology constraint set of CMeIE, use the Chinese BERT model to represent each relationship name as an embedding vector to obtain the semantic information of the relationship, that is, the relationship embedding C = {c 1 , c 2 , …, c l}, where l is the total number of relationships;

[0073] Step 2: Obtain the labeled Chinese medical information extraction dataset CMeIE (including the relationship names, head entities, and tail entity names and types of each medical text, as shown in Table 2, "text" refers to the medical text, "predicate" refers to the relationship name, "subject" and "subject_type" respectively refer to the name and type of the head entity, "object" and "object_type" respectively refer to the name and type of the tail entity,) as Data, and preprocess Data to obtain the head and tail positions of the head entity and the tail entity in each medical text;

[0074] Table 2

[0075]

[0076] Table 2 shows the labeled Chinese medical information extraction data.

[0077] Step 3: Based on the Flat_Lattice structure, perform lexical enhancement on Text and Data. Calculate the 4 relative distances between any two character (or word) vectors of each medical text to represent the possible cross, inclusion, or separation relationships between the character (or word) vectors, and obtain the character (or word) vectors and their relative distance matrices for each medical text. The specific process is as follows:

[0078] 3.1 For each medical text in Text and Data, use the Chinese BERT model to obtain their respective character vectors;

[0079] 3.2 Obtain the pre-trained Chinese biomedical word vectors, such as the Chinese biomedical word vectors (Chinese-Word2vec-Medicine) containing 278,256 biomedical-related vocabulary words and with a dimension of 512 obtained by Word2Vec training. Match each medical text in Text and Data with the word list of the Chinese biomedical word vectors respectively, identify the words with intersections with the word list for lexical enhancement, and obtain the word vectors of each medical text in Text and Data;

[0080] 3.3 Perform head and tail position encoding on the character vectors and word vectors of each medical text in Text and Data to obtain the start and end positions of the characters and words. Use the relative position encoding technology in Flat_Lattice to obtain the 4 relative distances i between any two character (or word) vectors x j and and and put them into the relative distance matrix:

[0081]

[0082] where head[i] and tail[i] represent the head and tail positions of the i-th character (or word) vector x i and head[j] and tail[j] represent the head and tail positions of the j-th character (or word) vector x j . represents the distance from the start position of x i to the start position of x j , represents the distance from the start position of x i to the end position of x j , represents the distance from the end position of x i to the start position of x i , represents the distance from the end position of x i to the end position of x j ;

[0083] Step 4: Take a batch of training data sets from Data, and input the word (or term) vectors Z and position encoding vectors R of its medical text into the Transformer-XL encoder, and output the word vectors H = {h 1 , h 2 , …, h n} after enhancing the medical text vocabulary. Here, n is the length of the medical text. The Transformer-XL encoder consists of two sub-layers: a self-attention layer and a feed-forward layer. After each sub-layer, a residual connection and layer normalization are connected. For any two word (or term) vectors x i and x i , the position encoding R ij between them is obtained by concatenating four relative distances and in the form of absolute position encoding and then passing through a fully connected layer with ReLU as the activation function:

[0084]

[0085] where W r are the parameters to be trained, and P d adopts absolute position encoding:

[0086]

[0087] where d refers to and k is the dimensional index inside the position encoding vector (k ∈ [0, (d model - ) / 2]), and d model = H × d head (d head is the dimension of each head of the multi-head attention mechanism, with a total of H heads);

[0088] The self-attention mechanism based on the position encoding vector R is as follows:

[0089] Attention(A * , V) = Softmax(A * )V,

[0090]

[0091] [Q, K, V] = E x [W q , W k , W v ,

[0092] where W q , W k,Z , W k,R , u, v, W k,W v All are parameters to be trained, A * The first two terms of which are the semantic interaction and position interaction between two words (or terms), and the last two terms are the global content bias and global position bias;

[0093] Step 5: Predict the relationship based on the relationship embedding C and the medical text word vectors H output by the Transformer-XL encoder to obtain a list of predicted relationships. The specific process includes self-attention mechanism, relationship attention mechanism, attention fusion mechanism and relationship prediction:

[0094] 5.1 Input the medical text word vectors H into two fully connected layers to obtain the self-attention value A (s) , where the first fully connected layer uses an activation function, and the second fully connected layer uses the softmax activation function. Calculate the medical text representation M (s) : (s)

[0095] A (s) = softmax(W 2 tanh(W 1 H)),

[0096] M (s) = A (s) H T ,

[0097] where, W 1 and W 2 are parameters to be trained;

[0098] 5.2 Calculate the relationship attention value A (l) and the medical text representation M (l) based on the relationship attention mechanism according to C and H:

[0099] A (l) = CH,

[0100] M (l) = A (l) H T ;

[0101] 5.3 Through the attention fusion mechanism, input M (s) and M (l) into a fully connected layer using the sigmoid activation function respectively to obtain α and β, and constrain α and β by α + β = 1 to fuse and obtain M:

[0102] α = sigmoid(M (s) W 3 ),

[0103] β = sigmoid(M (l) W​4 )

[0104] M = αM (s) + βM (l) ,

[0105] where W 3 and W 4 are parameters to be trained;

[0106] 5.4 Input M into two fully connected layers to obtain the predicted probability of the relationship label The first fully connected layer uses the ReLU activation function, and the second fully connected layer uses the sigmoid activation function:

[0107]

[0108] where, W 5 and W 6 are parameters to be trained. If is greater than the threshold 0.5, it is added to the predicted relationship list;

[0109] Step Six: Concatenate every two word vectors h i and h j output by the Transformer-XL encoder and perform a fully connected layer to obtain the character pair vector h ij :

[0110]

[0111] where the activation function used is tanh, and W h and b h are parameters to be trained;

[0112] Step Seven: Decode through the TPLinker decoder that fuses specific relationship embeddings to obtain the subject-predicate-object triple. Use EH-to-ET to mark the head and tail characters of the entity, SH-to-OH to mark the head characters of the head and tail entities of the relationship, and ST-to-OT to mark the tail characters of the head and tail entities of the relationship. Among them, the EH-to-ET, SH-to-OH, and ST-to-OT decoders are implemented by the same fully connected layer:

[0113]

[0114] where, represents the predicted value of the character pair h ij being marked, k q represents the embedding of the q-th relationship, and W t , b t are parameters to be trained. The activation function used is softmax. The specific decoding process is as follows:

[0115] 7.1) Decode EH-to-ET to obtain all entities and their starting characters in the medical text;

[0116] 7.2) For each relationship in the predicted relationship list, decode ST-to-OT to obtain the ending character pairs of the head and tail entities, store the ending character pairs and the relationship in set O. At the same time, decode SH-to-OH to obtain the starting character pairs of the head and tail entities, match the starting character pairs with the starting characters of all entities, and find the head and tail entities corresponding to the starting character pairs and store them in set S;

[0117] 7.3) Determine whether the ending character pairs of each pair of head and tail entities in S are in O. If so, then determine the triple as (head entity, relationship, tail entity);

[0118] Step Eight: Calculate the total loss function L and perform joint training through the backpropagation algorithm to obtain the joint extraction model:

[0119] L = L rel + L tp ,

[0120]

[0121]

[0122] where L rel is the loss function for relationship prediction, the true value of the q-th relationship the predicted value of the q-th relationship L tp is the loss function after adding relationship prediction. E, H, and T represent EH-to-ET, SH-to-OH, and ST-to-OT respectively, represents the predicted value of the character pair h ij being marked, y ijq represents the true value of the character pair h ij being marked, represents the probability that the character pair h ij is marked as y ijq when decoding the q-th relationship, represents the number of predicted relationships, is the number of head and tail entity types corresponding to the predicted relationships found according to the given ontology constraint set, that is, the number of predicted entity types;

[0123] Step Nine: Take a batch of validation data sets from Data, input the word (or term) vectors of its medical text and their relative distance matrices into the joint extraction model, and calculate the F 1 score of the joint extraction model:

[0124]

[0125] Where precision is the accuracy rate and recall is the recall rate;

[0126] Step 10: Repeat steps 4 to 9 until the predetermined F is exceeded. 1 Score, such as the F score of the predetermined CMeIE validation dataset 1 The score can be set to 0.65 to save the joint extraction model;

[0127] Step 11: Input the enhanced character (or word) vectors and their relative distance matrices of each medical text vocabulary of Text into the joint extraction model to obtain entity relationship triples (as shown in Table 1), which are stored in the graph database Neo4j as the knowledge graph of the Chinese medical information consultation system.

[0128] Table 1

[0129]

[0130] Table 1 shows the triple diagram of normal relations and overlapping relations (SEO and EPO) in Chinese medical text

[0131] Step 12: Input the user's question into the Chinese medical information consultation system. After parsing and keyword matching the question, use Cypher's match to match and query the Chinese medical knowledge graph, assemble the answer based on the returned knowledge, and give the query result of the question.

[0132] The present invention also includes a system for implementing a Chinese medical entity relationship joint extraction method of the present invention, including: a medical relationship embedding representation module, a head and tail position acquisition module of a head entity and a tail entity in a medical text, a medical text word vector and its relative distance calculation module, a word vector output module after vocabulary enhancement, a medical text relationship prediction module, a medical text character pair vector generation module, a subject-predicate-object triple output module, a joint extraction model training module, and an F-based joint extraction model. 1 Score calculation module, cyclic training joint extraction model module, medical text entity relationship acquisition module. The above modules correspond to the contents of step 1 to step 11 of the method of the present invention in sequence.

[0133] As described above, the specific implementation steps of this patent implementation make the present invention clearer. Within the spirit of the present invention and the protection scope of the claims, any modifications and changes made to the present invention fall within the protection scope of the present invention.

Claims

1. A method for jointly extracting Chinese medical entity relationships, characterized in that: It includes the following steps: Step 1: Prepare the Chinese medical text Text for entity relation extraction. According to the given ontology constraint set, the ontology constraint set includes relation names, head entity types, and tail entity types. Use the Chinese BERT model to represent each relation name as an embedding vector to obtain the semantic information of the relation, denoted as relation embedding C = {c 1 , c 2 ,..., c l}, where l is the total number of relations; Step 2: Obtain the labeled Chinese medical information extraction dataset Data. The Chinese medical information extraction dataset Data includes the relationship names, the names and types of the head entities and tail entities of each medical text. Preprocess Data to obtain the head and tail positions of the head entities and tail entities in each medical text; Step 3: Based on the Flat_Lattice structure, perform lexical enhancement on Text and Data, calculate the 4 relative distances between any two characters or word vectors in each medical text of them, and obtain the character or word vectors and their relative distance matrices of each medical text. The specific process is as follows: 3.1) Use the Chinese BERT model for each medical text of Text and Data to obtain their respective character vectors; 3.2) Obtain the pre-trained Chinese biomedical word vectors, match each medical text of Text and Data with the vocabulary of the Chinese biomedical word vectors respectively, identify the words with intersections with the vocabulary for lexical enhancement, and obtain the word vectors of each medical text of Text and Data; 3.3) Encode the start and end positions of each character vector and word vector in the medical text of Text and Data to obtain the start and end positions of characters and words, and use the relative position encoding technology in Flat_Lattice to obtain the 4 relative distances between any two character or word vectors x i and x j and put them into the relative distance matrix, where and represents the distance from the start position of x to the start position of x i ; represents the distance from the start position of x to the end position of x i ; represents the distance from the end position of x to the start position of x i ; represents the distance from the end position of x to the end position of x i ; ; Step 4: Take a batch of training datasets from Data, and input the word or word vector Z and position encoding vector R of its medical text into the Transformer-XL encoder to obtain the word vector H = {h 1 , h 2 , …, h n} after enhancing the medical text vocabulary. Here, n is the length of the medical text. The Transformer-XL encoder consists of two sub-layers: a self-attention layer and a feed-forward layer. After each sub-layer, there are residual connections and layer normalization. The position encoding R i between any two word or word vectors x j and x ij is obtained by concatenating four relative distances and in the form of absolute position encoding and then passing through a fully connected layer with ReLU as the activation function: Among them, W r is the parameter to be trained, P d adopts absolute position encoding, and d refers to and The self-attention mechanism based on the position encoding vector R is as follows: Attention(A * , V) = Softmax(A * )V, [Q, K, V] = E x [W q , W k , W v , Among them, W q , W k,Z , W k,R , u, v, W k , W v are all parameters to be trained; Step 5: Predict the relationship according to the relationship embedding C and the medical text character vectors H output by the Transformer-XL encoder to obtain a list of predicted relationships. The specific process is as follows: 5.1 Input H into two fully connected layers to obtain the self-attention value A (s) , where the first fully connected layer uses the tanh activation function, and the second fully connected layer uses the softmax activation function. Based on A (s) calculate the medical text representation M based on the self-attention mechanism (s) : A (s) = softmax(W 2 tanh(W 1 H)), M (s) = A (s) H T , where W 1 and W 2 are parameters to be trained; 5.2 Calculate the relational attention value A based on C and H (l) and the medical text representation M based on the relational attention mechanism (l) : A (l) =CH, M (l) = A (l) H T ; 5.3 Through the attention fusion mechanism, input M (s) and M (l) into a fully connected layer using the sigmoid activation function respectively to obtain α and β, and constrain α and β by α + β = 1, and fuse to obtain M: α = sigmoid(M (s) W 3 ) β = Sigmoid(M (l) W 4 ) M = αM (s) + βM (l) , where W 3 and W 4 are parameters to be trained; 5.4 Input M into two fully connected layers to obtain the predicted probability of the relationship label The first fully connected layer uses the ReLU activation function, and the second fully connected layer uses the sigmoid activation function: Among them, W 5 and W 6 are parameters to be trained. If is greater than the threshold of 0.5, it is added to the prediction relationship list; Step 6: Concatenate every two word vectors h of the medical text output by the Transformer-XL encoder i and h j and then perform a fully connected operation to obtain the character pair vector h ij : where the activation function used is tanh, W h and b h are the parameters to be trained; Step 7: Decode through the TPLinker decoder that fuses specific relationship embeddings to obtain the subject-predicate-object triples. Mark the head and tail characters of the entities with EH-to-ET, mark the head characters of the head and tail entities of the relationship with SH-to-OH, and mark the tail characters of the head and tail entities of the relationship with ST-to-OT. Among them, the EH-to-ET, SH-to-OH, and ST-to-OT decoders are implemented by the same fully connected layer: Among them, represents the character pair h ij the marked predicted value, k q represents the embedding of the q-th relationship, W t , b t are the parameters to be trained. The activation function used is softmax. The specific process is as follows: 7.1) Decode EH-to-ET to obtain all entities and their head characters in the medical text; 7.2) For each relationship in the list of predicted relationships, decode ST-to-OT to obtain the tail character pairs of the head and tail entities, store the tail character pairs and the relationships in the set O, and at the same time decode SH-to-OH to obtain the head character pairs of the head and tail entities, match the head character pairs with the head characters of all entities, and find the head and tail entities corresponding to the head character pairs and store them in the set S; 7.3) Judge whether the tail character pairs of each pair of head and tail entities in S are in O. If so, then determine that the triple is the head entity, relationship, and tail entity; Step 8: Calculate the total loss function L and perform joint training through the backpropagation algorithm to obtain the joint extraction model: L = L rel + L tp , Among which L rel is the loss function for relation prediction, and the true value of the q-th relation the predicted value of the q-th relation L tp is the loss function after adding relation prediction. E, H, and T represent EH-to-ET, SH-to-OH, and ST-to-OT respectively, represents the predicted value of the character pair h ij being marked, y ijq represents the true value of the character pair h ij being marked, represents that when decoding the q-th relation, the character pair h ij is marked as y ijq with a probability of, represents the number of predicted relations, is the number of head and tail entity types corresponding to the predicted relations found according to the given ontology constraint set, and is the number of predicted entity types; Step 9: Extract the validation dataset from Data, input the word or word vector of its medical text and its relative distance matrix into the joint extraction model, and calculate the F of the joint extraction model 1 score: where precision is the precision rate and recall is the recall rate; Step Ten: Repeat Steps Four to Nine until the predefined F 1 score is exceeded, and save the joint extraction model; Step 11: Input the character or word vectors and their relative distance matrices after lexical enhancement of each medical text of Text into the joint extraction model to obtain the entity relationship triples.

2. A system for implementing the method for jointly extracting Chinese medical entity relationships described in claim 1, characterized in that it includes: Medical relationship embedding representation module, head and tail position acquisition module for head entity and tail entity in medical text, medical text word vector and its relative distance calculation module, word vector output module after vocabulary enhancement, relationship prediction module for medical text, character pair vector generation module for medical text, subject-verb-object triple output module, joint extraction model training module, F of joint extraction model 1 score calculation module, loop training joint extraction model module, medical text entity relationship acquisition module.

Citation Information

Patent Citations

  • Chinese entity relationship extraction method based on character and word feature fusion of entity meaning items

    CN111291556A

  • Medical entity relation joint extraction method

    CN112818676A

Cited By

  • Medical data retrieval method and system and storage medium

    CN122112235A