A medical entity recognition and relationship extraction method based on deep learning
By applying deep learning-based methods in medical text analysis, using word embedding models and Attention mechanisms for entity recognition and relationship extraction, and combining with medical knowledge base for verification, the limitations of medical text analysis in the existing technology are solved, and the accuracy and practicality of the analysis are significantly improved.
Patent Information
- Application Number
- CN202411393643.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-08
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2044-10-08
AI Technical Summary
The prior art has limitations in entity recognition and relationship extraction of medical texts, especially when facing complex medical texts and multi-sense medical terms, the recognition rate is low, the misclassification rate is high, and it is difficult to deal with multi-level semantic associations and complex relationships.
Using a deep learning-based method, we use the word embedding model and multi-layer neural network model in the medical field, combining the Attention mechanism and a two-way long and short-term memory network to perform entity recognition and relationship extraction, and use the medical field knowledge base to verify and supplement entities and relationships.
It significantly improves the accuracy and practicality of medical text analysis, can effectively identify entities and relationships in complex medical texts, reduce misclassification and misidentification, and enhances the performance of the model in complex scenarios.
Smart Images

Figure CN119167938B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical technology, and in particular to a medical entity recognition and relationship extraction method based on deep learning. Background Art
[0002] In the existing technology, entity recognition and relationship extraction of medical texts mainly rely on traditional natural language processing technology and rule-based methods. However, traditional methods usually identify specific medical entities such as disease names, drugs and symptoms from medical texts by defining fixed rules or using simple statistical models, and extract the relationships between entities. Traditional methods have significant limitations when facing complex medical texts. The professionalism and complexity of medical texts lead to difficulties in entity recognition. Many medical terms are ambiguous or context-dependent. Common NLP technologies are prone to errors when processing these terms. In addition, the extensive use of abbreviations and terminology changes in medical texts exacerbate the difficulty of recognition, resulting in low recognition rates and high misclassification rates.
[0003] In terms of relationship extraction, traditional rule-based methods can usually only handle simple entity relationships, such as the basic correspondence between "drug-disease treatment" or "disease-induced symptoms". Such methods are unable to handle multi-level semantic associations and complex relationships, and cannot accurately extract complex causal relationships and time sequence relationships. With the increase in the amount of medical data and the complexity of text structure, the limitations of traditional methods are becoming more and more obvious. In addition, existing technologies rely on fixed rules and cannot self-optimize and learn according to the dynamic changes of data, which limits the adaptability of the system in actual application scenarios.
[0004] In addition, although some existing technologies attempt to introduce machine learning algorithms to improve the performance of entity recognition and relationship extraction, machine learning algorithms usually rely on large-scale labeled data. However, labeled data in the medical field is difficult to obtain and the labeling cost is high, which makes it difficult to effectively promote the model in practical applications. In addition, existing machine learning models are usually difficult to deeply integrate with the knowledge base in the medical field, and cannot fully utilize the professional knowledge in the medical field for entity recognition and relationship extraction, which affects the accuracy and practicality of the overall model. Summary of the invention
[0005] One purpose of the present invention is to propose a medical entity recognition and relationship extraction method based on deep learning, which improves the reliability and practicality of medical text analysis.
[0006] A medical entity recognition and relationship extraction method based on deep learning according to an embodiment of the present invention comprises the following steps:
[0007] S1. Obtain and preprocess the medical text dataset, perform word segmentation, stop word removal and standardization on the medical text dataset;
[0008] S2. Build a word embedding model in the medical field, use the terminology dictionary and corpus in the medical field to train the word embedding model, so that the word embedding model generates a high-dimensional vector representation of medical terms;
[0009] S3. Based on the word embedding model, a multi-layer neural network model using the Attention mechanism is constructed and trained;
[0010] S4, applying the trained multi-layer neural network model to the new medical text dataset to identify and classify entities in the medical text dataset;
[0011] S5. After identifying the entities, the bidirectional Attention mechanism is used to extract the relationships between the entities in the medical text dataset. The relationship extraction includes identifying the direct relationship between entities and the potential relationship inferred through context perception.
[0012] S6. Use the bidirectional Attention mechanism to process different types of relationships separately.
[0013] S7. Verify and supplement the identified entities and relationships by combining with the medical domain knowledge base;
[0014] S8. Output the finally identified and extracted medical entities and their relationships in a structured form.
[0015] Optionally, the S1 specifically includes:
[0016] S11. Obtain medical text dataset:
[0017] D={d1,d2,...,d n};
[0018] Among them, d i represents the i-th medical text, and n is the total number of medical texts;
[0019] S12, use the word segmentation algorithm to perform word segmentation on the medical text dataset D, and separate each medical text d i Represented as a sequence of words:
[0020] W i ={w i1 ,w i2 ,...,w im};
[0021] Among them, w ij is the jth word in the i-th medical text, and m is the medical text d iThe total number of words in the
[0022] S13, remove stop words and convert each word sequence W i The stop word set S in is filtered out to obtain a new word sequence W′ i ;
[0023] S14, word sequence W′ after removing stop words i Perform standardization, convert uppercase and lowercase letters, convert all words to lowercase, and restore the words in the word sequence to the stem form W″ i ;
[0024] S15. The preprocessed medical text dataset is expressed as:
[0025] D′={W″1,W″2,...,W″ n};
[0026] Among them, W″ i is the i-th preprocessed stemmed word sequence.
[0027] Optionally, the S2 specifically includes:
[0028] S21. Obtain a dictionary of medical terms:
[0029] V={v1,v2,...,v t};
[0030] Among them, v i represents the i-th medical term, and t is the total number of terms in the term dictionary;
[0031] S22. Collecting corpus in the medical field:
[0032] C={c1,c2,...,c p};
[0033] where c j represents the text segment in the jth medical field, and p is the total number of text segments in the corpus;
[0034] S23, constructing a word embedding model in the medical field based on the medical field terminology dictionary V and the medical field corpus C, and defining the co-occurrence matrix M(i, j) of the word embedding model as:
[0035]
[0036] Among them, f(v i ,e kl ) represents the term v i In the kth text segment, kl The degree of correlation, g(v j ,ekl+1 ) represents the term v j With the adjacent entity e kl+1 The correlation degree, h(c k ) is the weight factor of the segment, representing the importance of the segment in the entire corpus, q k For the text fragment c k The number of entities in ;
[0037] S24. Use the context-aware model in the medical field to train the word vector of the co-occurrence matrix M(i,j), and define the word vector update as:
[0038]
[0039] in, represents the vector representation of the i-th term at the t-th iteration, η is the learning rate, λ is the regularization term, Word embedding values representing medical entities;
[0040] S25. Use the standard knowledge base in the medical field to optimize the word vector and map the word vector with the standard encoding to obtain the optimized word vector
[0041]
[0042] Among them, α is the adjustment weight, δ(WSK k ,v i ) indicates medical term v i The same as the code in the Ministry of Health's disease classification coding standard WSK k The degree of match between WSK k is the vector representation corresponding to WSK encoding;
[0043] S26. Output the optimized word embedding model:
[0044]
[0045] Among them, each is the optimized high-dimensional vector of medical terms.
[0046] Optionally, the S3 specifically includes:
[0047] S31, receiving the preprocessed medical text dataset D', the input layer maps the medical text dataset into a high-dimensional vector sequence X through a word embedding model i ={x i1 ,x i2 ,...,x ik};
[0048] S32, through the global Attention mechanism, the input word vector sequence X i Perform context-aware processing and calculate the attention weight α ij , the attention weight is determined by the similarity of word vectors and the semantic association of medical terms:
[0049]
[0050] Among them, q i and k j are query vector and key vector respectively, K med (v i ,v j ) represents the domain knowledge association between medical terms, λ1 is the importance factor of controlling domain knowledge, and d k is the dimension of the key vector;
[0051] S33, through attention weight α ij For the input vector X i Perform weighted summation to generate a context-aware vector h that includes local context information in the word sequence and global semantic associations of medical terms i :
[0052]
[0053] Among them, K global (v ip ,x ij ) represents the contextual influence of the global medical knowledge base on the word vector, and λ2 is the importance factor controlling the global medical knowledge;
[0054] S34, using a bidirectional long short-term memory network to process the context-aware vector h i , the forward and backward long short-term memory networks capture the contextual information in the sequence respectively, and the feature extraction of the hidden layer is:
[0055]
[0056] in, and They are forward and backward long short-term memory networks, and are the forward and backward impact factors of medical terms, W f and W b is the weight matrix of LSTM, d f and d b are the dimensions of the forward and backward LSTM networks.
[0057] Optionally, the S5 specifically includes:
[0058] S51. After identifying the entities, a bidirectional Attention mechanism is constructed. The bidirectional Attention mechanism extracts the relationship between the identified entities and generates a relationship for each pair of identified entities (e i ,e j ) Calculate the context-based association weights and define the bidirectional attention weights and Represent the forward and backward attention weights respectively:
[0059]
[0060]
[0061] in, and denote the forward and backward attention scores, q i For entity e i The query vector is and Entity e j The forward and backward key vectors of ;
[0062] S52, based on forward and backward attention weights and Computational Entity i and e j The context-aware relationship vector r between ij :
[0063]
[0064] in, and Respectively represent entity e j The context feature vectors in the forward and backward passes, r ij is the relationship vector between two entities;
[0065] S53, after calculating the direct relationship of the entity, the context-aware mechanism is used to infer the potential association relationship based on the identified entity and context relationship, using the global medical knowledge base K global As additional information, define the potential association weight β ij :
[0066] β ij =λ5·K global (e i ,e j );
[0067] Among them, λ5 is the weight factor that controls the influence of the global medical knowledge base, K global (e i ,e j) represents entity e i and e j degree of association in the global medical knowledge base;
[0068] S54, context-aware direct relationship ij and potential association weight β ij , calculate the final relationship prediction value R between entities ij :
[0069] R ij =r ij +β ij ;
[0070] Among them, R ij For entity e i and e j The final relationship between them is expressed.
[0071] Optionally, the S7 specifically includes:
[0072] S71, the identified entity set E = {e1, e2, ..., e n} and entity relationship set R = {r1,r2,...,r m} and the medical field knowledge base K med To match, the knowledge base K med Includes standard medical terms, disease classification codes and drug databases, defines entity matching score S match (e i ):
[0073]
[0074] Among them, K k is the kth medical term or entity record in the knowledge base, δ(e i ,K k ) is a binary function, when the entity e i and the entity K in the knowledge base k When matching, δ(e i ,K k )=1, otherwise 0, l is the entity e in the knowledge base i The total number of entries that matched;
[0075] S72. Score S based on entity matching match (e i ) for low matching entity e i Perform additional verification from the knowledge base K med Retrieve entities with similar semantics and calculate the entity similarity score S sim (e i ,K k ):
[0076]
[0077] in, and Entity e i and entity K in the knowledge base k The word vector representation of , cos represents the cosine similarity between word vectors;
[0078] S73, calculate each pair of entity relationships r for the identified relationship set R j =(e i ,e k )’s relationship matching score S match (r j ):
[0079] S match (r j )=λ6·Rel med (e i ,e k );
[0080] Among them, Rel med (e i ,e k ) is the entity e in the medical knowledge base i and entity e k The vector representation of the relationship between them, λ6 is the adjustment coefficient, which controls the influence of the knowledge base on the relationship verification;
[0081] S74, the final relationship prediction value R generated by combining context perception ij Matching score S match (r j ) for comparison, when the final relationship prediction value R ij When the difference between the matching score and the matching score exceeds the preset threshold, the semantic information in the knowledge base is used to verify and supplement the relationship;
[0082] S75. Output the medical entity set and medical relationship set after verification and supplementation by the knowledge base:
[0083]
[0084]
[0085] in, and It is a medical entity and medical relationship that has been verified and supplemented.
[0086] Optionally, the S73 specifically includes:
[0087] S731, for each pair of entity relationships identified j =(e i ,e k ) From the medical field knowledge base K med Extract entity e i With entity e k The relationship vector Rel med (e i ,e k ):
[0088] Rel med (e i ,e k )=[rel1(e i ,e k ),rel2(e i ,e k ),...,rel n (e i ,e k )];
[0089] Among them, rel n (e i ,e k ) represents entity e i With entity e k The nth relation attribute in the knowledge base, Rel med (e i ,e k ) is a vector representation summarizing all relationship attributes;
[0090] S732, Relationship vector Rel extracted based on knowledge base med (e i ,e k ) and the identified entity relationship r j =(e i ,e k ) to match and calculate the relationship matching score S used to measure the consistency between the relationship in the knowledge base and the identified relationship match (r j )):
[0091]
[0092] in, represents the nth relation vector generated by context perception, λ6 is the adjustment coefficient, which controls the influence of the knowledge base on relation verification, N is the total number of relation vectors in the knowledge base, ∥·∥ represents the norm of the vector, and the scoring result is the cosine similarity between the two vectors;
[0093] S733, scoring the relationship matching S match (rj ) is compared with the preset matching threshold. If S match (r j ) is lower than the set threshold, the relevant entity relationship information is retrieved from the knowledge base again and the relationship r j Verify and supplement.
[0094] The beneficial effects of the present invention are:
[0095] (1) The present invention significantly improves the accuracy of entity recognition by introducing a word embedding model optimized based on the medical field and combining it with a professional dictionary and corpus training model of medical terms. The Attention mechanism is used to capture medical terms and their contextual information. The word embedding model effectively handles the polysemous words and abbreviations in medical texts, and solves the problems of misclassification and missed recognition that traditional NLP models are prone to when facing professional medical terms. At the same time, combined with a bidirectional long short-term memory network, the model can capture contextual information in the text, thereby further enhancing the semantic understanding ability of entity recognition and making the recognition results more accurate.
[0096] (2) The present invention proposes a relationship extraction method based on a bidirectional Attention mechanism and context awareness, which can automatically identify and infer complex entity relationships in medical texts. By combining global and local Attention mechanisms, it can not only identify the direct relationship between entities, but also infer implicit potential relationships through context information. Combined with the relationship vectors in the medical field knowledge base, the model can automatically verify and supplement when extracting complex semantic associations, effectively solving the problem of information loss in the existing technology when dealing with multi-level semantic relationships, and greatly improving the model's relationship extraction ability in complex scenarios.
[0097] (3) The present invention introduces a medical knowledge base to automatically verify and supplement the identified entities and relationships, thereby solving the problem in the prior art that medical knowledge cannot be fully utilized. Through matching scores and similarity calculations, the system can dynamically adjust the recognition results of low-matching entities and relationships, and combine domain knowledge for effective supplementary verification to ensure the accuracy of entities and relationships. Through the cross-validation mechanism of the knowledge base, the model can automatically correct errors in recognition, reduce erroneous recognition caused by missing data or semantic ambiguity, and further improve the reliability and practicality of medical text analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0098] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0099] Figure 1This is a flowchart of a medical entity recognition and relationship extraction method based on deep learning proposed by the present invention;
[0100] Figure 2 This is a detailed schematic diagram of the bidirectional Attention mechanism used in the medical entity recognition and relationship extraction method based on deep learning proposed in the present invention. DETAILED DESCRIPTION
[0101] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, which only illustrate the basic structure of the present invention in a schematic manner, and therefore only show the components related to the present invention.
[0102] refer to Figure 1-2 , a medical entity recognition and relationship extraction method based on deep learning, comprising the following steps:
[0103] S1. Obtain and preprocess the medical text dataset, perform word segmentation, stop word removal and standardization on the medical text dataset;
[0104] S2. Build a word embedding model in the medical field, use the terminology dictionary and corpus in the medical field to train the word embedding model, so that the word embedding model generates a high-dimensional vector representation of medical terms;
[0105] S3. Based on the word embedding model, build and train a multi-layer neural network model using the Attention mechanism;
[0106] S4, applying the trained multi-layer neural network model to the new medical text dataset to identify and classify entities in the medical text dataset;
[0107] S5. After identifying the entities, the bidirectional Attention mechanism is used to extract the relationships between the entities in the medical text dataset. The relationship extraction includes identifying the direct relationship between entities and the potential relationship inferred through context perception.
[0108] S6. Use the bidirectional Attention mechanism to process different types of relationships separately.
[0109] S7. Verify and supplement the identified entities and relationships by combining with the medical domain knowledge base;
[0110] S8. Output the finally identified and extracted medical entities and their relationships in a structured form.
[0111] In this implementation, S1 specifically includes:
[0112] S11. Obtain medical text dataset:
[0113] D={d1,d2,...,d n};
[0114] Among them, d i represents the i-th medical text, and n is the total number of medical texts;
[0115] S12, use the word segmentation algorithm to perform word segmentation on the medical text dataset D, and separate each medical text d i Represented as a sequence of words:
[0116] W i ={w i1 ,w i2 ,...,w im};
[0117] Among them, w ij is the jth word in the i-th medical text, and m is the medical text d i The total number of words in the
[0118] S13, remove stop words and convert each word sequence W i The stop word set S in is filtered out to obtain a new word sequence W′ i ;
[0119] S14, word sequence W′ after removing stop words i Perform standardization, convert uppercase and lowercase letters, convert all words to lowercase, and restore the words in the word sequence to the stem form W″ i ;
[0120] S15. The preprocessed medical text dataset is expressed as:
[0121] D′={W″1,W″2,...,W″ n};
[0122] Among them, W″ i is the i-th preprocessed stemmed word sequence.
[0123] In this implementation, S2 specifically includes:
[0124] S21. Obtain a dictionary of medical terms:
[0125] V={v1,v2,...,v t};
[0126] Among them, v i represents the i-th medical term, and t is the total number of terms in the term dictionary;
[0127] S22. Collecting corpus in the medical field:
[0128] C={c1,c2,...,cp};
[0129] where c j represents the text segment in the jth medical field, and p is the total number of text segments in the corpus;
[0130] S23. Construct a word embedding model in the medical field based on the medical terminology dictionary V and the medical corpus C, and define the co-occurrence matrix M(i,j) of the word embedding model as:
[0131]
[0132] Among them, f(v i ,e kl ) represents the term v i In the kth text segment, kl The degree of correlation, g(v j ,e kl+1 ) represents the term v j With the adjacent entity e kl+1 The correlation degree, h(c k ) is the weight factor of the segment, representing the importance of the segment in the entire corpus, q k For the text fragment c k The number of entities in ;
[0133] S24. Use the context-aware model in the medical field to train the word vector of the co-occurrence matrix M(i,j), and define the word vector update as:
[0134]
[0135] in, represents the vector representation of the i-th term at the t-th iteration, η is the learning rate, λ is the regularization term, Word embedding values representing medical entities;
[0136] S25. Use the standard knowledge base in the medical field to optimize the word vector and map the word vector with the standard encoding to obtain the optimized word vector
[0137]
[0138] Among them, α is the adjustment weight, δ(WSK k ,v i ) indicates medical term v i The same as the code in the Ministry of Health's disease classification coding standard WSK k The degree of match between WSK k is the vector representation corresponding to WSK encoding;
[0139] S26. Output the optimized word embedding model:
[0140]
[0141] Among them, each is the optimized high-dimensional vector of medical terms.
[0142] In this implementation, S3 specifically includes:
[0143] S31, receiving the preprocessed medical text dataset D', the input layer maps the medical text dataset into a high-dimensional vector sequence X through a word embedding model i ={x i1 ,x i2 ,...,x ik};
[0144] S32, through the global Attention mechanism, the input word vector sequence X i Perform context-aware processing and calculate the attention weight α ij , the attention weight is determined by the similarity of word vectors and the semantic association of medical terms:
[0145]
[0146] Among them, q i and k j are query vector and key vector respectively, K med (v i ,v j ) represents the domain knowledge association between medical terms, λ1 is the importance factor of controlling domain knowledge, and d k is the dimension of the key vector;
[0147] S33, through attention weight α ij For the input vector X i Perform weighted summation to generate a context-aware vector h that includes local context information in the word sequence and global semantic associations of medical terms i :
[0148]
[0149] Among them, K global (v ip ,x ij ) represents the contextual influence of the global medical knowledge base on the word vector, and λ2 is the importance factor controlling the global medical knowledge;
[0150] S34, using a bidirectional long short-term memory network to process the context-aware vector h i, the forward and backward long short-term memory networks capture the contextual information in the sequence respectively, and the feature extraction of the hidden layer is:
[0151]
[0152] in, and They are forward and backward long short-term memory networks, and are the forward and backward impact factors of medical terms, W f and W b is the weight matrix of LSTM, d f and d b are the dimensions of the forward and backward LSTM networks.
[0153] In this implementation, S5 specifically includes:
[0154] S51. After identifying the entities, a bidirectional Attention mechanism is constructed. The bidirectional Attention mechanism extracts the relationship between the identified entities and generates a relationship for each pair of identified entities (e i ,e j ) Calculate the context-based association weights and define the bidirectional attention weights and Represent the forward and backward attention weights respectively:
[0155]
[0156]
[0157] in, and denote the forward and backward attention scores, q i For entity e i The query vector is and Entity e j The forward and backward key vectors of ;
[0158] S52, based on forward and backward attention weights and Computational Entity i and e j The context-aware relationship vector r between ij :
[0159]
[0160] in, and Respectively represent entity e jThe context feature vectors in the forward and backward passes, r ij is the relationship vector between two entities;
[0161] S53, after calculating the direct relationship of the entity, the context-aware mechanism is used to infer the potential association relationship based on the identified entity and context relationship, using the global medical knowledge base K global As additional information, define the potential association weight β ij :
[0162] β ij =λ5·K global (e i ,e j );
[0163] Among them, λ5 is the weight factor that controls the influence of the global medical knowledge base, K global (e i ,e j ) represents entity e i and e j degree of association in the global medical knowledge base;
[0164] S54, context-aware direct relationship ij and potential association weight β ij , calculate the final relationship prediction value R between entities ij :
[0165] R ij =r ij +β ij ;
[0166] Among them, R ij For entity e i and e j The final relationship between them is expressed.
[0167] In this implementation, S7 specifically includes:
[0168] S71, the identified entity set E = {e1, e2, ..., e n} and entity relationship set R = {r1,r2,...,r m} and the medical field knowledge base K med To match, the knowledge base K med Includes standard medical terms, disease classification codes and drug databases, defines entity matching score S match (e i ):
[0169]
[0170] Among them, K kis the kth medical term or entity record in the knowledge base, δ(e i ,K k ) is a binary function, when the entity e i and the entity K in the knowledge base k When matching, δ(e i ,K k )=1, otherwise 0, l is the entity e in the knowledge base i The total number of entries that matched;
[0171] S72. Score S based on entity matching match (e i ) for low matching entity e i Perform additional verification from the knowledge base K med Retrieve entities with similar semantics and calculate the entity similarity score S sim (e i ,K k ):
[0172]
[0173] in, and Entity e i and entity K in the knowledge base k The word vector representation of , cos represents the cosine similarity between word vectors;
[0174] S73, calculate each pair of entity relationships r for the identified relationship set R j =(e i ,e k )’s relationship matching score S match (r j ):
[0175] S match (r j )=λ6·Rel med (e i ,e k );
[0176] Among them, Rel med (e i ,e k ) is the entity e in the medical knowledge base i and entity e k The vector representation of the relationship between them, λ6 is the adjustment coefficient, which controls the influence of the knowledge base on the relationship verification;
[0177] S74, the final relationship prediction value R generated by combining context perception ij Matching score S match (r j) for comparison, when the final relationship prediction value R ij When the difference between the matching score and the matching score exceeds the preset threshold, the semantic information in the knowledge base is used to verify and supplement the relationship;
[0178] S75. Output the medical entity set and medical relationship set after verification and supplementation by the knowledge base:
[0179]
[0180]
[0181] in, and It is a medical entity and medical relationship that has been verified and supplemented.
[0182] In this implementation manner, S73 specifically includes:
[0183] S731, for each pair of entity relationships identified j =(e i ,e k ) From the medical field knowledge base K med Extract entity e i With entity e k The relationship vector Rel med (e i ,e k ):
[0184] Rel med (e i ,e k )=[rel1(e i ,e k ),rel2(e i ,e k ),...,rel n (e i ,e k )];
[0185] Among them, rel n (e i ,e k ) represents entity e i With entity e k The nth relation attribute in the knowledge base, Rel med (e i ,e k ) is a vector representation summarizing all relationship attributes;
[0186] S732, Relationship vector Rel extracted based on knowledge base med (e i ,e k) and the identified entity relationship r j =(e i ,e k ) to match and calculate the relationship matching score S used to measure the consistency between the relationship in the knowledge base and the identified relationship match (r j )):
[0187]
[0188] in, represents the nth relation vector generated by context perception, λ6 is the adjustment coefficient, which controls the influence of the knowledge base on relation verification, N is the total number of relation vectors in the knowledge base, ∥·∥ represents the norm of the vector, and the scoring result is the cosine similarity between the two vectors;
[0189] S733, scoring the relationship matching S match (r j ) is compared with the preset matching threshold. If S match (r j ) is lower than the set threshold, the relevant entity relationship information is retrieved from the knowledge base again and the relationship r j Verify and supplement.
[0190] Embodiment 1:
[0191] From July 2023 to September 2023, a large tertiary hospital analyzed a large amount of historical case data it had accumulated and used the deep learning-based medical entity recognition and relationship extraction method in the present invention. The hospital's goal was to improve the efficiency of its electronic medical record system through automation technology, reduce the time doctors spend on text information processing, and improve the accuracy of diagnosis and treatment recommendations. During this period, the hospital selected 10,000 case data containing complex medical information as training and test sets for the model.
[0192] In the test on July 15, a patient's case description was: "The patient took metformin for type 2 diabetes, with poor blood sugar control, accompanied by blurred vision and weight loss." When processing this type of text, the traditional NLP model can identify the two entities of "diabetes" and "metformin", but cannot infer the correlation between blurred vision and weight loss as complications, resulting in incomplete extracted information. However, the method of the present invention, through its bidirectional Attention mechanism, combines context perception with the medical field knowledge base, accurately identifies "type 2 diabetes" as a disease entity and "metformin" as a drug entity, and automatically associates "blurred vision" and "weight loss" as complication symptoms through context inference.
[0193] In this test, the hospital selected 5,000 case data for entity recognition and relationship extraction tests, and conducted a comparative experiment between the method of the present invention and the traditional rule-based NLP method. The specific experimental data are as follows:
[0194] Table 1 Comparison of medical entity recognition and relationship extraction methods
[0195]
[0196] In actual cases, the system successfully identified a large number of important medical entities and inferred multiple complex relationships by processing case data sets in real time. In Example 1, a case description on August 2 was: "The patient has been taking aspirin for a long time and recently stopped taking the medication due to recurrence of gastric ulcer." Traditional methods can only identify the two entities "aspirin" and "gastric ulcer", while the method of the present invention can infer "aspirin" as a potential cause of "recurrence of gastric ulcer" through context, and generate a complex association report of "drug-symptom".
[0197] On August 10, a doctor in the hospital searched the case system for a patient who had been taking medication for high blood pressure for a long time. The case description was: "The patient took amlodipine and recently developed symptoms of headache and nausea." Traditional analysis methods can identify the two entities "high blood pressure" and "amlodipine", but cannot accurately extract "headache" and "nausea" as adverse drug reactions. The method of the present invention not only identifies the relationship between "high blood pressure" and "amlodipine" through a deep learning-based entity recognition model, but also infers that "headache" and "nausea" may be related to the drug through context perception. After verification by the medical knowledge base, the system confirms that these symptoms are closely related to the adverse reactions of "amlodipine", and generates a detailed drug-adverse reaction report to promptly remind doctors to adjust the patient's treatment plan.
[0198] The system generates the following report through automated analysis:
[0199] Entities: hypertension, amlodipine, headache, nausea;
[0200] Relationship: drug-treats disease, drug-induces adverse reactions;
[0201] Verification time: 2 seconds (quick verification based on knowledge base);
[0202] In order to demonstrate the effect more intuitively, the hospital conducted a comparative test on 1,000 case data containing complex symptom descriptions. In the test, the traditional rule-based NLP model could not accurately identify more than 50% of complications, while the method of the present invention could accurately extract more than 85% of complication symptoms. In Example 1, in a case described as "a patient took pantoprazole for chronic gastritis and recently developed stomach pain and acid reflux symptoms", the traditional method could not effectively extract the association between "stomach pain" and "acid reflux" and the drug, while the method of the present invention accurately identified and generated a drug-symptom association report.
[0203] In an emergency case handling in September 2023, the method of the present invention once again demonstrated its high efficiency in complex medical scenarios. The emergency department received a case of a patient with acute myocardial infarction, described as "the patient suddenly had chest pain and shortness of breath. The doctor injected atropine and the symptoms were relieved, but then blurred vision appeared." Traditional methods can only identify "chest pain" and "shortness of breath" as direct symptoms of myocardial infarction, while the method of the present invention, through context and knowledge base verification, not only identifies "blurred vision" as a potential adverse reaction of atropine, but also automatically generates a risk report on the association between drugs and symptoms, helping doctors to quickly adjust treatment plans and avoid more serious adverse reactions.
[0204] In this emergency treatment, the actual operation data of the method of the present invention are as follows:
[0205] Entity recognition time: 10 seconds, entity relationship extraction accuracy: 94.6%, adverse reaction detection rate: 92.4%, system generation of complete report time: 3 minutes.
[0206] Through a series of real-world scenario tests, the method of the present invention has demonstrated its excellent performance in medical entity recognition and relationship extraction. Compared with traditional NLP models, the method of the present invention is significantly superior to traditional methods in entity recognition accuracy, relationship extraction efficiency and processing time in complex medical scenarios. In emergency medical scenarios, the method of the present invention can promptly detect potential adverse reactions, accurately associate drugs with symptoms, and generate detailed analysis reports to help doctors make quick diagnostic decisions, thereby effectively improving the quality of medical services.
[0207] The present invention significantly improves the accuracy of entity recognition by introducing a word embedding model optimized based on the medical field and combining it with a professional dictionary and corpus training model of medical terms. The Attention mechanism is used to capture medical terms and their contextual information. The word embedding model effectively handles the problems of polysemous words and abbreviations in medical texts, and solves the problems of misclassification and missed recognition that are prone to occur in traditional NLP models when facing professional medical terms. At the same time, combined with a bidirectional long short-term memory network, the model can capture contextual information in the text, thereby further enhancing the semantic understanding ability of entity recognition and making the recognition results more accurate.
[0208] The present invention proposes a relationship extraction method based on a bidirectional Attention mechanism and context awareness, which can automatically identify and infer complex entity relationships in medical texts. By combining global and local Attention mechanisms, it can not only identify direct relationships between entities, but also infer implicit potential relationships through context information. Combined with the relationship vectors in the medical field knowledge base, the model can automatically verify and supplement when extracting complex semantic associations, effectively solving the problem of information loss in the prior art when processing multi-level semantic relationships, and greatly improving the model's relationship extraction capabilities in complex scenarios.
[0209] The present invention introduces a medical field knowledge base to automatically verify and supplement the identified entities and relationships, thereby solving the problem that the medical field knowledge cannot be fully utilized in the prior art. Through matching scores and similarity calculations, the system can dynamically adjust the recognition results of low-matching entities and relationships, and combine domain knowledge for effective supplementary verification to ensure the accuracy of entities and relationships. Through the cross-validation mechanism of the knowledge base, the model can automatically correct errors in recognition, reduce erroneous recognition caused by missing data or semantic ambiguity, and further improve the reliability and practicality of medical text analysis.
[0210] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A medical entity recognition and relationship extraction method based on deep learning, characterized in that: The steps include: S1. Obtain and preprocess the medical text dataset, perform word segmentation, stop word removal and standardization on the medical text dataset; S2. Build a word embedding model in the medical field, use the terminology dictionary and corpus in the medical field to train the word embedding model, so that the word embedding model generates a high-dimensional vector representation of medical terms; S3. Based on the word embedding model, a multi-layer neural network model using the Attention mechanism is constructed and trained; S4, applying the trained multi-layer neural network model to the new medical text dataset to identify and classify entities in the medical text dataset; S5. After identifying the entities, the bidirectional Attention mechanism is used to extract the relationships between the entities in the medical text dataset. The relationship extraction includes identifying the direct relationship between entities and the potential relationship inferred through context perception. S6. Use the bidirectional Attention mechanism to process different types of relationships separately. S7. Verify and supplement the identified entities and relationships by combining with the medical domain knowledge base; S8, output the finally identified and extracted medical entities and their relationships in a structured form; The S7 specifically includes: S71, the identified entity set E = {e1, e2, ..., e n } and entity relationship set R = {r1,r2,...,r m } and the medical field knowledge base K med To match, the knowledge base K med Includes standard medical terms, disease classification codes and drug databases, defines entity matching score S match (e i ): Among them, K k is the kth medical term or entity record in the knowledge base, δ(e i ,K k ) is a binary function, when the entity e i and the entity K in the knowledge base k When matching, δ(e i ,K k )=1, otherwise 0, l is the entity e in the knowledge base i The total number of entries that matched; S72, according to the entity matching score S match (e i ) for low matching entity e i Perform additional verification from the knowledge base K med Retrieve entities with similar semantics and calculate the entity similarity score S sim (e i ,K k ): in, and Entity e i and entity K in the knowledge base k The word vector representation of , cos represents the cosine similarity between word vectors; S73, calculate each pair of entity relationships r for the identified relationship set R j =(e i ,e k )’s relationship matching score S match (r j ): S match (r j )=λ6·Rel med (e i ,e k ); Among them, Rel med (e i ,e k ) is the entity e in the medical knowledge base i and entity e k The vector representation of the relationship between them, λ6 is the adjustment coefficient, which controls the influence of the knowledge base on the relationship verification; S74, the final relationship prediction value R generated by combining context perception ij Matching score S match (r j ) for comparison, when the final relationship prediction value R ij When the difference between the matching score and the matching score exceeds the preset threshold, the semantic information in the knowledge base is used to verify and supplement the relationship; S75. Output the medical entity set and medical relationship set after verification and supplementation by the knowledge base: in, and It is a medical entity and medical relationship that has been verified and supplemented.
2. According to the deep learning-based medical entity recognition and relationship extraction method of claim 1, it is characterized in that: The S1 specifically includes: S11. Obtain medical text dataset: D={d1,d2,...,d n }; Among them, d i represents the i-th medical text, and n is the total number of medical texts; S12, use the word segmentation algorithm to perform word segmentation on the medical text dataset D, and separate each medical text d i Represented as a sequence of words: IN i ={in i1 ,In i2 ,...,In im }; Among them, w ij is the jth word in the i-th medical text, and m is the medical text d i The total number of words in the S13, remove stop words and convert each word sequence W i The stop word set S in is filtered out to obtain a new word sequence W′ i ; S14, word sequence W′ after removing stop words i Perform standardization, convert uppercase and lowercase letters, convert all words to lowercase, and restore the words in the word sequence to the stem form W″ i ; S15. The preprocessed medical text dataset is expressed as: D′={W″1,W″2,...,W″ n }; Among them, W″ i is the i-th preprocessed stemmed word sequence.
3. According to the deep learning-based medical entity recognition and relationship extraction method of claim 1, it is characterized in that: The S2 specifically includes: S21. Obtain a dictionary of medical terms: V={v1,v2,...,v t }; Among them, v i represents the i-th medical term, and t is the total number of terms in the term dictionary; S22. Collecting corpus in the medical field: C={c1,c2,...,c p }; where c j represents the text segment in the jth medical field, and p is the total number of text segments in the corpus; S23, constructing a word embedding model in the medical field based on the medical field terminology dictionary V and the medical field corpus C, and defining the co-occurrence matrix M(i, j) of the word embedding model as: Among them, f(v i ,e kl ) represents the term v i In the kth text segment, kl The degree of correlation, g(v j ,e kl+1 ) represents the term v j With the adjacent entity e kl+1 The correlation degree, h(c k ) is the weight factor of the segment, representing the importance of the segment in the entire corpus, q k For the text fragment c k The number of entities in ; S24. Use the context-aware model in the medical field to train the word vector of the co-occurrence matrix M(i,j), and define the word vector update as: in, represents the vector representation of the i-th term at the t-th iteration, η is the learning rate, λ is the regularization term, Word embedding values representing medical entities; S25. Use the standard knowledge base in the medical field to optimize the word vector and map the word vector with the standard encoding to obtain the optimized word vector Among them, α is the adjustment weight, δ(WSK k ,v i ) indicates medical term v i The same as the code in the Ministry of Health's disease classification coding standard WSK k The degree of match between WSK k is the vector representation corresponding to WSK encoding; S26. Output the optimized word embedding model: Among them, each is the optimized high-dimensional vector of medical terms.
4. The method for medical entity recognition and relationship extraction based on deep learning according to claim 1, characterized in that: The S3 specifically includes: S31, receiving the preprocessed medical text dataset D', the input layer maps the medical text dataset into a high-dimensional vector sequence X through a word embedding model i ={x i1 ,x i2 ,...,x ik }; S32, through the global Attention mechanism, the input word vector sequence X i Perform context-aware processing and calculate the attention weight α ij , the attention weight is determined by the similarity of word vectors and the semantic association of medical terms: Among them, q i and k j are query vector and key vector respectively, K med (v i ,v j ) represents the domain knowledge association between medical terms, λ1 is the importance factor of controlling domain knowledge, and d k is the dimension of the key vector; S33, through attention weight α ij For the input vector X i Perform weighted summation to generate a context-aware vector h that includes local context information in the word sequence and global semantic associations of medical terms i : Among them, K global (v ip ,x ij ) represents the contextual influence of the global medical knowledge base on the word vector, and λ2 is the importance factor controlling the global medical knowledge; S34, using a bidirectional long short-term memory network to process the context-aware vector h i , the forward and backward long short-term memory networks capture the contextual information in the sequence respectively, and the feature extraction of the hidden layer is: in, and They are forward and backward long short-term memory networks, and are the forward and backward impact factors of medical terms, W f and W b is the weight matrix of LSTM, d f and d b are the dimensions of the forward and backward LSTM networks.
5. The method for medical entity recognition and relationship extraction based on deep learning according to claim 1, characterized in that: The S5 specifically includes: S51. After identifying the entities, a bidirectional Attention mechanism is constructed. The bidirectional Attention mechanism extracts the relationship between the identified entities and generates a relationship for each pair of identified entities (e i ,e j ) Calculate the context-based association weights and define the bidirectional attention weights and Represent the forward and backward attention weights respectively: in, and denote the forward and backward attention scores, q i For entity e i The query vector is and Entity e j The forward and backward key vectors of ; S52, based on forward and backward attention weights and Computational Entity i and e j The context-aware relationship vector r between ij : in, and Respectively represent entity e j The context feature vectors in the forward and backward passes, r ij is the relationship vector between two entities; S53, after calculating the direct relationship of the entity, the context-aware mechanism is used to infer the potential association relationship based on the identified entity and context relationship, using the global medical knowledge base K global As additional information, define the potential association weight β ij : b ij =λ5·K global (e i ,e j ); Among them, λ5 is the weight factor that controls the influence of the global medical knowledge base, K global (e i ,e j ) represents entity e i and e j degree of association in the global medical knowledge base; S54, context-aware direct relationship ij and potential association weight β ij , calculate the final relationship prediction value R between entities ij : R ij =r ij +β ij ; Among them, R ij For entity e i and e j The final relationship between .
6. The method for medical entity recognition and relationship extraction based on deep learning according to claim 1, characterized in that: The S73 specifically includes: S731, for each pair of entity relationships identified j =(e i ,e k ) From the medical field knowledge base K med Extract entity e i With entity e k The relationship vector Rel med (e i ,e k ): Relative med (And i ,And k )=[rel1(and i ,And k ),rel2(e i ,And k ),...,rel n (And i ,And k )]; Among them, rel n (e i ,e k ) represents entity e i With entity e k The nth relation attribute in the knowledge base, Rel med (e i ,e k ) is a vector representation summarizing all relationship attributes; S732, Relationship vector Rel extracted based on knowledge base med (e i ,e k ) and the identified entity relationship r j =(e i ,e k ) to match and calculate the relationship matching score S used to measure the consistency between the relationship in the knowledge base and the identified relationship match (r j )): in, represents the nth relation vector generated by context perception, λ6 is the adjustment coefficient, which controls the influence of the knowledge base on relation verification, N is the total number of relation vectors in the knowledge base, ∥·∥ represents the norm of the vector, and the scoring result is the cosine similarity between the two vectors; S733, scoring the relationship matching S match (r j ) is compared with the preset matching threshold. If S match (r j ) is lower than the set threshold, the relevant entity relationship information is retrieved from the knowledge base again and the relationship r j Verify and supplement.
Citation Information
Patent Citations
Composite neural network psychological medical knowledge graph construction method and system
CN115879546A