A Medical Entity Disambiguation Method Based on Bio-LinkBERT and Context Awareness

By adopting the Bio-LinkBERT and ELMo models in the medical entity disambiguation method, combining the cross-attention mechanism and the Bi GRU encoder, the problem of existing methods being difficult to capture cross-document dependencies and rich knowledge is solved, achieving higher accuracy and reliability.

CN115796184BActive Publication Date: 2025-06-27DALIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211674890.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-26
Publication Date
2025-06-27
Estimated Expiration
2042-12-26

AI Technical Summary

Technical Problem

Existing medical entity disambiguation methods are difficult to effectively capture cross-document dependencies and rich knowledge, cannot perform multi-hop reasoning, and the implicit knowledge in the medical knowledge base is difficult to reflect, resulting in insufficient performance and scalability.

Method used

Using a medical entity disambiguation method based on Bio-LinkBERT and context-aware, cross-document dependencies are extracted through Bio-LinkBERT, combined with the ELMo model encoding context, and using a cross-attention mechanism and a multi-layer Bi GRU encoder, the relationship and context information between mentioned and candidate entities are captured.

Benefits of technology

It improves the accuracy and reliability of medical entities' disambiguation, can effectively capture cross-document dependencies and context information, and solves the problems of multi-hop reasoning and implicit knowledge embodied.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0004017800500000052
    Figure BDA0004017800500000052
  • Figure BDA0004017800500000091
    Figure BDA0004017800500000091
  • Figure HDA0004017800510000011
    Figure HDA0004017800510000011
Patent Text Reader

Abstract

The present invention discloses a medical entity disambiguation method based on Bio-LinkBERT and context awareness, including: preprocessing all mentions in medical texts and entity names in the knowledge base; obtaining a candidate entity set for mention M through similarity; encoding the mentions in medical texts and candidate entities in the knowledge base using Bio-LinkBERT; using a cross-attention mechanism for the mentions and candidate entities to capture interaction relationships and obtain word representations of the mentions and candidate entities; inputting the above word representations into a BiGRU for further encoding to obtain word representations containing more semantic information; using a context awareness mechanism to obtain the correlation degree between the mention context and candidate entities according to the self-attention mechanism; obtaining the matching scores of each predicted candidate entity through a feed-forward neural network, and taking the candidate entity with the highest score as the target entity uniquely corresponding to the mention in the knowledge base. The present invention can extract the dependency relationships between medical documents and capture the context text clues of mentions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a medical entity disambiguation method based on Bio-LinkBERT and context awareness. Background Art

[0002] In recent years, medical technology has been continuously advancing, the number of medical texts has increased significantly, and the number of medical knowledge bases has also been growing rapidly. How to utilize the rich knowledge contained in these records to provide high-quality information to facilitate clinical decision-making is crucial. However, many different medical concepts may have very similar mentions, and if the ambiguity among them cannot be eliminated, it will lead to misinterpretation of the entire context, which will pose a great risk to medical-related decisions. Therefore, medical entity disambiguation (BioNED) is the key to correctly utilizing such knowledge bases.

[0003] BioNED is a task of eliminating the ambiguity of mentions by linking the mentions in medical texts to the corresponding entities in a medical knowledge base. Currently, the methods for solving BioNED mainly include rule-based methods, traditional machine learning-based methods, and deep learning-based methods. A rule-based entity disambiguation system refers to using artificially defined rules to simulate the text coherence between mentions and entities, and calculating string similarity by specifying a certain order or weight combination of these rules to perform the disambiguation task. These methods usually have extremely high accuracy, but the disadvantage is extremely low recall. A traditional machine learning-based entity disambiguation system will automatically learn the similarity between mentions and entities, and has a higher recall rate than rule-based methods. However, the disadvantage is that they cannot utilize semantic information to distinguish similar words, and in order to achieve higher accuracy, complex feature engineering needs to be used for calculation. A deep learning-based entity disambiguation system overcomes the dependence on feature engineering; for example, the medical entity disambiguation method based on the BERT model has achieved SOTA results on many medical benchmark datasets; but there are still some problems: they only model the current single document, so that although the word embeddings have context knowledge, they cannot capture cross-document dependencies and rich knowledge, nor can they perform multi-hop reasoning.

[0004] In traditional entity disambiguation tasks, mentions need to be linked to real entities in a general knowledge base that provides various information (such as entity names, entity descriptions, entity attributes, etc.), while medical knowledge bases only contain entity names, and a large amount of implicit knowledge is difficult to be reflected in sample data, and the available information is scarce. Therefore, how to improve the performance and scalability of the method has important theoretical significance and practical application value for medical entity disambiguation. Summary of the Invention

[0005] The object of the present invention is to provide a medical entity disambiguation method based on Bio-LinkBERT and context awareness, which can extract the dependency relationships between medical documents, model the interaction information between text mentions and target entities, capture the context text clues of mentions, and improve the performance of existing medical entity disambiguation.

[0006] To achieve the above object, the present application proposes a medical entity disambiguation method based on Bio-LinkBERT and context awareness, including:

[0007] Preprocess all mentions in the medical text and entity names in the knowledge base to unify the format for subsequent operations;

[0008] Obtain the candidate entity set of mention M through similarity to control the number of candidate entities;

[0009] Use Bio-LinkBERT to encode the mentions in the medical text and candidate entities in the knowledge base respectively, extract cross-document dependency relationships, and achieve multi-hop reasoning;

[0010] Use the cross-attention mechanism for the mentions and candidate entities to capture the relationships between mention-candidate entities and candidate entity-mentions, making the representations of mentions and candidate entities more attention-grabbing;

[0011] Encode the mentions and candidate representations processed by the cross-attention mechanism using a multi-layer Bi GRU encoder to obtain the word representations of the final mentions and the word representations of candidate entities; concatenate the two word representations to get the result output;

[0012] Use the context awareness mechanism to obtain the correlation degree between the mention context and candidate entities through the self-attention mechanism, fully mine the hidden information in the context text, and calculate the context score;

[0013] After concatenating the context score with the result output, input it into a feed-forward neural network to obtain the matching scores of each predicted candidate entity, and take the candidate entity with the highest score as the target entity uniquely corresponding to the mention in the knowledge base.

[0014] Further, preprocessing all mentions in the medical text and entity names in the knowledge base specifically includes:

[0015] Expand English abbreviations: Use the Ab3p toolkit to expand abbreviations;

[0016] Entity segmentation: Use SimConcept tools to segment composite entities;

[0017] Digital conversion: Manually create a digital dictionary based on the Wikipedia page, and replace different forms of numbers in the mentioned phrases with Arabic numerals;

[0018] Other processing: Remove the punctuation marks existing in the mention, and convert all characters to lowercase letters.

[0019] Furthermore, obtain the candidate entity set of mention M through similarity, specifically:

[0020] First, perform exact and fuzzy matching, that is, select candidate entities according to entity names that exactly match all letters in the mention or share multiple common characters with the mention; in addition, the present invention also considers the information of other mention phrases; specifically, if the current mention is an abbreviation or substring mentioned in other phrases, the original mention is merged and the candidate set of the mention is expanded.

[0021] Split mention M and candidate entity E into tokens, use the Levenshtein ratio and cosine similarity to calculate the similarity between the mention string and the candidate entity name, and finally select the top k entities with the highest scores as candidate entities. The specific implementation method is: calculate the Levenshtein ratio and select the top N (for example, N can take the value of 100) entity names with the highest scores; on this basis, considering the word order problem, calculate the similarity between the mention tokens and the entity name tokens and the similarity between the entity name tokens and the mention tokens at the same time, that is, the alignment cosine similarity; obtain the similarity score between the mention and the candidate entity name through the average value of the alignment cosine similarity;

[0022] Finally, each mention M has a candidate entity set C containing N / 2 (when N takes the value of 100, this value is 50) candidate entities M ={<id1,C1,score1>,<id2,C2,score2>,...,<id k ,C k ,score k >}, where id i is the candidate entity number, C i is the candidate entity name, and score i is the similarity score of the candidate entity.

[0023] Furthermore, use Bio-LinkBERT to encode the mentions in medical texts and the candidate entities in the knowledge base respectively, specifically:

[0024] Represent the mention tokens and the candidate entity tokens respectively using Bio-LinkBERT to obtain the corresponding word embeddings and Meanwhile, to overcome the problem of out-of-vocabulary (OOV) words, a bidirectional long short-term memory neural network (BiLSTM) is used to capture character-level features and obtain character embeddings. and Then, the character embeddings are concatenated with the word embeddings to finally obtain a word representation that contains both word-level and character-level information. and

[0025] Furthermore, a cross-attention mechanism is used for the mentioned and candidate entities to capture the relationships between the mention-candidate entity and the candidate entity-mention. Specifically:

[0026] Calculate the word representation through a cross-attention module and The interaction between them is calculated, so that the relationships between text features can be learned to obtain more accurate results. The present invention uses a bidirectional attention mechanism to calculate attention in two directions: from the mention to the candidate entity and from the candidate entity to the mention. These two attentions are obtained through a shared similarity matrix The similarity matrix S is obtained through the word representations and The element s tj in the matrix represents the similarity between the mention token i and the entity name token j. Then, S can be used to obtain the attentions in two directions: mention-to-candidate attention (M2CAtt) and candidate-to-mention attention (C2MAtt).

[0027] To obtain a word representation containing more information, the present invention encodes the mentions and candidate representations processed by the cross-attention mechanism using a multi-layer BiGRU encoder to obtain the word representation r i m of the final mention and the word representation r i c of the candidate entity; the word representations of the two are concatenated to obtain the result output.

[0028] Even further, a context-aware mechanism is used to obtain the correlation degree between the mention context and the candidate entity through a self-attention mechanism. Specifically:

[0029] Evaluate the relevance between the context and the candidate entity by calculating the context score: First, use an ELMo model containing two layers of bidirectional long short-term memory neural networks (BiLSTM) to encode the candidate entity and the mention context to obtain the candidate entity representation ctx Eand the initial mention context representation ctx' M To be able to select important keywords and ignore the influence of noise, the present invention uses an attention mechanism to assign weights to each token in the mention context, and then calculates the weighted sum to obtain the final mention context representation ctx M ; then the context score is calculated as ctx M and ctx E 's dot product, and splices it to the result vector output.

[0030] As a further step, after splicing the context score with the result output, it is input into a two-layer feed-forward neural network.

[0031] As a further step, positive samples are randomly selected from the given training set, and negative samples are selected from the candidate entities (excluding the target entity) generated in the candidate entity generation stage. This makes the negative samples very similar to the positive samples, forcing the present invention to disambiguate entities in a more fine-grained manner. The hinge loss is used as the loss function, and the purpose of the hinge loss function is to separate positive and negative sample pairs at a certain margin by optimizing the embedding space to ensure that positive sample pairs are close enough to each other and negative sample pairs are far enough away from each other.

[0032] The present invention adopts the above technical solutions. Compared with the prior art, the advantages are: using Bio-LinkBERT that can capture cross-document dependencies to encode mentions and entities, solving the multi-hop reasoning problem. At the same time, using ELMo to encode the context to obtain rich disambiguation knowledge implicit in the context, effectively improving the accuracy and reliability of medical entity disambiguation. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is a flowchart of a medical entity disambiguation method based on Bio-LinkBERT and context awareness in an embodiment of the present invention;

[0034] Figure 2 is a schematic diagram of an embedding layer in an embodiment of the present invention;

[0035] Figure 3 is a schematic diagram of a cross-attention layer in an embodiment of the present invention;

[0036] Figure 4 is a schematic diagram of a Bi-GRU layer in an embodiment of the present invention;

[0037] Figure 5 is a schematic diagram of a context encoding layer in an embodiment of the present invention;

[0038] Figure 6 is a medical entity disambiguation model diagram in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application, that is, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments.

[0040] Therefore, the detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application claimed, but merely represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present application.

[0041] Embodiment 1

[0042] As Figure 1 shown, this embodiment provides a medical entity disambiguation method based on Bio-LinkBERT and context awareness, including a data preprocessing step, a candidate entity generation step, and a candidate entity ranking step.

[0043] Step 1. Preprocess all mentions in the medical text and entity names in the knowledge base;

[0044] Specifically, extract the mention "CopperToxicosis" to be disambiguated from the text, and after data preprocessing, obtain the standardized data "coppertoxicosis".

[0045] Step 2. Obtain the candidate entity set of mention M through similarity to control the number of candidate entities;

[0046] Specifically, through the candidate generation step, calculate the candidate entity set of "coppertoxicosis" from the knowledge base: "D020149, manganese poisoning, 0.80687", "D002105, cadmium poisoning, 0.7961", "OMIM:215600, copper toxicosis idiopathic included, 0.78953", "D008630, mercury poisoning, 0.74249", "C53546, copper deficiency, 0.7309" and "D020261, arsenic poisoning, 0.73047".

[0047] Step 3. Use Bio-LinkBERT to encode the mentions in the medical text and the candidate entities in the knowledge base respectively, and extract cross-document dependencies;

[0048] As Figure 2 shown in the embedding layer, the mention "coppertoxicosis" and each candidate entity (taking the first candidate entity "D020149, manganese poisoning, 0.80687" as an example) are encoded using Bio-LinkBERT respectively to obtain word embeddings; run Bi LSTM on "copper toxicosis" and "manganese poisoning" to obtain character embeddings, and finally concatenate the word embeddings and character embeddings to obtain word representations.

[0049] Step 4. Use the cross-attention mechanism for the mentioned and candidate entities to capture the relationships between mention-candidate entities and candidate entity-mentions;

[0050] As Figure 3 shown in the cross-attention layer, input the word representations of "copper toxicosis" and "manganese poisoning" into the cross-attention module. First, calculate the similarity matrix where Then calculate the attention in two directions of "copper toxicosis"-"manganese poisoning" and "manganese poisoning"-"copper toxicosis" respectively: S α = softmax(row(S)), S β = softmax(max col (S)),

[0051] Step 5. Encode the mentions and candidate representations processed by the cross-attention mechanism using a multi-layer Bi GRU encoder to obtain the word representations of the final mentions and the word representations of the candidate entities; concatenate the two word representations to obtain the result output;

[0052] As Figure 4 shown in the Bi GRU layer, use Bi GRU to encode the word representations of "coppertoxicosis" and "manganesepoisoning" after passing through the cross-attention mechanism:

[0053] Finally, ri m With r i c Concatenate to obtain the result output.

[0054] Step 6. Use the context-aware mechanism to obtain the relevance between the mentioned context and the candidate entity through the self-attention mechanism, fully mine the hidden information in the context text, and calculate the context score.

[0055] Step 7. After concatenating the context score with the result output, input it into the feed-forward neural network to obtain the matching scores of each predicted candidate entity, and use the candidate entity with the highest score as the target entity uniquely corresponding to the mention in the knowledge base.

[0056] As Figure 5 Shown in the context encoding layer, use the ELMo model containing two layers of Bi LSTM to encode the candidate entity "manganese poisoning" and the mention context "by investigating the common autosomal recessive copper toxicosis in bedlington terriers we have identified a new locus involved in progressive liver disease" to obtain the candidate entity representation ctx E And the mention context representation ctx' M ; At the same time, use the attention mechanism to assign weights to each token in the context, and then calculate the weighted sum to obtain the mention context representation ctx M = weight · ctx' M ; Finally, calculate the context score as the dot product of ctx M And ctx E : ctx score (M, E) = ctx M · ctx E , and concatenate it to output: output = [output, ctx score

[0057] ​The final output is obtained using a two-layer fully connected neural network: Φ' = ReLU(W1·output + b1), Φ(M, E) = sigmoid(W2·Φ' + b2). The candidate entity with the highest score "OMIM: 215600, copper toxicosis idiopathic included, 0.6824" is regarded as the final linking result of the mention "copper toxicosis" in the knowledge base.

[0058] Optimization training is performed using the hinge loss function: where E + represents the positive sample, and E - represents the negative sample, and μ is the margin hyperparameter.

[0059] When training the disambiguation model, the parameter settings on the NCBI, ADR, and ShARe / CLEF datasets are as follows: the learning rate is 0.001 for all, the batch size is 128 for all, the decay rate is 0.05 for all, dropout(0.1) is used to prevent overfitting, the optimizer is Adam, and Recall and Accuracy are used as evaluation metrics.

[0060] In this embodiment, Bio-LinkBERT capable of learning cross-document dependencies is used to obtain the embedding representations of mentions and candidate entities, and character-level information is added to overcome the OOV problem. Then, a cross-attention mechanism is used for mentions and candidate entities to capture the relationships between mention-candidate entity and candidate entity-mention. A context-aware mechanism is also used to encode the context of the mention with ELMo, and the degree of relevance between the context and the candidate entity is obtained through the self-attention mechanism, obtaining clues for disambiguation in the context, which greatly improves the accuracy and reliability of the disambiguation model.

[0061] The foregoing description of the specific exemplary embodiments of the present invention is for purposes of illustration and exemplification. These descriptions are not intended to limit the invention to the precise forms disclosed, and obviously, many changes and variations are possible in light of the above teachings. The purpose of selecting and describing the exemplary embodiments is to explain the specific principles of the invention and its practical applications, so that those skilled in the art can implement and utilize the various different exemplary embodiments of the present invention, as well as various different selections and changes. The scope of the present invention is intended to be defined by the claims and their equivalents.

Claims

1. A medical entity disambiguation method based on Bio-LinkBERT and context awareness, characterized in that, Including: Preprocess all mentions in the medical text and entity names in the knowledge base; Obtain the candidate entity set of mention M through similarity to control the number of candidate entities; Use Bio-LinkBERT to encode the mentions in the medical text and candidate entities in the knowledge base respectively, and extract cross-document dependencies; Use a cross-attention mechanism for the mentioned and candidate entities to capture the relationship between mention-candidate entity and candidate entity-mention; Encode the mentions and candidate representations after being processed by the cross-attention mechanism using a multi-layer BiGRU encoder to obtain the word representations of the final mentions and candidate entities; concatenate the two word representations to get the result output; Use a context-aware mechanism to obtain the relevance between the mention context and candidate entities through a self-attention mechanism, fully mine the hidden information in the context text, and calculate the context score; After concatenating the context score with the result output, input it into a feed-forward neural network to obtain the matching scores of each predicted candidate entity, and use the candidate entity with the highest score as the target entity uniquely corresponding to the mention in the knowledge base.

2. The method for medical entity disambiguation based on Bio-LinkBERT and context awareness according to claim 1, wherein Preprocess all mentions in the medical text and entity names in the knowledge base, specifically including: Expand English abbreviations: Use the Ab3p toolkit to expand abbreviations; Entity segmentation: Use the SimConcept tool to segment composite entities; Digital conversion: Manually create a digital dictionary according to Wikipedia pages, and replace different forms of numbers in the mention phrase with Arabic numerals; Other processing: Delete the punctuation marks in the mention and convert all characters to lowercase letters.

3. The method for medical entity disambiguation based on Bio-LinkBERT and context awareness according to claim 1, wherein Obtain the candidate entity set of mention M through similarity, specifically: First, perform exact and fuzzy matching, that is, select candidate entities according to entity names that exactly match all letters in the mention or share multiple common characters with the mention; Then calculate the Levenshtein ratio and select the top N entity names with the highest scores; at the same time, calculate the similarity between mention tokens and entity name tokens and the similarity between entity name tokens and mention tokens, that is, the alignment cosine similarity; Obtain the similarity score between the mention and the candidate entity name through the average value of the alignment cosine similarity; Finally, each mention M has a candidate entity set containing N / 2 candidate entities.

4. The medical entity disambiguation method based on Bio-LinkBERT and context awareness according to claim 1, wherein Use Bio-LinkBERT to encode the mentions in the medical text and candidate entities in the knowledge base respectively, specifically: Represent the mention tokens and candidate entity tokens using Bio-LinkBERT respectively to obtain word embeddings; capture character-level features through a bidirectional long short-term memory neural network BiLSTM to obtain character embeddings; then concatenate the character embeddings with the word embeddings to obtain a word representation containing word-level and character-level information.

5. The medical entity disambiguation method based on Bio-LinkBERT and context awareness according to claim 1, wherein Use a cross-attention mechanism for the mentioned and candidate entities to capture the relationship between mention-candidate entity and candidate entity-mention, specifically: Obtain attention in two directions, from mentions to candidate entities and from candidate entities to mentions; these two attentions are obtained through a shared similarity matrix, where the elements in the matrix represent the similarity between mention tokens and entity name tokens; calculate the attentions in the two directions using the similarity matrix.

6. The method for medical entity disambiguation based on Bio-LinkBERT and context awareness according to claim 5, characterized in that, Use BiGRU to capture specific pre- or post- features in the mention-candidate entity, enhance semantic associations, and obtain a more informative word representation of the mention-candidate entity.

7. The method for medical entity disambiguation based on Bio-LinkBERT and context awareness according to claim 1, wherein Use a context-aware mechanism to obtain the relevance between the mention context and the candidate entity through the self-attention mechanism, specifically: Use the ELMo model with two layers of bidirectional long short-term memory neural networks (BiLSTM) to encode the candidate entity and the mention context, obtaining a candidate entity representation and a preliminary mention context representation; Assign weights to each token in the mention context through the self-attention mechanism, calculate the weighted sum to obtain the final mention context representation; then calculate the context score as the dot product of the candidate entity representation and the final mention context representation.

8. The method for medical entity disambiguation based on Bio-LinkBERT and context awareness according to claim 1, wherein After concatenating the context score with the result output, input it into a two-layer feed-forward neural network.

9. The method for medical entity disambiguation based on Bio-LinkBERT and context awareness according to claim 1, wherein Randomly select positive samples from the training set and negative samples from the generated candidate entities; use the hinge loss as the loss function, and the purpose of the hinge loss function is to optimize the embedding space.