Method, apparatus, and storage medium for normalizing biomedical entity mentions

By generating and expanding candidate concepts using a biomedical dictionary and determining semantic similarity, the method addresses the challenges of accurate biomedical entity mention normalization, enhancing the mapping process to canonical concepts.

JP7707949B2Active Publication Date: 2025-07-15FUJITSU LTD
View PDF 13 Cites 0 Cited by

Patent Information

Application Number
JP2022009594
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-02-04
Filing Date
2022-01-25
Publication Date
2025-07-15
Estimated Expiration
2042-01-25

AI Technical Summary

Technical Problem

Conventional entity mention normalization methods in biomedical literature face challenges in accurately determining the correct candidate concepts due to similarities in generated names and context ambiguity, leading to difficulties in correctly mapping biomedical entity mentions to their corresponding identifiers in knowledge graphs.

Method used

A method and apparatus that utilize a biomedical dictionary to generate candidate concepts, expand them based on related concepts, and determine semantic similarity to map mentions to the most similar concept, incorporating synonyms and superordinate concepts to improve accuracy.

Benefits of technology

Enhances the accuracy of determining canonical concepts in biomedical entity mentions by leveraging semantic similarity and hierarchical relationships, improving the mapping process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007707949000001
    Figure 0007707949000001
  • Figure 0007707949000002
    Figure 0007707949000002
  • Figure 0007707949000003
    Figure 0007707949000003
Patent Text Reader

Abstract

To provide a method for normalizing a biomedical entity mention, a device and a storage medium.SOLUTION: A method comprises the steps of: receiving, as a mention to be mapped, a biomedical entity mention, then searching a biomedical dictionary, for generating a candidate concept collection of the mention to be mapped; determining whether or not the same concept as the mention to be mapped is included in the candidate concept collection; expanding, when the same concept is not included, the candidate concept collection on the basis of a related concept collection acquired from the biomedical dictionary for each candidate concept, for updating the candidate concept collection; determining semantic similarity between each candidate concept in the updated candidate concept collection and the mention to be mapped, then acquiring semantic similarity collection; and mapping the mention to be mapped to the candidate concept which corresponds to maximum semantic similarity in the semantic similarity collection. According to the method, a device and a storage medium of the disclosure, it is possible to improve accuracy of determination of a normal concept.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to knowledge discovery, and more specifically, to a method, an apparatus, and a storage medium for normalizing biomedical entity mentions.

Background Art

[0002] With the rapid development of technologies in the biomedical field, various biomedical literature such as scientific and technical papers and patent documents are increasing day by day. This has promoted the development of text mining technology in the biomedical field. Biomedical terms mentioned in the literature are referred to as biomedical entity mentions. Text mining technology includes the normalization of biomedical entity mentions. The purpose of the normalization task of biomedical entity mentions is to determine the corresponding unique identifier in the knowledge graph of the entity mention in the biomedical literature and establish the association between the entity mention and the knowledge graph. Establishing this association is important in the technical research of the biomedical field.

[0003] Conventional entity mention normalization methods usually include two modules, namely candidate generation and candidate rearrangement. Conventional entity mention normalization methods have achieved good results in the normalization of biomedical entities, but there are still certain limitations. First, since the generated candidate names are similar, it is difficult to determine the correct candidate based only on the candidate names. Second, since the candidate and the mention are entities within the same field and information such as the context of the candidate is also similar, it is difficult to correctly rearrange the candidates even using the context information.

Summary of the Invention

Problems to be Solved by the Invention

[0004] The following provides a brief overview of the present disclosure to facilitate a basic understanding of the aspects of the present disclosure. It should be noted that this brief overview is not an exhaustive overview of the present disclosure, nor is it intended to specifically identify the key points or important parts of the present disclosure, nor is it intended to limit the scope of the present disclosure. Instead, it is for the purpose of simply explaining the concepts in a simple form as a preamble to the more detailed description to follow.

[0005] The present disclosure provides a method, an apparatus, and a storage medium for normalizing biomedical entity mentions.

Means for Solving the Problem

[0006] In one aspect of the present disclosure, there is provided a method for normalizing biomedical entity mentions, which is executed by a computer, the method including: receiving the mention to be mapped of the biomedical entity mention; searching a biomedical dictionary to generate a set of candidate concepts for the mention to be mapped; determining whether the set of candidate concepts includes a concept identical to the mention to be mapped; if the set of candidate concepts does not include a concept identical to the mention to be mapped, expanding the set of candidate concepts based on a set of related concepts obtained from the biomedical dictionary for each candidate concept in the set of candidate concepts, updating the set of candidate concepts, determining the semantic similarity between each candidate concept in the updated set of candidate concepts and the mention to be mapped to obtain a set of semantic similarities, and mapping the mention to be mapped to the candidate concept corresponding to the maximum semantic similarity in the set of semantic similarities, wherein the set of related concepts includes synonymous concepts of the corresponding candidate concept, and if the corresponding candidate concept has a corresponding superordinate concept, the set of related concepts further includes the corresponding superordinate concept and synonymous concepts of the corresponding superordinate concept.

[0007] In one aspect of the present disclosure, an apparatus for normalizing biomedical entity mentions, comprising: a memory storing instructions; and one or more processors communicatively coupled to the memory and configured to execute the instructions retrieved from the memory, the instructions causing the one or more processors to: receive a mention to be mapped as the biomedical entity mention; search a biomedical dictionary to generate a set of candidate concepts for the mention to be mapped; determine whether the set of candidate concepts contains a concept identical to the mention to be mapped; if the set of candidate concepts does not contain a concept identical to the mention to be mapped, expand the set of candidate concepts based on a set of related concepts retrieved from the biomedical dictionary for each candidate concept in the set of candidate concepts to update the set of candidate concepts, determine a semantic similarity degree between each candidate concept in the updated set of candidate concepts and the mention to be mapped to obtain a set of semantic similarity degrees, and map the mention to be mapped to a candidate concept corresponding to the maximum semantic similarity degree in the set of semantic similarity degrees, wherein the set of related concepts includes synonyms of the corresponding candidate concept, and if the corresponding candidate concept has a corresponding superordinate concept, the set of related concepts further includes the corresponding superordinate concept and synonyms of the corresponding superordinate concept.

[0008] In another aspect of the present disclosure, there is provided a computer-readable storage medium storing a program, the program causing a computer to perform steps of: receiving, as a mention to be mapped, a biomedical entity mention; searching a biomedical dictionary to generate a set of candidate concepts for the mention to be mapped; determining whether the set of candidate concepts includes a concept identical to the mention to be mapped; if the set of candidate concepts does not include a concept identical to the mention to be mapped, expanding the set of candidate concepts based on a set of related concepts obtained from the biomedical dictionary for each candidate concept in the set of candidate concepts to update the set of candidate concepts, determining a semantic similarity degree between each candidate concept in the updated set of candidate concepts and the mention to be mapped to obtain a set of semantic similarity degrees, and mapping the mention to be mapped to a candidate concept corresponding to the maximum semantic similarity degree in the set of semantic similarity degrees; wherein the set of related concepts includes synonyms of the corresponding candidate concept, and if the corresponding candidate concept has a corresponding superordinate concept, the set of related concepts further includes the corresponding superordinate concept and synonyms of the corresponding superordinate concept.

[0009] As at least one advantageous effect of the method, apparatus, and storage medium for normalizing biomedical entity mentions according to the present disclosure, according to the method, apparatus, and storage medium of the present disclosure, the accuracy of determining a canonical concept can be improved.

Brief Description of the Drawings

[0010] To more easily understand the above and other objects, features, and advantages of the present disclosure, embodiments of the present disclosure will be described below with reference to the drawings. It should be noted that the drawings are only for explaining the principle of the present disclosure. In the drawings, it is not necessary to draw the sizes and relative positions of each part according to a scale. The same reference numerals may represent the same features.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

DETAILED DESCRIPTION OF THE INVENTION

[0011] Exemplary embodiments of the present disclosure will be described below with reference to the drawings. For convenience of explanation, not all features of the actual embodiments are shown in the specification. It should be noted that when those skilled in the art implement the embodiments, they may make specific decisions to implement the embodiments, and these decisions may be changed according to the embodiments.

[0012] It should be noted that, for clarity of the present disclosure, only the components of the apparatus and / or the processing steps closely related to the present disclosure are shown in the drawings, and details not related to the present disclosure are omitted.

[0013] It should be noted that the present disclosure is not limited to the described embodiments for the following description with reference to the accompanying drawings. In this specification, when feasible, embodiments may be combined with each other, features of different embodiments may be replaced or utilized, or one or more features may be omitted in one embodiment.

[0014] The computer program code for performing the operations of each aspect of the exemplary embodiments disclosed herein may be written in any combination of one or more programming languages, which may include object-oriented programming languages such as Java, Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" programming language or similar programming languages.

[0015] The method of the present disclosure may be implemented by a circuit having a corresponding functional configuration. The circuit includes a circuit for a processor.

[0016] One aspect of the present disclosure provides a method for normalizing biomedical entity mentions executed by a computer. The following exemplarily describes the method with reference to FIG. 1.

[0017] FIG. 1 is a flowchart exemplarily showing a method 100 for normalizing biomedical entity mentions executed by a computer according to one embodiment of the present disclosure. The method 100 can output a normalization result for the biomedical entity mention related to the biomedical dictionary, for example, based on the input biomedical entity mention. The normalization result can indicate the corresponding canonical representation of the biomedical entity mention shown in the biomedical dictionary.

[0018] In step S101, a mention mm to be mapped of a biomedical entity mention is received. mm is, for example, a biomedical term mentioned in a biomedical document. A term extraction application may be used to extract biomedical terms appearing in an image or text of a biomedical document. An example of the mention mm to be mapped is "sporadic t-cell leukemia". The mention mm to be mapped is, for example, an English term or a term of other Indo-European language families.

[0019] In step S103, the biomedical dictionary D is searched to generate a candidate concept set Sc of mentions mm to be mapped. Sc = {mc[1], mc[2], …, mc[i], …, mc[i_max]}, where i_max is the number of elements in the candidate concept set. For example, using the full-text search engine Lucene tool, indexes are established for each concept and its identifier in the biomedical dictionary D (biomedical knowledge base). Next, the Lucene tool is used to perform a search with the indexes, and the top 20 search results are used as candidate concepts to form the candidate concept set. The upper limit of the elements in the candidate concept set may be restricted. For example, the upper limit is 20. If the number of retrieved candidate concepts has not reached the upper limit, the candidate concept set is constructed with all the retrieved candidate concepts. Each candidate concept mc[i] has an index related to the biomedical dictionary D, and this index is denoted as mc[i].ind. The concepts in the biomedical dictionary D are already recorded standardized biomedical terms. The biomedical dictionary D is, for example, an English dictionary or another Indo-European language dictionary.

[0020] In step S105, it is determined whether the candidate concept set Sc contains a concept identical to the mention to be mapped. For example, for the mention mm = "sporadic t-cell leukemia" to be mapped, if the candidate concept mc[i] = "sporadic t-cell leukemia" is included in Sc, it is determined that the candidate concept set Sc contains a concept identical to the mention to be mapped.

[0021] If the decision result is "NO", step S107 is executed. In step S107, for each candidate concept mc[i] in the candidate concept set Sc, the candidate concept set mc[i] is extended based on the related concept set Sr[i] obtained from the biomedical dictionary D to update the candidate concept set. Here, the related concept set includes synonymous concepts of the corresponding candidate concept. If the corresponding candidate concept has a corresponding upper concept, the related concept set further includes the corresponding upper concept and synonymous concepts of the corresponding upper concept. For example, the related concept set Sr[i] of the candidate concept mc[i] includes all the synonymous concepts sys[i][j] shown in the biomedical dictionary D of the candidate concept mc[i]. j may be from 1 to the number of synonymous concepts j_max. Also, if the candidate concept mc[i] has an upper concept, the related concept set Sr[i] further includes the upper concept mp[i] of the candidate concept mc[i] and all the synonymous concepts syp[i][k] shown in the biomedical dictionary D of the upper concept mp[i]. k may be from 1 to the number of synonymous concepts k_max. In the biomedical dictionary, there is a hierarchical relationship between concepts. This means that there is an upper concept or a lower concept of the concept, or both. Here, the lower concept includes information about the upper concept. For example, the concept "lymphoma, non - hodgkins" is a type of "lymphoma". Here, "lymphoma, non - hodgkins" is a lower concept of "lymphoma", and "lymphoma" is an upper concept of "lymphoma, non - hodgkins". In each concept item in the biomedical dictionary D, the synonymous concepts of the concept and its upper concept (if it exists) are recorded. If there are multiple upper concepts of mc[i], preferably, the related concept set Sr[i] includes these multiple upper concepts and further includes the synonymous concepts of each of these multiple upper concepts. Preferably, if the candidate concept mc[i] has at least one upper concept, the related concept set Sr[i] further includes the at least one upper concept and the synonymous concepts of each concept among the at least one upper concept.

[0022] In step S109, the semantic similarity sm[i] between each candidate concept mc[i] in the updated candidate concept set and the mention mm to be mapped is determined, and a semantic similarity set Ss is obtained.

[0023] In step S111, the mention mm to be mapped is mapped to the candidate concept mc[i_smax] corresponding to the maximum semantic similarity in the semantic similarity set Ss. Specifically, using the index mc[i_smax].ind of the candidate concept mc[i_smax] corresponding to the maximum semantic similarity in the semantic similarity set Ss in the biomedical dictionary D, the mention mm to be mapped is mapped to the corresponding concept mc[i_smax] in the biomedical dictionary D, indicating the canonical concept of the mention mm to be mapped. For example, the index mm.ind of the mention mm to be mapped in the biomedical dictionary D is set to mc[i_smax].ind, mm is mapped to the candidate concept mc[i_smax] corresponding to the maximum semantic similarity, and the canonical concept of the mention mm to be mapped is set to the candidate concept mc[i_smax] corresponding to the maximum semantic similarity.

[0024] If the determination result in step S105 is YES, step S113 is executed. In step S113, the mention mm to be mapped is mapped to the same candidate concept mc[i_s] in the candidate concept set Sc. For example, the index mm.ind of the mention mm to be mapped in the biomedical dictionary D is set to mc[i_s].ind, and mm is mapped to the same candidate concept mc[i_s].

[0025] In one embodiment, the step of searching the biomedical dictionary D to generate a set of candidate concepts for the mention to be mapped includes a step of preprocessing the mention mm to be mapped. Here, the preprocessing may include at least one of converting an abbreviation in the mention to be mapped to its formal name, replacing non-Arabic numerals in the mention to be mapped with Arabic numerals, and converting a plural mention to be mapped to a singular entity. Preferably, the preprocessing may further include converting the first character of the mention mm to be mapped to the corresponding lower case character if the first character of the mention mm to be mapped is in upper case.

[0026] In one embodiment of the present disclosure, candidate concepts may be extended based on word frequency and the mention to be mapped. FIG. 2 is a flowchart exemplarily showing a method 200 for diffusing candidate concepts according to one embodiment of the present disclosure. For each candidate concept (i.e., the selected candidate concept) in the set of candidate concepts, the method 200 may be executed, and the set of candidate concepts may be updated by expanding the set of candidate concepts based on the set of related concepts obtained from the biomedical dictionary.

[0027] In step S203, based on the word frequencies of the words in the related concept set Sr[i] of the selected candidate concept mc[i] in the candidate concept set, a word sequence Sw[i] sorted in descending order of word frequency is determined. When the selected candidate concept mc[i] has a superordinate concept mp[i], in the related concept set Sr composed of all the synonymous concepts (sys[i][1],…,sys[i][j],…,sys[i][j_max]) of the selected candidate concept, its superordinate concept mp[i], and all the synonymous concepts (syp[i][1],…,sys[i][k ],…,sys[i][k_max]) of the superordinate concept, the same word may appear in multiple concepts. When the selected candidate concept mc[i] has no superordinate concept, in the related concept set Sr composed of all the synonymous concepts (sys[i][1],…,sys[i][j],…,sys[i][j_max]) of the selected candidate concept, the same word may also appear in multiple concepts. Note that the number of superordinate concepts of mc[i] may be plural. It is also possible to select the first superordinate concept among the multiple superordinate concepts shown in D and construct Sr. The related concept set Sr preferably includes other superordinate concepts other than mc[i] and synonymous concepts of other superordinate concepts. The word frequency of each word (for example, a total of 30 different words) in the related concept set is determined by counting the number of times the word appears in the related concept set. Thereby, the original sequence So[i] (for example, the original sequence including 30 words) composed of the words sorted in descending order of word frequency can be determined. The word sequence Sw[i] may be set to be the same as the original sequence So[i], or when the number of words in the original sequence is more than N (N is a natural number, for example, N = 5), the word sequence Sw[i] may be set to the sequence composed of the first N words of the original sequence So[i]. Alternatively, the word sequence Sw[i] may include only the words whose word frequencies in the original sequence are greater than the word frequency threshold. In other words, the length (i.e., the number of words) of the word sequence Sw[i] may be restricted based on the word frequency threshold or the sequence length threshold.

[0028] In step S205, the word sequence Sw[i] is updated. For example, the word sequence Sw[i] is updated based on the mention mm to be mapped. Specifically, each word in mm may be checked to determine whether the word appears in the word sequence Sw[i]. For the words that appear, in the word sequence Sw[i], move the word forward and update the word sequence Sw[i]. For example, move the words that appear in the mention mm to be mapped in the word sequence Sw[i] before the first word of the word sequence Sw[i]. That is, set the appeared word as the first word of the word sequence. If there are multiple words that appear in the mention mm to be mapped in the word sequence Sw[i], move all of these words to the previous position of the word sequence Sw[i]. For example, execute the forward movement so that these words occupy the first position, the second position, and the third position of the word sequence (when there are three words that appear in the mention mm to be mapped in the word sequence Sw[i]). In one example, if the word being checked in the current mm is already the first word of the current word sequence Sw[i], proceed to check the next word in mm (that is, there is no need to execute the forward movement). For example, if Sw[i] is w1, w2, w3, w4, w5 and w4 appears in mm, execute the forward movement and Sw[i] is updated to w4, w1, w2, w3, w5.

[0029] In one example, the word sequence Sw[i] determined in step S203 may be the same as the original sequence So[i]. In step S203, after updating the word sequence Sw[i] based on the mention mm to be mapped, only the first N words of the word sequence Sw[i] are truncated to update the word sequence. That is, updating the word sequence may further include updating the word sequence based on the sequence length threshold.

[0030] The word sequence Sw[i] output in step S205 may be represented as w[i][1], w[i][2], …, w[i][n], …, w[i][n_max].

[0031] In step S207, an attempt is made to expand the selected candidate concept mc[i] by selecting a word from the word sequence Sw[i] based on the selected candidate concept mc[i] and adding it to the end of the selected candidate concept mc[i]. The sequence position pointer P is used to indicate the number of the word to be checked in the word sequence Sw[i] (the initial value is set to 1), and it may be checked whether the word indicated by the sequence position pointer P in the word sequence Sw[i] is in the selected candidate concept mc[i]. If the check result is "NO", the word is added to the end of the selected candidate concept mc[i], and P is updated to P + 1. For example, if mc[i] = "b-cell lymphoma" and the word to be added is "non-hodgkins", mc[i] after expansion is "b-cell lymphoma non-hodgkins". In one example, step S207 includes selecting the first word that does not appear in the selected candidate concept mc[i] in the word sequence Sw[i] and adding it to the end of the selected candidate concept to update the selected candidate concept. For example, if w[i][1] is not included in mc[i], mc[i] is updated to mc[i]+space+w[i][1]. Here, space represents a space. If the check result is "YES", P is updated to P + 1, and step S209 is executed (that is, the word indicated by P is not used for expansion).

[0032] In step S209, it is determined whether the number of words of the selected candidate concept mc[i] is greater than the length threshold Lth. If the determination result is "YES", the expansion of the selected candidate concept ends. The length threshold is determined, for example, in the manner of Lth = min{2 * the number of words of the mention to be mapped, M}. Here, M is the maximum number of words of the concepts in D, and min{} is the minimum parameter among the selection parameters. Here, "2" is just an example and may be adjusted according to experience.

[0033] In step S211, when word order is ignored, it is determined whether the selected candidate concept and the mention to be mapped are the same. If the determination result is "YES", the expansion of the selected candidate concept ends. In one variant, the same determination (S211) may be performed first, and then the word count determination (S209) may be performed.

[0034] In step S213, it is determined whether the check of the word sequence has advanced to the end of the word sequence. For example, this may be realized by checking whether the current sequence position pointer P has reached the end of the word sequence (whether P is equal to n_max + 1). When P = n_max + 1, it means that the check of all words in the word sequence Sw[i] is completed and the check has advanced to the end of the word sequence.

[0035] In step S215, it is determined whether there is a superordinate concept above the selected candidate concept mc[i]. This may be checked by referring to the dictionary D to see whether the dictionary D indicates a superordinate concept above the superordinate concept. If it is determined that there is no superordinate concept above the selected candidate concept mc[i], step S219 is executed, and for example, prompt information may be output to notify the user so that the user can manually determine the processing method. Preferably, if it is determined that there is no superordinate concept above the selected candidate concept mc[i], it may be set to end the expansion of the selected candidate concept.

[0036] In step S217, update the word sequence based on a higher-level superordinate concept. For example, update Sw[i] by constructing a sequence of words sorted in descending order of word frequency using the higher-level superordinate concept and its synonymous concepts. Similarly, the upper limit of the number of words included in the word sequence Sw[i] may be restricted based on a word frequency threshold or a sequence length threshold. For example, if the original word sequence contains 30 words, construct Sw[i] using the first 5 words. For example, Sw[i] may be constructed using only words with a word frequency greater than 1. Then, return to step S205.

[0037] That is, method 200 may include a step of conditionally expanding the selected candidate concept based on a higher-level superordinate concept of the selected candidate concept.

[0038] The following describes a method for obtaining a semantic similarity set according to the present disclosure.

[0039] The semantic similarity set may be represented as Ss = {sm[1], …, sm[i], … sm[i_max]}. Here, sm[i] represents the semantic similarity between the candidate concept mc[i] and the mention mm to be mapped.

[0040] In one example, sm[i] may be determined using a conventional semantic similarity calculation method. For example, determine the feature vectors of the corresponding two word strings, and directly calculate the similarity between these two feature vectors as the corresponding semantic similarity.

[0041] In one embodiment of the present disclosure, an attention matrix and a convolutional neural network model are used to determine each semantic similarity in the semantic similarity set Ss. Hereinafter, a method for determining the semantic similarity will be exemplarily described with reference to FIG. 3.

[0042] FIG. 3 is a flowchart exemplarily showing a method 300 for determining semantic similarity according to one embodiment of the present disclosure.

[0043] In step S301, an attention matrix A is determined based on the concept vector Fcv of the target candidate concept mc[i] and the mention vector Fmv of the mention mm to be mapped. The element a of the attention matrix A uv =match_score(Vwm[u],Vwc[v]), where Vwm[u] is the word vector of the u-th word of mm, Vwc[v] is the word vector of the v-th word of mc[i], and match_score(wm[u],wc[v]) represents the matching degree between the word wc[v] in the target candidate concept mc[i] and the word wc[u] in the mention mm to be mapped. Vwm[u] is a component of Fmv, and Vwc[v] is a component of Fcv. The matching degree between two words may be defined according to a predetermined rule. Generating word vectors based on words is a prior art (for example, generating word vectors through a language model such as Word2Vec), and the description thereof is omitted here.

[0044] In step S302, an attention feature vector Fma of the mention mm to be mapped and an attention feature vector Fca of the target candidate concept mc[i] are determined based on the attention matrix A. Here, F ma =W0·A T 、F ca =W1·A, where T represents a transpose transformation. W0 and W1 are parameters that need to be learned when training the model.

[0045] In step S303, a mention feature Fm of the mention mm to be mapped is generated based on the mention vector Fmv of the mention mm to be mapped and the attention feature vector Fma using a convolutional neural network layer CNN. In one example, Fma and Fmv are concatenated to form a feature Fm' (i.e., Fm' = Fma + Fmv), the feature Fm' is input into the convolutional neural network layer CNN, and after processing Fm' by CNN, Fm is output. This may be expressed as Fm = CNN(Fm').

[0046] In step S304, a convolutional neural network layer CNN is used to generate the concept feature Fc of the target candidate concept mc[i] based on the concept vector Fcv and the attention feature vector Fca of the target candidate concept mc[i]. In one example, Fca and Fcv are concatenated to form a feature Fc’ (i.e., Fc’ = Fca + Fcv), the feature Fc’ is input into the convolutional neural network layer CNN, and after Fc’ is processed by the CNN, Fc is output. This may be expressed as Fc = CNN(Fc’).

[0047] In step S305, at least one hidden layer is used to generate the depth feature Fd based on the mention feature Fm and the concept feature Fc. In one example, Fm and Fc are concatenated to form a feature Fd’ (i.e., Fd’ = Fm + Fc), the feature Fd’ is input into the hidden layer, and after Fd’ is processed by the hidden layer, Fd is output. This may be expressed as Fd = Hidden(Fd’).

[0048] In step S306, a Softmax layer is used to determine the semantic similarity sm[i] between the target candidate concept mc[i] and the mention mm to be mapped based on the depth feature Fd. In one example, the output of the Softmax layer is a two-dimensional vector, and one component in the two-dimensional vector is the semantic similarity between the target candidate concept and the mention to be mapped.

[0049] As can be understood by those skilled in the art, before determining semantic similarity using a convolutional neural network model, it is necessary to train the convolutional neural network model using samples and determine the parameters of the convolutional neural network model. At the stage of training the convolutional neural network model, for each entity mention with a normalized result in the training corpus, a set of candidate concepts is generated using the candidate concept generation method described in the present disclosure. When the candidate concept matches the labeled normalized result of the entity mention, the semantic similarity (one component of the output of the Softmax layer, corresponding to the similarity score) of the pair <mm, mc[i]> of the mention and the candidate is labeled as 1, and otherwise it is labeled as 0. At the test stage, sorting is performed using the similarity score, and the candidate concept with the highest similarity score is selected as the normalized result.

[0050] The present disclosure further provides a normalization device for biomedical entity mentions. The following exemplarily describes the device with reference to FIG. 4. FIG. 4 is a block diagram exemplarily showing a normalization device 400 for biomedical entity mentions according to one embodiment of the present disclosure. The device 400 includes a search unit 403, a determination unit 405, an update unit 407, an acquisition unit 409, and a mapping unit 411. The search unit 403 searches a biomedical dictionary and generates a set of candidate concepts for mentions to be mapped as biomedical entity mentions. The determination unit 405 determines whether the set of candidate concepts contains the same concept as the mention to be mapped. The update unit 407, when the set of candidate concepts does not contain the same concept as the mention to be mapped, expands the set of candidate concepts based on the set of related concepts obtained from the biomedical dictionary for each candidate concept in the set of candidate concepts, and updates the set of candidate concepts. Here, the set of related concepts includes synonyms of the corresponding candidate concept, and when the corresponding candidate concept has a corresponding superordinate concept, the set of related concepts further includes the corresponding superordinate concept and synonyms of the corresponding superordinate concept. The acquisition unit 409 determines the semantic similarity between each candidate concept in the updated set of candidate concepts and the mention to be mapped, and obtains a set of semantic similarities. The mapping unit 411 maps the mention to be mapped to the candidate concept corresponding to the maximum semantic similarity in the set of semantic similarities. The device 400 corresponds to method 100. For further configurations of the device 400, reference may be made to the description of the method of the present disclosure. For example, when the determination result is "NO", the mapping unit 411 maps the mention to be mapped to the same candidate concept in the set of candidate concepts.

[0051] The present disclosure further provides an apparatus for normalizing biomedical entity mentions. The following exemplarily describes the apparatus with reference to FIG. 5. FIG. 5 is a block diagram exemplarily showing a normalization apparatus 500 for biomedical entity mentions according to one embodiment of the present disclosure. The apparatus 500 includes a memory 501 in which instructions are stored, and one or more processors 503 that can communicate with the memory 501 and execute the instructions obtained from the memory 501. The instructions cause the one or more processors 503 to perform the steps of receiving, as a mention to be mapped, a biomedical entity mention; searching a biomedical dictionary to generate a set of candidate concepts for the mention to be mapped; determining whether the set of candidate concepts includes a concept identical to the mention to be mapped; if the set of candidate concepts does not include a concept identical to the mention to be mapped, expanding the set of candidate concepts based on a set of related concepts obtained from the biomedical dictionary for each candidate concept in the set of candidate concepts, updating the set of candidate concepts, determining a semantic similarity degree between each candidate concept in the updated set of candidate concepts and the mention to be mapped to obtain a set of semantic similarity degrees, and mapping the mention to be mapped to a candidate concept corresponding to the maximum semantic similarity degree in the set of semantic similarity degrees. Here, the set of related concepts includes synonyms of the corresponding candidate concept, and if the corresponding candidate concept has a corresponding superordinate concept, the set of related concepts further includes the corresponding superordinate concept and synonyms of the corresponding superordinate concept. The apparatus corresponds to the method for normalizing biomedical entity mentions of the present disclosure. For further configurations of the apparatus, reference may be made to the description of the method 100 of the present disclosure.

[0052] One aspect of the present disclosure provides a computer-readable storage medium storing a program. The program causes a computer to perform steps of receiving, as a mention to be mapped, a biomedical entity mention; searching a biomedical dictionary to generate a set of candidate concepts for the mention to be mapped; determining whether the set of candidate concepts includes a concept identical to the mention to be mapped; if the set of candidate concepts does not include a concept identical to the mention to be mapped, expanding the set of candidate concepts based on a set of related concepts obtained from the biomedical dictionary for each candidate concept in the set of candidate concepts to update the set of candidate concepts, determining a semantic similarity degree between each candidate concept in the updated set of candidate concepts and the mention to be mapped to obtain a set of semantic similarity degrees, and mapping the mention to be mapped to a candidate concept corresponding to the maximum semantic similarity degree in the set of semantic similarity degrees. Here, the set of related concepts includes synonyms of the corresponding candidate concept, and if the corresponding candidate concept has a corresponding superordinate concept, the set of related concepts further includes the corresponding superordinate concept and synonyms of the corresponding superordinate concept. For details of further configurations of the program, reference may be made to the description of the method for normalizing biomedical entity mentions of the present disclosure.

[0053] In one aspect of the present disclosure, an information processing apparatus is further provided.

[0054] FIG. 6 is a block diagram exemplarily showing an information processing apparatus according to an embodiment of the present disclosure. In FIG. 6, a central processing unit (CPU) 601 executes various processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage unit 608 to a random access memory (RAM) 603. In the RAM 603, data necessary for the CPU 601 to execute various processes is stored as needed.

[0055] The CPU 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output interface 605 is also connected to the bus 604.

[0056] The input unit 606 (including a keyboard, a mouse, etc.), the output unit 607 (including a display, such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.), the storage unit 608 (including, for example, a hard disk, etc.), and the communication unit 609 (including a network interface card, such as a LAN card, a modem, etc.) are connected to the input / output interface 605. The communication unit 609 executes communication processing via a network, such as the Internet.

[0057] Optionally, the driver 610 may be connected to the input / output interface 605. The removable medium 611 is, for example, a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., and is set up in the driver 610 as needed, and the computer program read from it is installed in the storage unit 608 as needed.

[0058] The CPU 601 may execute a program for normalizing biomedical entity mentions. The program can implement the functions of Method 100.

[0059] The technology according to the present disclosure includes determining a canonical concept of a biomedical entity mention based on extended same-level and higher-level concepts. Thereby, the accuracy of determining the canonical concept can be improved. Aspects of the present disclosure further include determining a canonical concept of a biomedical entity mention based on an attention matrix. Thereby, the accuracy of determining the canonical concept can be further improved.

[0060] As described above, according to the present disclosure, a principle of a method for normalizing biomedical entity mentions is provided. It should be noted that the effects of the present disclosure are not necessarily limited to the above effects, and in addition to the effects described in the above paragraphs, or instead of them, any of the effects shown in this specification can be obtained, or other effects can be understood from this specification.

[0061] The above has described specific embodiments of the present disclosure. However, those skilled in the art can make various changes (in the case of rows, combining or replacing the features of each embodiment), improvements or equivalents to the present disclosure within the gist and scope of the appended claims. These changes, improvements or equivalents belong to the protection scope of the present disclosure. For example, in Method 200, the following exemplary modifications may be made. After it is determined in step S207 that the expansion condition is not satisfied for the current P, if P+1 does not exceed the range of the word sequence, P is updated, the expansion is attempted again, and after the expansion attempt is successful, step S209 is executed. If P+1 exceeds the range of the word sequence, step S215 is executed.

[0062] The above has illustratively described a method for normalizing terms in the biomedical field. It should be noted that the above scheme may be used to normalize terms in other fields (for example, the chemical field) after simple adaptive adjustments are made (for example, selecting a dictionary for the corresponding field).

[0063] It should be noted that the terms "comprising" and "having" mean the presence of the features, elements, steps or members described in this specification, but do not exclude the presence or addition of one or more other features, elements, steps or members.

[0064] Furthermore, the methods of each embodiment of the present invention are not limited to being executed in the order of time described in the specification or shown in the drawings. They may be executed in other orders of time, or may be executed in parallel or independently. Therefore, the execution order of the methods described in this specification does not limit the technical scope of the present invention.

[0065] In addition, regarding the embodiments including the above-described respective embodiments, the following supplementary notes are further disclosed, but are not limited to these supplementary notes. (Supplementary Note 1) A method for normalizing biomedical entity mentions executed by a computer, comprising: Receiving the biomedical entity mention as a mention to be mapped; Searching a biomedical dictionary to generate a set of candidate concepts for the mention to be mapped; Determining whether the set of candidate concepts contains a concept identical to the mention to be mapped; If the set of candidate concepts does not contain a concept identical to the mention to be mapped, expanding the set of candidate concepts based on a set of related concepts obtained from the biomedical dictionary for each candidate concept in the set of candidate concepts to update the set of candidate concepts; determining a semantic similarity degree between each candidate concept in the updated set of candidate concepts and the mention to be mapped to obtain a set of semantic similarity degrees, and mapping the mention to be mapped to a candidate concept corresponding to the maximum semantic similarity degree in the set of semantic similarity degrees, and a method including: the set of related concepts includes synonyms of the corresponding candidate concept, when the corresponding candidate concept has a corresponding superordinate concept, the set of related concepts further includes the corresponding superordinate concept and synonyms of the corresponding superordinate concept. (Appendix 2) The step of searching a biomedical dictionary to generate a set of candidate concepts for the mention to be mapped includes performing preprocessing on the mention to be mapped, the preprocessing includes converting an abbreviation in the mention to be mapped to a formal name, replacing non-Arabic numerals in the mention to be mapped with Arabic numerals, and converting a plural mention to be mapped to a singular entity, and the method according to Appendix 1 including at least one of the above. (Appendix 3) The step of expanding the set of candidate concepts based on a set of related concepts obtained from the biomedical dictionary for each candidate concept in the set of candidate concepts to update the set of candidate concepts is Determining a word sequence sorted in descending order of word frequency based on the word frequencies of words in the associated concept set of the selected candidate concept in the candidate concept set; Updating the word sequence; Trying to expand the selected candidate concept by selecting a word from the word sequence based on the selected candidate concept and adding it to the end of the selected candidate concept, the method according to Appendix 1. (Appendix 4) The step of updating the word sequence is Setting, as the first word of the word sequence, the word that appears in the mention to be mapped in the word sequence, the method according to Appendix 3. (Appendix 5) The step of trying to expand the selected candidate concept by selecting a word from the word sequence and adding it to the end of the selected candidate concept is Updating the selected candidate concept by adding, to the end of the selected candidate concept, the first word that does not appear in the selected candidate concept in the word sequence, the method according to Appendix 3. (Appendix 6) The step of expanding the candidate concept set based on the associated concept set obtained from the biomedical dictionary for each candidate concept in the candidate concept set and updating the candidate concept set is Determining whether the number of words in the selected candidate concept is greater than a length threshold; If the number of words in the selected candidate concept is greater than the length threshold, ending the expansion of the selected candidate concept, the method according to Appendix 3. (Appendix 7) The step of expanding the candidate concept set based on the associated concept set obtained from the biomedical dictionary for each candidate concept in the candidate concept set and updating the candidate concept set is Determining whether the selected candidate concept and the mention to be mapped are the same when word order is ignored; When word order is ignored, when the selected candidate concept is the same as the mention to be mapped, the step of ending the expansion of the selected candidate concept, and the method according to Supplementary Note 3, including this step. (Supplementary Note 8) For each candidate concept in the candidate concept set, based on the set of related concepts obtained from the biomedical dictionary, the step of expanding the candidate concept set and updating the candidate concept set is as follows: Based on the hypernym of the selected candidate concept that is higher in hierarchy, the step of conditionally expanding the selected candidate concept, and the method according to Supplementary Note 3, including this step. (Supplementary Note 9) Using a convolutional neural network model to determine the similarity between each candidate concept in the updated candidate concept set and the mention to be mapped, the method according to Supplementary Note 1. (Supplementary Note 10) The convolutional neural network model determines an attention matrix based on the concept vector of the target candidate concept and the mention vector of the mention to be mapped, determines the attention feature vector of the mention to be mapped and the attention feature vector of the target candidate concept based on the attention matrix, uses a convolutional neural network layer to generate the mention features of the mention to be mapped based on the mention vector and the attention feature vector of the mention to be mapped, uses a convolutional neural network layer to generate the concept features of the target candidate concept based on the concept vector and the attention feature vector of the target candidate concept, uses at least one hidden layer to generate depth features based on the mention features and the concept features, uses a Softmax layer to determine the semantic similarity between the target candidate concept and the mention to be mapped based on the depth features, Each element of the attention matrix represents the degree of matching between a word in the target candidate concept and a word in the mention to be mapped, according to the method described in Appendix 9. (Appendix 11) The output of the Softmax layer is a two-dimensional vector, One component in the two-dimensional vector is the semantic similarity between the target candidate concept and the mention to be mapped, according to the method described in Appendix 10. (Appendix 12) An apparatus for normalizing biomedical entity mentions, A search unit that searches a biomedical dictionary to generate a set of candidate concepts for the mention to be mapped as biomedical entity mentions, A determination unit that determines whether the set of candidate concepts includes the same concept as the mention to be mapped, When the set of candidate concepts does not include the same concept as the mention to be mapped, an update unit that expands the set of candidate concepts based on a set of related concepts obtained from the biomedical dictionary for each candidate concept in the set of candidate concepts and updates the set of candidate concepts, An acquisition unit that determines the semantic similarity between each candidate concept in the updated set of candidate concepts and the mention to be mapped and obtains a set of semantic similarities, A mapping unit that maps the mention to be mapped to the candidate concept corresponding to the maximum semantic similarity in the set of semantic similarities, and includes, The set of related concepts includes synonyms of the corresponding candidate concept, When the corresponding candidate concept has a corresponding superordinate concept, the set of related concepts further includes the corresponding superordinate concept and synonyms of the corresponding superordinate concept, an apparatus. (Appendix 13) An apparatus for normalizing biomedical entity mentions, A memory in which instructions are stored, One or more processors that communicate with the memory and can execute the instructions obtained from the memory, and includes, The command causes the one or more processors to receive the biomedical entity mention as a mention to be mapped, search a biomedical dictionary to generate a set of candidate concepts for the mention to be mapped, determine whether the set of candidate concepts contains a concept identical to the mention to be mapped, if the set of candidate concepts does not contain a concept identical to the mention to be mapped, update the set of candidate concepts by expanding the set of candidate concepts based on a set of related concepts obtained from the biomedical dictionary for each candidate concept in the set of candidate concepts, determine a semantic similarity degree between each candidate concept in the updated set of candidate concepts and the mention to be mapped to obtain a set of semantic similarity degrees, and map the mention to be mapped to a candidate concept corresponding to the maximum semantic similarity degree in the set of semantic similarity degrees, and cause the execution of the steps, the set of related concepts includes synonyms of the corresponding candidate concept, when the corresponding candidate concept has a corresponding superordinate concept, the set of related concepts further includes the corresponding superordinate concept and synonyms of the corresponding superordinate concept, a device. (Appendix 14) A computer-readable storage medium storing a program, where the program causes a computer to receive the biomedical entity mention as a mention to be mapped, search a biomedical dictionary to generate a set of candidate concepts for the mention to be mapped, determine whether the set of candidate concepts contains a concept identical to the mention to be mapped, if the set of candidate concepts does not contain a concept identical to the mention to be mapped, For each candidate concept in the candidate concept set, expand the candidate concept set based on the set of related concepts obtained from the biomedical dictionary, and update the candidate concept set. Determine the semantic similarity between each candidate concept in the updated candidate concept set and the mention to be mapped to obtain a set of semantic similarities, and map the mention to be mapped to the candidate concept corresponding to the maximum semantic similarity in the set of semantic similarities. The set of related concepts includes synonymous concepts of the corresponding candidate concept. When the corresponding candidate concept has a corresponding superordinate concept, the set of related concepts further includes the corresponding superordinate concept and synonymous concepts of the corresponding superordinate concept.

Claims

1. A method for normalizing biomedical entity mentions, executed by a computer, comprising: receiving the mention to be mapped as the biomedical entity mention; searching a biomedical dictionary to generate a set of candidate concepts for the mention to be mapped; determining whether the set of candidate concepts contains a concept identical to the mention to be mapped; if the set of candidate concepts does not contain a concept identical to the mention to be mapped, expanding the set of candidate concepts based on a set of related concepts obtained from the biomedical dictionary for each candidate concept in the set of candidate concepts, updating the set of candidate concepts, determining a semantic similarity degree between each candidate concept in the updated set of candidate concepts and the mention to be mapped to obtain a set of semantic similarity degrees, and mapping the mention to be mapped to the candidate concept corresponding to the maximum semantic similarity degree in the set of semantic similarity degrees. The method further comprises: the set of related concepts includes synonyms of the corresponding candidate concept; if the corresponding candidate concept has a corresponding superordinate concept, the set of related concepts further includes the corresponding superordinate concept and synonyms of the corresponding superordinate concept.

2. The step of searching a biomedical dictionary to generate a set of candidate concepts for the mention to be mapped includes performing preprocessing on the mention to be mapped, wherein the preprocessing includes converting an abbreviation in the mention to be mapped to its formal name, replacing non-Arabic numerals in the mention to be mapped with Arabic numerals, and converting a plural mention to be mapped to a singular entity, at least one of which is included in the method according to Claim 1.

3. The step of expanding the set of candidate concepts based on a set of related concepts obtained from the biomedical dictionary for each candidate concept in the set of candidate concepts and updating the set of candidate concepts includes: determining a word sequence sorted in descending order of word frequency based on the word frequency of words in the set of related concepts of a selected candidate concept in the set of candidate concepts; updating the word sequence. Attempting to expand the candidate concept by selecting a word from the word sequence based on the selected candidate concept and adding it to the end of the selected candidate concept, the method according to claim 1, comprising the step of.

4. The step of updating the word sequence is The method according to claim 3, comprising the step of setting the word that appears in the mention to be mapped in the word sequence as the first word of the word sequence.

5. The step of attempting to expand the candidate concept by selecting a word from the word sequence and adding it to the end of the selected candidate concept is The method according to claim 3, comprising the step of updating the selected candidate concept by adding the first word that does not appear in the selected candidate concept in the word sequence to the end of the selected candidate concept.

6. For each candidate concept in the candidate concept set, the step of updating the candidate concept set by expanding the candidate concept set based on the set of related concepts obtained from the biomedical dictionary is Determining whether the number of words in the selected candidate concept is greater than a length threshold; If the number of words in the selected candidate concept is greater than the length threshold, the expansion of the selected candidate concept ends, the method according to claim 3, comprising the step of.

7. Using a convolutional neural network model to determine the similarity between each candidate concept in the updated candidate concept set and the mention to be mapped, the method according to claim 1.

8. The convolutional neural network model is Determining an attention matrix based on the concept vector of the target candidate concept and the mention vector of the mention to be mapped, Determining an attention feature vector of the mention to be mapped and an attention feature vector of the target candidate concept based on the attention matrix, Generating a mention feature of the mention to be mapped based on the mention vector and the attention feature vector of the mention to be mapped using a convolutional neural network layer, Generating a concept feature of the target candidate concept based on the concept vector and the attention feature vector of the target candidate concept using the convolutional neural network layer, Generate depth features based on the mentioned features and the concept features using at least one hidden layer, Determine the semantic similarity between the target candidate concept and the mention to be mapped based on the depth features using a Softmax layer, The method according to claim 7, wherein each element of the attention matrix represents the degree of matching between a word in the target candidate concept and a word in the mention to be mapped.

9. An apparatus for normalizing biomedical entity mentions, comprising: A memory storing instructions; One or more processors communicable with the memory and operative to execute the instructions retrieved from the memory, The instructions cause the one or more processors to: Receive the biomedical entity mention as a mention to be mapped; Search a biomedical dictionary to generate a set of candidate concepts for the mention to be mapped; Determine whether the set of candidate concepts contains a concept identical to the mention to be mapped; If the set of candidate concepts does not contain a concept identical to the mention to be mapped, Update the set of candidate concepts by expanding the set of candidate concepts based on a set of related concepts retrieved from the biomedical dictionary for each candidate concept in the set of candidate concepts; Determine the semantic similarity between each candidate concept in the updated set of candidate concepts and the mention to be mapped to obtain a set of semantic similarities, and Map the mention to be mapped to the candidate concept corresponding to the maximum semantic similarity in the set of semantic similarities. The set of related concepts includes synonyms of the corresponding candidate concept, The apparatus, wherein if the corresponding candidate concept has a corresponding superordinate concept, the set of related concepts further includes the corresponding superordinate concept and synonyms of the corresponding superordinate concept.

10. A computer-readable storage medium storing a program, the program causing a computer to: Receive the biomedical entity mention as a mention to be mapped; Search a biomedical dictionary to generate a set of candidate concepts for the mention to be mapped; Determine whether the set of candidate concepts contains a concept identical to the mention to be mapped; If the candidate concept set does not contain a concept identical to the mention to be mapped, for each candidate concept in the candidate concept set, expand the candidate concept set based on the associated concept set obtained from the biomedical dictionary to update the candidate concept set, determine the semantic similarity between each candidate concept in the updated candidate concept set and the mention to be mapped to obtain a semantic similarity set, and map the mention to be mapped to the candidate concept corresponding to the maximum semantic similarity in the semantic similarity set, and execute the steps, the associated concept set includes synonyms of the corresponding candidate concept, if the corresponding candidate concept has a corresponding superordinate concept, the associated concept set further includes the corresponding superordinate concept and synonyms of the corresponding superordinate concept, storage medium.

Citation Information

Patent Citations

  • Knowledge graph-based method for constructing foreign Chinese learning contents

    CN110008354A

  • Entity linking method, electronic device and computer equipment

    CN110569328A

  • Intelligent question answering method based on subject knowledge graph and convolutional neural network

    CN111324709A

  • Method and device for retrieval term extension and method and device for information retrieval

    JP1998207896A

  • Method and apparatus for identifying terminology

    JP2011513810A