A Medical Ontology Alignment Method Based on Point Set Registration

CN113139067BActive Publication Date: 2026-10-09HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202110475715.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-29
Publication Date
2026-10-09
Estimated Expiration
2041-04-29

AI Technical Summary

Technical Problem

[0007]本发明的目的是为了解决现有方法解决医学本体数据异构性的存在,导致医学知识图谱融合的准确率低,数据匹配精度差,数据处理量大的问题,而提出一种基于PointSet Registration的医学本体对齐方法

Benefits of technology

[0056] To address the issues of low accuracy, poor data matching precision, and large data processing volume in medical knowledge graph fusion, this invention proposes a medical ontology alignment method based on Point Set Registration. This method does not require the introduction of external knowledge and uses an unsupervised algorithm (steps one through six) in the alignment step. The algorithm is simple, easy to implement, highly reliable, and requires less data processing. Furthermore, unlike typical ontology matching algorithms with fixed concept embeddings, this invention introduces the Point Set Registration algorithm into the medical ontology matching process and iteratively updates the concept embedding representation to obtain concept embeddings that maximize the optimization target. This results in high data matching precision, high reliability, and rigorous interpretability, thereby improving the accuracy of medical knowledge graph fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113139067B_ABST
    Figure CN113139067B_ABST
Patent Text Reader

Abstract

The application relates to a medical ontology alignment method based on Point Set Registration. The application aims to solve the problems of low accuracy, poor data matching precision and large data processing amount of existing medical knowledge graph fusion. The process comprises the following steps: 1, obtaining a vector representation of a concept; 2, establishing a mixed Gaussian model; 3, obtaining a transformation relationship; 4, mapping the vector representation to the same vector space through the transformation relationship; 5, in the vector space, for a concept in a group of medical ontologies, if there is an embedded post-vector of a concept in another group of medical ontologies within a given threshold radius of the embedded post-vector corresponding to the concept, then the two groups of medical ontology objects have an alignment relationship; 6, judging whether a new alignment appears, if yes, generating a new triple positive example by using the new alignment relationship, and executing step 1, if not, outputting a result. The application is used in the field of medical knowledge graphs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a medical ontology alignment method based on Point Set Registration, belonging to the field of medical knowledge graphs. Background Technology

[0002] Ontology uses concepts and relationships between concepts to represent domain knowledge, providing support for applications such as semantic annotation, knowledge discovery and sharing, data integration, and decision-making. Because medical ontology is constructed in various ways and from diverse perspectives, heterogeneity arises between different medical ontology ontologies; that is, the same concept often has different contexts and not entirely the same meaning in different medical ontology ontologies.

[0003] To integrate such a large medical ontology, an automatic medical ontology matching tool becomes an essential solution. Medical ontology alignment is a crucial technique for addressing the heterogeneity of medical ontology data and plays a significant role in the fusion of medical knowledge graphs. Its research methods mainly fall into two categories:

[0004] (1) A medical ontology alignment method based on string similarity and logical rules.

[0005] (2) A medical ontology alignment method based on concept embedding.

[0006] However, when existing methods use concept embedding to solve the medical ontology alignment problem, the concept embedding is not further optimized during the model learning process. The alignment results generated in each iteration are not used in subsequent iterations, resulting in low accuracy of medical knowledge graph fusion, poor data matching precision, and large data processing volume. Summary of the Invention

[0007] The purpose of this invention is to address the problems of low accuracy, poor data matching precision, and large data processing volume in medical ontology fusion caused by the heterogeneity of medical ontology data in existing methods. Therefore, this invention proposes a medical ontology alignment method based on PointSet Registration.

[0008] The specific process of a medical ontology alignment method based on Point Set Registration is as follows:

[0009] Step 1: Embed each concept in the two sets of medical ontology datasets to obtain the vector representation of the concept;

[0010] Step 2: Establish a Gaussian mixture model based on Step 1;

[0011] Step 3: Use the EM algorithm to solve the Gaussian mixture model obtained in Step 2 to obtain the transformation relationship T between the two sets of medical ontology datasets.θ (y m );

[0012] Step 4: Map the vector representations of the two sets of medical ontology obtained in Step 1 to the same vector space through the transformation relationship in Step 3;

[0013] Step 5: In this vector space, for a certain concept in one set of medical ontologies, if there exists an embedded vector of a concept in another set of medical ontologies within a given threshold radius of the embedded vector corresponding to that concept, then the two sets of medical ontologies have an alignment relationship.

[0014] Step 6: Determine if a new alignment occurred in Step 5. If yes, generate a new positive triplet using the new alignment and execute Step 1; otherwise, output the result.

[0015] Preferably, in step one, each concept in the two sets of medical ontology datasets is embedded to obtain a vector representation of the concept; the specific process is as follows:

[0016] Using the TransE method, the triple relations contained in the medical ontology dataset FMA are used as input to embed each concept in the medical ontology dataset FMA, resulting in a vector representation X of the concept. N×D ;

[0017] Using the TransE method, the triple relations contained in the Medical Ontology Dataset (NCI) are used as input to embed each concept in the NCI dataset, resulting in a vector representation Y of the concept. M×D ;

[0018] X N×D and Y M×D The expression is:

[0019] X N×D =(x1,…,x n ,…,x N ) T ;

[0020] Y M×D =(y1,…,y m ,…,y M ) T ;

[0021] In the formula, x1 is the vector of the concept obtained by embedding the first concept in the medical ontology dataset FMA, and x n x is the vector of the concept obtained by embedding the nth concept in the medical ontology dataset FMA. NLet y1 be the vector of the concept obtained by embedding the Nth concept in the medical ontology dataset FMA, where T is the transpose, and y1 is the vector of the concept obtained by embedding the first concept in the medical ontology dataset NCI. m y is the vector of the concept obtained by embedding the m-th concept in the NCI medical ontology dataset. M D is the vector of the concept obtained by embedding the Mth concept in the Medical Ontology dataset NCI; D is the dimension of the vector obtained by embedding a single concept in the Medical Ontology dataset; N is the size of the Medical Ontology dataset FMA; M is the size of the Medical Ontology dataset NCI.

[0022] Preferably, in step two, a Gaussian mixture model is established based on step one; the specific process is as follows:

[0023] The probability density function of the Gaussian mixture model is expressed as follows:

[0024]

[0025] In the formula, p(m) is the prior probability of the m-th Stokes model, and p(x) is the prior probability of the m-th Stokes model. n |m) represents x given the m-th Gaussian model. n The conditional probability distribution of x n The vector of the concept obtained by embedding the nth concept in the medical ontology dataset FMA, where M is the size of the ontology dataset NCI;

[0026] In the first M terms:

[0027]

[0028] In the formula, σ 2 T is the covariance of each Gaussian model in the first M Gaussian models. θ (y m ) represents the transformation relationship; D represents the dimension of the vector obtained by embedding a single concept in the medical ontology;

[0029] Item M+1:

[0030]

[0031] In the formula, N is the size of the ontology dataset FMA.

[0032] Preferably, the prior probability of the m-th Gaussian mixture model is:

[0033]

[0034] In the formula, λ is the prior of the noise ratio.

[0035] Preferably, in step three, the Gaussian mixture model obtained in step two is solved using the EM algorithm to obtain the transformation relationship T between the two sets of medical ontology datasets. θ (y m The specific process is as follows:

[0036] The Q function of the EM algorithm is defined as:

[0037]

[0038] In the formula, p(m|x n Given x n Given the conditional probability distribution of the m-th Gaussian model, x n Let p(x) be the vector of the concept after embedding the nth concept in the medical ontology dataset FMA. n |m) represents x given the m-th Gaussian model. n The conditional probability distribution of , where θ is s, R and t;

[0039] Where s is the scaling factor, R is the rotation matrix, and t is the translation vector;

[0040] According to Bayes' theorem, with vector y m The Gaussian model with respect to vector x as the centroid n The posterior probability is p(m|x) n );

[0041] p(m|x) n Substituting the Q function of the EM algorithm, we obtain the transformation relationship T. θ (y m ).

[0042] Preferably, the step of applying Bayes' theorem to a vector y m The Gaussian model with respect to vector x as the centroid n The posterior probability p(m|x) n The expression for ) is:

[0043]

[0044] Among them, T θ (y m ) represents the transformation relationship, T θ (y i ) represents the transformation relationship, y i For the ontology dataset Y M×D The vector of the concept after embedding the i-th concept, P(i) is the prior probability of the m-th Gaussian model, and P(x) is the vector of the concept after embedding the i-th concept. n |i) is the value generated by the i-th Gaussian model after it is selected. n The probability of ; c is the replacement variable.

[0045] Preferably, the transformation relationship T θ (y m The expression for ) is:

[0046] T θ (y m )=sRy m +t

[0047] In the formula, s is the scaling factor, R is the rotation matrix, and t is the translation vector.

[0048] Preferably, the expression for the replacement variable c is:

[0049]

[0050] Preferably, the given threshold radius in step five is obtained by using cosine distance.

[0051] Preferably, the specific process of generating new positive triplet examples in step six is ​​as follows:

[0052] Using the concept of alignment, for (o1,o2)∈P, generate new relational triples according to the following formula:

[0053]

[0054] In this context, o1, t1, and h1 are concepts in FMA, r1 is a relation in FMA, S1 is the set of triple relations contained in the FMA ontology dataset, o2, t2, and h2 are concepts in NCI, r2 is a relation in NCI, and S2 is the set of triple relations contained in the NCI ontology dataset.

[0055] The beneficial effects of this invention are as follows:

[0056] To address the issues of low accuracy, poor data matching precision, and large data processing volume in medical knowledge graph fusion, this invention proposes a medical ontology alignment method based on Point Set Registration. This method does not require the introduction of external knowledge and uses an unsupervised algorithm (steps one through six) in the alignment step. The algorithm is simple, easy to implement, highly reliable, and requires less data processing. Furthermore, unlike typical ontology matching algorithms with fixed concept embeddings, this invention introduces the Point Set Registration algorithm into the medical ontology matching process and iteratively updates the concept embedding representation to obtain concept embeddings that maximize the optimization target. This results in high data matching precision, high reliability, and rigorous interpretability, thereby improving the accuracy of medical knowledge graph fusion. Attached Figure Description

[0057] Figure 1 This is a flowchart of the present invention. Detailed Implementation

[0058] Specific Implementation Method 1: The specific process of this implementation method for medical ontology alignment based on Point Set Registration is as follows:

[0059] Point Set Registration is used for point cloud registration.

[0060] Step 1: Embed each concept in the two sets of medical ontology datasets to obtain the vector representation of the concept;

[0061] Step 2: For the ontology alignment problem, establish a Gaussian mixture model based on Step 1;

[0062] Step 3: Use the EM algorithm to solve the Gaussian mixture model obtained in Step 2, and obtain the transformation relationship T between the two sets of medical ontology datasets (Medical ontology dataset FMA and Medical ontology dataset NCI). θ (y m );

[0063] Step 4: Transform the two sets of medical ontology vectors obtained in Step 1 (the vector representation of the concepts obtained in Step 1 by embedding each concept in the ontology datasets FMA and NCI using the TransE method) through the transformation relation T in Step 3. θ (y m (The transformation relationship T between the two sets of ontology datasets obtained in step three) θ (y m Mapped to the same vector space;

[0064] Using the Point Set Registration algorithm described in step three, we found the relationship T between the two sets of ontologies. R,t,s (y m Then, one of the ontologies is transformed and mapped to the space of the other ontology;

[0065] Step 5: In this vector space, for a certain concept in one set of medical ontologies, if there exists an embedded vector of a concept in another set of medical ontologies within a given threshold radius of the embedded vector (the vector representation of the concept is obtained by embedding the concept), then the two sets of medical ontologies have an alignment relationship.

[0066] Step 6: Determine if a new alignment occurred in Step 5. If yes, use the new alignment to generate a new positive triplet, restart the iteration, and execute Step 1. If no, output the result.

[0067] The medical entities involved in this invention include the Basic Anatomical Model (FMA), the National Cancer Institute (NCI), and the Systematic Nomenclature of Medical-Clinical Terminology (SNOMED CT).

[0068] Specific Implementation Method Two: This implementation method differs from Specific Implementation Method One in that: in step one, each concept in the two sets of medical ontology datasets is embedded to obtain a vector representation of the concept; the specific process is as follows:

[0069] Using the TransE method, the triple relations contained in the medical ontology dataset FMA are used as input to embed each concept (e.g., Monoblast) in the medical ontology dataset FMA, resulting in a vector representation X of the concept. N×D ;

[0070] Using the TransE method, the triple relations contained in the Medical Ontology Dataset NCI are used as input to embed each concept (e.g., Chondroblast) in the Medical Ontology Dataset NCI, resulting in the vector representation Y of the concept. M×D ;

[0071] X N×D and Y M×D The expression is:

[0072] X N×D =(x1,…,x n ,…,x N ) T ;

[0073] Y M×D =(y1,…,y m ,…,y M ) T ;

[0074] In the formula, x1 is the vector of the concept obtained by embedding the first concept in the medical ontology dataset FMA, and x n x is the vector of the concept obtained by embedding the nth concept in the medical ontology dataset FMA. N Let y1 be the vector of the concept obtained by embedding the Nth concept in the medical ontology dataset FMA, where T is the transpose, and y1 is the vector of the concept obtained by embedding the first concept in the medical ontology dataset NCI. m y is the vector of the concept obtained by embedding the m-th concept in the NCI medical ontology dataset. M D is the vector of the concept obtained by embedding the Mth concept in the Medical Ontology dataset NCI; D is the dimension of the vector obtained by embedding a single concept in the Medical Ontology dataset; N is the size of the Medical Ontology dataset FMA; M is the size of the Medical Ontology dataset NCI.

[0075] The other steps and parameters are the same as in Specific Implementation Method 1.

[0076] Specific Implementation Method Three: This implementation method differs from Specific Implementation Method One or Two in that: in step two, the ontology alignment problem is addressed by establishing a Gaussian mixture model based on step one; the specific process is as follows:

[0077] Y M×D In y1,…,y M The vector is considered as the centroid of the Gaussian mixture model, X N×D x1,…,x N Vectors can be viewed as points generated by a Gaussian mixture model;

[0078] Thus, the original ontology alignment problem was transformed into the problem of solving the parameters of the Gaussian mixture model;

[0079] The probability density function of the Gaussian mixture model is expressed as follows:

[0080]

[0081] In the formula, p(m) is the prior probability of the m-th Stokes model, and p(x) is the prior probability of the m-th Stokes model. n |m) represents x given the m-th Gaussian model. n The conditional probability distribution of x n The vector of the concept obtained by embedding the nth concept in the medical ontology dataset FMA, where M is the size of the ontology dataset NCI (y1 is the centroid of the first Gaussian model, y2 is the centroid of the second Gaussian model, and so on).

[0082] in:

[0083] In the first M terms:

[0084]

[0085] In the formula, σ 2 T is the covariance of each Gaussian model in the first M Gaussian models. θ (y m ) represents the transformation relationship; D represents the dimension of the vector obtained by embedding a single concept in the medical ontology;

[0086] The (M+1)th term is a uniformly distributed noise:

[0087]

[0088] In the formula, N is the size of the ontology dataset FMA.

[0089] Other steps and parameters are the same as in specific implementation method one or two.

[0090] Specific Implementation Method Four: This implementation method differs from Specific Implementation Methods One to Three in that the prior probability of the m-th Gaussian mixture model is:

[0091]

[0092] In the formula, λ is the prior of the noise ratio.

[0093] The other steps and parameters are the same as those in one of the specific implementation methods one to three.

[0094] Specific Implementation Method Five: This implementation method differs from Specific Implementation Methods One through Four in that, in step three, the EM algorithm is used to solve the Gaussian mixture model obtained in step two, thereby obtaining the transformation relationship T between the two sets of medical ontology datasets (Medical Ontology Dataset FMA and Medical Ontology Dataset NCI). θ (y m The specific process is as follows:

[0095] The Q function of the EM algorithm is defined as:

[0096]

[0097] In the formula, p(m|x n Given x n Given the conditional probability distribution of the m-th Gaussian model, x n Let p(x) be the vector of the concept after embedding the nth concept in the medical ontology dataset FMA. n |m) represents x given the m-th Gaussian model. n The conditional probability distribution of , where θ is s, R and t;

[0098] Where s is the scaling factor, R is the rotation matrix, and t is the translation vector;

[0099] According to Bayes' theorem, with vector y m The Gaussian model with respect to vector x as the centroid n The posterior probability is p(m|x) n );

[0100] p(m|x) n Substituting the Q function of the EM algorithm, we obtain the transformation relationship T. θ (y m ).

[0101] The other steps and parameters are the same as those in one of the specific implementation methods one to four.

[0102] Specific Implementation Method Six: This implementation method differs from Specific Implementation Methods One through Five in that: the step of using the Bayesian formula with vector y... mThe Gaussian model with respect to vector x as the centroid n The posterior probability p(m|x) n The expression for ) is:

[0103]

[0104] Among them, T θ (y m ) is y m To x n The transformation relationship, T θ (y i ) is y i To x n The transformation relationship, y i For the ontology dataset Y M×D The vector of the concept after embedding the i-th concept, P(i) is the prior probability of the m-th Gaussian model, and P(x) is the vector of the concept after embedding the i-th concept. n |i) is the value generated by the i-th Gaussian model after it is selected. n The probability of ; c is the replacement variable.

[0105] The other steps and parameters are the same as those in one of the specific implementation methods one to five.

[0106] Specific Implementation Method Seven: This implementation method differs from Specific Implementation Methods One through Six in that the transformation relationship T is... θ (y m The expression for ) is:

[0107] T θ (y m )=sRy m +t

[0108] In the formula, s is the scaling factor, R is the rotation matrix, and t is the translation vector.

[0109] The other steps and parameters are the same as those in one of the specific implementation methods one to six.

[0110] Specific Implementation Method Eight: This implementation method differs from one of Specific Implementation Methods One to Seven in that the expression for the replacement variable c is:

[0111]

[0112] The other steps and parameters are the same as those in any of the specific implementation methods one to seven.

[0113] Specific Implementation Method Nine: This implementation method differs from Specific Implementation Methods One to Eight in that the given threshold radius in step five is obtained by using cosine distance.

[0114] The other steps and parameters are the same as those in one of the specific implementation methods one to eight.

[0115] Specific Implementation Method Ten: This implementation method differs from Specific Implementation Methods One through Nine in that the method for generating new positive triplet examples in step six is ​​as follows:

[0116] Using alignment relations, a concept in a triplet of another ontology set S2 is replaced with a concept from an ontology set S1 to form a new positive triplet; the process is as follows:

[0117] For example, using the concept of alignment for (o1,o2)∈P, a new relation triplet is generated according to the following formula:

[0118]

[0119] In this context, o1, t1, and h1 are concepts in FMA, r1 is a relation in FMA (e.g., subclassof), S1 is the set of triple relations contained in the FMA ontology dataset, o2, t2, and h2 are concepts in NCI, r2 is a relation in NCI (e.g., Gene_Plays_Role_In_Process), and S2 is the set of triple relations contained in the NCI ontology dataset.

[0120] The other steps and parameters are the same as those in any of the specific implementation methods one to nine.

[0121] This invention may have other embodiments. Without departing from the spirit and essence of this invention, those skilled in the art can make various corresponding changes and modifications according to this invention, but these corresponding changes and modifications should all fall within the protection scope of the appended claims.

Claims

1. A medical ontology alignment method based on Point Set Registration, characterized in that: The specific process of the method is as follows: Step 1: Embed each concept in the two sets of medical ontology datasets to obtain the vector representation of the concept; Step 2: Establish a Gaussian mixture model based on Step 1; Step 3: Use the EM algorithm to solve the Gaussian mixture model obtained in Step 2 to obtain the transformation relationship between the two sets of medical ontology datasets. ; Step 4: Map the vector representations of the two sets of medical ontology obtained in Step 1 to the same vector space through the transformation relationship in Step 3; Step 5: In this vector space, for a certain concept in one set of medical ontologies, if there exists an embedded vector of a concept in another set of medical ontologies within a given threshold radius of the embedded vector corresponding to that concept, then the two sets of medical ontologies have an alignment relationship. Step 6: Determine if a new alignment occurred in Step 5. If yes, generate a new positive triplet using the new alignment and execute Step 1; otherwise, output the result. In step one, each concept in the two sets of medical ontology datasets is embedded to obtain a vector representation of the concept; the specific process is as follows: Using the TransE method, the triple relations contained in the Medical Ontology Dataset (FMA) are used as input to embed each concept in the FMA dataset, resulting in a vector representation of the concept. ; Using the TransE method, the triple relations contained in the Medical Ontology Dataset (NCI) are used as input to embed each concept in the NCI dataset, resulting in a vector representation of the concept. ; and The expression is: ; ; In the formula, This is the vector of the concept obtained by embedding the first concept in the medical ontology dataset FMA. This refers to the vector of the concept obtained by embedding the nth concept in the medical ontology dataset FMA. To analyze the first medical ontology dataset FMA The vector of concepts obtained by embedding 1 concept, where T is the transpose. This is the vector of the concept obtained by embedding the first concept in the NCI medical ontology dataset. This is the vector of the concept obtained by embedding the m-th concept in the NCI medical ontology dataset. Let be the vector of the concept obtained by embedding the Mth concept in the NCI medical ontology dataset; D is the dimension of the vector obtained by embedding a single concept in the medical ontology. The size of the medical ontology dataset FMA; The size of the Medical Ontology dataset NCI; In step two, a Gaussian mixture model is established based on step one; the specific process is as follows: The probability density function of the Gaussian mixture model is expressed as follows: In the formula, For the first Prior probabilities of the model For a given number In the case of The conditional probability distribution, The vector of the concept obtained by embedding the nth concept in the medical ontology dataset FMA, where M is the ontology dataset. Size; In the first M terms: In the formula, It is the covariance of each Gaussian model in the first M Gaussian models. It represents the transformation relationship; D is the dimension of the vector obtained by embedding a single concept in the medical ontology; Item M+1: In the formula, N is the ontology dataset. Size; The first The prior probabilities of the Gaussian mixture model are: In the formula, It is a priori about the noise ratio; In step three, the Gaussian mixture model obtained in step two is solved using the EM algorithm to obtain the transformation relationship between the two sets of medical ontology datasets. The specific process is as follows: The Q function of the EM algorithm is defined as: In the formula, For a given In the case of choosing the first The conditional probability distribution of a Gaussian model. This represents the vector of the concept after embedding the nth concept in the medical ontology dataset FMA. For a given number In the case of The conditional probability distribution, Let s, R, and t be the numbers; Where s is the scaling factor, R is the rotation matrix, and t is the translation vector; According to Bayes' theorem, using vectors The Gaussian model with centroid about vectors The posterior probability is ; Will Substituting the Q function of the EM algorithm, we obtain the transformation relationship. ; According to Bayes' theorem, the vector... The Gaussian model with centroid about vectors posterior probability The expression is: in, To transform the relationship, To transform the relationship, For ontology dataset The vector of the concept after embedding the i-th concept. For the first Prior probabilities of a Gaussian model To select the first After a Gaussian model is generated, the model produces The probability of; To replace variables; The transformation relationship The expression is: In the formula, s is the scaling factor, R is the rotation matrix, and t is the translation vector; The replacement variable The expression is; The given threshold radius in step five is obtained by using cosine distance; The specific process for generating new positive triplet examples in step six is ​​as follows: Using the concept of alignment Generate new relation triples according to the following formula: in, These are all concepts from FMA. For the relationship in FMA, The set of triple relations contained in the FMA ontology dataset. These are all concepts from NCI. For relationships in NCI, This is the set of triple relations contained in the NCI ontology dataset.