Method for predicting comorbidities using semantic profiling

The method addresses the limitation of existing comorbidity prediction methods by employing overrepresentation analysis, semantic profiling, and Pearson correlation to determine comorbidity accurately, even when no common genes are present, leveraging gene lists and functional gene sets.

JP2026031870AActive Publication Date: 2026-02-25GIL MEDICAL CENT +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024182627
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-12
Filing Date
2024-10-18
Publication Date
2026-02-25
Estimated Expiration
2044-10-18

AI Technical Summary

Technical Problem

Existing methods for predicting comorbidity between diseases using the Jaccard index or overlap coefficient fail to accurately determine the likelihood of comorbidity when there are no common genes associated with the diseases, and cannot account for shared biological mechanisms.

Method used

A method involving overrepresentation analysis, semantic profile calculation, and Pearson correlation coefficient calculation to predict comorbidity, utilizing gene lists and functional gene sets, with optional core subset extraction using mixture model regression analysis.

Benefits of technology

Accurately determines the presence or absence of comorbidity between diseases, even in the absence of common genes, by quantifying shared biological mechanisms through high-accuracy prediction systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026031870000007
    Figure 2026031870000007
  • Figure 2026031870000008
    Figure 2026031870000008
  • Figure 2026031870000009
    Figure 2026031870000009
Patent Text Reader

Abstract

To provide a prediction method and a prediction system for knowing the presence / absence of a comorbidity in a first disease and a second disease with high accuracy even when there is no common gene.SOLUTION: A method of predicting the likelihood of co-morbidity includes the steps of (a1) calculating k p-values using a genetic list of a first disease and k functional gene sets; (a2) calculating k p-values using a genetic list of a second disease and the k functional gene sets; (a3) calculating a first semantic profile using the k p-values calculated in (a1); (a4) calculating a second semantic profile using the k p-values calculated in (a2); and (a5) calculating a Pearson Correlation Coefficient using the first semantic profile and the second semantic profile.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a comorbidity prediction method, and more particularly to a prediction method for determining the likelihood of comorbidities in multiple diseases through semantic profiling. [Background technology]

[0002] This discussion is merely intended to provide background information regarding the present invention and may not constitute prior art.

[0003] Comorbidity refers to other diseases that occur in conjunction with the main disease. It is similar to the concept of complications, but differs from them in that the other diseases are not related to the main disease.

[0004] The Jaccard index or overlap coefficient (OC) method is used to indicate the degree of comorbidity. These methods use genes known to be highly associated with specific diseases to indicate the degree of comorbidity between two diseases, and predict the possibility of comorbidity between diseases based on genes associated with the onset of two diseases.

[0005] Specifically, the index value is determined based on the number of overlapping genes that exist in common between two diseases, which creates the problem that if there are no overlapping genes, it is not possible to calculate the probability of a comorbid disease.

[0006] On the other hand, even if there is no common disease gene, if there is a shared biological mechanism in the development of two diseases, comorbidity will occur between the two diseases. However, these methods have the problem of not being able to explain comorbidities for which no common disease gene is found. Summary of the Invention [Problem to be solved by the invention]

[0007] Therefore, an object of the present invention is to provide a prediction method and a prediction system that can determine with high accuracy whether or not there is a comorbid disease between disease A and disease B, even if there are no common genes.

[0008] The various problems that the present invention aims to solve are not limited to the above problems, and other problems not mentioned will be clearly understood by those skilled in the art from the following description. [Means for solving the problem]

[0009] According to one embodiment of the present invention, a method for predicting the possibility of comorbidity between a first disease and a second disease includes the steps of: (a1) an overrepresentation analysis module 10, 10' calculating k P values ​​using a gene list of a first disease and k functional gene sets; (a2) an overrepresentation analysis module 10, 10' calculating k P values ​​using a gene list of a second disease and the k functional gene sets; and (b1) a semantic profile calculation module 12, 12' calculating a first semantic profile (P A (b2) the semantic profile calculation module 12, 12' calculates a second semantic profile (P B ), and (c) the Pearson correlation coefficient calculation module 14 calculates the first semantic profile (P A ) and the second semantic profile (P B and calculating the Pearson correlation coefficient using

[0010] In addition, in the steps (a1) and (a2) according to an embodiment of the present invention, it is preferable that the P value is calculated according to Equation 1.

[0011] Furthermore, in the step (b1) according to an embodiment of the present invention, the first semantic profile (P A) is a k-th order vector calculated by substituting the k P values ​​calculated in the step (a1) into an exponential function, and in the step (b2), the second semantic profile (P B ) is preferably a k-th order vector calculated by substituting the k P values ​​calculated in step (a2) into an exponential function.

[0012] Furthermore, in step (c) according to one embodiment of the present invention, it is preferable that the Pearson correlation coefficient is calculated according to Equation 2, where n in Equation 2 is the number of functional gene sets.

[0013] Furthermore, the prediction method according to another embodiment of the present invention further includes, before step (a1), a step (a0) in which the core subset extraction module 16′ selects some genes from the first disease gene list and some genes from the second disease gene list.

[0014] Furthermore, in the prediction method according to another embodiment of the present invention, the core subset extraction module 16' selects some genes using mixture model regression analysis. [Effects of the Invention]

[0015] As described above, one embodiment and another embodiment of the present invention have the advantage that the presence or absence of comorbidities between disease A and disease B can be determined with high accuracy even if there are no common genes. [Brief explanation of the drawings]

[0016] [Figure 1] FIG. 1 is a block diagram of a prediction system according to one embodiment of the present invention. [Figure 2] FIG. 2 is a diagram illustrating the flow of data processed by a prediction system according to an embodiment of the present invention. [Figure 3] FIG. 2 is a block diagram of a prediction system according to another embodiment of the present invention. [Figure 4]1 is a flowchart of a prediction method using a prediction system according to one or other embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0017] Some embodiments of the present invention will be described in detail below with reference to the illustrative drawings. When referring to components in each drawing, the same components will be referred to by the same reference numerals whenever possible, even if they appear in different drawings. Furthermore, when describing the present invention, if it is determined that a detailed description of related known structures or functions would obscure the gist of the present invention, such a detailed description will be omitted.

[0018] In describing the components of the embodiments of the present invention, reference numerals such as 1, 2, i), ii), a), b) are used. These reference numerals are merely used to distinguish the components from other components, and do not limit the essence, order, or sequence of the components. In this specification, when a part "includes" or "has" a certain component, it does not mean that other components are excluded, but rather that the part may further include other components, unless otherwise specified.

[0019] In the present invention, the plurality of genes included in the disease gene list means a set of genes involved in the disease.

[0020] Explanation of the comorbidity prediction system Fig. 1 is a block diagram of a prediction system according to an embodiment of the present invention, and Fig. 2 is a diagram showing the flow of data processed by the prediction system according to an embodiment of the present invention.

[0021] A prediction system 1 according to an embodiment of the present invention will be described with reference to Figures 1 and 2. The prediction system 1 is configured to calculate a Pearson correlation coefficient and predict the presence or absence of a comorbidity. To this end, the prediction system 1 according to an embodiment of the present invention includes all or part of an overrepresentation analysis module 10, a semantic profile calculation module 12, and a Pearson correlation coefficient calculation module 14.

[0022] The overrepresentation analysis module 10 calculates a P-value in cooperation with multiple databases 2 and 3. Specifically, the overrepresentation analysis module 10 is configured to calculate a P-value using a gene list for disease A and a gene list for disease B stored in a gene list database 2 and a functional gene set stored in a functional gene set database 3.

[0023] Here, the gene list for disease A may include gene A1, gene B2, gene A3, gene A4, etc. Furthermore, the gene list for disease B may include gene B1, gene B2, gene B3, gene A4, etc. Note that the names A1, B1, etc. are used to indicate that the genes are distinct from one another.

[0024] The functional gene set also includes gene ontology (GO) or gene path. Gene ontology is a type of database consortium that, for the purpose of studying gene function, refers to a model structured for individual genes according to the biological processes, molecular functions, and intracellular and extracellular components to which the genes are related. A gene pathway is a biological database that represents the dynamic relationships or interactions between biological elements such as proteins, genes, and cells in a network format.

[0025] The over-representation analysis module 10 calculates the P-value by an over-representation analysis method.

[0026] The overrepresentation analysis method is a method of calculating a P-value using Table 1 and Equation 1.

[0027] [Table 1]

[0028]

number

[0029] In Table 1, a plus sign means that the gene is included in the disease list or functional gene set, and a minus sign means that the gene is not included. For example, a indicates the number of human genes that are included in the gene list for disease A and are included in Gene Ontology. b indicates the number of human genes that are not included in the gene list for disease A and are included in Gene Ontology. c indicates the number of human genes that are included in the gene list for disease A and are not included in Gene Ontology. d indicates the number of human genes that are not included in the gene list for disease A and are not included in Gene Ontology.

[0030] Note that if a+b+c+d=n, n represents the total number of genes in humans.

[0031] Using a to d and n, the P-value is calculated using Equation 1.

[0032] The overrepresentation analysis module 10 according to one embodiment of the present invention compares a disease gene list with multiple functional gene sets.

[0033] For example, when P values ​​are calculated using a gene list for disease A and a gene set including Gene Ontology 1 to Gene Ontology 1000, 1000 P values ​​for disease A are calculated. If the calculation process for disease A is repeated for disease B, 1000 P values ​​for disease B are calculated.

[0034] The 1000 calculated P values ​​for disease A and the 1000 calculated P values ​​for disease B are sent to the semantic profile calculation module 12 .

[0035] The semantic profile calculation module 12 receives the calculated 1000 P values ​​for disease A and 1000 P values ​​for disease B, and uses them to calculate a semantic profile.

[0036] Here, the semantic profile refers to the result of expressing 1000 calculated P values ​​for disease A and 1000 calculated P values ​​for disease B as vectors and substituting them into an exponential function.

[0037] The semantic profile is expressed as follows: It should be noted that the numbers shown in the present invention are only for the purpose of facilitating understanding, and the numbers will vary depending on the type of disease and the number and type of functional gene sets. Exp(P-vector A)=(1.0,2.7,1.8,1.1,…) Exp(P-vector B)=(1.3,2.1,2.6,2.7,…)

[0038] The Pearson correlation coefficient calculation module 14 calculates a semantic profile (hereinafter referred to as P A ) and semantic profile in disease B (hereafter referred to as P B ) is configured to calculate the Pearson's correlation coefficient (PCC). Here, the formula for calculating PCC using PA and PB is shown in Equation 2.

[0039]

number

[0040] In Formula 2, A and B each represent a disease, and n represents the number of functional gene sets used in calculating the P value. For example, if 1,000 gene ontologies are used as described above, n is 1,000.

[0041] PCC quantifies the degree of similarity between the semantic profiles of gene groups of two diseases. A positive PCC indicates a higher comorbidity rate, whereas a lower PCC indicates a lower comorbidity rate.

[0042] According to the prediction system 1 according to an embodiment of the present invention, if disease A and disease B share similar biological mechanisms, the semantic profile (P A ,P B ) has a relatively high PCC. Therefore, the prediction system 1 according to one embodiment of the present invention has an advantage in that it can determine the presence or absence of a comorbidity between disease A and disease B even if there are no common genes.

[0043] FIG. 3 is a block diagram of a prediction system according to another embodiment of the present invention.

[0044] As shown in FIG. 3, a prediction system 1′ according to another embodiment of the present invention further includes a core subset extraction module 16′ in addition to an overrepresentation analysis module 10′, a semantic profile calculation module 12′, and a Pearson correlation coefficient calculation module 14′.

[0045] The core subset extraction module 16' is configured to extract a core subset. Specifically, when two disease gene lists show similarity to each other, the P value can be calculated using only a portion of the genes in each list, rather than using the entire gene list for each disease. For this purpose, the core subset extraction module 16' extracts only a portion of the genes in the list.

[0046] When extracting the core subset, it is preferable to use a mixture model regression method.

[0047] FIG. 4 is a flowchart of a prediction method using a prediction system according to one or other embodiments of the present invention.

[0048] As shown in FIG. 4, the prediction system 1 according to the embodiment of the present invention can predict comorbidities in the following order:

[0049] The overrepresentation analysis modules 10, 10' perform overrepresentation analysis for disease A and disease B (S410).

[0050] Here, the overrepresentation analysis module 10, 10' works in conjunction with a plurality of databases 2, 3 to calculate P values.

[0051] Specifically, a P-value for a disease can be calculated using the gene list for each disease stored in the gene list database 2 and at least one functional gene set stored in the functional gene set database 3. The over-representation analysis module 10 calculates the P-value using an over-representation analysis method. The P-value is calculated using Table 1 and Equation 1.

[0052] The number of P-values ​​is determined by the number of functional gene sets. That is, if overrepresentation analysis is performed on n functional gene sets and the gene list of disease A, a total of n P-values ​​are determined. If the same process is performed on the gene list of disease B, n P-values ​​for disease B are calculated.

[0053] The semantic profile calculation module 12 calculates a semantic profile using the n P values ​​for disease A and the n P values ​​for disease B calculated in step S410 (S420).

[0054] Specifically, the semantic profile calculation module 12 substitutes the calculated n P values ​​for disease A into an exponential function to obtain an n-th order vector (P A ) and the n-th order vector (P B ) where the two vectors mentioned above are each referred to as a semantic profile according to the present invention.

[0055] The Pearson correlation coefficient calculation module 14 calculates the semantic profile P A , P B The Pearson correlation coefficient is calculated using (S430).

[0056] Here, the Pearson correlation coefficient is calculated using Equation 2. According to the present invention, by applying the value of the Pearson correlation coefficient to logistic regression analysis, a model is constructed to predict whether two diseases are in a comorbid disease relationship, and the constructed model is applied to clinical use to determine comorbid diseases. In order to train the prediction model, conventionally known comorbid disease relationships between two diseases are used as class labels.

[0057] Steps S410 to S430 are common to the prediction methods using the prediction systems 1 and 1' according to the one embodiment and other embodiments of the present invention.

[0058] A prediction method using a prediction system 1' according to another embodiment of the present invention further includes a core subset extraction step (S400) performed before step S410.

[0059] The core subset extraction module 16' extracts some genes from the gene list for disease A and some genes from the gene list for disease B. Note that the number of genes extracted from each list must be the same.

[0060] On the other hand, a preferred method for extracting a subset of genes from each list is to use mixed model regression analysis, which generates a randomly selected gene list and checks whether it yields a higher Pearson correlation coefficient than when using the entire gene set.

[0061] When a gene list with a higher Pearson correlation coefficient than the entire gene list is determined by the above method, the extracted gene list is considered to be a set of genes more related to the onset of comorbidities in the two diseases, and such a gene list is determined to be a core gene set.

[0062] Performance of the comorbidity prediction system Table 2 shows the accuracy of comorbidity prediction results using prediction systems 1 and 1' according to one embodiment and another embodiment of the present invention (where accuracy is greater than or equal to 0 and less than or equal to 1) and the accuracy of comorbidity prediction results using conventional methods.

[0063] [Table 2]

[0064] In Table 2, "n" indicates the number of disease pairs. "RR thres" indicates the relative risk threshold, "JI" indicates the results calculated using the Jaccard coefficient method, "OC" indicates the results calculated using the overlap coefficient method, and "Sab" indicates the results calculated using the separation measurement method proposed by Menche et al. That is, JI, OC, and Sab indicate the results obtained using the conventional method. Here, the relative risk is the result of calculating the degree of comorbidity in a disease based on data for one million people published by the National Cancer Institute. That is, Table 2 compares the degree to which the degree of comorbidity obtained from actual clinical data using relative risk can be predicted using gene sets alone.

[0065] "GOBP" indicates the results of a prediction method using Gene Ontology as a functional gene set in a prediction system 1 according to an embodiment of the present invention, "GOMF" indicates the results of a prediction method using GOMF as a functional gene set in a prediction system 1 according to an embodiment of the present invention, and "Reactome" indicates the results of a prediction method using the Reactome gene set as a functional gene set in a prediction system 1 according to an embodiment of the present invention. That is, GOBP, GOMR, and Reactome indicate the results of a method using the prediction system 1 according to an embodiment of the present invention.

[0066] "LR" indicates the results of a prediction method using a prediction system 1' according to another embodiment of the present invention.

[0067] As shown in Table 2, the results for GOBP, GOMF, Reactome, and LR are always higher than those of the conventional method, indicating that the method is superior to the conventional method.

[0068] The above description is merely an example of the technical concept of the present invention, and various modifications and variations may be made by a person skilled in the art without departing from the essential characteristics of the present invention. Therefore, the present embodiment is intended to illustrate the technical concept of the present invention, but is not intended to limit the technical concept of the present invention. The scope of protection of the present invention should be interpreted by the scope of the claims, and any technical concept within the scope equivalent thereto should be interpreted as being included in the present invention. [Explanation of symbols]

[0069] 1,1' Comorbidity Prediction System 2. Gene List Database 3. Functional Gene Set Database 10,10' Overrepresentation Analysis Module 12, 12' Semantic profile calculation module 14, 14' Pearson correlation coefficient calculation module 16' Core subset extraction module

Claims

1. 1. A method for predicting the likelihood of comorbidity between a first disease and a second disease, comprising: (a1) an overrepresentation analysis module (10, 10') calculates k P-values ​​using a first disease gene list and k functional gene sets; (a2) the overrepresentation analysis module (10, 10') calculates k P-values ​​using the second disease gene list and the k functional gene sets; (b1) The semantic profile calculation module (12, 12') generates a first semantic profile (P A ) (b2) The semantic profile calculation module (12, 12') generates a second semantic profile (P B ) (c) a Pearson correlation coefficient calculation module (14) for calculating the first semantic profile (P A ) and a second semantic profile (P B and calculating a Pearson correlation coefficient using method.

2. In the steps (a1) and (a2), the P value is calculated by Equation 1. The method of claim 1. [Equation 1] a: Number of human genes that are included in the disease gene list and the functional gene set b: Number of human genes that are not included in the disease gene list and are included in the functional gene set c: Number of human genes included in the disease gene list but not included in the functional gene set d: Number of human genes that are not included in the disease gene list and the functional gene set n: total number of human genes

3. In the step (b1), the first semantic profile (P A ) is a k-th order vector calculated by substituting the k P values ​​calculated in step (a1) into an exponential function, In the step (b2), the second semantic profile (P B ) is a k-th order vector calculated by substituting the k P values ​​calculated in step (a2) into an exponential function, The method of claim 1.

4. In the step (c), the Pearson correlation coefficient is calculated according to Formula 2, where n is the number of functional gene sets. The method of claim 3. [Equation 2]

5. carried out before the step (a1), (a0) the core subset extraction module (16') further includes a step of selecting some genes from the first disease gene list and some genes from the second disease gene list; The method according to any one of claims 1 to 4.

6. The core subset extraction module (16') selects a portion of genes using a mixture model regression analysis. The method of claim 5.

Citation Information

Patent Citations

  • Patient assistant for chronic diseases and co-morbidities

    US20190198174A1

  • Pharmacogenomics of Intergenic Single-Nucleotide Polymorphisms and in Silico Modeling for Precision Therapy

    US20200051661A1

  • Systems and methods for identification of clinically similar individuals, and interpretations to a target individual

    US20200395129A1

  • Treatment of cancer by risk stratification of patients based on comordidities

    US20220044764A1