Novel cancer predisposition genes derived from mendelian disease-associated gene group

By detecting pathogenic germline mutations in Mendelian disease-associated genes, the method effectively identifies novel cancer predisposition genes, addressing the limitations of current methods and enhancing cancer predisposition diagnosis.

WO2025135848A1PCT designated stage expired Publication Date: 2025-06-26SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/020749
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2024-12-19
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Current methods struggle to identify novel cancer predisposition genes (CPGs) due to low statistical power in detecting rare pathogenic germline variants and limited understanding of their functional contribution to cancer.

Method used

The method involves detecting pathogenic germline mutations in Mendelian disease-associated genes using a genetic analysis panel and composition capable of identifying such mutations in biological samples, thereby identifying novel CPGs.

Benefits of technology

This approach enables the identification of novel CPGs among Mendelian disease-related genes, providing valuable information for cancer predisposition diagnosis and potentially improving early detection and management of cancer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024020749_26062025_PF_FP_ABST
    Figure KR2024020749_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a novel cancer predisposition genes and, more specifically, to a method for providing information for cancer predisposition diagnosis, the method comprising a step of detecting pathogenic germline variants in novel cancer predisposition genes derived from a Mendelian disease-associated gene group, and to a composition, a kit, and a gene analysis panel for cancer predisposition diagnosis, comprising an agent capable of detecting the pathogenic germline variants in the cancer predisposition genes. The novel cancer predisposition genes of the present invention may be very usefully employed for cancer patient prevention and treatment by enabling management, diagnosis, and treatment prior to cancer onset or at an early stage of cancer onset by preemptively identifying cancer risk.
Need to check novelty before this filing date? Find Prior Art

Description

Novel cancer predisposition genes derived from a group of Mendelian disease-related genes

[0001] The present invention relates to a novel cancer predisposition gene (CPG), and more particularly, to a method for providing information for cancer predisposition diagnosis, comprising a step of detecting a pathogenic germline mutation of a novel cancer predisposition gene derived from a group of Mendelian disease-related genes, and a composition, kit, and genetic analysis panel for cancer predisposition diagnosis, comprising an agent capable of detecting a pathogenic germline mutation of the cancer predisposition gene.

[0002] It is well known that germline mutations increase the risk of cancer, and to date, approximately 140 cancer predisposition genes (CPGs) have been identified through various strategies, including genome-wide association analysis (Nature, 2014, 505: 302). However, many genes associated with cancer predisposition are believed to remain unidentified, likely due to the low statistical power to detect associations between rare pathogenic germline mutations and cancer risk.

[0003] In addition to the challenges of identifying CPGs, understanding their functional contributions to cancer remains limited, although they are known to be closely linked to fundamental cellular processes central to cancer cell function. For example, representative CPGs such as BRCA1 / 2, ATM, and TP53 can affect DNA repair processes and genome instability, disrupt cellular metabolism, and resist apoptosis. Furthermore, CPGs are typically monogenic and exhibit a high-penetrance phenotype in tumors, exhibiting bi-allelic inactivation. Data from large-scale cancer consortia suggest that CPGs contribute to cancer development through mechanisms such as immune evasion, epigenetic reprogramming, and persistent proliferative signals (Cancer Discov., 2022, 12: 31).

[0004] To systematically identify novel CPGs, we focused on Mendelian disease-associated genes obtained from OMIM (Online Mendelian Inheritance in Man; Nucleic Acids res., 2005, 33: D514), a catalog of human genes associated with genetic disorders primarily elucidated by genetic association studies. OMIM Mendelian disease-associated genes share aspects similar to CPGs in that they are monogenic factors exhibiting highly penetrant phenotypes. Furthermore, it has been reported that patients with genetic disorders associated with rare mutations in OMIM genes exhibit secondary phenotypes in adulthood. For example, patients with rare mutations in the GBA gene, which are associated with childhood Gaucher disease, develop Parkinson's disease in adulthood (Curr Opin Neurobiol, 2021, 72: 148). From these examples, the inventors of the present invention believed that among the Mendelian disease-related genes that were previously considered unrelated to cancer, there may be genes that can contribute to cancer through germline mutations.

[0005] Against this backdrop, the inventors of the present invention discovered novel CPGs by confirming the frequency of rare germline mutations in Mendelian disease-related genes in cancer patients, thereby completing the present invention.

[0006] One object of the present invention is to provide a method for providing information for diagnosing a predisposition to cancer, comprising the step of detecting pathogenic germline mutations in Mendelian disease-associated genes in a biological sample isolated from an individual.

[0007] Another object of the present invention is to provide a cancer predisposition diagnostic composition, kit and genetic analysis panel comprising an agent capable of detecting pathogenic germline mutations of Mendelian disease-related genes from a biological sample isolated from an individual.

[0008] This is explained in detail as follows. Meanwhile, each description and embodiment disclosed in the present invention can also be applied to each other description and embodiment. In other words, all combinations of the various elements disclosed in the present invention fall within the scope of the present invention. Furthermore, the scope of the present invention should not be considered limited by the specific descriptions described below.

[0009] One aspect of the present invention for achieving the above object is a method for providing information for diagnosing a predisposition to cancer, comprising the step of detecting a pathogenic germline mutation of a Mendelian disease-related gene in a biological sample isolated from an individual.

[0010] The present inventors aimed to identify novel cancer predisposition genes (CPGs) among Mendelian disease-associated genes. Specifically, we sought to (1) identify new CPGs by identifying genes with a high frequency of pathogenic germline mutations in cancer patients compared to controls from the Online Mendelian Inheritance in Man gene (OMIM) database, and (2) functionally classify these new CPGs. After case-control analysis, we considered (1) the frequency of biallelic gene inactivation and (2) tissue-wide gene expression profiles to define various classes of OMIM genes as potential CPGs.

[0011] CPGs, such as BRCA1 / 2, TP53, and APC, are often involved in essential cellular processes and are therefore ubiquitously and nonspecifically expressed in many tissues (Nature 505, 302-308 (2014)). Furthermore, they tend to follow the 'two-hit' hypothesis, which states that pathogenic germline mutations occur on one allele of a gene, and a 'second hit' occurs somatically on the second allele, increasing cancer risk.

[0012] However, exceptions to the widespread tissue expression observed in several CPGs have been found, and furthermore, various mechanisms of action beyond the standard CPG model have been proposed (Trends in Genetics 37, 433-443 (2021)). For example, the α-1-antitrypsin gene SERPINA1 is preferentially expressed in the liver (RNA level) or in the lung, kidney, and gastrointestinal tract (protein level) (Proc Natl Acad Sci US A114, E10244-e10253 (2017)), yet pathogenic variants in it increase susceptibility to breast cancer (Oncotarget 6, 25815-25827 (2015)). Furthermore, some CPGs show that a single genetic variant is sufficient to increase cancer risk. For example, second allele inactivation of the Fanconi anemia-associated genes (FANCA, FANCC, FANCG; Nature Reviews Cancer 18, 168-185 (2018)), which play important roles in DNA repair and are associated with a high risk of cancers including acute myeloid leukemia and squamous cell carcinoma, is not as widespread as that seen in BRCA1 / 2-associated cancers (Cancer Res 67, 9591-9596 (2007)). We hypothesized that heterozygous pathogenic germline mutation carriers may also develop cancer through mechanisms beyond the classical double-hit hypothesis.

[0013] The present inventors hypothesized that rare pathogenic germline mutations in OMIM genes, such as those associated with genetic insufficiency in inherited diseases, would also increase the risk of tumor development. To verify this hypothesis, the inventors performed a comprehensive analysis using large-scale cancer genetic data obtained from the Pan-Cancer Analysis of Whole Genomes Network (PCAWG) and data from the 1000 Genomes Project as a normal control. Our results suggest that various OMIM genes, previously unknown as CPGs, also harbor pathogenic germline mutations that predispose individuals in the general population to cancer. Furthermore, the inventors analyzed the tumorigenic mechanism of the PAH gene, which showed the highest frequency of pathogenic germline mutations among the OMIM genes, and discovered a novel, previously unknown mechanism of action: OMIM genes increase cancer risk by regulating metabolic and immune response-related mechanisms.

[0014] In the present invention, the term "cancer predisposition" means that a specific individual is at a higher risk of developing cancer compared to a normal individual. Even if a clinically cancer-free individual is not currently in a state, an individual with cancer predisposition is at a higher potential, tendency, or risk of developing cancer in the future compared to a normal individual. Furthermore, in the present invention, cancer predisposition includes all cases where cancer has not been diagnosed based on clinical symptoms but already has cancer cells. In the present invention, a "cancer predisposition gene (CPG)" refers to a gene that can be judged to be at a higher risk of developing cancer in the future if the individual has a pathogenic germline mutation in a specific gene. To date, approximately 140 CPGs have been identified, including BRCA1 / 2, ATM, and TP53. However, the present inventors focused on Mendelian disease-related genes to identify new CPGs.

[0015] As used herein, the term "Mendelian disorder" refers to a genetic disease that follows Mendel's inheritance pattern and is caused by a single mutation in the DNA structure, resulting in a pathological outcome. Specifically, in the present invention, "Mendelian disorder" refers to a disease other than cancer. In other words, cancer is excluded from the Mendelian disorders mentioned herein.

[0016] In the present invention, the term "Mendelian disorder-related genes" refers to genes that undergo mutations and cause Mendelian diseases, and genetic information can be obtained, for example, from the Online Mendelian Inheritance in Man (OMIM) database (https: / / www.omim.org / ). In the definition of "Mendelian disease" in the present invention, genes known to cause cancer among OMIM are not included in the category of "Mendelian disease-related genes." In other words, the novel CPG of the present invention refers to a "Mendelian disease-related gene" known to cause genetic disorders other than cancer when germline mutations occur, and is a gene that was not previously known as a CPG, but was confirmed in the present invention to have the potential to cause cancer as a secondary phenotype. In the present invention, the terms "Mendelian disease-related gene" and "OMIM gene" are used interchangeably, and are used with a meaning clearly distinct from existing CPGs such as BRCA1 / 2, ATM, and TP53.

[0017] <h2 style=";text-align:left;direction:ltr">상기 멘델 질환-관련 유전자는 AAGAB, ABCA4, ABCA7, ABCC6, ABCG2, ABCG5, ACADM, ACADVL, ACPT, ADK, ADSSL1, AGBL1, AHSG, AIRE, ALB, ALPL, AMPD1, AMPD3, ANLN, ANO5, ANO6, ASB10, ASL, ASPM, ASPN, ATP5A1, ATP6AP1, ATP7B, BBS1, BCHE, BEST1, C7, C8A, C9, CALCR, CD207, CD36, CD96, CERKL, CFHR5, CFTR, CHIT1, CHST8, CLEC4M, COL9A1, CPT2, CRB1, CSF2RB, CYP1B1, CYP24A1, DAG1, DFNA5, DHCR7, DMGDH, DNAH11, DNAH17, DNAH5, DOCK6, EEF2, EPHX2, EYS, FAM111A, FCN3, FIG4, FLG2, FTCD, FUT2, FYCO1, G6PC, G6PD, GABRA1, GALNS, GCKR, GJB4, GNRHR, GPNMB, GYG1, HABP2, HAVCR2, HBB, HEXB, HGD, HMCN1, HS6ST2, IFIH1, IL12RB1, IRAK3, KANK1, KEL, KIAA0586, KRT74, KRT8, LDB3, LIPH, LRP2, LRRK2, LZTR1, MAD1L1, MC1R, MEFV, MIB1, MMAA, MMP19, MMP20, MOCOS, MPO, MS4A2, MUC7, MVK, MYH8, MYO18B, MYOC, NAGA, NIPAL4, NPC2, NR1H4, OBSL1, OCA2, OGG1, OTOG, OTOGL, PADI4, PARK2, PAX4, PCCB, PCDH15, PCK2, PER3, PEX1, PHYH, PHYKPL, PKHD1, PLEKHG2, PLOD3, POLG, PREPL, PRSS12, RNASEL, RP1, SAMD9L, SI, SKIV2L, SLC12A3, SLC22A12, SLC24A4, SLC26A8,One or more genes selected from the group consisting of SLC3A1, SLC7A9, SLCO2A1, SLFN14, SNRPN, SPG7, SPTB, SUGCT, TBC1D8B, TCHH, TET2, TEX11, TGFBI, TGIF1, TGM5, THBD, TLR5, TMEM67, TMPRSS3, TNFRSF13B, TRPM4, TTC21A, TTN, TUBB3, TYRP1, UCP3, UPB1, USP45, VPS13B, VWA3B, WARS2, and WDR52. If one or more genes among the above gene groups have a pathogenic germline mutation, the individual may be judged to have a high risk of developing cancer.

[0018] In the present invention, the term "pathogenic germline variant" refers to a mutation that causes loss of original function or alteration, degradation, or hyperactivity due to addition / deletion / substitution of one or more sequences in a wild-type gene, and specifically, it may be a premature protein truncation variant (PTV; Tier 1) or a "pathogenic" or "likely pathogenic" mutation according to the clinical standard ACMG / AMP criteria.

[0019] The above immature PTV is a mutation in which the full-length sequence of the protein is not translated but a truncated form of the protein is translated, and may be, for example, a splice donor / acceptor site mutation, a frameshift indel, or a stop codon gain / loss mutation, but is not limited thereto, and is included in the scope of the present invention as long as it satisfies the definition of a pathogenic germline mutation of the present invention. For example, pathogenic germline mutations of the PAH gene may be, but are not limited to, F55L, I65T, N167I, R169H, V177L, V190A, Y204C, R241C, R241H, V245A, R261Q, T278I, Y356*, V388M, A403V, ​​R408W, Y414C, Q419R.

[0020] The above "pathogenic or likely pathogenic variants according to clinical standard ACMG / AMP criteria" include all variants defined as 'Pathogenic', 'Likely pathogenic', 'association', and 'risk factors' based on clinical evidence developed by ACMG (American College of Medical Genetics and Genomic) of ClinVar (https: / www.ncbi.nlm.nih.gov / clinvar / ), and may further include all potential variants that have not been identified to date but may meet the above criteria. That is, in the present invention, a pathogenic germline variant can be viewed as a variant that causes a Mendelian disease by causing a loss of the function of a protein encoded by the gene due to the genetic mutation, and has the potential to secondarily cause cancer.

[0021] In a specific embodiment of the present invention, the cancer is pan-cancer, and the Mendelian disease-associated genes are AAGAB, ABCG5, ACADM, ADSSL1, AGBL1, ALB, ALPL, AMPD3, ANO6, ASPN, ATP5A1, ATP6AP1, ATP7B, BEST1, C7, C8A, CD207, CD36, CD96, CSF2RB, CYP1B1, CYP24A1, DNAH11, DNAH17, DNAH5, FTCD, FUT2, FYCO1, G6PC, G6PD, HABP2, HS6ST2, IFIH1, IRAK3, KRT74, KRT8, LRRK2, MEFV, MMP19, MVK, NPC2, NR1H4, OTOG, PARK2, PAX4, PCCB, PEX1, PHYH, It may be, but is not limited to, PHYKPL, PKHD1, PLEKHG2, PLOD3, PREPL, PRSS12, SAMD9L, SI, SLC12A3, SLC22A12, SLC26A8, SLCO2A1, SLFN14, TBC1D8B, TEX11, TGIF1, TLR5, TNFRSF13B, VWA3B, or WDR52.

[0022] In another specific embodiment of the present invention, the cancer may be Biliary-AdenoCA, and the Mendelian disease-associated gene may be, but is not limited to, HBB.

[0023] In another specific embodiment of the present invention, the cancer is Bone-Osteosarc and the Mendelian disease-associated gene may be, but is not limited to, EYS.

[0024] In another specific embodiment of the present invention, the cancer is Breast-AdenoCA, and the Mendelian disease-associated gene may be, but is not limited to, ABCA7, ACADM, AHSG, ASB10, ASPM, C7, DHCR7, EYS, FCN3, GCKR, LZTR1, MIB1, NAGA, PHYKPL, RNASEL, SLCO2A1, TCHH, TEX11, TMPRSS3, TTN, TYRP1, UCP3, or UPB1.

[0025] In another specific embodiment of the present invention, the cancer is CNS-Medullo, and the Mendelian disease-associated gene may be, but is not limited to, CD96, CHST8, CYP24A1, DNAH17, FAM111A, GJB4, GNRHR, HABP2, HAVCR2, LRRK2, MAD1L1, MMP20, MYH8, NPC2, PRSS12, RNASEL, SAMD9L, or USP45.

[0026] In another specific embodiment of the present invention, the cancer is CNS-PiloAstro, and the Mendelian disease-associated gene may be, but is not limited to, C7, C9, CHST8, SAMD9L, or SNRPN.

[0027] In another specific embodiment of the present invention, the cancer is ColoRect-AdenoCA, and the Mendelian disease-associated gene may be, but is not limited to, PCK2.

[0028] In another specific embodiment of the present invention, the cancer is Eso-AdenoCA, and the Mendelian disease-associated gene may be, but is not limited to, ABCA7, ABCC6, ASL, GYG1, SKIV2L, or TTN.

[0029] In another specific embodiment of the present invention, the cancer is Head-HCC, and the Mendelian disease-associated gene may be, but is not limited to, FIG4, OGG1, PHYKPL, or SLC7A9.

[0030] In another specific embodiment of the present invention, the cancer is Kidney-ChRCC, and the Mendelian disease-associated gene may be, but is not limited to, ANO5.

[0031] In another specific embodiment of the present invention, the cancer is Kidney-RCC, and the Mendelian disease-associated gene may be, but is not limited to, ACADM, ANLN, CD207, CHST8, CLEC4M, EEF2, FLG2, HGD, MYOC, NPC2, OTOGL, PHYKPL, SLC26A4, SUGCT, TTC21A, or WARS2.

[0032] In another specific embodiment of the present invention, the cancer is Liver-HCC, and the Mendelian disease-associated genes are ABCA4, ABCG2, ACPT, ADK, ADSSL1, AGBL1, ALB, AMPD3, ANO5, ATP6AP1, C7, CALCR, CD207, CD36, CERKL, CFHR5, CPT2, DAG1, DMGDH, DNAH5, EYS, HABP2, HEXB, HMCN1, IL12RB1, IRAK3, KANK1, KRT74, LDB3, LIPH, LRP2, LRRK2, MEFV, MPO, MUC7, OBSL1, OTOG, PAX4, PCCB, PCDH15, PKHD1, PLOD3, PREPL, RP1, SLC22A12, SLC24A4, SLC26A8, SLFN14, This may be, but is not limited to, SPG7, SPTB, TEX11, TGFBI, or VWA3B.

[0033] In another specific embodiment of the present invention, the cancer is Lung-AdenoCA, and the Mendelian disease-associated gene may be, but is not limited to, OCA2.

[0034] In another specific embodiment of the present invention, the cancer is Lung-SCC, and the Mendelian disease-associated gene may be, but is not limited to, FTCD or PCK2.

[0035] In another specific embodiment of the present invention, the cancer is Lymph-BNHL, and the Mendelian disease-associated gene may be, but is not limited to, AGBL1, C7, CFTR, GPNMB, HAVCR2, PEX1, SI, SLC3A1, TRPM4, or WARS2.

[0036] In another specific embodiment of the present invention, the cancer is Lymph-CLL, and the Mendelian disease-associated gene may be, but is not limited to, ABCC6, CFHR5, GABRA1, KIAA0586, MYO18B, RNASEL, SLC3A1, or TCHH.

[0037] In another specific embodiment of the present invention, the cancer is Ovary-AdenoCA, and the Mendelian disease-associated gene may be, but is not limited to, CFHR5, CFTR, CHST8, MAD1L1, NPC2, SI, TMEM67, or VPS13B.

[0038] In another specific embodiment of the present invention, the cancer is Panc-AdenoCA, and the Mendelian disease-associated gene may be, but is not limited to, ACADM, ACADVL, AIRE, AMPD1, BCHE, C9, CD36, CHIT1, CHST8, CRB1, CYP24A1, DFNA5, DNAH5, EPHX2, FCN3, FTCD, G6PD, LRRK2, MMAA, MMP19, MS4A2, NIPAL4, OCA2, OGG1, PADI4, POLG, SI, SUGCT, TCHH, TEX11, or THBD.

[0039] In another specific embodiment of the present invention, the cancer is Panc-Endocrine, and the Mendelian disease-associated gene may be, but is not limited to, AGBL1, PCDH15, or SI.

[0040] In another specific embodiment of the present invention, the cancer is Prost-AdenoCA, and the Mendelian disease-associated gene may be, but is not limited to, CFHR5, DOCK6, G6PD, GALNS, HABP2, KRT8, MAD1L1, MMP19, MOCOS, OGG1, PER3, SLC12A3, TCHH, TET2, or TGM5.

[0041] In another specific embodiment of the present invention, the cancer is Skin-Melanoma, and the Mendelian disease-associated gene may be, but is not limited to, BBS1, DNAH11, NPC2, OCA2, or WARS2.

[0042] In another specific embodiment of the present invention, the cancer is Stomach-AdenoCA, and the Mendelian disease-associated gene may be, but is not limited to, ATP7B, COL9A1, CYP24A1, HAVCR2, IFIH1, KEL, or MYH8.

[0043] In another specific embodiment of the present invention, the cancer is Uterus-AdenoCA, and the Mendelian disease-associated gene may be, but is not limited to, GNRHR, MC1R, MPO, or TUBB3.

[0044] In the present invention, the term "subject" refers to a mammal including a human, and "biological sample" refers to any biological sample obtained from the subject, including, but not limited to, blood, serum, plasma, lymph, saliva, sputum, mucus, urine or feces isolated from the subject whose cancer predisposition is to be confirmed, and samples capable of confirming germline genetic mutations.

[0045] In the present invention, the step of detecting a pathogenic germline mutation of the Mendelian disease-related gene may be polymerase chain reaction (PCR), Sanger sequencing, microarray, direct sequencing, next generation sequencing, targeted exome sequencing, read-depth sequencing, or whole genome sequence assembly, but is not particularly limited to the method as long as it can detect a desired pathogenic germline mutation in the target gene.

[0046] Another aspect of the present invention is a composition for diagnosing cancer predisposition, comprising an agent capable of detecting pathogenic germline mutations of Mendelian disease-related genes from a biological sample isolated from an individual.

[0047] Pathogenic germline mutations in biological samples isolated from individuals and Mendelian disease-related genes are as described above.

[0048] Agents capable of detecting the above pathogenic germline mutations may be, but are not limited to, mutation-specific primers, probes, antisense nucleic acids, aptamers, or antibodies.

[0049] As used herein, the term "primer" refers to a nucleic acid capable of forming base pairs with a complementary template and serving as a starting point for copying the template strand. The sequence of the primer need not be exactly the same as that of the template, but should be sufficiently complementary to hybridize with the template. The primer can initiate DNA synthesis in the presence of polymerization reagents and four different nucleoside triphosphates in an appropriate buffer and temperature. PCR conditions and the lengths of the sense and antisense primers can be applied or modified based on those known in the art.

[0050] As used herein, the term "probe" refers to a substance capable of specifically binding to a target substance to be detected within a sample, and through said binding, the presence of the target substance within the sample can be specifically confirmed. The probe may be manufactured in the form of an oligonucleotide probe, a single-stranded DNA probe, a double-stranded DNA probe, an RNA probe, etc., but is not limited thereto. The selection of the probe and hybridization conditions may be applied or modified based on those known in the art.

[0051] As used herein, the term "antisense nucleic acid" refers to a nucleic acid-based molecule that has a complementary nucleotide sequence to a target genetic variant and can form a dimer therewith. The antisense nucleic acid may be the nucleotide, a fragment thereof, or a complementary one thereof, and may be of an appropriate length to increase detection specificity.

[0052] In addition, the agent capable of detecting the pathogenic germline mutation of the present invention includes an agent capable of detecting not only a gene but also a mutant protein translated therefrom, and may be, for example, an antibody (monoclonal antibody, polyclonal antibody, chimeric antibody, fragment or mutant thereof) or an aptamer capable of specifically binding to the mutant protein, but is not limited thereto.

[0053] Another aspect of the present invention is a cancer predisposition diagnostic kit comprising the composition.

[0054] Specifically, the kit may further include, in addition to the composition, one or more other components, solutions, or devices suitable for the analytical method. For example, the kit may be, but is not limited to, an RT-PCR kit, a DNA chip kit, an ELISA (Enzyme-linked immunosorbent assay) kit, a protein chip kit, or a rapid kit.

[0055] Another aspect of the present invention is a panel for genetic analysis for providing information necessary for cancer predisposition diagnosis, comprising the composition.

[0056] The above genetic analysis panel is a genetic mutation testing method in which mutations for multiple target genes are composed of one panel, and may be based on NGS, but is not limited thereto.

[0057] The embodiments of the present invention may be modified in various ways, and the scope of the present invention is not limited to the embodiments described below. Furthermore, the embodiments of the present invention are provided to more completely explain the present invention to those skilled in the art. Furthermore, throughout the specification, the term "including" a certain component does not exclude other components, but rather implies the inclusion of other components, unless specifically stated otherwise.

[0058] The novel cancer predisposition gene of the present invention can be very useful in preventing and treating cancer patients, as it enables management, diagnosis, and treatment before or at the early stage of cancer onset by identifying the risk of cancer onset in advance.

[0059] Figure 1 shows the population composition of PCAWG and 1KG. AFR: African, AMR: American, EAS: East Asian, EUR: European, and SAS: South Asian.

[0060] Figure 2 is a schematic diagram showing gene classification by four characteristics. (i) Germline enrichment: accumulation of pathogenic germline variants in patients compared to controls, (ii) Two-hit preference: excess of germline variants in samples with LOH compared to samples without LOH, (iii) Tissue expression: number of tissues showing high expression levels among 20 different tissues, and (iv) Tissue-specificity: number of high-expressing tissues with a normalized expression level greater than 2.

[0061] Figure 3 shows a scree plot for determining the optimal number of clusters to use in the trimmed k-means clustering algorithm.

[0062] Figure 4 schematically illustrates the case-control analysis. Principal components analysis (PCA) using common variants was performed to adjust the populations of cancer patients and controls. After collecting rare pathogenic germline variants from the case-control population, the first two PC values ​​used to stratify the population using a linear regression model were used to test the clustering of germline variants in cancer patients compared to controls.

[0063] Figure 5 shows the population composition distribution and principal component analysis (PCA) results using germline variation in PCAWG and 1KG (1000 Genomes). The PCA results are displayed on the left for each population (blue: PCAWG-ICGC, green: PCAWG-TCGA, gray: 1KG). The right side shows the variance ratio for each configuration after PCA. The first two configurations account for approximately 90% of the total variance in the solution.

[0064] Figure 6 is a graph showing genes having at least three or more carriers identified in the present invention, including CPG (cancer predisposition gene), SOD (somatic driver gene), and OMIM gene (Online Mendelian Inheritance in Man gene).

[0065] Figure 7 shows the clustering of rare pathogenic germline mutations in cancer patients compared to controls in the gene set (overall: Tier 1 + 2) (**P< 0.01, ***P< 0.001).

[0066] Figure 8 shows the distribution of Log2 OR (odds ratio) in case-control analysis in the four gene sets and four subgroups of OMIM genes (*P<0.05, **P<0.01, ***P<0.001). The length of each line is 1.5 times the interquartile range.

[0067] Figure 9 shows the clustering of pathogenic germline mutations in disease classes by pan-cancer and single cancer types using a linear regression model after population adjustment. The size of the circles represents the degree of clustering of pathogenic germline mutations in cancer types compared to controls, and the color represents the disease class. Only single cancer types with >20 samples are shown. The numbers in parentheses indicate the number of samples tested for each cancer type (left) and the number of OMIM genes tested for each disease (bottom). The number of diseases detected in each cancer type (P<0.05) is indicated by the bar on the right.

[0068] Hereinafter, the composition and effects of the present invention will be described in more detail through examples. These examples are intended solely to illustrate the present invention and are not intended to limit the scope of the present invention.

[0069] Example 1: Obtaining tumor sequencing data from PCAWG (Pan-Cancer Analysis of Whole Genomes)

[0070] The final consensus set of merged germline mutation calls from the variant call format (VCF) files generated by the PCAWG Consortium was obtained from the International Cancer Genome Consortium portal (https: / / dcc.icgc.org / releases / PCAWG). This dataset represents 2,642 high-quality samples from 2,834 donor samples, excluding 192 samples with potential technical issues (Nature, 2020, 578: 82). The germline mutation calling file is divided into two germline VCF files: one for 1,823 donor data obtained from a non-US International Cancer Genome Consortium (ICGC), and the other for 819 donor data obtained from The Cancer Genome Atlas (TCGA). Each VCF file was merged using bcftools-1.9 following the method recommended by the ICGC DCC (data coordination center) (Bioinformatics (Oxford, England), 2011, 27: 2987). The final set of PCAWG germline calls contained data from 2,642 donors organized into four subgroups, as shown in Figure 1 : European / American (1,904 donors; 72.1%), East Asian (396 donors; 15.0%), African (125 donors; 4.7%), and South Asian (36 donors; 1.4%). This also included 181 donors (6.9%) for whom ethnicity information was unknown.Germline variant calling was originally performed by the ICGC staff using six variant callers that combined single-nucleotide variants (SNVs), indels, and structural variations (SVs): GATK HaplotypeCaller (bioRxiv, 2018: 201178), FreeBayes (arXiv preprint arxiv, 2012: 1207.3907), Real Time Genomics (RTG) (bioRxiv, 2015: 023754), Delly (Bioinformatics, 2012, 28: i333), TraFiC mobile element insertion caller (https: / gitlab.com / mobilegenomes / TraFiC), and Eagle2 (Nature Genetics, 2016, 48: 1443).

[0071] Example 2: Control exogenomes from the 1000 Genomes Project

[0072] The 1000 Genomes Project (1KG; Phase III high-coverage whole-exome sequences) VCF files of normal controls were collected from the FTP server (ftp: / / ftp.1000genomes.ebi.ac.uk / vol1 / ftp / release / 20130502 / ). The data include information on 2,504 individuals generated from four subgroups (European / American, N= 850; African, N= 661; East Asian, N= 504; South Asian, N= 489; Fig. 1 ). The data generation and variant calling methods are described in detail in the 1KG flagship paper (Nature, 2015, 526: 68).

[0073] Example 3: Clinical Information

[0074] Clinical and histological annotation information for PCAWG donors was obtained from the ICGC portal (https: / / dcc.icgc.org / releases / PCAWG). The PCAWG data consist of information from 1,462 males (55.3%) and 1,180 females (44.7%). The Pan-cancer population comprised 37 tumor types: Biliary-AdenoCA (Billary tract adenocarcinoma), Bladder-TCC (Bladder-Transitional cell carcinoma), Bone-Benign (Bone-benign), Bone-Epith (Bone-epithelioid), Bone-Osteosarc (Bone-Sarcoma / bone), Breast-AdenoCa (Breast-Adenocarcinoma), Breast-DCIS (Breast-In situ adenocarcinoma), Breast-LobularCA (Breast-Lobular carcinoma), Cervix-AdenoCA (Cervix-Adenocarcinoma), Cervix-SCC (Cervix-Squamous cell carcinoma), CNS-GBM (Central nervous systems Diffuse glioma), CNS-Medullo (Central nervous systems) Medulloblastoma), CNS-Oligo (Central nervous system) Oligodendroglioma), CNS-PiloAstro (Central nervous systems Non-diffuse gliomo), ColoRect-AdenoCA (Colon / Rectum Adenocarcinoma), Eso-AdenoCA (Esophagus-Adenocarcinoma), Head-SCC (Head / Neck Squamous cell carcinoma),Kidney-ChRCC (Kidney Renal cell carcinoma distal tubles), Kidney-RCC (Kidney Renal cell carcinoma), Liver-HCC (Liver-Hepatocellular carcinoma), Lung-AdenoCA (Lung-Adenocarcinoma), Lung-SCC (Lung-Squamous cell carcinoma), Lymph-BNHL (Lymphoid Mature B-cell lymphoma), Lymph-CLL (Lymphoid Chronic) lymphocytic leukemia), Myeloid-AML (Myeloid-Acute myeloid leukemia), Myeloid-MDS (Myeloid Myelodysplastic syndrome), Myeloid-MPN (Myeloid-Myeloproliferative neoplasm), Ovary-AdenoCA (Ovary-Adenocarcinoma), Panc-AdenoCA (Pancreas-Adenocarcinoma), Panc-Endocrine (Pancreas-Neuroendocrine tumor), Prost-AdenoCA (Prostate-Adenocarcinoma), Skin-Melanoma (Skin-Melanoma), SoftTissue-Leiomyo (Bone / Soft tissue Sarcoma), SoftTissue-Liposarc (Bone / Soft tissue Liposarcoma), Stomach-AdenoCA (Stomach Adenocarcinoma), Thy-AdenoCA (Thyroid Adenocarcinoma), Uterus-AdenoCA (Uterus Adenocarcinoma)), and samples for checking germline mutations are mainly non-tumor samples obtained from blood, as well as tissue adjacent to the primary site or bone marrow,Includes samples obtained from other sites, such as lymph nodes. Demographic information for the 1KG control sample is available on the 1000 Genomes portal (https: / / www.internationalgenome.org / data).

[0075] Example 4: Mutant Annotation and Filtering Steps

[0076] To generate a consistent functional annotation set, we edited the VCF files of PCAWG (N= 2,642) and 1KG (N= 2,504). First, we filtered out variants with low mappability by removing variants within the ENCODE / DUKE (Genome Research, 2014, 24: 2022) and DAC blacklist (Scientific Reports, 2019, 9: 9354) regions, and selected only variants within the ENCODE / CRG GEM mappable regions (75mers; PLOS ONE, 2012, 7: e30377). Next, we annotated the variants collected in the VCF files using ANNOVAR (version 2014 Apr 14). Among the data derived from Annovar's filter-based functional annotation, we used (i) the clinical and phenotypic effects of variants from ClinVar (accessed on 18 June 2019), and (ii) minor allele frequency (MAF) of variants across eight populations (African / African Americans (AFR), South Asians (SAS), East Asians (EAS), South Americans / Mixed Race Americans (AMR), Non-Finnish Europeans (NFE), Finnish Europeans (FIN), Ashkenazi Jews (ASJ), and Others (OTH)) and across all ethnicities from the gnomAD (Genome Aggregation Database) exogenome (v2.1.1) (Nature, 2020, 581: 434). Simultaneously, the gene-based information of the standard transcriptome was interpreted using the Variant Effect Predictor (VEP)-release-96 (Genome Biology, 2016, 17: 122).Among the VEP analysis data, we used the results of protein-coding variants (synonymous, missense, stop gain and loss, splice site, frameshift indel, and in-frame indel). Next, we removed all variants flagged as likely technical artifacts in the gnomAD exogenous genome, including excessive heterozygosity at the variant site (InbreedingCoeff), variants with zero allele counts after filtering out low-confidence genotypes (ACO), and random forest filtering threshold failures (RF).

[0077] Example 5: Principal component analysis (PCA) using common variants

[0078] Many germline variants identified from whole-genome sequencing data or genome-wide association studies were expected to vary by ethnicity. To assess the potential confounding effect of population stratification, PCA was performed at 90% genotyping coverage using PLINK version 2.0 (Am J Hum Genet, 2007; 81: 559) for common germline variants (excluding synonymous variants) with a MAF ≥ 5% across all ethnic groups in the PCAWG and 1KG, and within each ethnic group (four subgroups) in the PCAWG / 1KG.

[0079] Example 6: Rare pathogenic germline mutations

[0080] Rare variants are defined as variants with a frequency of less than 0.5% across all ethnicities and less than 1% in each of the eight subgroups in the gnomAD exogenous genome. For variants not present in the gnomAD exogenous genome, we checked the number of detections within the populations (PCAWG, 1KG separately) across all variants to remove possible technical artifact variants, and removed samples with variants detected in more than 1% of the PCAWG or 1KG populations (9.4% of total variants; 139,678 of 1,485,345).

[0081] Pathogenic variants from the collected rare variants are defined as follows: (1) Tier 1 variants (potentially deleterious) are defined as premature protein-truncating variants (PTVs), including splice donor / acceptor site variants, frameshift indels, and stop codon gain / loss variants, and are not designated as benign or likely benign in ClinVar (http: / www.ncbi.nlm.nih.gov / clinvar / ; Nucleic Acids Res, 2014, 42: D980). (2) Tier 2 variants are defined in ClinVar as variants with clinical evidence developed by the American College of Medical Genetics and Genomics (ACMG) that are classified as ‘pathogenic,’ ‘likely pathogenic,’ ‘associated,’ and ‘risk factors.’

[0082] Example 7: Gene set classification

[0083] OMIM (Online Mendelian Inheritance in Man) database via Gene Map (genemap2.txt) and Morbid Map (mim2gene.txt) We retrieved genes with mutations known to cause clinical diseases and phenotypes from the OMIM database (https: / / www.omim.org / ; acquired Jan 21, 2020). From the OMIM database, we identified 17,076 disease-associated genes for 5,392 diseases represented in the Gene Map (cytogenetic location of genes) and Morbid Map (cytogenetic location of diseases). Next, we selected 5,460 disease-associated genes associated with at least one phenotype (genetic disease) from the genemap2.txt file for further analysis. We then cross-checked the genes represented in the mim2gene.txt file and annotated them to the Ensembl gene names using the annotables package in R (https: / / github.com / stephenturner / annotables), resulting in 4,095 unique gene groups.

[0084] Next, we collected a total of 152 known germline CPGs, including 114 from a recently published review article (Nature, 2014, 505: 302), 11 from the Cancer Gene Census-Germline (http: / cancer.sanger.ac.uk / census / ), 12 from a literature search (Cell, 2018, 173: 355), and 15 genes from the St. Jude PCGP germline study (New England Journal of Medicine, 2015, 373: 2336). Furthermore, to collect high-confidence somatic driver genes (SODs), we first collected 678 genes from three data sources: IntOGen (Nature Reviews Cancer, 2020, 20: 555), MutSig (Nature, 2014, 505: 495), and MutPan (Nature Genetics, 2020, 52: 208). To obtain a non-redundant gene set, we first defined a set of CPGs (N = 152), and then removed 143 genes overlapping with CPGs, resulting in 3,952 OMIM genes. Subsequently, we removed 421 genes overlapping with CPGs or OMIM genes, resulting in 257 SODs. Next, 53 CPGs for which no clinically pathogenic variants were found in TCGA-based germline variant analysis (Cell, 2018, 173: 355) were excluded from the final CPG set.

[0085] The final 4,308 genes selected are defined as input genes, including 99 CPGs, 207 SODs, and 3,952 OMIM genes [2,007 autosomal-recessive (AR) genes, 1,194 autosomal-dominant (AD) genes, 252 AD-AR genes (i.e., variants with dual effects of either dominant or recessive), 211 X-linked genes, 53 somatic genes, 3 digenic genes, 5 isolated genes, 13 multifactorial genes, 2 Y-linked genes, and 212 genes with unknown heritable phenotypes that were finally selected based on unique genetic annotations] (Table 1).

[0086]

[0087]

[0088]

[0089]

[0090]

[0091]

[0092]

[0093] From all human genes in PCAWG and 1KG, 5,598 genes were collected as a control set, excluding CPG, SOD, OMIM genes, 196 disease-specific / cancer-associated genes (https: / diseases.jensenlab.org / Downloads), 100 proteins and 126 DNA repair genes whose characteristics were not identified in the Ensembl gene descriptions annotated using biomaRt packages in R (Nature Protocols, 2009, 4: 1184).

[0094] Example 8: Statistical Analysis of Pathogenic Germline Variant Enrichment

[0095] A Burden test analysis was performed to compare the frequency of pathogenic germline variants within genes between cancer patients (PCAWG) and controls (1KG). This method was designed to estimate the clustering of cancer risk at the gene level by abbreviating potential rare pathogenic germline variants using a generalized linear regression model (GLM) using the stats package in R. The model was performed for each gene-tissue pair (>20 samples from pan-cancer and 26 single cancer types) as follows:

[0096] glm (N~Germline Variants+PC1 + PC2,family=''binomial'')

[0097] Where: N = patient group (1) or control group (0), Germline Variants = number of samples with rare pathogenic germline variants for each gene-tissue pair.

[0098] PC values ​​obtained from PCAWG and 1KG PCA analyses for the control group were used as inputs for regression analysis. Regression coefficients and P-values ​​were calculated for individual gene-tissue pairs using the summary function in R. Higher coefficient values ​​indicate stronger aggregation of pathogenic germline variants compared to the control group.

[0099] Example 9: Clustering of pathogenic germline mutations in disease classes and signaling pathways.

[0100] To systematically estimate the abundance of pathogenic germline variants in disease classes compared to controls, we collected clinically relevant disease diagnostic data from Centogene (https: / / www.centogene.com / diagnostics / ngspanels.html) and Blueprint Genetics (https: / / blueprintgenetics.com / tests / panels / ) and compiled 10 disease classes (cardiovascular, endocrine, epilepsy, hematology, immune, liver, lung, metabolism, muscle / skeletal, and neurological) containing 96 heritable diseases. Among 3,952 OMIM genes, 784 genes were mapped to the 10 disease classes containing 91 OMIM diseases.

[0101] We collected 186 KEGG (Kyoto Encyclopedia of Genes and Genomes) signaling pathways from the Molecular Signatures Database (MsigDB version 7.0). To exclude potential overlaps between KEGG pathways, we performed gene set overrepresentation analysis (GSET) by incorporating network-based gene weights using Gene Ontology functional link enrichment or gene sets (LEGO V2.0). Finally, we collected 70 KEGG signaling pathways containing 1,721 OMIM-related genes.

[0102] Aggregate analyses were performed for each disease class or pathway according to cancer type as follows:

[0103] glm (N~Pathway or Disease classes+PC1+PC2,family=''binomial'')

[0104] Here, N = patient group (1) or control group (0), Pathway or Disease classes = number of samples harboring rare pathogenic germline variants for each disease class or pathway-related gene.

[0105] PC values ​​obtained from PCA analysis of PCAWG and 1KG were used as input values ​​for the regression model.

[0106] Example 10: Obtaining copy number variation data from PCAWG

[0107] Genome-level copy-number alteration (CNA) data for 2,642 samples were obtained from the ICGC data portal (final_consensus_passonly.snv_mnv_indel.icgc.public.maf and final_consensus_passonly.snv_mnv_indel.tcga.controlled.maf). This provides a summarized intersection of genomic region calls from six copy-number variant callers (ABSOLUTE, ACEseq, Battenberg, cloneHD, JaBbA, and Sclust), implemented as described by Dentro et al. (Cell, 2021, 184: 2239). From CNA data, the copy number of loss of heterozygosity (LOH) when the minor allele is 0 is defined according to the PCAWG paper (Nature Communications, 2020, 11: 3400).

[0108] Example 11: Two-hit analysis

[0109] We tested the double-hit hypothesis at the gene level across cancer types by identifying an excess of pathogenic germline mutations in samples with LOH compared to those without LOH events. Analysis was performed for each gene-tissue pair as follows:

[0110] glm (N~Germline Variants+PC1 + PC2,family=''binomial'')

[0111] Here, N = samples with LOH (1) or no LOH event (0), Germline Variants = number of samples with rare pathogenic germline variants for each gene-tissue pair.

[0112] To control for population structure within the cancer population, PC values ​​obtained from PCA analysis for only the PCAWG samples were used as inputs for regression analysis. Regression coefficients and P-values ​​for individual gene-tissue pairs were calculated using the summary function in R. Higher coefficient values ​​indicate a stronger clustering of pathogenic germline mutations in samples with LOH compared to samples without LOH.

[0113] Example 12: PCA Analysis for Gene Classification

[0114] To classify OMIM genes and CPGs, we focused on genes with a frequency of rare germline mutations greater than 1.5-fold in both PCAWG and TCGA compared to controls (1KG). We then examined four features: (i) germline clustering [prevalence of germline mutations in cancer populations compared to normal controls (case-control analysis by regression model; Log2 odds ratio)], (ii) double-hit preference [excess pathogenic germline mutations in samples with LOH events compared to samples without LOH events by regression model; Log2 odds ratio], (iii) tissue expression [number of tissues with expression levels higher than the mean expression value for each gene across 20 tissue types in the Human Protein Atlas (HPA)], and (iv) tissue specificity [the extent to which a gene's expression level is broadly or specifically expressed across cancer types] (Fig. 2). Tissue-specific high expression is defined as a case where the normalized expression level (NX) is greater than 2 after checking the distribution of expression levels across genes in HPA. The above four values ​​were applied as input values ​​for PCA-based clustering analysis using the prcomp function of the stats package in R. The tclust package in R (https: / cran.r-project.org / web / packages / tclust / index.html) was used to classify individual input genes by the top three PC values, and the classified genes were classified into four clusters including outliers (Fig. 3).

[0115] Example 13: Independent cancer cohorts from TCGA

[0116] For independent validation, TCGA germline variants (PCA.r1.TCGAbarcode.merge.tnSwapCorrected.10389.vcf) for 10,389 samples were downloaded from the National Cancer Institute Genomic Data Commons (NCI GDC) Legacy Archive (dbGaP phs000178). The data included 7,573 Europeans (72.9%), 901 African Americans (8.7%), 651 Asians (6.3%), and 1,264 other or unknown races (12.2%). For further analysis, 782 samples overlapping with the PCAWG study were excluded, resulting in a final collection of 9,607 TCGA samples.

[0117] Example 14: Data Information

[0118] This study reanalyzed the PCAWG whole genome sequencing (retrieved from http: / / dcc.icgc.org / pcawg / ), the TCGA whole genome sequencing (https: / / cghub.ucsc.edu / ), and 1000 Genomes (http: / / www.internationalgenome.org / ). All data sets are available upon request from the ICGC Data Access Compliance Office (DACO; http: / / icgc.org / daco), the TCGA Data Access Committe (DAC) via dbGaP, and the study authors for 1000 Genomes. The rare variant data set was obtained from the gnomAD Browser version 2.1.1 (https: / / gnomad.broadinstitute.org / ).

[0119] Experimental Example 1: A System for Measuring the Accumulation of Germline Mutations in Cancer

[0120] To systematically estimate whether rare pathogenic germline variants in OMIM genes contribute to cancer development, we designed a statistical method to test for an excess of rare pathogenic germline variants in cancer patients compared to normal control samples.

[0121] The approach was to collect only highly reliable pathogenic variants. Variants with a minor allele frequency (MAF) <0.5% in the gnomAD (Genome Aggregation Database) exogenous genome were first screened as potentially pathogenic. Next, the following three sets of potentially pathogenic variants were considered, each with a rigorous variant curation step:

[0122] (i) Tier 1 - Protein truncating mutations (PTVs, global splicing mutations, frameshift indels and nonsense mutations),

[0123] (ii) Tier 2 - Variants previously reported as clinically significant in the ClinVar database using clinical standard ACMG / AMP criteria (designated as “pathogenic” or “likely pathogenic”);

[0124] (iii) Tier 1 + 2 Total - Tier 1 and Tier 2 combined to estimate the maximum contribution of potential pathogenic variants to cancer.

[0125] To determine whether selected pathogenic variants contribute to cancer risk, single nucleotide variants (SNVs) and indels were aggregated for each gene. Case-control analysis confirmed an excess of pathogenic variants in genes from cancer patients compared to non-cancer control patients. Burden test analysis was used to control for population structure using the first two PC values ​​based on common variants in the case-control samples (Figures 4 and 5).

[0126] Experimental Example 2: Confirming the Clustering of Rare Germline Variants in Mendelian Disease-Associated Genes in Pan Am

[0127] Two datasets were prepared for which individual-level genomic sequences were available. The first dataset is a database of multi-ancestral studies involving 2,642 cancer patients across four subpopulations of the Pan-Cancer Analysis of Whole Genomes Network (PCAWG). The second dataset contains the exomes of 2,504 normal control individuals across four subpopulations of the 1000 Genomes Project (1KG) (Figure 1). For quality control, both the PCAWG and 1KG databases included only variants in regions with sufficient sequencing coverage. Pathogenic variants were identified in 7,012 genes from 2,543 genes with at least one pathogenic variant, and the frequency of pathogenic variants (Tier 1 or Tier 2) varied widely across samples and cancer types. To increase statistical power, regression analyses were restricted to genes with at least three carriers.

[0128] In the pan-cancer analysis, known CPGs (N= 52, Figure 6) showed significant enrichment of Tier 1 variants in the cancer patient group compared to the control group (median Log2 odds ratio (OR) for Tier 1 enrichment in cancer patients versus controls = 1.25, P= 1.07 × 10 -3 by Wilcoxon rank sum test; Figs. 7 and 8) were confirmed. Above all, six CPGs (11.5%) were significantly clustered for Tier 1 in cancer patients compared to controls individually (P<0.05 in pan-cancer case-control analysis).

[0129] Specifically, BRCA2 showed the highest cancer prevalence compared to the control group (Log2 OR = 4.21, FDR = 4.51 × 10-4 ), and other major CPGs include BRCA1 (Log2 OR = 3.01, FDR = 1.07 × 10 -2 ),FAH(Log2 OR = 3.33,P= 5.33 × 10 -2 ),ATM(Log2 OR = 2.32,FDR= 6.60 × 10 -2 ),BUB1B(Log2 OR = 3.14,P= 8.69 × 10 -2 ), and GJB2 (Log2 OR = 1.28, P = 8.71 × 10 -2 ) was confirmed.

[0130] Next, we examined whether the Mendelian disease-related genes in OMIM (OMIM genes; N = 1,721), which contribute to genetic diseases but whose contribution to cancer risk was not known at all, showed an accumulation of germline mutations in cancer patients compared to control samples (Fig. 6).

[0131] As a result, it was confirmed that OMIM genes had a significant level of Tier 1 mutations in the cancer patient group compared to the control group (median Log2OR = 0.65, P = 1.07 × 10 -3 by Wilcoxon rank sum test, Fig. 8). The above clustering was more pronounced when OMIM was divided into subgroups, specifically, in the case of OMIM-autosomal recessive genes (OMIM-AR), median Log2 OR = 0.69, P = 4.08 × 10 -5 (by Wilcoxon rank sum test), and for OMIM-autosomal dominant gene (OMIM-AD), median Log2 OR = 0.63, P = 4.08 × 10 -5 (Wilcoxon rank sum test) was confirmed (Fig. 8).

[0132] These results suggest that rare pathogenic germline mutations in multiple OMIM genes may contribute to increased cancer risk.

[0133] Next, we integrated three different data sets and used SOD (somatic driver gene; N = 95), a gene known to increase cancer risk through frequent somatic mutations, as a reference gene set (Fig. 6). No mutational enrichment was observed in these genes in cancer (median Log2OR = -0.29, Fig. 8).

[0134] Meanwhile, in tier 2 variants, CPG showed strong clustering in Pan-Cancer patients in PCAWG (median Log2 OR = 0.99). OMIM genes also showed high clustering (median Log2 OR = 0.31), but SOD did not have any clinically proven pathogenic variants in ClinVar. Overall, tier 1 and tier 2 variants showed similar clustering results in Pan-Cancer patients. CPG showed the strongest clustering (median Log2 OR = 1.06; P = 4.66 × 10 -4 by Wilcoxon rank sum test), OMIM-AR (median Log2 OR = 0.46;P= 1.18 × 10 -2 by Wilcoxon rank sum test) and OMIM-AD (median Log2 OR = 0.60;P= 4.67 × 10 -2 OMIM genes (median Log2 OR = 0.51; P = 1.53 × 10 by Wilcoxon rank sum test) -2 (by Wilcoxon rank sum test) was next.

[0135] OMIM genes that showed statistical significance in the entire set of variants in Tier 1 + 2 for the Pan-Am patient population are shown in Table 2 below.

[0136] GeneLog2 ORP-valueAAGAB3.420.007ABCG52.350.049ACADM1.570.012ADSSL12.010.022AGBL10.750.045ALB3.510.025ALPL2.940.018AMPD32.150.009ANO62.640.036ASPN2.70.029ATP5A11.70.039ATP6AP13.420.026ATP7B1.390.015BEST12.230.01C71.550.006C8A2.580.022CD2071.39< 0.001CD360.750.038CD963.670.017CSF2RB1.980.017CYP1B11.990.045CYP24A12.130.007DNAH112.110.036DNAH172.41< 0.001DNAH51.430.034FTCD2.43< 0.001FUT21.740.001FYCO14.080.016G6PC3.340.028G6PD5.3< 0.001HABP22.070.026HS6ST23.230.03IFIH12.680.014IRAK31.950.029KRT741.610.012KRT81.50.024LRRK21.220.02MEFV1. 870.01MMP191.570.034MVK3.050.044NPC21.670.012NR1H43.660.038OTOG1.990.002PAH0.880.031PARK21.710.036PAX42.86< 0.001PCCB1.910.024PEX13.460.031PHYH2.20.012PHYKPL1.350.024PKHD11.490.026PLEKHG22.870.004PLOD33.70.02PREPL2.180.013P RSS121.750.009SAMD9L1.940.033SI1.470.021SLC12A31.590.047SLC22A121.680.009SLC26A82.560.005SLCO2A12.690.019SLFN142.66< 0.001TBC1D8B3.330.034TEX112.3< 0.001TGIF12.480.043TLR51.570.044TNFRSF13B2.720.017VWA3B1.420.008WDR522.90.031

[0137] Experimental Example 3: Identification of Clusters of Rare Germline Mutations in the OMIM Gene in a Single Cancer

[0138] Next, a case-control analysis was performed for each of the nine cancers with more than 100 samples (95.3% of the pan-cancer patients; Figure 9).

[0139] At the single-gene level, 109 genes (6 CPG and 103 OMIM) were identified as significantly associated with at least one cancer type and pan-cancer (FDR < 0.2).

[0140] Specifically, for CPG, as expected, BRCA1 was associated with ovarian adenocarcinoma (Ovary-AdenoCA; Log2OR = 6.91, FDR = 9.63 × 10 -6 ) and breast adenocarcinoma (Breast-AdenoCA; Log2OR = 4.81, P = 4.76 × 10 -3 ), and BRCA2 was associated with ovarian adenocarcinoma (Ovary-AdenoCA; Log2OR = 4.70, P = 6.93 × 10 -5 ), pancreatic adenocarcinoma (Panc-AdenoCA; Log2OR = 4.12, P = 6.45 × 10 -4 ) and breast adenocarcinoma (Log2OR = 3.76, P = 9.03 × 10 -3 ) were significantly accumulated.

[0141] Among the OMIM genes, the following results were confirmed: In renal cell carcinoma (Kidney-RCC), EEF2 (Log2OR = 5.38, FDR = 1.40 × 10 -3 ) was identified as the top gene-tumor pair showing a significant level of clustering. In addition, PEX1 (Log2OR = 5.13, P = 3.07 × 10 -3 ),FLG2(Log2OR = 5.04, P = 3.68 × 10 -3) is associated with renal cell carcinoma (Kidney-RCC), DHCR7 (Log2OR = 4.81, P = 4.76 × 10 -3 ) also showed a significant level of aggregation in breast cancer (Breast-AdenoCA).

[0142] OMIM genes with significant levels of pathogenic germline mutations identified in each cancer type are summarized in Tables 3 to 25 below.

[0143] Biliary-AdenoCAGeneLog2 ORP-valueHBB2.97< 0.001

[0144] Bone-OsteosarcGeneLog2 ORP-valueEYS4.8< 0.001

[0145] Breast-AdenoCAGeneLog2 ORP-valueABCA71.640.014ACADM1.910.042AHSG2.870.007ASB101.640.043ASPM3.620.003C720.033DHCR74.56< 0.001EYS2.440.007FCN32.340.019GCKR2.730.005LZTR13.260.001MIB13.220.001NAGA2.470.027PHYKPL2.930.001RNASEL1.920.027SLCO2A15.13< 0.001TCHH2.150.035TEX113.46< 0.001TMPRSS32.160.022TTN1.80.018TYRP14.66< 0.001UCP33.01< 0.001UPB13.340.003

[0146] CNS-MedulloGeneLog2 ORP-valueCD9621.19< 0.001CHST83.66< 0.001CYP24A14.260.002DNAH174.670.009FAM111A21.19< 0.001GJB43.590.014GNRHR1.810.047HABP23.80.022HAVCR23.460.033LRRK23.95< 0.001MAD1L13.330.016MMP202.390.01MYH83.180.002NPC23.530.004PRSS124.49< 0.001RNASEL4.31< 0.001SAMD9L4.850.003USP452.190.002

[0147] CNS-PiloAstroGeneLog2 ORP-valueC73.290.01C93.220.005CHST84.39< 0.001SAMD9L5.550.001SNRPN2.040.018

[0148] ColoRect-AdenoCAGeneLog2 ORP-valuePCK22.940.001

[0149] Eso-AdenoCAGeneLog2 ORP-valueABCA73.320.034ABCC63.090.017ASL3.860.011GYG13.110.023SKIV2L5.130.006TTN5.480.003

[0150] Head-SCCGeneLog2 ORP-valueFIG44.33< 0.001OGG14.71< 0.001PHYKPL3.82< 0.001SLC7A92.80.002

[0151] Kidney-ChRCCGeneLog2 ORP-valueANO53.94< 0.001

[0152] Kidney-RCCGeneLog2 ORP-valueACADM2.530.006ANLN5.150.003CD2072.170.019CHST82.220.028CLEC4M2.230.021EEF27.43< 0.001FLG24.51< 0.001HGD2.130.032MYOC1.910.045NPC22.870.011OTOGL4.010.002PAH1.830.021PHYKPL2.60.011SLC26A42.930.006SUGCT2.470.019TTC21A5.88< 0.001WARS22.170.018

[0153] Liver-HCCGeneLog2 ORP-valueABCA42.380.013ABCG21.430.018ACPT5.670.006ADK4.060.032ADSSL12.370.045AGBL11.150.037ALB16.9< 0.001AMPD33.30.013ANO52.930.006ATP6AP116.9< 0.001C72.470.046CALCR7.21< 0.001CD2071.96< 0.001CD361.350.031CERKL10.660.014CFHR52.740.011CPT25.40.003DAG12.570.031DMGDH2.590.025DNAH52.990.009EYS1.750.029HABP25.67< 0.001HEXB2.970.027HMCN13.25< 0.001IL12RB15.89< 0.001IRAK33.610.025KANK12.620.036KRT741.550.046LDB316.86< 0.001LIPH2.30.038LRP24.290.002LRRK21.520.039MEFV2.260.019MPO2.40.011MUC71.670.032OBSL12.80.009OTOG3.770.002PAX43.11< 0.001PCCB4.190.013PCDH151.790.027PKHD14.010.002PLOD316.86< 0.001PREPL2.820.014RP130.043SLC22A121.810.021SLC24A416.92< 0.001SLC26A810.560.021SLFN143.04< 0.001SPG73.650.028SPTB1.960.047TEX114.760.004TGFBI2.530.014VWA3B1.980.002

[0154] Lung-AdenoCAGeneLog2 ORP-valueOCA23.52< 0.001

[0155] Lung-SCCGeneLog2 ORP-valueFTCD8.940.001PCK22.630.007

[0156] Lymph-BNHLGeneLog2 ORP-valueAGBL11.510.048C73.280.014CFTR2.820.022GPNMB8.11< 0.001HAVCR23.420.036PAH2.270.026PEX143.73< 0.001SI6.370.001SLC3A14.230.01TRPM43.480.008WARS22.780.012

[0157] Lymph-CLLGeneLog2 ORP-valueABCC63.480.019CFHR54.40.015GABRA157.73< 0.001KIAA05862.280.031MYO18B3.550.008RNASEL3.240.049SLC3A14.090.022TCHH4.020.008

[0158] Ovary-AdenoCAGeneLog2 ORP-valueCFHR52.460.007CFTR2.570.007CHST82.490.01MAD1L13.320.003NPC23.230.003PAH2.180.004SI3.77< 0.001TMEM673.260.003VPS13B5.59< 0.001

[0159] Panc-AdenoCAGeneLog2 ORP-valueACADM1.890.038ACADVL3.280.016AIRE4.690.001AMPD13.050.001BCHE1.980.019C91.90.048CD363< 0.001CHIT12.360.023CHST81.910.03CRB14.33< 0.001CYP24A12.880.008DFNA55.89< 0.001DNAH53.30.008EPHX24.730.001FCN32.630.025FTCD6.99< 0.001G6PD7.72< 0.001LRRK22.890.012MMAA5.46< 0.001MMP193.570.011MS4A23.790.002NIPAL43.430.007OCA21.620.041OGG12.660.036PADI44.16< 0.001POLG3.150.014SI3.270.003SUGCT2.110.039TCHH2.390.012TEX1130.02THBD6.130.001

[0160] Panc-EndocrineGeneLog2 ORP-valueAGBL11.650.039PAH1.940.034PCDH154.36< 0.001SI4.88< 0.001

[0161] Prost-AdenoCAGeneLog2 ORP-valueCFHR52.390.04DOCK64.70.014G6PD12.58< 0.001GALNS2.250.033HABP23.250.027KRT82.330.015MAD1L12.840.018MMP195.45< 0.001MOCOS5.92< 0.001OGG12.980.039PAH1.790.013PER36.22< 0.001SLC12A32.620.026TCHH2.250.038TET22.820.035TGM52.140.032

[0162] Skin-MelanomaGeneLog2 ORP-valueBBS14.490.02DNAH1166.810.035NPC24.260.013OCA24.5< 0.001WARS24.1< 0.001

[0163] Stomach-AdenoCAGeneLog2 ORP-valueATP7B4.56< 0.001COL9A111.49< 0.001CYP24A15.56< 0.001HAVCR26.37< 0.001IFIH17.73< 0.001KEL3.79< 0.001MYH82.090.013

[0164] Thy-AdenoCAGeneLog2 ORP-valuePAH2.97< 0.001

[0165] Uterus-AdenoCAGeneLog2 ORP-valueGNRHR3.44< 0.001MC1R3.19< 0.001MPO3.38< 0.001TUBB33.19< 0.001

[0166] From the above description, those skilled in the art will understand that the present invention can be implemented in other specific forms without altering its technical spirit or essential characteristics. In this regard, it should be understood that the embodiments described above are illustrative in all respects and not restrictive. The scope of the present invention should be interpreted as encompassing all changes or modifications derived from the meaning and scope of the following claims and their equivalent concepts, rather than the detailed description above.

Claims

1. A step of detecting a pathogenic germline mutation of a Mendelian disease-related gene in a biological sample isolated from an individual, The above Mendelian disease-related genes are one or more genes selected from the group consisting of ATP7B, CD36 and SI. A method of providing information for diagnosing a predisposition to cancer.

2. In paragraph 1, The cancer is hepatocellular carcinoma or pancreatic adenocarcinoma, and the Mendelian disease-related gene is CD36. A method of providing information for diagnosing a predisposition to cancer.

3. In paragraph 1, The cancer is mature B-cell lymphoma, ovarian adenocarcinoma, pancreatic adenocarcinoma or pancreatic neuroendocrine tumor, and the Mendelian disease-related gene is SI. A method of providing information for diagnosing a predisposition to cancer.

4. In paragraph 1, The above cancer is gastric cancer, and the Mendelian disease-related gene is ATP7B. A method of providing information for diagnosing a predisposition to cancer.

5. In paragraph 1, The above pathogenic germline mutations are immature protein truncation mutations, or pathogenic or pseudo-pathogenic mutations according to clinical standard ACMG / AMP criteria. A method of providing information for diagnosing a predisposition to cancer.

6. In paragraph 5, The above immature protein truncation mutations are splice donor / acceptor site mutations, frameshift indels, or stop codon gain / loss mutations. A method of providing information for diagnosing a predisposition to cancer.

7. In paragraph 1, The above biological samples include blood, serum, plasma, lymph, saliva, sputum, mucus, urine or stool, A method of providing information for diagnosing a predisposition to cancer.

8. In paragraph 1, The step of detecting the pathogenic germline mutation of the above Mendelian disease-related gene is a method selected from the group consisting of polymerase chain reaction (PCR), Sanger sequencing, microarray, direct sequencing, next generation sequencing, targeted exome sequencing, read-depth sequencing, and whole genome sequence assembly. A method of providing information for diagnosing a predisposition to cancer.

9. A preparation capable of detecting pathogenic germline mutations of Mendelian disease-related genes from a biological sample isolated from an individual, The above Mendelian disease-related genes are one or more genes selected from the group consisting of ATP7B, CD36 and SI. A composition for diagnosing cancer.

10. In paragraph 9, The agent capable of detecting the above pathogenic germline mutation is a mutation-specific primer, probe, antisense nucleic acid, aptamer or antibody. A composition for diagnosing cancer.

11. A composition comprising a composition according to Article 9 or 10; Kit for cancer predisposition diagnosis.

12. A composition comprising a composition according to Article 9 or 10, A panel for genetic analysis to provide information needed for cancer predisposition diagnosis.

Citation Information

Patent Citations

  • CD36 mutant gene and method for determining disease caused by abnormal lipid metabolism and diagnostic kit therefor

    JP2001149082A

  • Detection of mutations in ATP7B gene using Real-Time PCR

    KR1020140140386A

  • KR20220061737A