Marker and kit for detecting gastric cancer, and use thereof
By detecting differentially methylated regions and gene methylation levels specific to gastric cancer, and combining this with a machine learning model, the problem of insufficient sensitivity and specificity in existing gastric cancer detection methods has been solved, achieving efficient and low-cost early gastric cancer screening.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BGI GENOMICS CO LTD
- Filing Date
- 2024-10-28
- Publication Date
- 2026-05-07
AI Technical Summary
Existing gastric cancer detection methods lack sensitivity and specificity, failing to meet the need for accurate screening of early gastric cancer. Furthermore, endoscopic examinations are costly, require large equipment, and have low patient acceptance.
By detecting gastric cancer-specific differentially methylated regions, including 28 differentially methylated regions, and combining the methylation levels of the ELMO1 and KCNA3 genes, primers and probes were used for detection. A cancer risk prediction model was constructed using machine learning to achieve accurate auxiliary diagnosis of gastric cancer.
It improves the accuracy of gastric cancer diagnosis, reduces testing costs, simplifies the testing process, enables early detection of gastric cancer, and increases patient acceptance.
Smart Images

Figure PCTCN2024127717-FTAPPB-I100001 
Figure PCTCN2024127717-FTAPPB-I100002 
Figure PCTCN2024127717-FTAPPB-I100003
Abstract
Description
Biomarkers, reagent kits and their applications for detecting gastric cancer Technical Field
[0001] This invention relates to the field of biomedicine, specifically to biomarkers, reagent kits, and their applications for detecting gastric cancer. Background Technology
[0002] Gastric cancer is one of the cancers with the highest incidence and mortality rates worldwide. Endoscopic examination and biopsy are currently the gold standard for diagnosing gastric cancer; however, endoscopic screening requires extensive equipment and endoscopists, and the examination is relatively expensive and somewhat painful, leading to poor patient acceptance. Protein biomarker-based detection methods (serum PG, G-17, and Hp antibodies) have lower than expected sensitivity and specificity, failing to meet the needs of early gastric cancer diagnosis. In my country, the early diagnosis and treatment rate of gastric cancer is very low, only about 20%, with most cases already in advanced stages. The overall 5-year survival rate is less than 50%, and early gastric cancer detection could increase the 5-year survival rate to over 90%. Therefore, establishing an accurate, simple, and economical method for early gastric cancer screening is of great significance.
[0003] When cells in the human body rupture or die, they release their DNA into the circulatory system, which is called cell-free DNA (cfDNA). Similarly, when tumor cells rupture or die, they release circulating tumor DNA (ctDNA), which carries the genetic information of the tumor cells. By detecting ctDNA mixed with cfDNA and analyzing the mutations and epigenetic information it carries, the likelihood of the subject having cancer can be inferred.
[0004] Scientists have developed numerous high-throughput sequencing technologies with high molecular weight utilization, making it possible to detect trace mutation signals in cell-free DNA in plasma, thus driving the development of precise early cancer screening. Scientists initially searched for biomarkers suitable for early cancer screening from gene mutations; however, research showed that using mutation signals alone for early cancer screening had limited effectiveness. Therefore, scientists began exploring early cancer screening at the epigenetic level. DNA methylation is an important gene expression regulatory mechanism that can regulate gene expression and silencing, playing a significant role in the occurrence and development of tumors. Abnormal methylation of cancer-related genes often occurs in the early stages of cancer development; therefore, DNA methylation signals are considered potential biomarkers for early cancer screening.
[0005] Looking at companies both domestically and internationally focused on early cancer screening technology research: Grail primarily employs cfDNA targeted sequencing, WGS, and WGBS, performing whole-genome sequencing on a large number of cancer and non-cancer control samples to uncover tumor-specific mutations and methylation molecular markers. This strategy allows for a comprehensive study of the tumor genome map, but the enormous cost of high-depth whole-genome sequencing is beyond the reach of most research institutions. Guardant Health focuses on liquid biopsy technology, using highly sensitive detection techniques for early cancer screening. However, liquid biopsy technology still has many limitations for early cancer screening, such as: extremely weak early tumor mutation signals, the presence of some gene mutations in multiple different cancer types, and significant interference from clonal hematopoiesis with ctDNA detection. Therefore, using mutations alone as molecular markers has limited effectiveness. Genecast combines mutation and protein marker detection; this research also shows that applying multi-omics detection can effectively improve detection performance. Kunyuan Gene and Benchmark Medical focus on the detection of methylation. Changes in DNA methylation often occur at multiple sites simultaneously, thus having higher sensitivity than gene mutations at a single site. Moreover, the tissue specificity of DNA methylation signals makes early screening of pan-cancer tumors possible. Therefore, methylation is an ideal molecular marker for early cancer screening.
[0006] Currently, there are many studies on early cancer screening based on DNA methylation both domestically and internationally. For example, Dr. Dennis Lo, a renowned clinical molecular biology expert in Hong Kong and hailed as the "founder of non-invasive prenatal DNA testing," published a study in 2019 on the use of low-depth WGBS sequencing to detect methylation and copy number variation (CNA) in urinary cfDNA. This method achieved a sensitivity of 93.5% (specificity 95.8%) for bladder cancer detection. Another example is the clinical validation study on methylation biomarkers for gastric cancer published by Anderson, BW, et al. in 2018. This study first identified candidate DNA methylation biomarkers related to gastric cancer from the DNA methylome, and then used methylation-specific PCR (MSP) to test a large number of samples, ultimately obtaining a panel containing three markers (ELMO1, ZNF569, C13orf18). The sensitivity of this method reached 86% (specificity 95%, CI 71-95%). A growing body of research reports has demonstrated the enormous potential of DNA methylation biomarkers in early cancer screening. Based on this extensive research, developing a methylation-based, "convenient" detection method will accelerate the process of translating methylation-based early cancer screening into clinical applications.
[0007] The ELMO1 gene encodes a protein involved in cell motility, promoting phagocytosis and cell migration. The KCNA3 gene encodes a member 3 of the potassium voltage-gated channel subfamily A, which mediates voltage-dependent potassium permeability of excitable membranes.
[0008] Invention disclosure
[0009] The technical problem to be solved by the present invention is to provide a more accurate auxiliary diagnostic method for gastric cancer, characterized by using the detection of differential methylation markers for gastric cancer to assist in the diagnosis of gastric cancer.
[0010] This invention claims protection for a group of differentially methylated regions (DMRs).
[0011] The differentiated methylation region group claimed in this invention includes all or part of the following 28 differentiated methylation regions (as shown in Table 3):
[0012] (A1) is located at positions 23847173 to 23848069 on chromosome 16;
[0013] (A2) is located at positions 58497626 to 58497874 on chromosome 16;
[0014] (A3) is located at positions 3785618 to 3786199 on chromosome 19;
[0015] (A4) is located at positions 133464016 to 133464255 on chromosome 12;
[0016] (A5) is located at positions 182321660 to 182322797 on chromosome 2;
[0017] (A6) is located at positions 27940318 to 27940640 on chromosome 17;
[0018] (A7) is located at positions 49814781 to 49815692 on chromosome 7;
[0019] (A8) is located at positions 37487610 to 37489066 on chromosome 7;
[0020] (A9) is located at positions 87229481 to 87230462 on chromosome 7;
[0021] (A10) is located at positions 132382356 to 132383095 on chromosome 9;
[0022] (A11) is located at positions 29106420 to 29107080 on chromosome 13;
[0023] (A12) is located at positions 102068552 to 102068996 on chromosome 13;
[0024] (A13) is located at positions 123301050 to 123302058 on chromosome 11;
[0025] (A14) is located at positions 114695178 to 114695581 on chromosome 1;
[0026] (A15) is located at positions 67344498 to 67345106 on chromosome 8;
[0027] (A16) is located at positions 752063 to 752203 on chromosome 7;
[0028] (A17) is located at positions 29337988 to 29338749 on chromosome 2;
[0029] (A18) is located at positions 111217001 to 111217968 on chromosome 1;
[0030] (A19) is located at positions 7450242 to 7451419 on chromosome 10;
[0031] (A20) is located at positions 36909233 to 36910064 on chromosome 19;
[0032] (A21) is located at positions 48936828 to 48938691 on chromosome 15;
[0033] (A22) is located at positions 68546355 to 68547096 on chromosome 2;
[0034] (A23) is located at positions 61808614 to 61809996 on chromosome 20;
[0035] (A24) is located at positions 108507504 to 108508116 on chromosome 1;
[0036] (A25) is located at positions 58238586 to 58239135 on chromosome 19;
[0037] (A26) is located at positions 58951350 to 58951895 on chromosome 19;
[0038] (A27) is located at positions 144694307 to 144695311 on chromosome 2;
[0039] (A28) is located at positions 129693355 to 129694176 on chromosome 3;
[0040] The physical locations of the 28 differentially methylated regions were determined based on alignment with the human whole genome sequence (version hg19).
[0041] The differentiated methylation region group includes, but is not limited to, selecting a subset of (A1)-(A28) from the previous text, small-scale replacement, or small-scale addition.
[0042] Furthermore, the group of differentiated methylation regions may consist of all or part of the 28 differentiated methylation regions shown in (A1)-(A28) above.
[0043] In one embodiment of the invention, the group of differentiated methylation regions comprises all of the 28 differentiated methylation regions shown in (A1)-(A28) above.
[0044] In another embodiment of the present invention, the group of differentiated methylation regions consists of the two differentiated methylation regions shown in (A8)-(A18) above.
[0045] This invention also claims protection for a reagent kit.
[0046] The kit claimed in this invention includes substances for detecting the group of differentially methylated regions described above.
[0047] Furthermore, the kit includes substances for detecting the methylation level of the ELMO1 gene and / or substances for detecting the methylation level of the KCNA3 gene.
[0048] Furthermore, the kit includes substances for detecting the nucleotide sequence of the gastric cancer-specific methylated region of the ELMO1 gene (SEQ ID No. 115) and / or substances for detecting the nucleotide sequence of the gastric cancer-specific methylated region of the KCNA3 gene (SEQ ID No. 116).
[0049] The substance used to detect the methylation level of the ELMO1 gene can be primer pair A and probe A; the substance used to detect the methylation level of the KCNA3 gene can be primer pair B and probe B.
[0050] The primer pair A consists of two single-stranded DNAs as shown in any one of SEQ ID No. 3 to SEQ ID No. 52; the probe A is a single-stranded DNA as shown in any one of SEQ ID No. 104 to SEQ ID No. 108;
[0051] The primer pair B consists of two single-stranded DNAs as shown in any one of SEQ ID No. 53 to SEQ ID No. 102; the probe B is a single-stranded DNA as shown in any one of SEQ ID No. 109 to SEQ ID No. 114;
[0052] The kit has any of the following uses: diagnosing or screening for cancer; providing early warning of cancer before clinical symptoms appear; or differentiating or assisting in the differentiation between cancer and benign lesions.
[0053] Further, primer pair A consists of two single-stranded DNAs as shown in SEQ ID No. 19 and SEQ ID No. 20; probe A is the single-stranded DNA as shown in SEQ ID No. 107. Primer pair B consists of two single-stranded DNAs as shown in SEQ ID No. 77 and SEQ ID No. 78; probe B is the single-stranded DNA as shown in SEQ ID No. 113.
[0054] Furthermore, the kit VI also contains primer pair C and probe C for amplifying the internal reference gene ACTB; primer pair C consists of two single-stranded DNAs as shown in SEQ ID No. 1 and SEQ ID No. 2; probe C is the single-stranded DNA as shown in SEQ ID No. 103.
[0055] In the kit described above, the non-cancer patients can be healthy controls or have benign lesions. The cancers are stomach cancer, liver cancer, colorectal cancer, lung cancer, pancreatic cancer, prostate cancer, esophageal cancer, or urothelial carcinoma.
[0056] In one embodiment of the present invention, the cancer is gastric cancer; the benign lesion is a benign lesion of the stomach, such as gastritis, gastric ulcer, gastric polyp, etc.
[0057] This invention also claims protection for any of the following applications:
[0058] I) Application of the ELMO1 gene and / or KCNA3 gene as methylation markers in any of the following (C1)-(C6);
[0059] II) The use of substances for detecting the methylation level of the ELMO1 gene and / or KCNA3 gene in any of the following (C1)-(C6);
[0060] III) The application of the aforementioned differentially methylated region group as a methylation marker in any of the following (C1)-(C6);
[0061] IV) Application of substances used to detect the methylation level of the aforementioned differentially methylated region groups in any of the following (C1)-(C6);
[0062] V) The combination of substance and medium is used in any of the following (C1)-(C6); said substance is a substance for detecting the differentially methylated region group described above; said medium stores a method for constructing and using a cancer risk prediction model;
[0063] VI) The application of a medium containing methods for constructing and using cancer risk prediction models in any of the following (C1)-(C6);
[0064] (C1) Prepare products for the diagnosis or auxiliary diagnosis of cancer;
[0065] (C2) Diagnosing or assisting in the diagnosis of cancer;
[0066] (C3) Prepare products for the early warning of cancer before clinical symptoms appear;
[0067] (C4) Providing early warning of cancer before clinical symptoms appear;
[0068] (C5) To prepare products for distinguishing or assisting in the distinction between cancer and benign lesions;
[0069] (C6) To differentiate or assist in differentiating between cancer and benign lesions;
[0070] In V) and VI), the method for constructing and using the cancer risk prediction model includes the following steps:
[0071] (D1) Construct a training set to obtain methylation level data for the differentially methylated region groups described above, including samples from known cancer patients and known non-cancer patients.
[0072] (D2) A cancer risk prediction model is constructed using machine learning, and then the cancer risk prediction model is used to diagnose or assist in the diagnosis of cancer and / or provide early warning of cancer before clinical symptoms and / or differentiate or assist in the differentiation of cancer and benign lesions.
[0073] The ELMO1 gene may include not only its coding region but also its promoter region. Similarly, the KCNA3 gene may include not only its coding region but also its promoter region.
[0074] The methylation level of the ELMO1 gene can be the methylation level of all or part of the CpG sites in the DNA fragment shown in SEQ ID No. 115. The methylation level of the KCNA3 gene can be the methylation level of all or part of the CpG sites in the DNA fragment shown in SEQ ID No. 116.
[0075] Furthermore, the substances used to detect the methylation level of the ELMO1 gene are primer set A and probe A; primer set A consists of two single-stranded DNAs as shown in any one of SEQ ID No. 3 to SEQ ID No. 52; probe A is any one of the single-stranded DNAs shown in SEQ ID No. 104 to SEQ ID No. 108.
[0076] Furthermore, the primer set A consists of two single-stranded DNAs shown in SEQ ID No. 19 and SEQ ID No. 20; the probe A is the single-stranded DNA shown in SEQ ID No. 107.
[0077] Furthermore, the substances used to detect the methylation level of the KCNA3 gene are primer set B and probe B; primer set B consists of two single-stranded DNAs as shown in any one of SEQ ID No. 53 to SEQ ID No. 102; probe B is any one of the single-stranded DNAs shown in SEQ ID No. 109 to SEQ ID No. 114.
[0078] Furthermore, the primer set B consists of two single-stranded DNAs shown in SEQ ID No. 77 and SEQ ID No. 78; the probe B is the single-stranded DNA shown in SEQ ID No. 113.
[0079] In the applications described above, all of the machine learning methods can be random forests.
[0080] In the applications described above, all of the samples can be samples from which DNA can be extracted.
[0081] Furthermore, the sample may be plasma, tissue, saliva, urine, or / and feces.
[0082] In the applications described above, all methods for obtaining the methylation level data include, but are not limited to, detection after bisulfite conversion, restriction enzyme digestion technology, methylation-specific PCR (MS-PCR), pyrosequencing, high-throughput sequencing, third-generation sequencing, or single-molecule sequencing.
[0083] In the application described above, all of the non-cancer patients can be either healthy controls or patients with benign lesions.
[0084] In the applications described above, all of the cancers mentioned can be stomach cancer, liver cancer, intestinal cancer, lung cancer, pancreatic cancer, prostate cancer, esophageal cancer, or urothelial carcinoma.
[0085] In one embodiment of the present invention, the cancer is gastric cancer; the benign lesion is a benign lesion of the stomach, such as gastritis, gastric ulcer, gastric polyp, etc.
[0086] This invention also claims protection for a data processing apparatus having any of the following functions:
[0087] (E1) Diagnosing or assisting in the diagnosis of cancer;
[0088] (E2) Provides early warning of cancer before clinical symptoms appear.
[0089] (E3) To differentiate or help differentiate between cancer and benign lesions;
[0090] The data processing device includes unit X and unit Y;
[0091] The unit X is used to establish a cancer risk prediction model, including a data acquisition module and a data analysis and processing module;
[0092] The data acquisition module is used to collect methylation level data for the differentially methylated region group described above from known cancer patient samples and known non-cancer patient samples.
[0093] The data analysis and processing module can use the methylation level data of the differential methylation region group mentioned above, collected by the data acquisition module from known cancer patient samples and known non-cancer patient samples, as a training set to construct a cancer risk prediction model based on the principle of machine learning.
[0094] The unit Y is capable of diagnosing or assisting in the diagnosis of cancer and / or providing early warning of cancer before clinical symptoms appear, and or differentiating or assisting in the differentiation of cancer from benign lesions, based on the cancer risk prediction model and methylation level data of the differentially methylated regions group described above from the test subject sample.
[0095] This invention also claims protection for a system.
[0096] The system claimed in this invention is system A or system B;
[0097] System A includes:
[0098] (F1) Reagents and / or instruments used to detect the methylation level of the differentially methylated region group described above;
[0099] (F2) A data processing device, the data processing device comprising unit X and unit Y;
[0100] The unit X is used to establish a cancer risk prediction model, including a data acquisition module A and a data analysis and processing module;
[0101] The data acquisition module A is used to acquire methylation level data for the differentially methylated region group described above from known cancer patient samples and known non-cancer patient samples obtained by (F1) detection.
[0102] The data analysis and processing module can use the methylation level data of the differential methylation region group mentioned above, collected by the data acquisition module from known cancer patient samples and known non-cancer patient samples, as a training set to construct a cancer risk prediction model based on the principle of machine learning.
[0103] The unit Y is capable of diagnosing or assisting in the diagnosis of cancer and / or providing early warning of cancer before clinical symptoms appear, and or differentiating or assisting in the differentiation of cancer from benign lesions, based on the cancer risk prediction model and methylation level data of the differentially methylated regions group described above from the test subject sample.
[0104] System B includes:
[0105] (G1) The kit and real-time quantitative PCR instrument described above;
[0106] (G2) Device, the device includes a data acquisition module B, a threshold storage module, a data comparison module, and a data processing and conclusion output module;
[0107] The data acquisition module B is configured to acquire real-time fluorescence quantitative PCR amplification result data from the sample obtained by the test subject through (G1) detection;
[0108] The threshold storage module is configured to store threshold A, threshold B, and threshold C; threshold A is the threshold value of the Ct value of the ACTB gene; threshold B is the threshold value of the Ct value of the ELMO1 gene; and threshold C is the threshold value of the Ct value of the KCNA3 gene.
[0109] The data comparison module is configured to receive real-time quantitative PCR amplification results data of the sample from the test subject sent by the data acquisition module, and to call the threshold A, threshold B, and threshold C stored in the threshold storage module. Then, it compares the Ct value of the ACTB gene of the test subject with the threshold A, compares the Ct value of the ELMO1 gene of the test subject with the threshold B, and compares the Ct value of the KCNA3 gene of the test subject with the threshold C.
[0110] The data processing and conclusion output module is configured to receive the comparison results sent from the data comparison module, and then output the conclusion as follows:
[0111] If the Ct value of the ACTB gene in the test subject is greater than the threshold A, the result is deemed unreliable; when the Ct value of the ACTB gene is less than or equal to the threshold A, the result is deemed reliable, and further determination is made as follows:
[0112] If the Ct value of the ELMO1 gene of the test subject is greater than the threshold B, and / or the Ct value of the KCNA3 gene of the test subject is greater than the threshold C, then the test subject is determined to be or a candidate non-cancer patient, or the test subject is a low-risk cancer patient; otherwise, the test subject is determined to be or a candidate cancer patient, or the test subject is a high-risk cancer patient.
[0113] In one embodiment of the present invention, the threshold A, the threshold B, and the threshold C are all 40.
[0114] This invention also claims protection for a computer-readable storage medium.
[0115] The computer-readable storage medium claimed in this invention stores a computer program for performing the following steps:
[0116] Methylation level data for the differentially methylated region group as described in claim 1 were collected from known cancer patient samples and known non-cancer patient samples;
[0117] The methylation level data of the differential methylation region groups mentioned above, obtained from known cancer patient samples and known non-cancer patient samples, were used as the training set to construct a cancer risk prediction model based on the principle of machine learning.
[0118] Based on the aforementioned cancer risk prediction model and methylation level data from the subject sample for the differentially methylated regions group described above, it is possible to diagnose or assist in the diagnosis of cancer and / or provide early warning of cancer before clinical symptoms appear and / or differentiate or assist in the differentiation of cancer from benign lesions.
[0119] The machine learning method mentioned above can be the random forest method.
[0120] In the above text, the sample may be a sample from which DNA can be extracted.
[0121] Furthermore, the sample may specifically be plasma, tissue, saliva, urine, and / or feces.
[0122] The methods for obtaining the methylation level data mentioned above include, but are not limited to, detection after bisulfite conversion, restriction enzyme digestion technology, methylation-specific PCR (MS-PCR), pyrosequencing, high-throughput sequencing, third-generation sequencing, or single-molecule sequencing.
[0123] In the above text, the non-cancer patients mentioned may be healthy controls or patients with benign lesions;
[0124] The cancers mentioned above may be stomach cancer, liver cancer, intestinal cancer, lung cancer, pancreatic cancer, prostate cancer, esophageal cancer, or urothelial carcinoma.
[0125] In one embodiment of the present invention, the cancer is gastric cancer; the benign lesion is a benign lesion of the stomach, such as gastritis, gastric ulcer, gastric polyp, etc.
[0126] This invention also claims protection for the use of the aforementioned reagent kit, data processing apparatus or system, or computer-readable storage medium in any of the following:
[0127] (C1) Prepare products for the diagnosis or auxiliary diagnosis of cancer;
[0128] (C2) Diagnosing or assisting in the diagnosis of cancer;
[0129] (C3) Prepare products for the early warning of cancer before clinical symptoms appear;
[0130] (C4) Providing early warning of cancer before clinical symptoms appear;
[0131] (C5) To prepare products for distinguishing or assisting in the distinction between cancer and benign lesions;
[0132] (C6) To differentiate or help differentiate between cancer and benign lesions.
[0133] The cancers mentioned may be stomach cancer, liver cancer, intestinal cancer, lung cancer, pancreatic cancer, prostate cancer, esophageal cancer, or urothelial carcinoma.
[0134] In one embodiment of the present invention, the cancer is gastric cancer; the benign lesion is a benign lesion of the stomach, such as gastritis, gastric ulcer, gastric polyp, etc.
[0135] This invention also claims protection for any of the following methods:
[0136] Method I: A method for diagnosing or assisting in the diagnosis of cancer may include the following steps: analyzing the methylation status of a group of differentially methylated regions described above from a sample of a subject, thereby enabling the diagnosis or assisting in the diagnosis of cancer.
[0137] Furthermore, the method may include the following steps:
[0138] (H1) Construct a training set to obtain methylation level data for the differentially methylated region groups described above, including samples from known cancer patients and known non-cancer patients.
[0139] (H2) A cancer risk prediction model is constructed using machine learning, and then the cancer risk prediction model is used to diagnose or assist in the diagnosis of cancer.
[0140] Method II: A method for early warning of cancer before clinical symptoms, which may include the following steps: analyzing the methylation status of differentially methylated regions from a sample of a subject, as described above, thereby enabling early warning of cancer before clinical symptoms.
[0141] Furthermore, the method includes the following steps:
[0142] (I1) Construct a training set to obtain methylation level data for the differentially methylated region groups described above, including samples from known cancer patients and known non-cancer patients.
[0143] (I2) A cancer risk prediction model is constructed using machine learning, and then the cancer risk prediction model is used to provide early warning of cancer before clinical symptoms appear.
[0144] Method III: A method for distinguishing or assisting in the distinction between cancer and benign lesions may include the following steps: analyzing the methylation status of differentially methylated regions from a sample of a subject, in relation to the differentially methylated regions described above, thereby distinguishing or assisting in the distinction between cancer and benign lesions.
[0145] Furthermore, the method includes the following steps:
[0146] (J1) Construct a training set to obtain methylation level data for the differentially methylated region groups described above, including samples from known cancer patients and samples from known corresponding benign lesion patients;
[0147] (J2) A cancer risk prediction model is constructed using machine learning, and then the cancer risk prediction model is used to distinguish or assist in distinguishing between cancer and benign lesions.
[0148] Method IV: A method for diagnosing or assisting in the diagnosis of cancer may include the steps of detecting the methylation levels of the ELMO1 gene and / or the KCNA3 gene in a sample from a subject, thereby enabling the diagnosis or assisting in the diagnosis of cancer.
[0149] Method V: A method for providing early warning of cancer before clinical symptoms, which may include the steps of detecting the methylation levels of the ELMO1 and / or KCNA3 genes in a sample from a subject, thereby enabling early warning of cancer before clinical symptoms.
[0150] Method VI: A method for distinguishing or assisting in the distinction between cancer and benign lesions may include the following steps: detecting the methylation levels of the ELMO1 gene and / or KCNA3 gene in a sample from a subject, thereby distinguishing or assisting in the distinction between cancer and benign lesions.
[0151] In the method described above, the methylation level of the ELMO1 gene is the methylation level of all or part of the CpG sites in the DNA fragment shown in SEQ ID No. 115. The methylation level of the KCNA3 gene is the methylation level of all or part of the CpG sites in the DNA fragment shown in SEQ ID No. 116.
[0152] In the method described above, detecting the methylation levels of the ELMO1 gene and / or KCNA3 gene in the sample from the subject can be performed according to a method including the following steps:
[0153] (K1) DNA was extracted from the sample and then methylated;
[0154] (K2) DNA obtained by real-time quantitative PCR amplification of (K1) after methylation conversion;
[0155] The primer pair A for the ELMO1 gene used in the real-time quantitative PCR amplification consists of two single-stranded DNAs as shown in any one of SEQ ID No. 3 to SEQ ID No. 52, and the probe A is any one of the single-stranded DNAs shown in SEQ ID No. 104 to SEQ ID No. 108.
[0156] The primer pair B for the KCNA3 gene amplification used in the real-time quantitative PCR amplification consists of two single-stranded DNAs as shown in any one of SEQ ID No. 53 to SEQ ID No. 102, and the probe B is any one of the single-stranded DNAs shown in SEQ ID No. 109 to SEQ ID No. 114.
[0157] Furthermore, the primer pair A for the ELMO1 gene used in the real-time quantitative PCR amplification consists of two single-stranded DNAs as shown in SEQ ID No. 19 and SEQ ID No. 20; the probe A is the single-stranded DNA shown in SEQ ID No. 107; the primer pair B for the KCNA3 gene used in the real-time quantitative PCR amplification consists of two single-stranded DNAs as shown in SEQ ID No. 77 and SEQ ID No. 78; the probe B is the single-stranded DNA shown in SEQ ID No. 113.
[0158] Further, in step (K2), the internal reference used in the real-time quantitative PCR amplification is the ACTB gene, the primer pair C used to amplify the ACTB gene consists of two single-stranded DNAs as shown in SEQ ID No. 1 and SEQ ID No. 2, and the probe C is the single-stranded DNA as shown in SEQ ID No. 103.
[0159] Furthermore, after step (K2), the following steps may also be included:
[0160] (K3) The results of the real-time quantitative PCR amplification described in (K2) shall be judged as follows:
[0161] When the Ct value of the ACTB gene is greater than 40, the result is considered unreliable; when the Ct value of the ACTB gene is ≤40, the result is considered reliable, and further determination is made as follows:
[0162] When the Ct value of the ELMO1 gene and / or the Ct value of the KCNA3 gene are greater than 40, the subject is determined to be a non-cancer patient or a candidate for non-cancer patient, or a low-risk cancer patient; otherwise, the subject is determined to be a cancer patient or a candidate for non-cancer patient, or a high-risk cancer patient.
[0163] In the method described above, the methods for analyzing the methylation state and obtaining the methylation level data include, but are not limited to, detection after bisulfite conversion, restriction enzyme digestion technology, methylation-specific PCR (MS-PCR), pyrosequencing, high-throughput sequencing, third-generation sequencing, or single-molecule sequencing.
[0164] In the method described above, the sample is a sample from which DNA can be extracted.
[0165] Furthermore, the samples include, but are not limited to, plasma, serum, blood, tissue, saliva, urine, and feces.
[0166] In the methods described above, the non-cancer patients may be healthy controls or have benign lesions. The cancers may be stomach cancer, liver cancer, intestinal cancer, lung cancer, pancreatic cancer, prostate cancer, esophageal cancer, or urothelial carcinoma.
[0167] In one embodiment of the present invention, the cancer is gastric cancer; the benign lesion is a benign lesion of the stomach, such as gastritis, gastric ulcer, gastric polyp, etc.
[0168] In this invention, the method for analyzing the methylation level of DMR is related to computer software and / or computer hardware, including but not limited to determining the methylation state of DMR, comparing the methylation state of DMR, determining the Ct value, determining the specificity and / or sensitivity of the analyte or label, calculating the ROC curve and related AUC, sequence analysis, etc.
[0169] In this invention, the gastroscopy examination of the healthy controls showed no abnormalities.
[0170] In a specific embodiment of the present invention, the cancer mentioned above is gastric cancer. Further, the gastric cancer may be gastric cancer of different pathological subtypes (including adenocarcinoma, adenosquamous carcinoma, signet ring cell carcinoma, neuroendocrine carcinoma, mucinous adenocarcinoma, etc.) and / or at different stages (including stages I-IV).
[0171] In a specific embodiment of the present invention, the methylation level data mentioned above includes methylation rate or Ct value. Attached Figure Description
[0172] Figure 1 is a heatmap of 630 DMRs (Disseminated Methylation Regulators) identified based on 59 pairs of gastric cancer tissues and adjacent normal tissues. Gastric cancer tissues are on the left, and adjacent normal tissues are on the right. Each cell represents the methylation rate of the DMR at that site for that sample, ranging from 0 to 1. The closer the methylation rate is to 0, the darker the color. Among them, 596 are hypermethylated DMRs and 34 are hypomethylated DMRs.
[0173] Figure 2 is a heatmap of the methylation rate of 28 DMRs in sample set I (174 gastric cancer patients and 201 healthy individuals). The horizontal axis represents the samples, with gastric cancer patients on the far left and healthy individuals on the far right.
[0174] Figure 3 shows the performance of the gastric cancer methylation model based on 28 DMRs in sample set I. The horizontal axis represents the false positive rate (1-specificity), and the vertical axis represents the sensitivity.
[0175] Figure 4 shows the ROC curves of gastric cancer samples for each gene detection in Example 3.
[0176] Figure 5 shows the ROC curve of gastric cancer samples detected by the combined detection of ELMO1 and KCNA3 genes in Example 3.
[0177] In Figures 4 and 5, when plotting ROC curves for binary classification, patients are divided into gastric cancer patients and non-gastric cancer patients (non-gastric cancer patients include patients with benign gastric lesions, interfering samples, and healthy individuals).
[0178] The best way to implement an invention
[0179] The following examples are provided to better understand the present invention, but do not limit the invention. Unless otherwise specified, the experimental methods in the following examples are conventional methods. Unless otherwise specified, the experimental materials used in the following examples were purchased from conventional biochemical reagent stores. All quantitative experiments in the following examples were performed in triplicate, and the results were averaged.
[0180] Example 1: Discovery of Gastric Cancer-Specific Methylation Markers
[0181] To screen for biomarkers of specific methylation in gastric cancer, whole genome bisulfite sequencing (WGBS) was performed on genomic DNA from gastric cancer tissues and adjacent normal tissues of 59 pairs of gastric cancer patients. Analysis of the data identified 630 differentially methylated regions (DMRs) that may be associated with the development and progression of gastric cancer (Figure 1).
[0182] 1. Tissue sample extraction
[0183] DNA was extracted from 59 pairs of gastric cancer and adjacent normal tissue samples using the DNeasy Blood & Tissue Kit (Qiagen, #69506). Quantification was performed using the Qubit 3.0 system (Invitrogen, USA).
[0184] 2. DNA fragmentation
[0185] Take 200 ng of extracted DNA and add 1 ng of Unmethylated lambda DNA (PROMEGA, #D1521) for subsequent CU transformation quality control. Use an ultrasonic fragmentation device to break the DNA fragments, and then use AMPure XP (AGENCOURT, #A63882) to select DNA fragments so that the DNA fragment size is concentrated at around 160 bp.
[0186] 3. Library Construction
[0187] Library construction was performed using the KAPA HyperPlus Library Preparation Kit (KAPA, #KK8510). For specific instructions, please refer to the kit manual. However, the adapters used in the adapter ligation step and the PCR primers used in the PCR step were replaced with adapters and primers suitable for the MGISEQ platform.
[0188] 4. Bisulfite Conversion
[0189] The constructed DNA library was subjected to bisulfite conversion using the EZ DNA Metlylation-Gold Kit (Zymo Research, #D5006).
[0190] 5. Library hybrid capture
[0191] Hybridization, capture, and elution were performed using the Seq Cap EZ Hybridization and Wash Kit (ROCHE, 5634253001) and the Seq Cap Epi CpGiant Enrichment Kit (ROCHE, 7138911001). Because the sequencing instrument used was an MGI platform, the blocks used in the hybridization process had to be the blocks corresponding to the MGI platform.
[0192] 6. High-throughput sequencing
[0193] PE100 sequencing was performed using MGISEQ-2000 (MGI).
[0194] 7. Data Analysis
[0195] The sequencing data underwent quality control, filtering, alignment, and calculation to determine the methylation rate of each CpG site. By comparing the differences in methylation rates of various CpG sites between gastric cancer tissue and adjacent normal tissue, a hierarchical Bayesian approach was used to identify 630 differentially methylated regions (DMRs), i.e., gastric cancer-specific methylation regions, between gastric cancer tissue and adjacent normal tissue.
[0196] Figure 1 shows a heatmap of 630 DMRs (Diagnosis Modulations) found based on 59 pairs of gastric cancer tissues and adjacent normal tissues. Gastric cancer tissues are on the left, and adjacent normal tissues are on the right. Each cell represents the DMR methylation rate at that site for that sample, ranging from 0 to 1. The closer the methylation rate is to 0, the darker the color.
[0197] Example 2: Validation of gastric cancer-specific methylation markers
[0198] The subjects in this study are referred to as sample set I, which includes 174 newly diagnosed gastric cancer patients who have not received treatment (excluding the 59 gastric cancer patients involved in Example 1) and 201 healthy controls. The age and gender information of the two groups are shown in Table 1. Among them, the gastric cancer patients cover different stages, different sites of onset, and different pathological subtypes, as shown in Table 2.
[0199] Table 1. Age and gender information of the two groups of samples
[0200] Table 2. Information on different stages, sites of onset, and pathological subtypes of gastric cancer patients.
[0201] This invention performs targeted methylation high-throughput sequencing on plasma cell-free cellular DNA (cfDNA) from 174 gastric cancer patients and 201 healthy individuals from an independent set (i.e., sample set I) to select and validate methylation markers for gastric cancer.
[0202] 1. Plasma sample extraction
[0203] cfDNA was extracted from plasma samples from gastric cancer patients and healthy individuals using the MagPure Circulating DNA Maxi Kit (MAGEN, #12917PJ-100). Quantification was performed using a Qubit 3.0 system (Invitrogen, USA). 10 ng of cfDNA was collected, and 0.05 ng of fragmented and screened unmethylated lambda DNA down to approximately 160 bp was added for subsequent CU transformation quality control.
[0204] 2. Library Construction
[0205] Library construction was performed using the KAPA HyperPlus Library Preparation Kit (KAPA, #KK8510). For specific instructions, please refer to the kit manual. However, the adapters used in the adapter ligation step and the PCR primers used in the PCR step were replaced with adapters and primers suitable for the MGISEQ platform.
[0206] 3. Bisulfite Conversion
[0207] The constructed DNA library was subjected to bisulfite conversion using the EZ DNA Metlylation-Gold Kit (Zymo Research, #D5006).
[0208] 4. Library hybrid capture
[0209] Hybridization, capture, and elution were performed using the Seq Cap EZ Hybridization and Wash Kit (ROCHE, 5634253001) and the Seq Cap Epi CpGiant Enrichment Kit (ROCHE, 7138911001). Because the sequencing instrument used was an MGI platform, the blocks used in the hybridization process had to be the blocks corresponding to the MGI platform.
[0210] 5. High-throughput sequencing
[0211] PE100 sequencing was performed using MGISEQ-2000 (MGI).
[0212] 6. Data Analysis
[0213] The sequencing data underwent quality control, filtering, alignment, and calculation to determine the methylation rate of each CpG site. By comparing the methylation rates of various CpG sites in the plasma of gastric cancer patients and healthy individuals, 28 high-methylation DMRs were selected based on feature importance and chosen as biomarkers for gastric cancer screening (Table 3). A 10-fold cross-validation based on random forest was used, and the model performance was evaluated using the validation set at each fold. The average sensitivity of this set of 28 gastric cancer-specific biomarkers was 0.954, the specificity was 0.985, and the AUC was 0.992.
[0214] Table 3. 28 biomarkers that can be used for gastric cancer screening.
[0215] Note: The physical locations (chromosome, start point, and end point) in the table are determined based on alignment with the human whole genome sequence (version number hg19).
[0216] Figure 2 shows a heatmap of methylation rates in 28 DMR regions across 174 gastric cancer patients and 201 healthy individuals. The horizontal axis represents the sample, with gastric cancer patients on the far left and healthy individuals on the far right.
[0217] Figure 3 shows the performance of the gastric cancer methylation model based on 28 DMRs in sample set I. The horizontal axis represents the false positive rate (1-specificity), and the vertical axis represents the sensitivity.
[0218] Example 3: Evaluation of the ability of ELMO1 and KCNA3 gene methylation to detect gastric cancer.
[0219] The study subjects in this embodiment are referred to as sample set II, which includes 258 gastric cancer patients, 29 patients with benign gastric lesions (including gastritis, gastric ulcers, gastric polyps, etc.), 3 interfering samples (samples from patients with other malignant tumors other than gastric cancer, including 2 cases of colorectal cancer and 1 case of gastrointestinal stromal tumor), and 89 healthy individuals. This excludes the 174 gastric cancer patients and 201 healthy controls involved in Example 2.
[0220] The age and gender information of each group of samples is shown in Table 4. Among them, gastric cancer patients cover different stages, different sites of onset, and different pathological subtypes, as shown in Table 5.
[0221] Table 4. Age and gender information of each group of samples
[0222] Table 5. Information on different stages, sites of onset, and pathological subtypes of gastric cancer patients.
[0223] The main objective of this embodiment is to verify the performance of the gastric cancer-specific biomarker genes (each gene in Table 1) in detecting gastric cancer methylation in an independent validation set (i.e., sample set II). The test sample was cell-free DNA from plasma. Finally, from these 28 DMRs, the ELMO1 and KCNA3 genes with better performance were selected as methylation biomarkers for detecting gastric cancer. The nucleotide sequence of the gastric cancer-specific methylation region of the ELMO1 gene is shown in SEQ ID No. 115; the nucleotide sequence of the gastric cancer-specific methylation region of the KCNA3 gene is shown in SEQ ID No. 116.
[0224] 1. Plasma sample extraction
[0225] cfDNA was extracted from plasma samples from gastric cancer patients and healthy individuals using the MagPure Circulating DNA Maxi Kit (MAGEN, #12917PJ-100). Quantification was performed using the Qubit 3.0 system (Invitrogen, USA).
[0226] 2. Bisulfite Conversion
[0227] Take 10 ng of cfDNA and perform bisulfite conversion using the EZ DNA Metlylation-Gold Kit (Zymo Research, #D5006).
[0228] 3. Real-time quantitative PCR (qPCR)
[0229] cfDNA converted from bisulfite was amplified by qPCR.
[0230] The entire reaction system contained 5 μl of 10×PCR buffer (Novizan, China), 2 units of Taq polymerase (Novizan, China), 2.5 mM dNTPs (Novizan, China), 2 μl (10 pmol / μl) of PCR primers, 1 μl (10 pmol / μl) of probe, and cfDNA converted to bisulfite. The primer and probe sequences used for detecting the ELMO1 and KCNA3 genes, as well as the internal reference gene, are shown in Table 6.
[0231] Table 6. Primers and probes for detecting ELMO1 and KCNA3 genes and internal reference genes.
[0232] The PCR reaction was performed under the following conditions: pre-denaturation at 95°C for 3 minutes, followed by 40 cycles of denaturation at 95°C for 30 seconds, annealing at 60°C for 30 seconds, and extension at 72°C for 30 seconds.
[0233] The detection was performed using a Macroblock 96S qPCR instrument. The baseline and threshold were set to default. Fluorescence signals were collected at the end of each cycle, and a final extension was performed at 72°C for 5 minutes.
[0234] In this embodiment, the primers and probes used for detecting the ELMO1 gene in subsequent performance evaluation are SEQ ID No. 19, SEQ ID No. 20, and SEQ ID No. 107; the primers and probes used for detecting the KCNA3 gene are SEQ ID No. 77, SEQ ID No. 78, and SEQ ID No. 113; and the primers and probes used for detecting the internal reference ACTB gene are SEQ ID No. 1, SEQ ID No. 2, and SEQ ID No. 103.
[0235] Other biomarker gene primers and probes use the same logical design and are not shown.
[0236] 4. Results Analysis
[0237] The following examples illustrate the use of primers for amplifying SEQ ID No. 115 of the ELMO1 gene (SEQ ID No. 19 and SEQ ID No. 20), with probe SEQ ID No. 107, and primers for amplifying SEQ ID No. 116 of the KCNA3 gene (SEQ ID No. 77 and SEQ ID No. 78), with probe SEQ ID No. 113:
[0238] Methylation-specific multiplex fluorescent PCR was used to detect the methylation status of target genes. This technique is based on the principle of bisulfite conversion. After extraction of cell-free DNA from plasma samples, bisulfite conversion occurs, converting unmethylated cytosine to uracil, while methylated cytosine remains unchanged. Subsequent detection uses Bis-DNA as a template. During the detection process, specific primers are used to specifically amplify the methylated target gene, and probes labeled with different fluorescence report the amplification signal. Primers designed for conserved regions are used to detect the internal control gene ACTB, and probes labeled with the corresponding fluorescence report the amplification signal. Detection of the internal control gene can be used to monitor whether the sample Bis-DNA quantity is sufficient and whether the pretreatment process is up to standard.
[0239] The threshold for Ct values of each gene is 40.
[0240] Among them, when the Ct value of the ACTB gene is greater than 40, it is judged as a quality control failure (the result is unreliable);
[0241] When the Ct value of a gastric cancer marker gene is greater than 40, it is considered negative (i.e., not a gastric cancer patient); otherwise, it is considered positive (i.e., a gastric cancer patient).
[0242] ROC curves for each biomarker gene Ct were plotted using SPSS software. The ROC curves are shown in Figure 4. The performance and area under the curve (AUC) of each gene are shown in Table 7.
[0243] Table 7. Performance and area under curve of each gene
[0244] Among them, the ELMO1 and KCNA3 genes exhibited the best diagnostic performance. To further improve detection performance, the Ct values of the two genes were integrated for analysis, achieving an AUC of 0.867. When the following criteria were applied: Ct(ELMO1) ≤ 40 and Ct(KCNA3) ≤ 40, the gene was considered positive; if either gene was positive, the sample was considered positive. Under these criteria, the specificity for non-gastric cancer samples was 68.7%, while the sensitivity for gastric cancer detection reached 83.7% (Figure 5).
[0245] The above results demonstrate that the present invention, through the combined detection of multiple genes, achieves better detection performance than single-gene detection. Furthermore, the combination of KCNA3 and ELMO1 genes exhibits optimal performance in the detection of gastric cancer.
[0246] Industrial applications
[0247] This invention provides 28 methylation biomarkers (differential methylation regions, DMRs) that can be used to detect gastric cancer (Table 1). From these 28 DMRs, the ELMO1 and KCNA3 genes were selected as high-performance methylation biomarkers for gastric cancer detection. By detecting these DMRs and the methylation levels of the ELMO1 and / or KCNA3 genes, and analyzing the obtained data, the likelihood of a subject developing gastric cancer can be predicted, thereby achieving the goal of screening for gastric cancer in the general population or high-risk groups. This invention provides an accurate, simple, and economical method for early gastric cancer screening, which can improve the detection rate of gastric cancer, especially early-stage gastric cancer, in high-risk groups and general health check-up populations, thereby improving the survival rate of gastric cancer patients and saving significant medical expenses and reducing the medical burden.
Claims
1. Differentially methylated region group, containing all or part of the following 28 differentially methylated regions: (A1) is located at positions 23847173 to 23848069 on chromosome 16; (A2) is located at positions 58497626 to 58497874 on chromosome 16; (A3) is located at positions 3785618 to 3786199 on chromosome 19; (A4) is located at positions 133464016 to 133464255 on chromosome 12; (A5) is located at positions 182321660 to 182322797 on chromosome 2; (A6) is located at positions 27940318 to 27940640 on chromosome 17; (A7) is located at positions 49814781 to 49815692 on chromosome 7; (A8) is located at positions 37487610 to 37489066 on chromosome 7; (A9) is located at positions 87229481 to 87230462 on chromosome 7; (A10) is located at positions 132382356 to 132383095 on chromosome 9; (A11) is located at positions 29106420 to 29107080 on chromosome 13; (A12) is located at positions 102068552 to 102068996 on chromosome 13; (A13) is located at positions 123301050 to 123302058 on chromosome 11; (A14) is located at positions 114695178 to 114695581 on chromosome 1; (A15) is located at positions 67344498 to 67345106 on chromosome 8; (A16) is located at positions 752063 to 752203 on chromosome 7; (A17) is located at positions 29337988 to 29338749 on chromosome 2; (A18) is located at positions 111217001 to 111217968 on chromosome 1; (A19) is located at positions 7450242 to 7451419 on chromosome 10; (A20) is located at positions 36909233 to 36910064 on chromosome 19; (A21) is located at positions 48936828 to 48938691 on chromosome 15; (A22) is located at positions 68546355 to 68547096 on chromosome 2; (A23) is located at positions 61808614 to 61809996 on chromosome 20; (A24) is located at positions 108507504 to 108508116 on chromosome 1; (A25) is located at positions 58238586 to 58239135 on chromosome 19; (A26) is located at positions 58951350 to 58951895 on chromosome 19; (A27) is located at positions 144694307 to 144695311 on chromosome 2; (A28) is located at positions 129693355 to 129694176 on chromosome 3; The physical locations of the 28 differentially methylated regions were determined based on the hg19 alignment of the human whole genome sequence.
2. The differentiated methylation region set according to claim 1, characterized in that: The group of differentially methylated regions consists of all or part of the 28 differentially methylated regions shown in claims 1 (A1)-(A28); Furthermore, the group of differentially methylated regions consists of the two differentially methylated regions shown in (A8) and / or (A18).
3. A kit comprising a substance for detecting the group of differentially methylated regions as described in claim 1 or 2; Furthermore, the kit includes substances for detecting ELMO1 gene methylation levels and / or substances for detecting KCNA3 gene methylation levels; or Furthermore, the kit includes substances for detecting the nucleotide sequence of the ELMO1 gene gastric cancer-specific methylated region and / or substances for detecting the nucleotide sequence of the KCNA3 gene gastric cancer-specific methylated region.
4. The reagent kit according to claim 3, characterized in that: The substances used to detect the methylation level of the ELMO1 gene are primer pair A and probe A; the substances used to detect the methylation level of the KCNA3 gene are primer pair B and probe B. The primer pair A consists of two single-stranded DNAs as shown in any one of SEQ ID No. 3 to SEQ ID No. 52; the probe A is a single-stranded DNA as shown in any one of SEQ ID No. 104 to SEQ ID No. 108; The primer pair B consists of two single-stranded DNAs as shown in any one of SEQ ID No. 53 to SEQ ID No. 102; the probe B is a single-stranded DNA as shown in any one of SEQ ID No. 109 to SEQ ID No.
114.
5. The kit according to claim 3 or 4, characterized in that: The primer pair A consists of two single-stranded DNAs as shown in SEQ ID No. 19 and SEQ ID No. 20; the probe A is the single-stranded DNA as shown in SEQ ID No.
107. The primer pair B consists of two single-stranded DNAs as shown in SEQ ID No. 77 and SEQ ID No. 78; the probe B is the single-stranded DNA as shown in SEQ ID No.
113.
6. The kit according to any one of claims 3-5, characterized in that: The kit also contains primer pair C and probe C for amplifying the internal reference gene ACTB; primer pair C consists of two single-stranded DNAs as shown in SEQ ID No. 1 and SEQ ID No. 2; probe C is the single-stranded DNA as shown in SEQ ID No.
103.
7. The kit according to any one of claims 3-6, characterized in that: The kit has any of the following uses: diagnosing or screening for cancer; providing early warning of cancer before clinical symptoms appear; or differentiating or assisting in the differentiation between cancer and benign lesions. and / or The non-cancer patients were either healthy controls or had benign lesions; The cancers mentioned are stomach cancer, liver cancer, intestinal cancer, lung cancer, pancreatic cancer, prostate cancer, esophageal cancer, or urothelial carcinoma; Furthermore, the cancer is stomach cancer; the benign lesion is a benign lesion of the stomach.
8. Any of the following applications: I) Application of the ELMO1 gene and / or KCNA3 gene as methylation markers in any of the following (C1)-(C6); II) Substances used to detect the methylation levels of the ELMO1 and / or KCNA3 genes are listed in (C1). -(C6) Any of the following applications; III) The use of the differentiated methylation region set as a methylation marker as described in claim 1 or 2 in any of the following (C1)-(C6); IV) Use of a substance for detecting the methylation level of the differentially methylated region group as described in claim 1 or 2 in any of the following (C1)-(C6); V) The use of a combination of substance and medium in any of the following (C1)-(C6); said substance is a substance for detecting the group of differentially methylated regions as described in claim 1 or 2; said medium stores a method for constructing and using a cancer risk prediction model; VI) The application of a medium containing methods for constructing and using cancer risk prediction models in any of the following (C1)-(C6); (C1) Prepare products for the diagnosis or auxiliary diagnosis of cancer; (C2) Diagnosing or assisting in the diagnosis of cancer; (C3) Prepare products for the early warning of cancer before clinical symptoms appear; (C4) Providing early warning of cancer before clinical symptoms appear; (C5) To prepare products for distinguishing or assisting in the distinction between cancer and benign lesions; (C6) To differentiate or assist in differentiating between cancer and benign lesions; In V) and VI), the method for constructing and using the cancer risk prediction model includes the following steps: (D1) Construct a training set to obtain methylation level data for the differentially methylated region group as described in claim 1 or 2, including samples from known cancer patients and known non-cancer patients. (D2) A cancer risk prediction model is constructed using machine learning, and then the cancer risk prediction model is used to diagnose or assist in the diagnosis of cancer and / or provide early warning of cancer before clinical symptoms and / or differentiate or assist in the differentiation of cancer and benign lesions.
9. The application according to claim 8, characterized in that: The methylation level of the ELMO1 gene is the methylation level of all or part of the CpG sites in the DNA fragment shown in SEQ ID No. 115; and / or The methylation level of the KCNA3 gene is the methylation level of all or part of the CpG sites in the DNA fragment shown in SEQ ID No.
116.
10. The application according to claim 8 or 9, characterized in that: The substances used to detect the methylation level of the ELMO1 gene are primer set A and probe A; primer set A consists of two single-stranded DNAs as shown in any one of SEQ ID No. 3 to SEQ ID No. 52; probe A is any one of the single-stranded DNAs shown in SEQ ID No. 104 to SEQ ID No. 108; Furthermore, the primer set A consists of two single-stranded DNAs shown in SEQ ID No. 19 and SEQ ID No. 20; the probe A is the single-stranded DNA shown in SEQ ID No. 107; and / or The substances used to detect the methylation level of the KCNA3 gene are primer set B and probe B; primer set B consists of two single-stranded DNAs as shown in any one of SEQ ID No. 53 to SEQ ID No. 102; probe B is any one of the single-stranded DNAs shown in SEQ ID No. 109 to SEQ ID No. 114; Furthermore, the primer set B consists of two single-stranded DNAs shown in SEQ ID No. 77 and SEQ ID No. 78; the probe B is the single-stranded DNA shown in SEQ ID No.
113.
11. The application according to claim 8, characterized in that: The machine learning method mentioned is the random forest method.
12. The application according to claim 8, characterized in that: The sample is one from which DNA can be extracted; Furthermore, the sample is plasma, tissue, saliva, urine, and / or feces.
13. The application according to claim 8, characterized in that: The method for obtaining the methylation level data is selected from post-bisulfite conversion detection, restriction enzyme digestion technology, methylation-specific PCR, pyrosequencing, high-throughput sequencing, third-generation sequencing, or single-molecule sequencing.
14. The application according to any one of claims 8-13, characterized in that: The non-cancer patients were either healthy controls or had benign lesions; The cancers mentioned are stomach cancer, liver cancer, intestinal cancer, lung cancer, pancreatic cancer, prostate cancer, esophageal cancer, or urothelial carcinoma; Furthermore, the cancer is stomach cancer; the benign lesion is a benign lesion of the stomach.
15. A data processing apparatus having any of the following functions: (E1) Diagnosing or assisting in the diagnosis of cancer; (E2) Providing early warning of cancer before clinical symptoms appear; (E3) To differentiate or help differentiate between cancer and benign lesions; The data processing device includes unit X and unit Y; The unit X is used to establish a cancer risk prediction model, including a data acquisition module A and a data analysis and processing module; The data acquisition module is used to acquire methylation level data for the differential methylation region group as described in claim 1 or 2 from known cancer patient samples and known non-cancer patient samples. The data analysis and processing module A can use the methylation level data of the differential methylation region group as described in claim 1 or 2 collected by the data acquisition module from known cancer patient samples and known non-cancer patient samples as a training set to construct a cancer risk prediction model based on the principle of machine learning. The unit Y is capable of diagnosing or assisting in the diagnosis of cancer and / or providing early warning of cancer before clinical symptoms and / or differentiating or assisting in the differentiation of cancer from benign lesions, based on the cancer risk prediction model and methylation level data from the subject sample for the differentially methylated region group as described in claim 1 or 2.
16. The system, characterized in that: The system is either system A or system B; System A includes: (F1) Reagents and / or instruments for detecting the methylation level of the differentially methylated region group as described in claim 1 or 2; (F2) A data processing device, the data processing device comprising unit X and unit Y; Unit X is used to establish a cancer risk prediction model, including data acquisition module A and data analysis. Processing module; The data acquisition module A is used to acquire methylation level data for the differential methylation region group as described in claim 1 or 2, obtained from known cancer patient samples and known non-cancer patient samples using (F1) detection. The data analysis and processing module can use the methylation level data of the differential methylation region group as described in claim 1 or 2 collected by the data acquisition module from known cancer patient samples and known non-cancer patient samples as a training set to construct a cancer risk prediction model based on the principle of machine learning. The unit Y is capable of diagnosing or assisting in the diagnosis of cancer and / or providing early warning of cancer before clinical symptoms and / or differentiating or assisting in the differentiation of cancer and benign lesions based on the cancer risk prediction model and methylation level data of the differential methylation region group according to claim 1 or 2 from the sample of the test subject. System B includes: (G1) The kit and real-time quantitative PCR instrument as described in any one of claims 3-7; (G2) Device, the device includes a data acquisition module B, a threshold storage module, a data comparison module, and a data processing and conclusion output module; The data acquisition module B is configured to acquire real-time fluorescence quantitative PCR amplification result data from the sample obtained by the test subject through (G1) detection; The threshold storage module is configured to store threshold A, threshold B, and threshold C; threshold A is the threshold value of the Ct value of the ACTB gene; threshold B is the threshold value of the Ct value of the ELMO1 gene; and threshold C is the threshold value of the Ct value of the KCNA3 gene. The data comparison module is configured to receive real-time quantitative PCR amplification results data of the sample from the test subject sent by the data acquisition module, and to call the threshold A, threshold B, and threshold C stored in the threshold storage module. Then, it compares the Ct value of the ACTB gene of the test subject with the threshold A, compares the Ct value of the ELMO1 gene of the test subject with the threshold B, and compares the Ct value of the KCNA3 gene of the test subject with the threshold C. The data processing and conclusion output module is configured to receive the comparison results sent from the data comparison module, and then output the conclusion as follows: If the Ct value of the ACTB gene in the test subject is greater than the threshold A, the result is deemed unreliable; when the Ct value of the ACTB gene is less than or equal to the threshold A, the result is deemed reliable, and further determination is made as follows: If the Ct value of the ELMO1 gene of the test subject is greater than the threshold B, and / or the Ct value of the KCNA3 gene of the test subject is greater than the threshold C, then the test subject is determined to be or a candidate non-cancer patient, or the test subject is a low-risk cancer patient; otherwise, the test subject is determined to be or a candidate cancer patient, or the test subject is a high-risk cancer patient.
17. The system according to claim 16, characterized in that: The thresholds A, B, and C are all 40.
18. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program for performing the following steps: Methylation level data for the differentially methylated region group as described in claim 1 or 2 were collected from known cancer patient samples and known non-cancer patient samples; The methylation level data of the differential methylation region group as described in claim 1 or 2 obtained from known cancer patient samples and known non-cancer patient samples are used as the training set to construct a cancer risk prediction model based on the principle of machine learning. Based on the cancer risk prediction model and methylation level data from the subject sample for the differentially methylated region group as described in claim 1 or 2, it is possible to diagnose or assist in the diagnosis of cancer and / or provide early warning of cancer before clinical symptoms and / or differentiate or assist in the differentiation of cancer from benign lesions.
19. Any of the following methods: Method I: A method for diagnosing or assisting in the diagnosis of cancer, comprising the steps of: analyzing the methylation status of a group of differentially methylated regions according to claim 1 or 2 from a sample of a subject, thereby achieving the diagnosis or assisting in the diagnosis of cancer; Furthermore, method I includes the following steps: (H1) Construct a training set to obtain methylation level data for the differentially methylated region group as described in claim 1 or 2, including samples from known cancer patients and known non-cancer patients; (H2) A cancer risk prediction model is constructed using machine learning, and then the cancer risk prediction model is used to diagnose or assist in the diagnosis of cancer. Method II: A method for early warning of cancer before clinical symptoms, comprising the steps of: analyzing the methylation status of a group of differentially methylated regions as described in claim 1 or 2 from a sample of a subject, thereby enabling early warning of cancer before clinical symptoms; Furthermore, method II includes the following steps: (I1) Construct a training set to obtain methylation level data for the differentially methylated region group as described in claim 1 or 2, including samples from known cancer patients and known non-cancer patients; (I2) A cancer risk prediction model is constructed using machine learning, and then the cancer risk prediction model is used to provide early warning of cancer before clinical symptoms appear. Method III: A method for distinguishing or assisting in the distinction between cancer and benign lesions, comprising the following steps: analyzing the methylation status of a differentially methylated region group as described in claim 1 or 2 from a sample of a subject, thereby distinguishing or assisting in the distinction between cancer and benign lesions; Furthermore, method III includes the following steps: (J1) Construct a training set to obtain methylation level data for the differentially methylated region group as described in claim 1 or 2, including samples from known cancer patients and samples from known corresponding benign lesion patients; (J2) A cancer risk prediction model is constructed using machine learning, and then the cancer risk prediction model is used to distinguish or assist in distinguishing between cancer and benign lesions; Method IV: A method for diagnosing or assisting in the diagnosis of cancer, comprising the steps of: detecting the methylation levels of the ELMO1 gene and / or KCNA3 gene in a sample from a subject, thereby achieving diagnosis or assisting in the diagnosis. Diagnosing cancer; Method V: A method for early warning of cancer before clinical symptoms, comprising the steps of: detecting the methylation levels of the ELMO1 gene and / or KCNA3 gene in a sample from a subject, thereby enabling early warning of cancer before clinical symptoms; Method VI: A method for distinguishing or assisting in the distinction between cancer and benign lesions, comprising the following steps: detecting the methylation levels of the ELMO1 gene and / or KCNA3 gene in a sample from a subject, thereby distinguishing or assisting in the distinction between cancer and benign lesions.
20. The method according to claim 19, characterized in that: The methylation level of the ELMO1 gene is the methylation level of all or part of the CpG sites in the DNA fragment shown in SEQ ID No. 115; and / or The methylation level of the KCNA3 gene is the methylation level of all or part of the CpG sites in the DNA fragment shown in SEQ ID No.
116.
21. The method according to claim 19 or 20, characterized in that: The method for detecting the methylation level of the ELMO1 gene and / or KCNA3 gene in the sample from the subject is selected from bisulfite conversion detection, restriction enzyme digestion technology, methylation-specific PCR, pyrosequencing, high-throughput sequencing, third-generation sequencing or single-molecule sequencing.
22. The method according to any one of claims 19-21, characterized in that: The methylation levels of the ELMO1 and / or KCNA3 genes in the sample from the subject were detected according to a method comprising the following steps: (K1) DNA was extracted from the sample and then methylated; (K2) DNA obtained by real-time quantitative PCR amplification of (K1) after methylation conversion; The primer pair A for the ELMO1 gene used in the real-time quantitative PCR amplification consists of two single-stranded DNAs as shown in any one of SEQ ID No. 3 to SEQ ID No. 52, and the probe A is any one of the single-stranded DNAs shown in SEQ ID No. 104 to SEQ ID No.
108. The primer pair B for the KCNA3 gene amplification used in the real-time quantitative PCR amplification consists of two single-stranded DNAs as shown in any one of SEQ ID No. 53 to SEQ ID No. 102, and the probe B is any one of the single-stranded DNAs shown in SEQ ID No. 109 to SEQ ID No.
114.
23. The method according to claim 22, characterized in that: The primer pair A for the ELMO1 gene used in the real-time quantitative PCR amplification consists of two single-stranded DNAs as shown in SEQ ID No. 19 and SEQ ID No. 20; the probe A is the single-stranded DNA shown in SEQ ID No.
107. The primer pair B for the KCNA3 gene used in the real-time quantitative PCR amplification consists of two single-stranded DNAs as shown in SEQ ID No. 77 and SEQ ID No. 78; the probe B is the single-stranded DNA shown in SEQ ID No.
113.
24. The method according to claim 22 or 23, characterized in that: In step (K2), the internal reference used in the real-time quantitative PCR amplification is the ACTB gene. The primer pair C used to amplify the ACTB gene consists of two single-stranded DNAs as shown in SEQ ID No. 1 and SEQ ID No. 2, and the probe C is the single-stranded DNA as shown in SEQ ID No.
103.
25. The method according to any one of claims 22-24, characterized in that: Following step (K2), the following steps are also included: (K3) The results of the real-time quantitative PCR amplification described in (K2) shall be judged as follows: When the Ct value of the ACTB gene is greater than 40, the result is considered unreliable; when the Ct value of the ACTB gene is ≤40, the result is considered reliable, and further determination is made as follows: When the Ct value of the ELMO1 gene and / or the Ct value of the KCNA3 gene are greater than 40, the subject is determined to be or a candidate for non-cancer patients, or the subject is a low-risk cancer patient. Conversely, the test subject is determined to be a cancer patient or a candidate for cancer, or a high-risk cancer patient.
26. The method according to any one of claims 19-25, characterized in that: The sample is one from which DNA can be extracted; Furthermore, the samples are plasma, tissue, saliva, urine, and feces.
27. The method according to any one of claims 19-26, characterized in that: The methods for analyzing the methylation status and obtaining the methylation level data are selected from post-bisulfite conversion detection, restriction enzyme digestion technology, methylation-specific PCR, pyrosequencing, high-throughput sequencing, third-generation sequencing, or single-molecule sequencing.
28. The method according to any one of claims 19-27, characterized in that: The non-cancer patients were either healthy controls or had benign lesions; The cancers mentioned are stomach cancer, liver cancer, intestinal cancer, lung cancer, pancreatic cancer, prostate cancer, esophageal cancer, or urothelial carcinoma; Furthermore, the cancer is stomach cancer; the benign lesion is a benign lesion of the stomach.