Differentially methylated region combination and use

Through differentiated methylation region groups and machine learning models, the problem of insufficient sensitivity in traditional early diagnosis of pancreatic cancer was solved, high-accuracy early pancreatic cancer detection and warning was achieved, and the specificity and sensitivity of diagnosis were improved.

WO2025194302A1PCT designated stage Publication Date: 2025-09-25BGI GENOMICS CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/082172
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-18
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

Traditional early diagnosis technologies for pancreatic cancer lack sensitivity and are difficult to accurately detect early pancreatic cancer. Imaging examinations and tumor marker tests have limitations, low specificity, and are prone to false positives.

Method used

The differentially methylated region group (DMR) is used to detect the methylation level in blood or other samples using a bisulfite reagent, and a cancer risk prediction model is constructed in combination with machine learning methods for the diagnosis and early warning of pancreatic cancer.

Benefits of technology

It has improved the accuracy of early diagnosis of pancreatic cancer, especially the detection rate of early pancreatic cancer, which is significantly better than traditional methods, reduces the false positive rate, and improves the survival rate and medical efficiency of pancreatic cancer patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure PCTCN2024082172-FTAPPB-I100001
    Figure PCTCN2024082172-FTAPPB-I100001
  • Figure PCTCN2024082172-FTAPPB-I100002
    Figure PCTCN2024082172-FTAPPB-I100002
  • Figure PCTCN2024082172-FTAPPB-I100003
    Figure PCTCN2024082172-FTAPPB-I100003
Patent Text Reader

Abstract

Provided are a differentially methylated region combination and a use. Provided are 13 differentially methylated regions for diagnosing or assisting in diagnosing cancers. Provided is an accurate, convenient, and economical means for early screening of pancreatic cancer, which can increase the detection rate of pancreatic cancer, especially early pancreatic cancer, in populations at high risk of pancreatic cancer and populations receiving general physical examinations, thereby increasing the survival rate of patients with pancreatic cancer, saving the medical expenditure, and reducing the medical burden.
Need to check novelty before this filing date? Find Prior Art

Description

Differentially methylated region combination and application Technical Field

[0001] The present invention relates to the field of biomedicine, and in particular to differentially methylated region combinations and applications. Background Art

[0002] Traditional early detection techniques for pancreatic cancer mainly include imaging examinations and tumor marker detection.

[0003] Imaging tests, including ultrasound, CT scans, and MRI, are commonly used to screen for early pancreatic cancer. Ultrasound is a noninvasive, non-radioactive test that can visualize the morphology and structure of the pancreas, but its detection rate for early pancreatic cancer is low. CT scans and MRIs can provide more detailed images of the pancreas, but their detection rate for small tumors is also limited.

[0004] Tumor marker testing assesses pancreatic cancer risk by detecting specific proteins or other molecules in the blood. Commonly used pancreatic cancer markers include CA19-9 and CEA. However, tumor marker test results may be affected by other factors, lack sensitivity for early-stage pancreatic cancer, and have low specificity for pancreatitis, hepatitis, and obstructive diseases. Therefore, they cannot be used as the sole basis for the diagnosis of pancreatic cancer.

[0005] Traditional pancreatic cancer early detection technologies have certain limitations in early diagnosis. Therefore, more sensitive and accurate early pancreatic cancer screening technologies, such as early pancreatic cancer detection methods based on multi-omics technologies, are currently being researched and developed.

[0006] Traditional pancreatic cancer early detection technologies have limitations in early diagnosis. Early pancreatic cancer often has no obvious symptoms, and the pancreas is located deep within the body, making it difficult to accurately visualize with imaging tests. Broad-spectrum tumor markers such as CEA and CA19-9 have low specificity and are prone to false positives. They also suffer from insufficient sensitivity for early pancreatic cancer and low specificity for pancreatitis, hepatitis, and obstructive diseases.

[0007] Invention Disclosure

[0008] The technical problem to be solved by the present invention is to provide a more accurate method for diagnosing pancreatic cancer, especially an early diagnosis method, and a corresponding methylation marker.

[0009] The present invention claims a set of differentially methylated regions (DMRs).

[0010] The group of differentially methylated regions claimed in the present invention includes all or part of the following 13 differentially methylated regions (as shown in Table 1):

[0011] (A1) located at positions 111217374-111217715 on chromosome 1;

[0012] (A2) located at positions 248020350-248020956 on chromosome 1;

[0013] (A3) located at positions 12490772-12490958 on chromosome 10;

[0014] (A4) located at positions 37005986-37006281 on chromosome 13;

[0015] (A5) located at positions 65701719-65701941 on chromosome 15;

[0016] (A6) located at positions 124510649-124510809 on chromosome 3;

[0017] (A7) located at positions 15672531-15672777 on chromosome 1;

[0018] (A8) located at positions 153518054-153518476 on chromosome 1;

[0019] (A9) is located at positions 92907747-92908353 on chromosome 5;

[0020] (A10) is located at positions 49813377-49813880 on chromosome 7;

[0021] (A11) is located at positions 170629389-170630397 on chromosome 1;

[0022] (A12) is located at positions 15568312-15568674 on chromosome 19;

[0023] (A13) is located at positions 16318860-16323551 on chromosome 17.

[0024] The physical locations of the 13 differentially methylated regions were determined based on the alignment of the human whole genome sequence (version number is hg19).

[0025] The differentially methylated region group includes but is not limited to selecting the aforementioned subsets (A1)-(A13), small-scale replacements or small-scale additions, etc.

[0026] Furthermore, the differentially methylated region group may consist of all or part of the 13 differentially methylated regions shown in (A1) to (A13) above.

[0027] In one embodiment of the present invention, the differentially methylated region group consists of all of the 13 differentially methylated regions shown in (A1) to (A13) above.

[0028] The present invention also claims the use of the above-mentioned differentially methylated region group as a methylation marker in any of the following:

[0029] (B1) preparing products for diagnosing or assisting in the diagnosis of cancer;

[0030] (B2) diagnosis or assistance in diagnosis of cancer;

[0031] (B3) preparing products for early warning of cancer before clinical symptoms;

[0032] (B4) Warning of cancer before clinical symptoms.

[0033] The present invention also claims the use of a substance for detecting the methylation level of the differentially methylated region group described above in any of the following:

[0034] (B1) preparing products for diagnosing or assisting in the diagnosis of cancer;

[0035] (B2) diagnosis or assistance in diagnosis of cancer;

[0036] (B3) preparing products for early warning of cancer before clinical symptoms;

[0037] (B4) Warning of cancer before clinical symptoms.

[0038] The substance used to detect the aforementioned differentially methylated region group may include a bisulfite reagent.

[0039] The present invention also claims the use of the combination of the substance and the medium in any of the following:

[0040] (B1) preparing products for diagnosing or assisting in the diagnosis of cancer;

[0041] (B2) diagnosis or assistance in diagnosis of cancer;

[0042] (B3) preparing products for early warning of cancer before clinical symptoms;

[0043] (B4) Warning of cancer before clinical symptoms.

[0044] The substance is a substance used to detect the differentially methylated region group described above;

[0045] Furthermore, the substance used to detect the aforementioned differentially methylated region group may include a bisulfite reagent.

[0046] The medium stores a method for constructing and using a cancer risk prediction model;

[0047] The method for constructing and using the cancer risk prediction model comprises the following steps:

[0048] (C1) obtaining methylation level data of the differentially methylated region group described in the first aspect of the present invention for known cancer patient samples and known non-cancer patient samples as training samples;

[0049] (C2) constructing a cancer risk prediction model using the training samples, and then utilizing the cancer risk prediction model to diagnose or assist in the diagnosis of cancer and / or to provide early warning of cancer before clinical symptoms occur.

[0050] The number of the known cancer patients and the known non-cancer patients may be more than 50.

[0051] The present invention also protects the use of a medium storing a method for constructing and using a cancer risk prediction model in any of the following:

[0052] (B1) preparing products for diagnosing or assisting in the diagnosis of cancer;

[0053] (B2) diagnosis or assistance in diagnosis of cancer;

[0054] (B3) preparing products for early warning of cancer before clinical symptoms;

[0055] (B4) Warning of cancer before clinical symptoms.

[0056] The method for constructing and using the cancer risk prediction model comprises the following steps:

[0057] (C1) obtaining methylation level data of the differentially methylated region group described in the first aspect of the present invention for known cancer patient samples and known non-cancer patient samples as training samples;

[0058] (C2) constructing a cancer risk prediction model using the training samples, and then utilizing the cancer risk prediction model to diagnose or assist in the diagnosis of cancer and / or to provide early warning of cancer before clinical symptoms occur.

[0059] The number of the known cancer patients and the known non-cancer patients may be more than 50.

[0060] In the applications described above, all of the cancer risk prediction models constructed using the training samples may adopt a machine learning method; the machine learning method may be a random forest method.

[0061] In the applications described above, all of the samples can be samples from which DNA can be extracted.

[0062] Furthermore, the sample may be plasma, tissue, saliva, urine and / or feces.

[0063] In the aforementioned applications, all methods for obtaining the methylation level data include, but are not limited to, bisulfite conversion, endonuclease digestion technology, methylation-specific PCR (MS-PCR), pyrosequencing, high-throughput sequencing, third-generation sequencing, or single-molecule sequencing.

[0064] In the aforementioned applications, all of the non-cancer patients may be healthy controls.

[0065] In the aforementioned applications, all of the cancers may be pancreatic cancer, liver cancer, colorectal cancer, gastric cancer, esophageal cancer or lung cancer.

[0066] The present invention also claims protection for a kit, which may be specifically the following kit I, kit II or kit III:

[0067] Kit 1, comprising:

[0068] (a1) a bisulfite reagent; and

[0069] (a2) A control nucleic acid comprising a sequence from the aforementioned differentially methylated region group and having a methylation status associated with a non-cancer patient.

[0070] Kit II, containing:

[0071] (b1) a bisulfite reagent; and

[0072] (b2) a control nucleic acid comprising a sequence from the aforementioned differentially methylated region group and having a methylation status associated with a cancer patient.

[0073] Kit III, containing substances for detecting the differentially methylated region group described above and a medium storing methods for constructing and using a cancer risk prediction model;

[0074] Furthermore, the substance used to detect the aforementioned differentially methylated region group may include a bisulfite reagent.

[0075] The method for constructing and using the cancer risk prediction model comprises the following steps:

[0076] (C1) obtaining methylation level data of the differentially methylated region group described in the first aspect of the present invention for known cancer patient samples and known non-cancer patient samples as training samples;

[0077] (C2) constructing a cancer risk prediction model using the training samples, and then utilizing the cancer risk prediction model to diagnose or assist in the diagnosis of cancer and / or to provide early warning of cancer before clinical symptoms occur.

[0078] The number of the known cancer patients and the known non-cancer patients may be more than 50.

[0079] The present invention also claims protection for a device having any of the following functions:

[0080] (D1) Diagnosis or assistance in diagnosis of cancer;

[0081] (D2) Warning of cancer before clinical symptoms.

[0082] The device comprises:

[0083] M1, a cancer risk prediction model construction module: configured to construct a cancer risk prediction model using the methylation level data of the differentially methylated region group described in the first aspect of the above-mentioned known cancer patient samples and known non-cancer patient samples as training samples;

[0084] M2, a data receiving module: configured to receive sample data, wherein the sample data is the methylation level data of the differentially methylated region group described in the first aspect of the subject sample;

[0085] M3, data analysis and processing module: configured to obtain results based on the cancer risk prediction model and the sample data; the results are the results of diagnosing or assisting in the diagnosis of cancer and / or the results of warning of cancer before clinical symptoms.

[0086] The number of the known cancer patients and the known non-cancer patients may be more than 50.

[0087] The present invention also claims a system.

[0088] The system claimed in the present invention comprises:

[0089] (E1) Reagents and / or instruments for detecting the methylation levels of the differentially methylated region group described above;

[0090] Furthermore, the reagent for detecting the aforementioned differentially methylated region group may include a bisulfite reagent.

[0091] (E2) a device comprising:

[0092] M1, a cancer risk prediction model construction module: configured to construct a cancer risk prediction model using the methylation level data of the differentially methylated region group described in the first aspect of the above-mentioned known cancer patient samples and known non-cancer patient samples as training samples;

[0093] M2, a data receiving module: configured to receive sample data, wherein the sample data is the methylation level data of the differentially methylated region group described in the first aspect of the subject sample;

[0094] M3, data analysis and processing module: configured to obtain results based on the cancer risk prediction model and the sample data; the results are the results of diagnosing or assisting in the diagnosis of cancer and / or the results of warning of cancer before clinical symptoms.

[0095] The number of the known cancer patients and the known non-cancer patients may be more than 50.

[0096] The present invention also claims protection for a computer-readable storage medium.

[0097] The computer-readable storage medium claimed in the present invention stores a computer program, which, when executed by a processor, implements the following steps:

[0098] Using the methylation level data of the differentially methylated region group described in the first aspect of the above-mentioned known cancer patient samples and known non-cancer patient samples as training samples, a cancer risk prediction model is constructed;

[0099] receiving sample data, wherein the sample data is the methylation level data of the differentially methylated region group described in the first aspect above of the subject sample;

[0100] Based on the cancer risk prediction model and the sample data, a result is obtained; the result is a result of diagnosing or assisting in the diagnosis of cancer and / or a result of warning of cancer before clinical symptoms occur.

[0101] The number of the known cancer patients and the known non-cancer patients may be more than 50.

[0102] The computer-readable storage medium refers to a carrier for storing data, which may be a tape, disk, floppy disk, optical disk, magneto-optical disk, ROM, PROM, VCD, DVD, hard disk, flash memory, USB flash drive, CF card, SD card, MMC card, SM card, memory stick (Memory Stick) or xD card, etc.

[0103] In the above, a machine learning method may be used when constructing the cancer risk prediction model using the training samples; the machine learning method may be a random forest method.

[0104] In the above, the sample may be a sample from which DNA can be extracted.

[0105] Furthermore, the sample may specifically be plasma, tissue, saliva, urine and / or feces.

[0106] In the above, the methods for obtaining the methylation level data include but are not limited to bisulfite conversion, endonuclease digestion technology, methylation-specific PCR (MS-PCR), pyrosequencing, high-throughput sequencing, third-generation sequencing or single-molecule sequencing, etc.

[0107] In the above, the non-cancer patients may be healthy controls.

[0108] In the above, the cancer may be pancreatic cancer, liver cancer, colorectal cancer, gastric cancer, esophageal cancer or lung cancer.

[0109] The present invention also claims protection for the use of the above-mentioned kit or data processing device or system or computer-readable storage medium in any of the following:

[0110] (B1) preparing products for diagnosing or assisting in the diagnosis of cancer;

[0111] (B2) diagnosis or assistance in diagnosis of cancer;

[0112] (B3) preparing products for early warning of cancer before clinical symptoms;

[0113] (B4) Warning of cancer before clinical symptoms.

[0114] The cancer may be pancreatic cancer, liver cancer, colorectal cancer, gastric cancer, esophageal cancer or lung cancer.

[0115] The present invention also claims a method for diagnosing or assisting in the diagnosis of cancer.

[0116] The method for diagnosing or assisting in the diagnosis of cancer claimed in the present invention may include the following steps:

[0117] N1. Constructing a cancer risk prediction model: Using the methylation level data of the differentially methylated region group described in the first aspect of the above-mentioned known cancer patient samples and known non-cancer patient samples as training samples, constructing a cancer risk prediction model;

[0118] N2. Data reception: receiving sample data, wherein the sample data is the methylation level data of the differentially methylated region group described in the first aspect above of the subject sample;

[0119] N3. Data processing and result output: Based on the cancer risk prediction model and the sample data, obtain and output the result; the result is the result of diagnosing or assisting in the diagnosis of cancer.

[0120] The number of the known cancer patients and the known non-cancer patients may be more than 50.

[0121] In the above, a machine learning method may be used when constructing the cancer risk prediction model using the training samples; the machine learning method may be a random forest method.

[0122] In the above, the sample may be a sample from which DNA can be extracted.

[0123] Furthermore, the sample may specifically be plasma, tissue, saliva, urine and / or feces.

[0124] In the above, the methods for obtaining the methylation level data include but are not limited to bisulfite conversion, endonuclease digestion technology, methylation-specific PCR (MS-PCR), pyrosequencing, high-throughput sequencing, third-generation sequencing or single-molecule sequencing, etc.

[0125] In the above, the non-cancer patients may be healthy controls.

[0126] In the above, the cancer may be pancreatic cancer, liver cancer, colorectal cancer, gastric cancer, esophageal cancer or lung cancer.

[0127] The present invention also claims a method for early warning of cancer before clinical symptoms.

[0128] The method for early warning of cancer before clinical symptoms are present as claimed in the present invention may comprise the following steps:

[0129] N1. Constructing a cancer risk prediction model: Using the methylation level data of the differentially methylated region group described in the first aspect of the above-mentioned known cancer patient samples and known non-cancer patient samples as training samples, constructing a cancer risk prediction model;

[0130] N2. Data reception: receiving sample data, wherein the sample data is the methylation level data of the differentially methylated region group described in the first aspect above of the subject sample;

[0131] N3. Data processing and result output: Based on the cancer risk prediction model and the sample data, obtain and output the results; the results are the results of warning cancer before clinical symptoms occur.

[0132] The number of the known cancer patients and the known non-cancer patients may be more than 50.

[0133] In the above, a machine learning method may be used when constructing the cancer risk prediction model using the training samples; the machine learning method may be a random forest method.

[0134] In the above, the sample may be a sample from which DNA can be extracted.

[0135] Furthermore, the sample may specifically be plasma, tissue, saliva, urine and / or feces.

[0136] In the above, the methods for obtaining the methylation level data include but are not limited to bisulfite conversion, endonuclease digestion technology, methylation-specific PCR (MS-PCR), pyrosequencing, high-throughput sequencing, third-generation sequencing or single-molecule sequencing, etc.

[0137] In the above, the non-cancer patients may be healthy controls.

[0138] In the above, the cancer may be pancreatic cancer, liver cancer, colorectal cancer, gastric cancer, esophageal cancer or lung cancer.

[0139] The present invention also claims protection for a computer device.

[0140] The computer device claimed in the present invention includes a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the steps of the above method.

[0141] The present invention also claims protection for a computer program product.

[0142] The computer program product claimed for protection by the present invention comprises a computer program, characterized in that: when the computer program is executed by a processor, the steps of the method described above are implemented.

[0143] The computer program product may be a software product that primarily implements its solution through a computer program.

[0144] In the present invention, the analysis method of DMR methylation level is related to computer software and / or computer hardware, including but not limited to determining the methylation status of DMR, comparing the methylation status of DMR, generating a methylation standard curve, determining the Ct value, calculating the methylation rate of DMR, determining the specificity and / or sensitivity of the analysis or marker, calculating the ROC curve and related AUC, sequence analysis, etc.

[0145] In the present invention, the healthy control specifically refers to a physical examination sample without abnormality.

[0146] In a specific embodiment of the present invention, the cancer described above is pancreatic cancer. Furthermore, the pancreatic cancer may be pancreatic cancer of different clinical types and / or different stages, including pancreatic cancer samples of different stages including 0, I, II, III, and IV (TNM staging, AJCC, 8th edition), and pancreatic cancer of different pathological types including ductal adenocarcinoma, other epithelial pancreatic cancers, neuroendocrine tumors, undifferentiated carcinomas, and stromal cancers. The methylation markers of the present invention have a good predictive effect on early-stage (0, I, II) pancreatic cancer.

[0147] In the present invention, the cancer risk prediction model is a random forest model constructed based on selected features. The feature selection process uses a 10-fold cross-validation based on random forests, followed by the Boruta algorithm and recursive feature elimination method, and performs feature selection and validation based on importance, necessity, and robustness.

[0148] In a specific embodiment of the present invention, the aforementioned methylation level data is a methylation rate, that is, the methylation level of the aforementioned differentially methylated region group is the ratio of methylated cytosines in CpGs to all cytosines in the corresponding DMR region.

[0149] In the present invention, when the sample is plasma, the detection object is cfDNA. BRIEF DESCRIPTION OF THE DRAWINGS

[0150] Figure 1 shows a heatmap of the regional methylation rates of 13 DMRs in 113 pancreatic cancer patients and 84 healthy controls from Example 1. The horizontal axis represents the samples, with pancreatic cancer patients on the far left, followed by healthy controls. The vertical axis represents the DMRs, with hypermethylated DMRs at the top and hypomethylated DMRs at the bottom. Each cell represents the DMR methylation rate at that site in the corresponding sample, ranging from 0 to 1. The closer the methylation rate is to 0, the darker the color.

[0151] Figure 2 shows the performance of the pancreatic cancer methylation model based on 13 DMRs in the validation set of Example 1. The horizontal axis represents the false positive rate (1-specificity), and the vertical axis represents the sensitivity.

[0152] FIG3 is a heat map of the methylation rates of the methylation markers considered in the present invention (the 13 DMRs in Table 1) in an independent validation set (Set II).

[0153] FIG4 shows the AUC of the pancreatic cancer risk model of the present invention in the independent validation set (set II).

[0154] FIG5 is a comparison of the performance of the pancreatic cancer risk model of the present invention and the traditional CA19-9 test in diagnosing pancreatic cancer in Example 2.

[0155] FIG6 shows the performance of the pancreatic cancer risk model of the present invention and the traditional CA19-9 test in early detection of pancreatic cancer at different stages in Example 2.

[0156] Best Mode for Carrying Out the Invention

[0157] The following examples are provided to facilitate a better understanding of the present invention, but are not intended to limit the present invention. The experimental methods in the following examples, unless otherwise specified, are conventional methods. The test materials used in the following examples, unless otherwise specified, were purchased from conventional biochemical reagent stores. The quantitative tests in the following examples were all repeated three times, and the results were averaged.

[0158] Example 1: Discovery of 13 DMRs that can be used as pancreatic cancer screening markers

[0159] This example describes the discovery and validation of differentially methylated regions (DMRs) associated with the development and progression of pancreatic cancer and markers therein that can be used as markers for pancreatic cancer detection and screening.

[0160] First, targeted bisulfite sequencing (TBS) was performed on DNA extracted from 32 pairs of frozen pancreatic cancer tissues and corresponding adjacent cancer tissues. Through data analysis and calculation, DMRs related to the occurrence and development of pancreatic cancer were identified.

[0161] Second, targeted methylation high-throughput sequencing was performed on plasma cfDNA from independent cohorts of pancreatic cancer patients and healthy controls. A pancreatic cancer risk model was constructed using machine learning methods to identify DMRs that could serve as biomarkers for pancreatic cancer screening.

[0162] Study subjects and samples: Samples included frozen pancreatic cancer and adjacent tissues, plasma from pancreatic cancer patients, and plasma from healthy individuals.

[0163] The study subjects in this example, designated Set I, included 113 pancreatic cancer patients and 84 healthy controls. Inclusion criteria were age 18 years or older, gender, availability of peripheral blood, voluntary participation, and signed informed consent. Healthy controls were considered to be those undergoing physical examinations with no reported abnormalities. Pancreatic cancer patients were considered to have received a first diagnosis of pancreatic cancer, had primary lesions, had no prior history of malignancy, and had not received neoadjuvant therapy or chemotherapy. Sample characteristics of the pancreatic cancer patients in Set I are shown in Table 1.

[0164] Table 1. Sample characteristics of pancreatic cancer patients in Set I

[0165] 1. DNA Preparation

[0166] 1. Tissue sample extraction

[0167] DNA was extracted from 32 pairs of pancreatic cancer tissues and corresponding adjacent tissues using DNeasy Blood & Tissue Kit (Qiagen, #69506).

[0168] 2. DNA fragmentation

[0169] 200 ng of extracted DNA was taken and 1 ng of unmethylated lambda DNA (PROMEGA, #D1521) was added as a quality control for subsequent CU conversion. The DNA was sheared using an ultrasonic shearer and then fragmented using AMPure XP (AGENCOURT, #A63882) to select the DNA fragments so that the DNA fragment size was concentrated around 160 bp.

[0170] 3. Plasma sample extraction

[0171] cfDNA was extracted from pancreatic cancer and healthy human plasma samples using the MagPure Circulating DNA Maxi Kit (MAGEN, #12917PJ-100). 10 ng of cfDNA was added to 0.05 ng of unmethylated lambda DNA, sheared and filtered to approximately 160 bp.

[0172] 2. Library Construction

[0173] The DNA fragments were end-repaired and A-tailed at the 3' end using Klenow (3'-5'exo-) (ENZYMATICS, #P7010-LC-L). T4 DNA Ligase (Rapid) (ENZYMATICS, #L6030-600,000) was reacted at 16°C for one hour. MGISEQ adapters containing methyl modifications and barcode sequences were ligated to both ends of the DNA fragments, and the ligation products were purified using AMPure XP. Finally, the ligation products were amplified using KAPA HiFi HotStart ReadyMix (KAPA, #KK2602), and the PCR products were purified using AMPure XP to complete library construction.

[0174] 3. Bisulfite Conversion

[0175] The constructed DNA library was bisulfite converted using EZ DNA Metlylation-Gold Kit (Zymo Research, #D5006).

[0176] 4. Library hybridization capture

[0177] Hybridization, capture, and elution were performed using the Seq Cap EZ Hybridization and Wash Kit (ROCHE, 5634253001) and Seq Cap Epi CpGiant Enrichment Kit (ROCHE, 7138911001). Since an MGI platform sequencing instrument was used, the blocks used in the hybridization process must be those corresponding to the MGI platform.

[0178] 5. Sequencing

[0179] PE100 sequencing was performed using MGISEQ-2000 (MGI).

[0180] VI. Construction of a Pancreatic Cancer Risk Model

[0181] In a sample set I of 113 pancreatic cancer patients and 84 healthy controls, a 10-fold random forest-based cross-validation algorithm was used to model hypermethylated DMRs. Feature selection was performed based on importance, necessity, and robustness using the Boruta algorithm and recursive feature elimination, identifying biomarkers for pancreatic cancer detection and screening. The sample set was divided into 10 folds. In each fold of the 10-fold cross-validation, 90% of the samples were used as a training set (for model building) and 10% as a validation set (for model validation). The test samples in each fold were different. Modeling was performed using the DMR methylation rate (i.e., the ratio of methylated cytosines in CpGs to all cytosines within the DMR region) calculated from the depth of single CpG sites and the number of methylated cytosines obtained through targeted high-throughput sequencing of cfDNA. Model performance was evaluated using the validation set from each fold to confirm the robustness of the algorithm and the selected features. Finally, a predictive model was constructed based on the selected DMRs in sample set I. The threshold is 0.5. A value greater than 0.5 is considered a high risk of pancreatic cancer, otherwise it is considered a low risk.

[0182] 7. Results

[0183] Targeted methylome sequencing was performed on DNA extracted from 32 pairs of pancreatic cancer tissues and corresponding adjacent tissues. The sequencing data were analyzed and the methylation rate of each CpG site was calculated. By comparing the methylation rates of each CpG site in pancreatic cancer tissues and adjacent tissues, a hierarchical Bayesian method was used to identify 1,173 differentially methylated regions that may be associated with the development and progression of pancreatic cancer, of which 538 were hypermethylated DMRs.

[0184] In this example, targeted high-throughput sequencing was performed on cfDNA from 113 pancreatic cancer patients and 84 healthy controls from Sample Set 1. Using a 10-fold random forest-based cross-validation approach, followed by the Boruta algorithm and recursive feature elimination, DMR screening was performed, resulting in 13 DMRs that could serve as pancreatic cancer screening markers (Table 2). Model performance was evaluated using validation sets from each fold, resulting in an average AUC of 0.941 for the validation set.

[0185] Figure 1 shows a heat map of the regional methylation rates of 13 DMRs in 113 pancreatic cancer patients and 84 healthy controls. The horizontal axis represents the samples, with pancreatic cancer patients on the far left, followed by healthy controls. The vertical axis represents the DMRs, with hypermethylated DMRs at the top and hypomethylated DMRs at the bottom. Each cell represents the DMR methylation rate at that site in the corresponding sample, ranging from 0 to 1. The closer the methylation rate is to 0, the darker the color. As can be seen from the figure, the discovered DMRs clearly discriminate between pancreatic cancer patients and healthy controls.

[0186] Figure 2 shows the performance of the pancreatic cancer methylation model based on 13 DMRs in the validation set of this example. The horizontal axis represents the false positive rate (1-specificity), and the vertical axis represents the sensitivity. As can be seen from the figure, the methylation model has good diagnostic ability for the samples in this example.

[0187] Table 2. 13 DMRs that can be used as early detection markers for pancreatic cancer

[0188] Note: The physical positions (chromosomes, starting and ending positions) in the table are determined based on the alignment of the human whole genome sequence (version number is hg19).

[0189] Example 2: Validation of 13 DMRs in pancreatic cancer screening

[0190] The main purpose of this embodiment is to verify the performance of these 13 DMRs in screening pancreatic cancer in an independent validation set (denoted as Set II), which includes plasma free DNA from 142 pancreatic cancer patients and 125 healthy people. The inclusion criteria are over 18 years old, regardless of gender, with peripheral blood available, voluntary participation and signed informed consent. Among them, healthy people are those who have not reported any abnormalities in the physical examination population, and pancreatic cancer patients are those who have been diagnosed with pancreatic cancer for the first time, with pancreatic cancer as the primary lesion, no previous history of malignant tumors, and have not received neoadjuvant therapy or chemotherapy. The sample characteristic information of pancreatic cancer patients in Set II is shown in Table 3. Sample Set II does not contain any samples from Sample Set I in Example 1.

[0191] Table 3. Sample characteristics of pancreatic cancer patients in Set II

[0192] 1. Methods

[0193] The performance of these 13 DMRs (Table 1) in screening pancreatic cancer was verified in sample set II, directly validating the corresponding model constructed in Example 1. The specific operation was completed as described in Example 1.

[0194] 2. Results

[0195] Figure 3 shows a heat map of the methylation rates of the methylation markers considered in this analysis (13 DMRs in Table 1) in an independent validation set (set II). As can be seen from the figure, there is a clear distinction between pancreatic cancer and healthy subjects.

[0196] Figure 4 shows the AUC of the pancreatic cancer methylation rate model in the independent validation set (Set II), with an AUC of 0.96. As can be seen from the figure, the methylation model has good judgment ability in the independent validation set.

[0197] Figure 5 shows a comparison of the performance of this methylation model and the traditional CA19-9 test for pancreatic cancer diagnosis. As can be seen from the figure, the sensitivity of this model is significantly improved compared to the existing CA19-9 test (P<0.05).

[0198] Figure 6 shows the performance of this methylation model and traditional CA19-9 testing for early detection of pancreatic cancer at different stages. As can be seen from the figure, this model outperforms traditional CA19-9 testing for early detection of pancreatic cancer at all stages, especially for early-stage pancreatic cancer (stages I and II) (P<0.05).

[0199] Table 4 shows the performance of the pancreatic cancer risk model in an independent validation set (Set II). The model achieved a sensitivity of 88.7% in pancreatic cancer samples and a specificity of 95.2% in healthy controls. These results demonstrate that the model has excellent discriminatory power for pancreatic cancer patients and healthy controls.

[0200] Table 5 shows the performance of the pancreatic cancer risk model (methylation model) of the present invention and the traditional CA19-9 test (which is the clinical standard, with >37.0 U / ml clinically considered high, i.e., positive in this article; otherwise, negative) for early detection of pancreatic cancer at different stages. For stage I pancreatic cancer, the sensitivity of the methylation model was 82.3%, while the sensitivity of the traditional CA19-9 test was 54.9%. The results showed that the performance of the methylation model was significantly improved compared to the traditional CA19-9 test at each stage (P < 0.05).

[0201] Table 4. Performance of the pancreatic cancer risk model of the present invention in the independent validation set (set II)

[0202] Note: The blank space in the table indicates that this box is not applicable, that is, only pancreatic cancer has sensitivity, and healthy people have specificity.

[0203] Table 5. Performance of the pancreatic cancer risk model of the present invention and the traditional CA19-9 test in early detection of pancreatic cancer at different stages

[0204] Industrial Applications

[0205] The present invention provides 13 DMRs that can be used for pancreatic cancer screening (Table 1). By detecting the methylation levels of these DMRs and analyzing the resulting data, the likelihood of pancreatic cancer in the subject can be predicted, thereby achieving the goal of screening for pancreatic cancer in the general population or in high-risk populations. This invention provides an accurate, simple, and economical method for early screening of pancreatic cancer, which can improve the detection rate of pancreatic cancer, especially early-stage pancreatic cancer, in high-risk populations and the general population undergoing physical examinations, thereby improving the survival rate of pancreatic cancer patients and significantly saving medical expenses and reducing the burden of medical care.

Claims

1. A differentially methylated region group containing all or part of the following 13 differentially methylated regions: (A1) located at positions 111217374-111217715 on chromosome 1; (A2) located at positions 248020350-248020956 on chromosome 1; (A3) located at positions 12490772-12490958 on chromosome 10; (A4) located at positions 37005986-37006281 on chromosome 13; (A5) located at positions 65701719-65701941 on chromosome 15; (A6) located at positions 124510649-124510809 on chromosome 3; (A7) located at positions 15672531-15672777 on chromosome 1; (A8) located at positions 153518054-153518476 on chromosome 1; (A9) is located at positions 92907747-92908353 on chromosome 5; (A10) is located at positions 49813377-49813880 on chromosome 7; (A11) is located at positions 170629389-170630397 on chromosome 1; (A12) is located at positions 15568312-15568674 on chromosome 19; (A13) is located at positions 16318860-16323551 on chromosome 17; The physical locations of the 13 differentially methylated regions were determined based on the human whole genome sequence hg19 alignment.

2. The differentially methylated region panel according to claim 1, wherein: The differentially methylated region group consists of all or part of the 13 differentially methylated regions shown in (A1) to (A13) of claim 1.

3. Use of the differentially methylated region group according to claim 1 or 2 as a methylation marker in any of the following: (B1) preparing products for diagnosing or assisting in the diagnosis of cancer; (B2) diagnosis or assistance in diagnosis of cancer; (B3) preparing products for early warning of cancer before clinical symptoms; (B4) Warning of cancer before clinical symptoms.

4. Use of a substance for detecting the methylation level of the differentially methylated region group according to claim 1 or 2 in any of the following: (B1) preparing products for diagnosing or assisting in the diagnosis of cancer; (B2) diagnosis or assistance in diagnosis of cancer; (B3) preparing products for early warning of cancer before clinical symptoms; (B4) Warning of cancer before clinical symptoms.

5. The use of a combination of substances and media in any of the following: (B1) preparing products for diagnosing or assisting in the diagnosis of cancer; (B2) diagnosis or assistance in diagnosis of cancer; (B3) preparing products for early warning of cancer before clinical symptoms; (B4) Early warning of cancer before clinical symptoms; The substance is a substance for detecting the differentially methylated region group according to claim 1 or 2; The medium stores a method for constructing and using a cancer risk prediction model; The method for constructing and using the cancer risk prediction model comprises the following steps: (C1) obtaining methylation level data of the differentially methylated region group according to claim 1 for known cancer patient samples and known non-cancer patient samples as training samples; (C2) constructing a cancer risk prediction model using the training samples, and then utilizing the cancer risk prediction model to diagnose or assist in the diagnosis of cancer and / or to provide early warning of cancer before clinical symptoms occur.

6. Use of a medium storing methods for constructing and using a cancer risk prediction model in any of the following: (B1) preparing products for diagnosing or assisting in the diagnosis of cancer; (B2) diagnosis or assistance in diagnosis of cancer; (B3) preparing products for early warning of cancer before clinical symptoms; (B4) Early warning of cancer before clinical symptoms; The method for constructing and using the cancer risk prediction model comprises the following steps: (C1) obtaining methylation level data of the differentially methylated region group according to claim 1 for known cancer patient samples and known non-cancer patient samples as training samples; (C2) constructing a cancer risk prediction model using the training samples, and then utilizing the cancer risk prediction model to diagnose or assist in the diagnosis of cancer and / or to provide early warning of cancer before clinical symptoms occur.

7. The use according to claim 5 or 6, characterized in that: The machine learning method is the random forest method.

8. The use according to any one of claims 5 to 7, characterized in that: The sample is a sample from which DNA can be extracted.

9. The use according to claim 8, characterized in that: The sample is plasma, tissue, saliva, urine or / and feces.

10. The use according to any one of claims 5 to 9, characterized in that: The method for obtaining the methylation level data is bisulfate conversion, endonuclease digestion technology, methylation-specific PCR, pyrophosphate sequencing, high-throughput sequencing, third-generation sequencing or single-molecule sequencing.

11. The use according to any one of claims 5 to 10, characterized in that: The non-cancer patients were healthy controls.

12. The use according to any one of claims 3 to 11, characterized in that: The cancer is pancreatic cancer, liver cancer, colorectal cancer, gastric cancer, esophageal cancer or lung cancer.

13. A kit comprising: (a1) a bisulfite reagent; and (a2) A control nucleic acid comprising a sequence from the differentially methylated region group according to claim 1 or 2, and having a methylation status associated with a non-cancer patient.

14. A kit comprising: (b1) a bisulfite reagent; and (b2) a control nucleic acid comprising a sequence from the differentially methylated region group according to claim 1 or 2 and having a methylation status associated with a cancer patient.

15. A kit comprising a substance for detecting the differentially methylated region group according to claim 1 or 2 and a medium storing a method for constructing and using a cancer risk prediction model; The method for constructing and using the cancer risk prediction model comprises the following steps: (C1) obtaining methylation level data of the differentially methylated region group according to claim 1 for known cancer patient samples and known non-cancer patient samples as training samples; (C2) constructing a cancer risk prediction model using the training samples, and then utilizing the cancer risk prediction model to diagnose or assist in the diagnosis of cancer and / or to provide early warning of cancer before clinical symptoms occur.

16. A device having any of the following functions: (D1) Diagnosis or assistance in diagnosis of cancer; (D2) Early warning of cancer before clinical symptoms; The device comprises: M1, a cancer risk prediction model construction module: configured to construct a cancer risk prediction model using the methylation level data of the differentially methylated region group according to claim 1 of known cancer patient samples and known non-cancer patient samples as training samples; M2, a data receiving module: configured to receive sample data, wherein the sample data is the methylation level data of the differentially methylated region group according to claim 1 of the subject sample; M3, data analysis and processing module: configured to obtain results based on the cancer risk prediction model and the sample data; the results are the results of diagnosing or assisting in the diagnosis of cancer and / or the results of warning of cancer before clinical symptoms.

17. System, including: (E1) a reagent and / or an apparatus for detecting the methylation level of the differentially methylated region group according to claim 1 or 2; (E2) The device of claim 16.

18. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the following steps are implemented: constructing a cancer risk prediction model using the methylation level data of the differentially methylated region group according to claim 1 of known cancer patient samples and known non-cancer patient samples as training samples; receiving sample data, wherein the sample data is methylation level data of the differentially methylated region group according to claim 1 of a subject sample; Based on the cancer risk prediction model and the sample data, a result is obtained; the result is a result of diagnosing or assisting in the diagnosis of cancer and / or a result of warning of cancer before clinical symptoms occur.

19. The kit according to claim 15, the data processing device according to claim 16, the system according to claim 17, or the computer-readable storage medium according to claim 18, characterized in that: A machine learning method is used when constructing the cancer risk prediction model using the training samples; Furthermore, the machine learning method is a random forest method.

20. The kit, data processing device, system, or computer-readable storage medium according to any one of claims 15 to 19, wherein: The sample is a sample from which DNA can be extracted.

21. The kit, data processing device, system, or computer-readable storage medium according to claim 20, wherein: The sample is plasma, tissue, saliva, urine or / and feces.

22. The kit, data processing device, system, or computer-readable storage medium according to any one of claims 15 to 21, wherein: The method for obtaining the methylation level data is bisulfate conversion, endonuclease digestion technology, methylation-specific PCR, pyrophosphate sequencing, high-throughput sequencing, third-generation sequencing or single-molecule sequencing.

23. The kit, data processing device, system, or computer-readable storage medium according to any one of claims 13 to 22, wherein: The non-cancer patients were healthy controls.

24. The kit, data processing device, system, or computer-readable storage medium according to any one of claims 13 to 23, wherein: The cancer is pancreatic cancer, liver cancer, colorectal cancer, gastric cancer, esophageal cancer or lung cancer.

25. Use of the kit, data processing device, system, or computer-readable storage medium according to any one of claims 13 to 24 in any of the following: (B1) preparing products for diagnosing or assisting in the diagnosis of cancer; (B2) diagnosis or assistance in diagnosis of cancer; (B3) preparing products for early warning of cancer before clinical symptoms; (B4) Warning of cancer before clinical symptoms.

26. The use according to claim 25, characterized in that: The cancer is pancreatic cancer, liver cancer, colorectal cancer, gastric cancer, esophageal cancer or lung cancer.

27. A method for diagnosing or assisting in the diagnosis of cancer, comprising the following steps: N1. Constructing a cancer risk prediction model: using the methylation level data of the differentially methylated region group described in claim 1 of known cancer patient samples and known non-cancer patient samples as training samples to construct a cancer risk prediction model; N2. Data reception: receiving sample data, wherein the sample data is the methylation level data of the differentially methylated region group according to claim 1 of the subject sample; N3. Data processing and result output: Based on the cancer risk prediction model and the sample data, obtain and output the result; the result is the result of diagnosing or assisting in the diagnosis of cancer.

28. A method for early warning of cancer before clinical symptoms occur, comprising the following steps: N1. Constructing a cancer risk prediction model: using the methylation level data of the differentially methylated region group described in claim 1 of known cancer patient samples and known non-cancer patient samples as training samples to construct a cancer risk prediction model; N2. Data reception: receiving sample data, wherein the sample data is the methylation level data of the differentially methylated region group according to claim 1 of the subject sample; N3. Data processing and result output: Based on the cancer risk prediction model and the sample data, Obtain and output a result; the result is a result of early warning of cancer before clinical symptoms.

29. The method according to claim 27 or 28, characterized in that: A machine learning method is used when constructing the cancer risk prediction model using the training samples; Furthermore, the machine learning method is a random forest method.

30. The method according to any one of claims 27 to 29, characterized in that: The sample is a sample from which DNA can be extracted.

31. The method according to claim 30, wherein: The samples are plasma, tissue, saliva, urine, and feces.

32. The method according to any one of claims 27 to 31, characterized in that: The method for obtaining the methylation level data is bisulfate conversion, endonuclease digestion technology, methylation-specific PCR, pyrophosphate sequencing, high-throughput sequencing, third-generation sequencing or single-molecule sequencing.

33. The method according to claim 27 or 28, characterized in that: The non-cancer patients were healthy controls.

34. The method according to any one of claims 27 to 33, wherein: The cancer is pancreatic cancer, liver cancer, colorectal cancer, gastric cancer, esophageal cancer or lung cancer.

35. A computer device comprising a memory, a processor, and a computer program stored in the memory, wherein: The processor executes the computer program to implement the steps of the method according to any one of claims 27 to 34.

36. A computer program product comprising a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 27 to 34 are implemented.

Citation Information

Patent Citations

  • Detecting neoplasm

    CN105143465A

  • Pancreatic cancer diagnosis related DNA methylation marker and application thereof

    CN115491421A

  • Detecting pancreatic neuroendocrine tumors

    CN115551880A

  • Pancreatic cancer biomarker, nucleic acid product and kit

    CN116656810A

  • Method for the identification of the origin of a cancer of unknown primary origin by methylation analysis

    US20160017430A1