Methylation markers for differentiating pancreatitis and pancreatic cancer and use thereof

CN115491411BActive Publication Date: 2026-10-09SINGLERA GENOMICS (SHANGHAI) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202110680924.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-18
Publication Date
2026-10-09
Estimated Expiration
2041-06-18

AI Technical Summary

Technical Problem

然而,大多数这些研究未取得有效的结果

Benefits of technology

[0150] Based on the DNA methylation biomarkers of this invention, patients with pancreatic cancer and chronic pancreatitis can be effectively distinguished. This invention provides a diagnostic model of the relationship between cfDNA methylation biomarker methylation levels and pancreatic cancer based on high-throughput plasma cfDNA methylation sequencing. This model has the advantages of non-invasive detection, safe and convenient detection, high throughput, and high detection specificity. Based on the optimal sequencing volume obtained by this invention, detection costs can be effectively controlled while achieving good detection performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0003122793530000191
    Figure BDA0003122793530000191
  • Figure BDA0003122793530000201
    Figure BDA0003122793530000201
  • Figure BDA0003122793530000211
    Figure BDA0003122793530000211
Patent Text Reader

Abstract

The present application relates to methylation markers for identifying pancreatitis and pancreatic cancer and application thereof. Specifically, the present application provides use of methylation markers and detection reagents thereof for preparing a kit for identifying pancreatic cancer and chronic pancreatitis. Based on sequencing data analysis of DNA methylation of plasma samples of patients with malignant pancreatic cancer and patients with chronic pancreatitis, the inventors identify markers for identifying samples of pancreatic cancer and chronic pancreatitis. Based on the markers or marker groups of the present application, pancreatic cancer and chronic pancreatitis patients can be effectively distinguished with high sensitivity and specificity, providing a new method for early identification of pancreatic cancer. The detection process of the present application is non-invasive and safe.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a methylation marker for differentiating pancreatitis and pancreatic cancer and its application, used to differentiate between patients with pancreatic cancer and pancreatitis, and belongs to the field of molecular biomedical technology. Background Technology

[0002] Pancreatic cancer (e.g., pancreatic ductal adenocarcinoma) is one of the deadliest diseases in the world. The 5-year relative survival rate is 9%, which drops to only 3% for patients with distant metastases. A major reason for the high mortality rate is the still limited availability of methods for early detection of pancreatic cancer, which is crucial for patients undergoing surgical resection. Currently, carbohydrate antigen 19-9 (CA19-9) is the most commonly used clinical serum biomarker for the auxiliary detection of pancreatic cancer, achieving a sensitivity of 79-90% and a specificity of 75-90% in symptomatic patients before resection. However, several large population studies have demonstrated that CA19-9 is ineffective in detecting pancreatic cancer in asymptomatic individuals due to its low positive predictive value, essentially ruling it out for early screening of pancreatic cancer (Kim et al., 2004) (Chang et al., 2006; Homma & Tsuchiya, 1991; Kim et al., 2004; Satake, Takeuchi, Homma, & Ozaki, 1994). Endoscopic ultrasound-guided fine-needle aspiration (EUS-FNA) is another commonly used method to obtain pathological diagnosis without open surgery, but it is invasive and requires clear imaging evidence, which usually indicates that pancreatic cancer has progressed. During tumorigenesis and development, the DNA methylation patterns and levels of malignant cell genomic DNA undergo profound changes. Some tumor-specific DNA methylations have been shown to occur early in tumorigenesis and may become "driving factors" in tumorigenesis.

[0003] Typical early symptoms of pancreatic cancer, including abdominal and back pain, diarrhea, weight loss, and jaundice, are not specific and may be associated with other gastrointestinal disorders. These complications are particularly common in the diagnosis of chronic pancreatitis, especially since patients with chronic pancreatitis have a significantly higher long-term risk of developing pancreatic cancer. Therefore, accurate differential diagnosis between pancreatic cancer and chronic pancreatitis is crucial for screening patients with chronic pancreatitis for pancreatic cancer. However, the current accuracy rate for differential diagnosis between chronic pancreatitis and pancreatic cancer is 65% or lower, indicating significant room for improvement.

[0004] Circulating tumor DNA (ctDNA) molecules originate from apoptotic or necrotic tumor cells and carry tumor-specific DNA methylation markers from early-stage malignant tumors. In recent years, they have been investigated as promising new targets for developing non-invasive early screening tools for various cancers. However, most of these studies have not yielded effective results. Existing research indicates that the proportion of ctDNA in the plasma DNA of patients with early-stage tumors is very small, and chronic inflammation can affect the accuracy of DNA methylation markers (Abbosh et al., 2017). Therefore, identifying stable and consistent specific markers from plasma DNA to differentiate between chronic pancreatitis and pancreatic cancer presents a significant challenge. Summary of the Invention

[0005] This invention provides a method for detecting DNA methylation in patient plasma samples, and using the differential methylation level data analysis of the detection results to distinguish between pancreatic cancer patients and pancreatitis patients, thereby achieving the goal of non-invasive and precise diagnosis of pancreatic cancer with higher accuracy and lower cost.

[0006] Specifically, the first aspect of the present invention provides an isolated nucleic acid molecule from a mammal, said nucleic acid molecule being a methylation marker associated with the identification of pancreatic cancer and pancreatitis, said nucleic acid molecule having a sequence comprising (1) a sequence selected from one or more of the following sequences or variants having at least 70% identity with them: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, wherein the methylation sites in said variants are not mutated, (2) a complementary sequence to (1), and (3) a treated sequence of (1) or (2), said treatment converting unmethylated cytosine into bases with a lower binding affinity to guanine than cytosine.

[0007] In one or more embodiments, the methylation sites are continuous CpG.

[0008] In one or more embodiments, the methylation marker may be any one or more CpG sites in the sequence region.

[0009] In one or more embodiments, the nucleic acid molecule is used as an internal standard or control for detecting the DNA methylation level of a corresponding sequence in a sample.

[0010] In one or more embodiments, the pancreatic cancer is pancreatic ductal adenocarcinoma.

[0011] In one or more embodiments, the pancreatitis is chronic pancreatitis.

[0012] A second aspect of the present invention provides a reagent for detecting DNA methylation, the reagent comprising a reagent for detecting the methylation level of a DNA sequence or fragment thereof in a sample of a test subject, or the methylation state or level of one or more CpG dinucleotides in the DNA sequence or fragment thereof, the DNA sequence being selected from one or more (e.g., at least two) or all of the following gene sequences, or sequences within 20 kb upstream or downstream thereof: SIX3, TLX2, CILP2.

[0013] In one or more embodiments, the DNA sequence includes a sense strand or an antisense strand.

[0014] In one or more embodiments, the fragment length is 1-1000bp, preferably 1-700bp.

[0015] In one or more embodiments, the fragment contains at least one CpG dinucleotide.

[0016] In one or more embodiments, the DNA sequence is selected from one or more (e.g., at least two) or all of the following sequences or their complementary sequences: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, or variants thereof having at least 70% identity, wherein the methylation sites in the variants are not mutated.

[0017] In one or more embodiments, the reagent is a primer molecule that hybridizes with the DNA sequence or a fragment thereof. The primer molecule amplifies the DNA sequence or a fragment thereof. In one or more embodiments, the primer sequence is methylation-specific or non-specific. The primer molecule is at least 9 bp.

[0018] In one or more embodiments, the reagent is a probe molecule that hybridizes with the DNA sequence or a fragment thereof. In one or more embodiments, the probe further contains a detectable. In one or more embodiments, the detectable is a 5' fluorescent reporter group and a 3' labeled quencher group. In one or more embodiments, the fluorescent reporter gene is selected from Cy5, FAM, and VIC. Preferably, the probe sequence contains MGB (Minor Groove Binder) or LNA (Locked Nucleic Acid). The probe molecule is at least 12 bp.

[0019] In one or more embodiments, the reagent comprises the nucleic acid molecules described in the first aspect of this document.

[0020] In one or more embodiments, the sample is derived from a mammal, preferably a human.

[0021] A third aspect of the present invention provides a medium containing a DNA sequence or a fragment thereof and / or its methylation information, said DNA sequence being (i) selected from one or more (e.g., at least two) or all of the following gene sequences, or sequences within 20 kb upstream or downstream of them: SIX3, TLX2, CILP2, or (ii) a treated sequence of (i) wherein the treatment converts unmethylated cytosine into bases with a lower binding affinity to guanine than cytosine.

[0022] In one or more embodiments, the DNA sequence is (i) a gene sequence selected from any of the following groups, or a sequence within 20 kb upstream or downstream of it: (1) SIX3, TLX2; (2) SIX3, CILP2; (3) TLX2, CILP2; (4) SIX3, TLX2, CILP2, or (ii) a treated sequence of (i) wherein the treatment converts unmethylated cytosine into bases with a lower binding affinity to guanine than cytosine.

[0023] In one or more embodiments, the medium is used to compare with gene methylation sequencing data to determine the presence, content, and / or methylation level of nucleic acid molecules containing the sequence or fragment.

[0024] In one or more embodiments, the DNA sequence includes a sense strand or an antisense strand.

[0025] In one or more embodiments, the fragment length is 1-1000bp, preferably 1-700bp.

[0026] In one or more embodiments, the fragment contains at least one CpG dinucleotide.

[0027] In one or more embodiments, the DNA sequence is selected from one or more (e.g., at least two) or all of the following sequences or their complementary sequences: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, or variants thereof having at least 70% identity, wherein the methylation sites in the variants are not mutated.

[0028] In one or more embodiments, the methylation information includes information related to cytosine that may be methylated in the sequence of the nucleic acid molecule. Preferably, the cytosine that may be methylated is the C in CpG. In one or more embodiments, the methylation information is the location of a methylation site (such as a CpG dinucleotide) in the nucleic acid molecule.

[0029] In one or more embodiments, the medium is a carrier printed with the DNA sequence or fragments thereof and / or its methylation information, including cards, such as paper, plastic, metal, or glass cards.

[0030] In one or more embodiments, the medium is a computer-readable medium storing the sequence and / or its methylation information and a computer program that, when executed by a processor, performs the following steps: comparing the methylation sequencing data of a sample with the sequence to obtain the presence, abundance, and / or methylation level of nucleic acid molecules containing the sequence in the sample. The presence, abundance, and / or methylation level of nucleic acid molecules containing the sequence are used to differentiate between pancreatic cancer and pancreatitis.

[0031] In another aspect, the present invention also provides (a) and / or (b) the use in the preparation of a kit for differentiating pancreatic cancer and pancreatitis.

[0032] (a) A reagent or apparatus for determining the methylation level of a DNA sequence or fragment thereof in a sample of an object, or the methylation state or level of one or more CpG dinucleotides in the DNA sequence or fragment thereof.

[0033] (b) A treated nucleic acid molecule of the DNA sequence or a fragment thereof, wherein the treatment converts unmethylated cytosine into bases that have a lower binding affinity to guanine than cytosine.

[0034] The DNA sequence is selected from one or more (e.g., at least two) or all of the following gene sequences, or sequences within 20kb upstream or downstream of them: SIX3, TLX2, CILP2.

[0035] In one or more embodiments, the DNA sequence comprises a gene sequence selected from any of the following groups: (1) SIX3, TLX2; (2) SIX3, CILP2; (3) TLX2, CILP2; (4) SIX3, TLX2, CILP2.

[0036] In one or more embodiments, the DNA sequence includes a sense strand or an antisense strand.

[0037] In one or more embodiments, the fragment length is 1-1000bp, preferably 1-700bp.

[0038] In one or more embodiments, the fragment contains at least one CpG dinucleotide.

[0039] In one or more embodiments, the DNA sequence is selected from one or more (e.g., at least two) or all of the following sequences or their complementary sequences: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, or variants thereof having at least 70% identity, wherein the methylation sites in the variants are not mutated.

[0040] In one or more embodiments, the DNA sequence comprises a sequence selected from any of the following groups or a complementary sequence thereof: (1) SEQ ID NO:1, SEQ ID NO:2, (2) SEQ ID NO:1, SEQ ID NO:3, (3) SEQ ID NO:2, SEQ ID NO:3, (4) SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3.

[0041] In one or more embodiments, the nucleic acid molecule is the nucleic acid molecule described in the first aspect of this document.

[0042] In one or more embodiments, the reagent comprises primer molecules and / or probe molecules.

[0043] In one or more embodiments, the reagent comprises primer molecules that hybridize with the DNA sequence or a fragment thereof. The primer molecules amplify the DNA sequence or a fragment thereof. In one or more embodiments, the primer sequence is methylation-specific or non-specific. The primer molecule is at least 9 bp.

[0044] In one or more embodiments, the reagent is a probe molecule that hybridizes with the DNA sequence or a fragment thereof. In one or more embodiments, the probe further contains a detectable. In one or more embodiments, the detectable is a 5' fluorescent reporter group and a 3' labeled quencher group. In one or more embodiments, the fluorescent reporter gene is selected from Cy5, FAM, and VIC. Preferably, the probe sequence contains MGB (Minor Groove Binder) or LNA (Locked Nucleic Acid). The probe molecule is at least 12 bp.

[0045] In one or more embodiments, the reagent comprises the medium described in any of the embodiments herein.

[0046] In one or more embodiments, the kit is a non-invasive diagnostic kit.

[0047] In one or more embodiments, the kit is an auxiliary diagnostic kit.

[0048] In one or more embodiments, the object is a mammal, preferably a human.

[0049] In one or more embodiments, the object is an object diagnosed with pancreatitis (e.g., chronic pancreatitis).

[0050] In one or more embodiments, the sample is derived from mammalian tissue, cells, or body fluids, such as pancreatic tissue or blood, preferably a fine-needle aspiration biopsy or plasma.

[0051] In one or more embodiments, the sample comprises genomic DNA or cfDNA.

[0052] In one or more embodiments, the DNA sequence is transformed, wherein unmethylated cytosine is converted into bases with a lower binding affinity to guanine than cytosine. The transformation is performed using an enzymatic method, preferably deaminase treatment, or the transformation is performed using a non-enzymatic method, preferably treatment with bisulfite, acid sulfite, or metabisulfite, or a combination thereof.

[0053] In one or more embodiments, the DNA sequence is treated with a methylation-sensitive restriction endonuclease.

[0054] In one or more embodiments, the kit further includes PCR reaction reagents. Preferably, the PCR reaction reagents include DNA polymerase, PCR buffer, dNTPs, and Mg2+.

[0055] In one or more embodiments, the kit further includes other reagents for detecting DNA methylation, said other reagents being selected from one or more of the following methods: bisulfite-based PCR (e.g., methylation-specific PCR), DNA sequencing (e.g., bisulfite sequencing, whole-genome methylation sequencing, simplified methylation sequencing), methylation-sensitive restriction endonuclease assays, quantitative fluorescence assays, methylation-sensitive high-resolution melting curve assays, chip-based methylation mapping, and mass spectrometry (e.g., mass spectrometry of flight). Preferably, said other reagents are selected from one or more of the following: bisulfite, bisulfite, acid sulfite, or metabisulfite or derivatives thereof, methylation-sensitive or insensitive restriction endonucleases, enzyme digestion buffers, fluorescent dyes, fluorescence quenchers, fluorescent reporters, exonucleases, alkaline phosphatases, internal standards, and controls.

[0056] In one or more embodiments, the PCR reaction solution comprises Taq DNA polymerase, PCR buffer, dNTPs, KCl, MgCl2, and (NH4)2SO4. Preferably, the Taq DNA polymerase is a hot-start Taq DNA polymerase. Preferably, the final concentration of Mg2+ is 1.0-10.0 mM.

[0057] In one or more embodiments, the diagnosis includes: comparing with a control sample or calculating a score, and differentiating between pancreatic cancer and pancreatitis based on the score. In one or more embodiments, the calculation is performed by constructing a support vector machine model.

[0058] Another aspect of the present invention provides a method for differentiating pancreatic cancer from pancreatitis, comprising:

[0059] (1) The methylation level of the DNA sequence or fragment thereof in the sample of the test subject, or the methylation status or level of one or more CpG dinucleotides in the DNA sequence or fragment thereof, wherein the DNA sequence is selected from one or more or all of the following gene sequences: SIX3, TLX2, CILP2,

[0060] (2) Compare with the control sample, or calculate the score.

[0061] (3) Differentiate between pancreatic cancer and pancreatitis based on the scoring.

[0062] In one or more embodiments, the DNA sequence includes a sense strand or an antisense strand.

[0063] In one or more embodiments, the fragment length is 1-1000bp, preferably 1-700bp.

[0064] In one or more embodiments, the fragment contains at least one CpG dinucleotide.

[0065] In one or more embodiments, the method further includes DNA extraction and / or quality control prior to step (1).

[0066] In one or more embodiments, the DNA sequence is selected from one or more of the following sequences or their complementary sequences: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, or variants thereof having at least 70% identity, wherein the methylation sites in the variants are not mutated.

[0067] In one or more embodiments, step (1) includes performing the detection using the nucleic acid molecules, primer molecules, probe molecules and / or media described herein.

[0068] In one or more embodiments, the detection includes, but is not limited to: bisulfite-based PCR (e.g., methylation-specific PCR), DNA sequencing (e.g., bisulfite sequencing, whole-genome methylation sequencing, simplified methylation sequencing), methylation-sensitive restriction endonuclease assays, quantitative fluorescence assays, methylation-sensitive high-resolution melting curve assays, chip-based methylation mapping analysis, and mass spectrometry (e.g., mass spectrometry of flight).

[0069] In one or more embodiments, the detection is DNA sequencing. In one or more embodiments, the sequencing depth of the DNA sequencing is greater than or equal to 5M, preferably 7M, 11M, 13M, or 15M.

[0070] In one or more embodiments, the sample is derived from mammalian tissue, cells, or body fluids, such as pancreatic tissue or blood. The mammal is preferably human. In one or more embodiments, the sample is a fine-needle aspiration biopsy. In one or more embodiments, the sample is plasma.

[0071] In one or more embodiments, the sample comprises genomic DNA or cfDNA.

[0072] In one or more embodiments, the DNA sequence is transformed, wherein unmethylated cytosine is converted into bases that do not bind to guanine. The transformation is performed using an enzymatic method, preferably deaminase treatment, or the transformation is performed using a non-enzymatic method, preferably treatment with bisulfite, acid sulfite, or metabisulfite, or a combination thereof.

[0073] In one or more embodiments, the DNA sequence is treated with a methylation-sensitive restriction endonuclease.

[0074] In one or more implementations, the score in step (2) is calculated by constructing a support vector machine model.

[0075] In one or more embodiments, step (3) includes: comparing the change in methylation level of the subject sample with a control sample, and determining whether the subject has pancreatic cancer or pancreatitis based on whether the methylation level reaches a threshold.

[0076] In one or more implementations, step (3) includes determining whether the subject has pancreatic cancer or pancreatitis based on whether the score reaches a threshold.

[0077] Another aspect of the present invention provides a kit for differentiating pancreatic cancer and pancreatitis, comprising:

[0078] (a) A reagent or apparatus for determining the methylation level of a DNA sequence or fragment thereof in a sample of an object, or the methylation state or level of one or more CpG dinucleotides in said DNA sequence or fragment thereof, and

[0079] Option (b) is a treated nucleic acid molecule of the DNA sequence or a fragment thereof, wherein the treatment converts unmethylated cytosine into bases that have a lower binding affinity to guanine than cytosine.

[0080] The DNA sequence is selected from one or more (e.g., at least two) or all of the following gene sequences, or sequences within 20kb upstream or downstream of them: SIX3, TLX2, CILP2.

[0081] In one or more embodiments, the DNA sequence is selected from one or more (e.g., at least two) or all of the following sequences or their complementary sequences: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, or variants thereof having at least 70% identity, wherein the methylation sites in the variants are not mutated.

[0082] In one or more embodiments, the kit is suitable for the use described in any of the embodiments herein.

[0083] In one or more embodiments, the nucleic acid molecule is the nucleic acid molecule described in the first aspect of this document.

[0084] In one or more embodiments, the reagent comprises primer molecules and / or probe molecules.

[0085] In one or more embodiments, the reagent comprises primer molecules that hybridize with the DNA sequence or a fragment thereof. The primer molecules amplify the DNA sequence or a fragment thereof. In one or more embodiments, the primer sequence is methylation-specific or non-specific. The primer molecule is at least 9 bp.

[0086] In one or more embodiments, the reagent is a probe molecule that hybridizes with the DNA sequence or a fragment thereof. In one or more embodiments, the probe further contains a detectable. In one or more embodiments, the detectable is a 5' fluorescent reporter group and a 3' labeled quencher group. In one or more embodiments, the fluorescent reporter gene is selected from Cy5, FAM, and VIC. Preferably, the probe sequence contains MGB (Minor Groove Binder) or LNA (Locked Nucleic Acid). The probe molecule is at least 12 bp.

[0087] In one or more embodiments, the reagent comprises the medium described in any of the embodiments herein.

[0088] In one or more embodiments, the kit is a non-invasive diagnostic kit.

[0089] In one or more embodiments, the object is a mammal, preferably a human.

[0090] In one or more embodiments, the sample is derived from mammalian tissue, cells, or body fluids, such as pancreatic tissue or blood. In one or more embodiments, the sample is a fine-needle aspiration biopsy. In one or more embodiments, the sample is plasma.

[0091] In one or more embodiments, the sample comprises genomic DNA or cfDNA.

[0092] In one or more embodiments, the DNA sequence is transformed, wherein unmethylated cytosine is converted into bases with a lower binding affinity to guanine than cytosine. The transformation is performed using an enzymatic method, preferably deaminase treatment, or the transformation is performed using a non-enzymatic method, preferably treatment with bisulfite, acid sulfite, or metabisulfite, or a combination thereof.

[0093] In one or more embodiments, the DNA sequence is treated with a methylation-sensitive restriction endonuclease.

[0094] In one or more embodiments, the kit further includes PCR reaction reagents. Preferably, the PCR reaction reagents include DNA polymerase, PCR buffer, dNTPs, and Mg2+.

[0095] In one or more embodiments, the kit further includes reagents for detecting DNA methylation, said reagents being selected from one or more of the following methods: bisulfite-based PCR (e.g., methylation-specific PCR), DNA sequencing (e.g., bisulfite sequencing, whole-genome methylation sequencing, simplified methylation sequencing), methylation-sensitive restriction endonuclease assays, quantitative fluorescence assays, methylation-sensitive high-resolution melting curve assays, chip-based methylation mapping, and mass spectrometry (e.g., mass spectrometry of flight). Preferably, the reagents are selected from one or more of the following: bisulfite and its derivatives, methylation-sensitive or insensitive restriction endonucleases, enzyme digestion buffers, fluorescent dyes, fluorescence quenchers, fluorescent reporter agents, exonucleases, alkaline phosphatases, internal standards, and controls.

[0096] In another aspect, the present invention provides an apparatus for differentiating pancreatic cancer from pancreatitis, the apparatus comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor, when executing the program, performs the following steps:

[0097] (1) Obtain the methylation level of a DNA sequence or fragment thereof in a sample of the subject, or the methylation status or level of one or more CpGs in the DNA sequence or fragment thereof, wherein the DNA sequence is selected from one or more or all of the following gene sequences: SIX3, TLX2, CILP2,

[0098] (2) Compare with the control sample, or calculate the score, and

[0099] (3) Differentiate between pancreatic cancer and pancreatitis based on the scoring.

[0100] In one or more embodiments, step (1) further includes a step of obtaining DNA, such as DNA extraction and / or quality control.

[0101] In one or more embodiments, the DNA sequence is selected from one or more of the following sequences or their complementary sequences: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, or variants thereof having at least 70% identity, wherein the methylation sites in the variants are not mutated.

[0102] In one or more embodiments, step (1) includes detecting the methylation level of the sequence in the sample using the nucleic acid molecules, primer molecules, probe molecules, and / or media described herein. In one or more embodiments, the detection includes, but is not limited to: bisulfite-based PCR (e.g., methylation-specific PCR), DNA sequencing (e.g., bisulfite sequencing, whole-genome methylation sequencing, simplified methylation sequencing), methylation-sensitive restriction endonuclease assays, quantitative fluorescence assays, methylation-sensitive high-resolution melting curve assays, chip-based methylation mapping, and mass spectrometry (e.g., mass spectrometry of flight). In one or more embodiments, the detection is DNA sequencing. Preferably, the sequencing depth of the DNA sequencing is greater than or equal to 5M, more preferably 7M, 11M, 13M, or 15M.

[0103] In one or more embodiments, the sample is derived from mammalian tissue, cells, or body fluids, such as pancreatic tissue or blood. The mammal is preferably human. In one or more embodiments, the sample is a fine-needle aspiration biopsy. In one or more embodiments, the sample is plasma.

[0104] In one or more embodiments, the sample comprises genomic DNA or cfDNA.

[0105] In one or more embodiments, the sequence is transformed, wherein unmethylated cytosine is converted into a base that does not bind to guanine. The transformation is performed using an enzymatic method, preferably deaminase treatment, or the transformation is performed using a non-enzymatic method, preferably treatment with bisulfite, acid sulfite, or metabisulfite, or a combination thereof.

[0106] In one or more embodiments, the DNA sequence is treated with a methylation-sensitive restriction endonuclease.

[0107] In one or more implementations, the score in step (2) is calculated by constructing a support vector machine model.

[0108] In one or more embodiments, step (3) includes: comparing the methylation level of the subject sample with that of a control sample, and identifying the subject as having pancreatic cancer when the methylation level meets a threshold.

[0109] In one or more implementations, step (3) includes identifying the subject as having pancreatic cancer when the score meets a threshold. Attached Figure Description

[0110] Figure 1 This is a flowchart of the technical solution of one embodiment of the present invention.

[0111] Figure 2 This is the ROC curve of the pancreatic cancer prediction model in the training and testing groups, distinguishing between chronic pancreatitis and pancreatic cancer.

[0112] Figure 3 This is the distribution of prediction scores for the pancreatic cancer prediction model across different groups.

[0113] Figure 4 This represents the methylation levels of three methylation markers in the training group.

[0114] Figure 5 This represents the methylation levels of three methylation markers in the test group.

[0115] Figure 6 This is the ROC curve of the pancreatic cancer prediction model in samples that were negative by traditional methods (i.e., CA19-9 measurement value less than 37) for diagnosing pancreatic cancer. Detailed Implementation

[0116] This invention explores the relationship between DNA methylation and pancreatic cancer and pancreatitis, particularly chronic pancreatitis. The aim is to improve the accuracy of non-invasive diagnosis of pancreatic cancer by utilizing DNA methylation levels as a biomarker for differentiating between pancreatic cancer and chronic pancreatitis using a non-invasive method.

[0117] The inventors have discovered that the differentiation between pancreatic cancer and pancreatitis (e.g., chronic pancreatitis) is associated with the methylation levels of one, two, or three genes selected from the following or sequences within 20 kb upstream or downstream of them: SIX3, TLX2, CILP2. In one or more embodiments, the differentiation between pancreatic cancer and pancreatitis is associated with the methylation levels of genes selected from any of the following groups: (1) SIX3, TLX2; (2) SIX3, CILP2; (3) TLX2, CILP2; (4) SIX3, TLX2, CILP2. This invention provides nucleic acid molecules containing one or more CpGs of the above-mentioned genes or fragments thereof.

[0118] In this article, the term "gene" includes both coding and non-coding sequences of the gene in question on the genome. Non-coding sequences include introns, promoters, and regulatory elements or sequences.

[0119] Furthermore, the differentiation between pancreatic cancer and pancreatitis is associated with the methylation level of any one of the following segments or two or all three segments randomly selected from: SEQ ID NO:1 in the SIX3 gene region, SEQ ID NO:2 in the TLX2 gene region, and SEQ ID NO:3 in the CILP2 gene region.

[0120] In some implementations, the identification of pancreatic cancer and pancreatitis is associated with the methylation level of sequences selected from any of the following groups or their complementary sequences: (1) SEQ ID NO:1, SEQ ID NO:2, (2) SEQ ID NO:1, SEQ ID NO:3, (3) SEQ ID NO:2, SEQ ID NO:3, (4) SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3.

[0121] The “sequences relevant to the differentiation of pancreatic cancer and pancreatitis” mentioned in this article include the above three genes, sequences within 20kb upstream or downstream of them, the above three sequences (SEQ ID NO:1-3) or their complementary sequences.

[0122] The locations of the above three sequences in human chromosomes are as follows: SEQ ID NO:1:chr2, 45028785-45029307; SEQ ID NO:2:chr2, 74742834-74743351; SEQ ID NO:3:chr19, 19650745-19651270. In this paper, the base numbers of each sequence and methylation site correspond to the reference genome HG19.

[0123] In one or more embodiments, the nucleic acid molecule described herein is a fragment selected from one or more genes selected from SIX3, TLX2, CILP2; the fragment is 1bp-1kb in length, preferably 1bp-700bp; the fragment contains one or more methylation sites in the chromosomal region of the corresponding gene. The methylation sites in the genes or their fragments mentioned in this article include, but are not limited to: 45028802, 45028816, 45028832, 45028839, 45028956, 45028961, 45028965, 45028973, 45029004, 45029017, 45029035, 45029046, 45029057, 45029060, 45029063, 45029065, 45029071, 45029106, 45029112, 45029117, 45029128, and 45029. 146,45029176,45029179,45029184,45029189,45029192,45029195,45029218,45029226,45029228,45029231,45029235,45029263,45029273,45029285,45029288,45029295,74742838,74742840,74742844,74742855,74742879,74742882,74742891,74742913,747 42922,74742925,74742942,74742950,74742953,74742967,74742981,74742984,74742996,74743004,74743006,74743009,74743011,74743015,74743021,74743035,74743056,74743059,74743061,74743064,74743068,74743073,74743082,74743084,74743101,7 4743108,74743111,74743119,74743121,74743127,74743131,74743137,74743139,74743141,74743146,74743172,74743174,74743182,74743186,74743191,74743195,74743198,74743207,74743231,74743234,74743241,74743243,74743268,74743295,74743301,74743306,74743318,74743321,74743325,74743329,74743333,74743336,74743343,74743346; chr19 of 19650766,19650791,19650796,19650822,19650837,19650839,19650874,1965 0882, 19650887, 19650893, 19650895, 19650899, 19650907, 19650917, 19650955, 19650978, 19650981, 19650995, 19650997, 19651001, 19651008, 19651020, 19651028, 19651041, 196510 53, 19651059, 19651062, 19651065, 19651071, 19651090, 19651101, 19651109, 19651111, 19651113, 19651121, 19651123, 19651127, 19651133, 19651142, 19651144, 19651151, 1965116 The bases at the methylation sites listed above that have not undergone mutations correspond to HG19 in the reference genome. (List of bases is missing from the original text.)

[0124] In one or more embodiments, the nucleic acid molecule length is 1bp-1000bp, 1bp-900bp, 1bp-800bp, or 1bp-700bp. The nucleic acid molecule length can be any of the above-mentioned ranges.

[0125] In this document, methods for detecting DNA methylation are well known in the art, such as bisulfite-conversion-based PCR (e.g., methylation-specific PCR, MSP), DNA sequencing, whole-genome methylation sequencing, simplified methylation sequencing, methylation-sensitive restriction endonuclease assays, quantitative fluorescence assays, methylation-sensitive high-resolution melting curve assays, chip-based methylation mapping, and mass spectrometry. In one or more embodiments, the detection includes detecting any strand at a gene or site.

[0126] Therefore, this invention relates to reagents for detecting DNA methylation. Reagents used in the above-described methods for detecting DNA methylation are well known in the art. In detection methods involving DNA amplification, reagents for detecting DNA methylation include primers. The primer sequences are methylation-specific or non-specific. Preferably, the primer sequences may include non-methylation-specific blocking sequences. Blocking sequences can improve the specificity of methylation detection. Reagents for detecting DNA methylation may also include probes. Typically, the 5' end of the probe sequence is labeled with a fluorescent reporter group, and the 3' end is labeled with a quencher group. Exemplarily, the probe sequence contains MGB (Minor Groove Binder) or LNA (Locked Nucleic Acid). MGB and LNA are used to increase Tm values, increase the specificity of the analysis, and improve the flexibility of probe design.

[0127] The term "primer" as used in this article refers to a nucleic acid molecule with a specific nucleotide sequence that guides the synthesis of nucleotides at the initiation of nucleotide polymerization. Primers are typically two artificially synthesized oligonucleotide sequences. One primer is complementary to one DNA template strand at one end of the target region, and the other primer is complementary to the other DNA template strand at the other end of the target region. Their function is to serve as the initiation point for nucleotide polymerization. Primers are usually at least 9 bp. Artificially designed primers are widely used in polymerase chain reaction (PCR), qPCR, sequencing, and probe synthesis. Typically, primers are designed to amplify products with lengths of 1-2000 bp, 10-1000 bp, 30-900 bp, 40-800 bp, 50-700 bp, or at least 150 bp, at least 140 bp, at least 130 bp, or at least 120 bp.

[0128] The term "variant" or "mutant" in this document refers to a polynucleotide whose nucleic acid sequence is altered compared to a reference sequence by the insertion, deletion, or substitution of one or more nucleotides while retaining its ability to hybridize with other nucleic acids. A mutant described in any embodiment of this document comprises a nucleotide sequence having at least 70%, preferably at least 80%, preferably at least 85%, preferably at least 90%, preferably at least 95%, preferably at least 97% sequence identity with a reference sequence and retaining the biological activity of the reference sequence. Sequence identity between two aligned sequences can be calculated using, for example, NCBI's BLASTn. A mutant also includes a nucleotide sequence having one or more mutations (insertions, deletions, or substitutions) in the reference sequence and its nucleotide sequence while still retaining the biological activity of the reference sequence. The multiple mutations typically refer to 1-10, for example, 1-8, 1-5, or 1-3. Substitution can be between purine nucleotides and pyrimidine nucleotides, or between purine nucleotides or pyrimidine nucleotides. Substitution is preferably conserved. For example, in the art, conserved substitution with nucleotides of similar or comparable properties generally does not alter the stability and function of the polynucleotide. Conservative substitutions include, for example, the interchange of (A and G) between purine nucleotides and the interchange of (T or U and C) between pyrimidine nucleotides. Therefore, replacing one or more sites with residues from the same source in the polynucleotides of this invention will not substantially affect their activity. Furthermore, methylation sites (e.g., consecutive CG) in the variants of this invention are not mutated. That is, the method of this invention detects the methylation status of methylable sites in the corresponding sequence; mutations can occur at bases in non-methylable sites. Typically, methylation sites are consecutive CpG dinucleotides.

[0129] As described herein, base conversions can occur between DNA or RNA bases. The terms "conversion," "cytosine conversion," or "CT conversion" used herein refer to the process of treating DNA using non-enzymatic or enzymatic methods to convert unmodified cytosine bases (C) into bases with a lower binding affinity to guanine (e.g., uracil bases (U)). Non-enzymatic or enzymatic methods for performing cytosine conversions are well known in the art. Exemplarily, non-enzymatic methods include treatment with conversion reagents such as bisulfites, acid sulfites, or metabisulfites, such as calcium bisulfite, sodium bisulfite, potassium bisulfite, ammonium bisulfite, sodium disulfite, potassium disulfite, and ammonium disulfite. Exemplarily, enzymatic methods include deaminase treatment. The converted DNA may optionally be purified. DNA purification methods suitable for use herein are well known in the art.

[0130] This invention also provides a methylation detection kit for differentiating pancreatic cancer and pancreatitis. The kit includes the primers and / or probes described herein for detecting the methylation levels of sequences related to the differentiation of pancreatic cancer and pancreatitis discovered by the inventors. The kit may also contain nucleic acid molecules described herein, particularly those described in the first aspect, as internal standards or positive controls.

[0131] The term "hybridization" as used in this article primarily refers to nucleic acid sequence pairing under stringent conditions. An exemplary stringent condition is hybridization followed by membrane washing in a solution of 0.1×SSPE (or 0.1×SSC) and 0.1% SDS at 65°C.

[0132] In addition to the primers, probes, and nucleic acid molecules described above, the kit also contains other reagents required for the detection of DNA methylation. Exemplarily, these other reagents for detecting DNA methylation may include one or more of the following: bisulfite and its derivatives, PCR buffer, polymerase, dNTPs, primers, probes, methylation-sensitive or non-methylation-sensitive restriction endonucleases, enzyme digestion buffers, fluorescent dyes, fluorescence quenchers, fluorescent reporter agents, exonucleases, alkaline phosphatase, internal standards, and controls.

[0133] The kit may further include converted positive standards, wherein unmethylated cytosine is converted into bases that do not bind to guanine. The positive standards may be fully methylated. The kit may also include PCR reaction reagents. Preferably, the PCR reaction reagents include Taq DNA polymerase, PCR buffer, dNTPs, and Mg2+.

[0134] The present invention also provides a method for differentiating pancreatic cancer and pancreatitis, comprising: (1) detecting the methylation level of the pancreatic cancer and pancreatitis differentiation-related sequences described herein in the sample of the subject; (2) comparing with a control sample, or calculating a score; and (3) differentiating pancreatic cancer and pancreatitis of the subject based on the score. Typically, the method further includes, prior to step (1): extraction of sample DNA, quality control, and / or conversion of unmethylated cytosine on the DNA into bases that do not bind to guanine.

[0135] The subjects are, for example, patients diagnosed with or previously diagnosed with pancreatitis. That is, in one or more embodiments, the method identifies pancreatic cancer in patients diagnosed with chronic pancreatitis (including those previously diagnosed). Of course, the method of the present invention is not limited to the above-described subjects and can also be used to directly diagnose or differentiate pancreatitis or pancreatic cancer in undiagnosed subjects.

[0136] In a specific implementation, step (1) includes: treating genomic DNA or cfDNA with a transformation reagent to convert unmethylated cytosine into a base (e.g., uracil) that has a lower binding affinity to guanine; performing PCR amplification using primers suitable for amplifying the transformed sequences of the pancreatic cancer and pancreatitis identification-related sequences described herein; and determining the methylation level of at least one CpG by the presence or absence of the amplification product or by sequence identification (e.g., probe-based PCR detection or DNA sequencing identification).

[0137] Alternatively, step (1) may also include: treating genomic DNA or cfDNA with a methylation-sensitive restriction endonuclease; performing PCR amplification using primers suitable for amplifying sequences having at least one CpG in the pancreatic cancer and pancreatitis identification-related sequences described herein; and determining the methylation level of at least one CpG by the presence or absence of the amplification product.

[0138] The term "methylation level" as used herein refers to the relationship between the methylation levels of any number and any position of CpGs in the sequence in question. This relationship can be the result of addition or subtraction of methylation level parameters (e.g., 0 or 1) or calculations using mathematical algorithms (e.g., mean, percentage, fraction, proportion, degree, or calculations using mathematical models), including but not limited to methylation level metrics, methylation haplotype ratios, or methylation haplotype loadings. The term "methylation status" indicates the methylation of a specific CpG site, typically including methylated or unmethylated sites (e.g., methylation status parameter 0 or 1).

[0139] In one or more embodiments, the methylation level of the target sample increases or decreases when compared with a control sample. When the methylation marker level meets a certain threshold, pancreatic cancer is identified; otherwise, chronic pancreatitis is identified. Alternatively, mathematical analysis can be performed on the methylation level of the tested gene to obtain a score. For the tested sample, if the score is greater than the threshold, the result is considered positive, i.e., pancreatic cancer; otherwise, it is considered negative, i.e., pancreatitis. Conventional mathematical analysis methods and processes for determining thresholds are known in the art; an exemplary method is a support vector machine (SVM) mathematical model. For example, for differential methylation markers, an SVM is constructed on the training set samples, and the model is used to statistically analyze the accuracy, sensitivity, and specificity of the detection results, as well as the area under the predictive value characteristic curve (ROC) (AUC), and to statistically analyze the predicted scores for the test set samples. In an implementation of the SVM, the scoring threshold is 0.897; a score greater than 0.897 indicates that the subject is considered a pancreatic cancer patient, otherwise, a chronic pancreatitis patient.

[0140] In a preferred embodiment, the model training process is as follows: First, differentially methylated regions are obtained based on the methylation level of each site, and a differentially methylated region matrix is ​​constructed. For example, a methylation data matrix can be constructed from the methylation level data of a single CpG dinucleotide position in the HG19 genome using software such as samtools; then, SVM model training is performed.

[0141] An exemplary SVM model training process is as follows:

[0142] a) Use the sklearn package (v0.23.1) of Python software (v3.6.9) to build the training mode of the cross-validation training model. Command line: model=SVR().

[0143] b) Using the sklearn package (v0.23.1), input the data matrix and build an SVM model, model.fit(x_train,y_train), where x_train represents the training set data matrix and y_train represents the phenotypic information of the training set.

[0144] Typically, during model construction, pancreatic cancer is encoded as 1, and pancreatitis as 0. In this invention, the threshold is set to 0.897 using Python software (v3.6.9) and the sklearn package (v0.23.1). The constructed model ultimately uses a threshold of 0.897 to distinguish samples.

[0145] In this document, the samples are derived from mammalian subjects, preferably humans. Samples can be derived from any organ (e.g., pancreas), tissue (e.g., epithelial tissue, connective tissue, muscle tissue, and nerve tissue), cell, or body fluid (e.g., blood, plasma, serum, tissue fluid, urine). Generally, the sample is acceptable as long as it contains genomic DNA or cfDNA (circulating free DNA or cell-free DNA). cfDNA, also known as circulating cell-free DNA or cell-free DNA, is a fragment of degraded DNA released into the plasma. Exemplarily, the sample is a pancreatic cancer biopsy, preferably a fine-needle aspiration biopsy. Alternatively, the sample may be plasma or cfDNA.

[0146] This article also relates to methods for obtaining methylation haplotype ratios associated with pancreatic cancer and pancreatitis. Taking methylation data obtained from methylation-targeted sequencing (MethylTitan) as an example, the process of screening and testing biomarker sites is as follows: raw paired-end sequencing reads – readings are merged to obtain merged single-end reads – adapters are removed to obtain adapter-removed reads – Bismark is aligned to the human DNA genome to form a BAM file – samtools extracts the CpG site methylation level of each read to form a haplotype file – the proportion of C site methylation haplotype ratios is statistically analyzed to form a meth file – MHF (Methylated Haplotype Fraction) methylation values ​​are calculated – Coverage 200 filter sites are used to form a meth.matrix matrix file – filter sites according to NA values ​​greater than 0.1 – samples are pre-divided into training and test sets – for each haplotype in the training set, a logistic regression model is constructed for the phenotype, and the regression P-value of each methylation haplotype ratio is selected – the methylation haplotype ratio with the most significant P-value in each MethylTitan amplification region is selected to represent the methylation haplotype ratio level of that region and modeled using a support vector machine – the results of the training set (ROC plot) are generated and the model is used to predict the test set for validation. Specifically, the method for obtaining the methylation haplotype ratio associated with pancreatic cancer includes the following steps: (1) obtaining plasma samples from patients with pancreatic cancer or pancreatitis to be tested, extracting cfDNA, and performing library construction and sequencing using the MethylTitan method to obtain sequencing reads; (2) preprocessing the sequencing data, including adapter removal and splicing of the sequencing data generated by the sequencer; (3) aligning the preprocessed sequencing data to the HG19 reference genome sequence of the human genome to determine the position of each fragment. The data in step (2) can be obtained from paired-end 150bp sequencing on the Illumina sequencing platform. Adapter removal in step (2) involves removing the sequencing adapters at the 5' and 3' ends of the two paired-end sequencing data respectively, as well as low-quality base removal after adapter removal. Splicing in step (2) involves merging the paired-end sequencing data to restore the original library fragment. This allows for better alignment and accurate location of sequencing fragments. For example, the length of the sequencing library is about 180bp, and paired-end 150bp can completely cover the entire library fragment. Step (3) includes: (a) converting the HG19 reference genome data to CT and GA respectively, constructing two sets of converted reference genomes, and constructing alignment indexes for the converted reference genomes respectively; (b) converting the merged sequencing sequence data to CT and GA respectively; (c) aligning the converted reference genome sequences respectively, and finally summarizing the alignment results to determine the position of the sequencing data in the reference genome.

[0147] Furthermore, the method for obtaining methylation haplotype ratios associated with pancreatic cancer and pancreatitis also includes (4) calculating MHF; (5) constructing a methylation haplotype ratio MHF data matrix; and (6) constructing a logistic regression model for each methylation haplotype ratio based on sample grouping. Step (4) involves obtaining the methylation haplotype ratio status and sequencing depth information at the location of the HG19 reference genome based on the alignment results obtained in step (3). Step (5) involves merging the methylation haplotype ratio status and sequencing depth information data into a data matrix. Among them, each data point with a depth less than 200 is treated as a missing value, and the missing values ​​are filled using the K nearest neighbor (KNN) method. Step (6) involves statistically modeling each location in the above matrix using logistic regression and screening for haplotypes with significant regression coefficients between the two groups.

[0148] The term "multiple" as used herein refers to any integer. Preferably, "multiple" in "one or more" can be any integer greater than or equal to 2, including 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60 or more.

[0149] The beneficial effects of this invention are:

[0150] Based on the DNA methylation biomarkers of this invention, patients with pancreatic cancer and chronic pancreatitis can be effectively distinguished. This invention provides a diagnostic model of the relationship between cfDNA methylation biomarker methylation levels and pancreatic cancer based on high-throughput plasma cfDNA methylation sequencing. This model has the advantages of non-invasive detection, safe and convenient detection, high throughput, and high detection specificity. Based on the optimal sequencing volume obtained by this invention, detection costs can be effectively controlled while achieving good detection performance.

[0151] Example

[0152] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments. In the following embodiments, experimental methods without specific conditions are generally performed according to the methods described under conventional conditions.

[0153] Example 1: Methylation-targeted sequencing to screen differentially methylation sites in pancreatic cancer

[0154] The inventors collected blood samples from a total of 94 pancreatic cancer patients and 25 patients with chronic pancreatitis. All enrolled patients signed informed consent forms. The pancreatic cancer patients had a history of pancreatitis diagnosis. Sample information is shown in the table below.

[0155]

[0156]

[0157] Plasma DNA methylation sequencing data were obtained using the MethylTitan method, and DNA methylation classification markers were identified. The process is as follows:

[0158] 1. Extraction of plasma cfDNA samples

[0159] A 2ml whole blood sample was collected from the patient using a streck blood collection tube. Plasma was separated by centrifugation within 3 days and then transferred to the laboratory. cfDNA was extracted using the QIAGEN QIAamp Circulating Nucleic Acid Kit according to the instructions.

[0160] 2. Sequencing and Data Preprocessing

[0161] 1) The library was sequenced using an Illumina Nextseq 500 sequencer with paired ends.

[0162] 2) The Pear (v0.6.0) software merges the paired-end sequencing data of the same fragment of 150bp from the Illumina Hiseq X10 / Nextseq 500 / Nova seq sequencer into a single sequence with a minimum overlap length of 20bp and a minimum length of 30bp after merging.

[0163] 3) Trim_galore v 0.6.0 and cutadapt v1.8.1 software were used to remove adapters from the merged sequencing data. The adapter sequence “AGATCGGAAGAGCAC” was removed from the 5' end of the sequence, and bases with sequencing quality values ​​lower than 20 at both ends were removed.

[0164] 3. Sequencing data alignment

[0165] The reference genome data used in this article came from the UCSC database (UCSC:HG19, http: / / hgdownload.soe.ucsc.edu / goldenPath / hg19 / bigZips / hg19.fa.gz).

[0166] 1) First, HG19 was transformed into cytosine to thymine (CT) and adenine to guanine (GA) using Bismark software, and the transformed genomes were indexed using Bowtie2 software.

[0167] 2) Perform CT and GA conversion on the preprocessed data as well.

[0168] 3) Use Bowtie2 software to align the transformed sequences to the transformed HG19 reference genome. The minimum seed sequence length is 20, and mismatches in the seed sequence are not allowed.

[0169] 4. Calculation of MHF

[0170] For each target region HG19 CpG site, the methylation status corresponding to each site is obtained based on the alignment results above. In this paper, the nucleotide number of the site corresponds to the nucleotide position number of HG19. A target methylation region may have multiple methylation haplotypes. For each methylation haplotype within the target region, this value needs to be calculated. An example of the MHF calculation formula is as follows:

[0171]

[0172] Where i represents the target methylation region, h represents the target methylation haplotype, and N i N represents the number of reads located in the target methylation region. i,h This indicates the number of reads containing the target methylated haplotype.

[0173] 5. Methylation data matrix

[0174] 1) Merge the methylation sequencing data of each sample in the training set and the test set into a data matrix, and perform missing value processing on each site with a depth of less than 200.

[0175] 2) Remove sites with a missing value ratio higher than 10%.

[0176] 3) For missing values ​​in the data matrix, the KNN algorithm is used to impute the missing data.

[0177] 6. Identify characteristic methylation regions by grouping training set samples.

[0178] 1) For each methylation segment, a logistic regression model is constructed for the phenotype. For each amplified target region, the methylation segment with the most significant regression coefficient is selected to form a candidate methylation segment.

[0179] 2) Randomly divide the training set into ten parts and perform incremental feature selection with tenfold cross-validation.

[0180] 3) The candidate methylation segments in each region are sorted from largest to smallest according to the significance of the regression coefficient. One methylation segment is added at a time to predict the test data.

[0181] 4) In step 3), use the 10 datasets generated in step 2) to calculate the AUC 10 times each time, and take the average of the 10 calculations. If the AUC of the training data increases, retain the candidate methylation region as a feature methylation region; otherwise, discard it.

[0182] 5) Take the feature combination corresponding to the median of the average AUC under different feature counts in the training set as the final determined feature methylation segment combination.

[0183] The distribution of the selected characteristic methylation biomarkers in HG19 is as follows: SEQ ID NO:1 in the SIX3 gene region, SEQ ID NO:2 in the TLX2 gene region, and SEQ ID NO:3 in the CILP2 gene region. The levels of these methylation biomarkers increased or decreased in the cfDNA of pancreatic cancer patients (Table 1). The sequences of the three biomarker regions are shown in SEQ ID NO:1-3. The methylation levels of all CpG sites in each biomarker region can be obtained by MethylTitan sequencing. The mean methylation level of all CpG sites in each region, as well as the methylation status of individual CpG sites, can be used as biomarkers for diagnosing pancreatic cancer.

[0184] Table 1: Methylation levels of DNA methylation biomarkers in the training set

[0185] SEQ ID NO:1 chr2:45028785-45029307 0.843731054 0.909570522 SEQ ID NO:2 chr2:74742834-74743351 0.953274962 0.978544302 SEQ ID NO:3 chr19:19650745-19651270 0.408843665 0.514101315

[0186] Table 2 shows the methylation levels of methylation markers in the test group of pancreatic cancer and chronic pancreatitis. As can be seen from the table, the distribution of methylation levels of methylation markers differs significantly between the pancreatic cancer and chronic pancreatitis populations, demonstrating good discriminatory power.

[0187] Table 2: Methylation levels of DNA methylation biomarkers in the test set

[0188] SEQ ID NO:1 chr2:45028785-45029307 0.843896661 0.86791556 SEQ ID NO:2 chr2:74742834-74743351 0.926459851 0.954493044 SEQ ID NO:3 chr19:19650745-19651270 0.399831579 0.44918572

[0189] Table 3 lists the correlation (Pearson correlation coefficient) between the methylation level of 10 random CpG sites or combinations in each selected marker and the methylation level of the entire marker, as well as the corresponding significance p-value. It can be seen that the methylation status or level of a single CpG site or a combination of multiple CpG sites in the marker is significantly correlated with the methylation level of the entire region (p<0.05), and the correlation coefficients are all above 0.8, indicating strong or very strong correlation. This shows that a single CpG site or a combination of multiple CpG sites in the marker also has a good distinguishing effect, just like the entire marker.

[0190] Table 3: Correlation between methylation levels of random CpG sites or combinations of sites in the three biomarkers and the overall methylation level of the biomarkers.

[0191]

[0192]

[0193] Example 2: Predictive performance of a single methylation marker

[0194] To verify the ability of a single methylation biomarker to distinguish between pancreatitis and pancreatic cancer, the methylation level of the single methylation biomarker was used to validate its predictive performance.

[0195] First, the methylation levels of the three methylation markers were used individually in the training set samples to determine the threshold, sensitivity, and specificity for distinguishing between pancreatic cancer and pancreatitis. Then, the threshold was used to statistically analyze the sensitivity and specificity of the test set samples. The results are shown in Table 4 below, which shows that a single marker can also achieve good distinguishing performance.

[0196] Table 4: Predictive performance of 56 methylation markers

[0197] SEQ ID NO:1 training set 0.8870 0.7937 0.8824 0.8850 SEQ ID NO:1 test set 0.6532 0.7742 0.3750 0.8850 SEQ ID NO:2 training set 0.8497 0.6508 0.8824 0.9653 SEQ ID NO:2 test set 0.6210 0.8065 0.5000 0.9653 SEQ ID NO:3 training set 0.8301 0.4286 0.8824 0.3984 SEQ ID NO:3 test set 0.6694 0.5806 0.6250 0.3984

[0198] Example 3: Constructing a classification prediction model

[0199] To validate the potential of a classifier for pancreatic cancer-chronic pancreatitis patients using DNA methylation markers (such as methylation haplotype ratios), a support vector machine disease classification model was constructed based on a combination of three DNA methylation markers in the training group. The classification prediction performance of this set of DNA methylation markers was then validated in the test group. The training and test groups were divided proportionally, with 80 cases (samples 1-80) in the training group and 39 cases (samples 80-119) in the test group.

[0200] Support vector machine models were built on the training set using the discovered DNA methylation markers for both groups of samples.

[0201] 1) Divide the samples into two parts in advance, one part for training the model and the other part for testing the model.

[0202] 2) To explore the potential of using methylation biomarkers for pancreatic cancer identification, a disease classification system based on gene biomarkers was developed. An SVM model was trained using the methylation biomarker levels in the training set. The specific training process is as follows:

[0203] a) Use the sklearn package (v0.23.1) of Python software (v3.6.9) to build the training mode of the cross-validation training model. Command line: model=SVR().

[0204] b) Using the sklearn package (v0.23.1), input the methylation numerical matrix and build an SVM model, model.fit(x_train,y_train), where x_train represents the methylation numerical matrix of the training set and y_train represents the phenotypic information of the training set.

[0205] During model construction, pancreatic cancer was coded as 1 and chronic pancreatitis as 0. The threshold was set to 0.897 by default using the sklearn package (v0.23.1). The final model also used a scoring threshold of 0.897 to distinguish between pancreatic cancer and pancreatitis. The prediction scores of the two models for the training set samples are shown in Table 5.

[0206] Table 5: Model prediction scores on the training set

[0207]

[0208]

[0209] Example 4: Classification Prediction Model Testing

[0210] Blood samples from the aforementioned pancreatic cancer and pancreatitis patients were subjected to MethylTitan sequencing. Based on the characteristic methylation marker signals in the sequencing results, PCA, clustering, and other classification analyses were performed.

[0211] Based on the methylation biomarker group of this invention, predictions were made on the test set according to the model established by SVM in Example 3. A prediction function was used to predict the test set, and the output was the prediction result (disease probability: the default scoring threshold is 0.897; a score greater than 0.897 indicates the subject is considered a pancreatic cancer patient, otherwise a chronic pancreatitis patient). The test group sample consisted of 57 cases (samples 118-174), and the calculation process is as follows:

[0212] Command line:

[0213] test_pred=model.predict(test_df)

[0214] Where test_pred represents the prediction score obtained by the SVM prediction model constructed in Example 3 for the test set samples, model represents the SVM prediction model constructed in Example 3, and test_df represents the test set data.

[0215] The predicted scores for the test group are shown in Table 6, and the ROC curves are as follows: Figure 2 As shown, the distribution of predicted scores is as follows: Figure 3 As shown, the area under the overall AUC of the test group is 0.847. In the training set, the model achieves a sensitivity of 88.9% when the specificity is 88.2%; in the test set, the sensitivity reaches 74.2% when the specificity is 87.5%. This demonstrates that the SVM model built from the selected variables exhibits good discrimination.

[0216] Figure 4 and Figure 5 The distribution of the three methylation markers in the training and testing groups were shown separately. It can be found that the differences of the methylation markers in the plasma of patients with pancreatitis and patients with pancreatic cancer are relatively stable.

[0217] Table 6: Prediction scores of the model on the test set

[0218] Sample 81 Chronic pancreatitis 0.610488911 Sample 101 pancreatic cancer 15.62766141 Sample 82 pancreatic cancer 0.912018264 Sample 102 pancreatic cancer 0.909976179 Sample 83 pancreatic cancer 0.870225426 Sample 103 pancreatic cancer 0.92289051 Sample 84 pancreatic cancer 0.897368929 Sample 104 pancreatic cancer 1.823319531 Sample 85 pancreatic cancer 1.491556374 Sample 105 pancreatic cancer 0.913625979 Sample 86 pancreatic cancer 0.99785215 Sample 106 pancreatic cancer 0.730447081 Sample 87 pancreatic cancer 0.909901733 Sample 107 pancreatic cancer 0.900701224 Sample 88 pancreatic cancer 0.955726751 Sample 108 Chronic pancreatitis 0.893221308 Sample 89 pancreatic cancer 0.96582068 Sample 109 Chronic pancreatitis 0.899073184 Sample 90 pancreatic cancer 0.910414113 Sample 110 Chronic pancreatitis 0.783284566 Sample 91 pancreatic cancer 0.850903621 Sample 111 Chronic pancreatitis 0.725251615 Sample 92 pancreatic cancer 0.916651697 Sample 112 pancreatic cancer 0.893141436 Sample 93 Chronic pancreatitis 0.904231501 Sample 113 pancreatic cancer 1.354991317 Sample 94 pancreatic cancer 0.764872522 Sample 114 pancreatic cancer 0.817727331 Sample 95 pancreatic cancer 1.241367038 Sample 115 pancreatic cancer 1.079401681 Sample 96 Chronic pancreatitis 0.897789105 Sample 116 pancreatic cancer 0.969607597 Sample 97 Chronic pancreatitis 0.852404121 Sample 117 pancreatic cancer 0.878877727 Sample 98 pancreatic cancer 1.068601129 Sample 118 pancreatic cancer 0.911801452 Sample 99 pancreatic cancer 3.715591125 Sample 119 pancreatic cancer 0.934497862 Sample 100 pancreatic cancer 0.920532374

[0219] Example 5: Predictive effect for patients with negative tumor markers

[0220] Based on the methylation biomarker group of the present invention, patients who were negative for tumor marker CA19-9 (<37) were identified according to the model established by SVM in Example 3.

[0221] The predicted scores for the test group are shown in Table 7, and the ROC curves are as follows: Figure 6 As shown, the constructed SVM model can achieve good results even for patients who cannot be distinguished by the traditional tumor marker CA19-9.

[0222] Table 7: CA19-9 Measurement Values ​​and SVM Model Predicted Scores

[0223]

[0224]

[0225] This study investigated the differences between the plasma methylation levels of methylation markers in plasma cfDNA and those in patients with chronic pancreatitis and pancreatic cancer, identifying three DNA methylation markers with significant differences. Based on these DNA methylation markers, a risk prediction model for malignant pancreatic cancer was established using support vector machines. This model effectively distinguishes between patients with pancreatic cancer and those with chronic pancreatitis, exhibiting high sensitivity and specificity, and is suitable for screening and diagnosing pancreatic cancer in patients with chronic pancreatitis. sequence list <110> Jiangsu Kunyuan Biotechnology Co., Ltd. <120> Methylation marker for differentiating pancreatitis and pancreatic cancer and use thereof <130> 214175 <160> 3 <170> SIPOSequenceListing 1.0 <210> 1 <211> 523 <212> DNA <213> Homo sapiens <400> 1 taatttatgg aatccaccgt cacactctct ccgagcagcc agctccccgc ttaacgggga 60 aattgaagca gacagccttt gtctaaacac ttcttttgcc cagaatatct taattttcct 120 atttgaatgt ttaataaggt ttggggtgca gcagcttcct tttaattgtg acggtgcggc 180 cgcttgggcg tgatcccttg gctggggctg cagggggccc gtcctccagg ggcgcagagg 240 gaaggaccag cgtttccaag ccgggctctg gccgccggcg cgagagcgag gccaaggtct 300 gggggcagtt cagggggacc ccgaagtcgg gacggcccag aaacgctttg cccacagcca 360 ccgccctttc ctttgtgagt ttccccaaag ccgtcggtgc gacccggcgc cgactctcct 420 cctcttctcc ctgcgagggc ccgcgccgcc cgggcccagt cctgggggat agatccctcg 480 gggcccaacg gctgggccac cgccggtctc cggccactgc tgc 523 <210> 2 <211> 518 <212> DNA <213> Homo sapiens <400> 2 aagccgcgca cgtccttctc ccgctcacag gtgctggagt tggagcggcg cttcctgcgc 60 cagaagtacc tggcctctgc ggagagggcg gcgctggcca aggccttgcg catgaccgac 120 gcacaggtca aaacgtggtt ccagaaccga cgcaccaagt ggcggtgagg cgcggcgcgg 180 gcgagggcgg actggggttc ccgagcaggg cctggtgaga agcgacgcgg cgggcgcccc 240 gctgaccccg cgtctccctc ccttaggcgc cagacggcgg aggagcgcga ggccgagcgg 300 caccgcgcgg gccggctgct cctgcatctg cagcaggacg cgttgccacg gccgctgcgg 360 ccgccgctgc ccccggaccc tctctgcctg cacaactcgt cgctcttcgc gctgcagaac 420 ctgcagccct gggccgagga caacaaagtg gcttcagtgt ccgggctcgc ctcggtggtg 480 tgagcgacgc ccgtccgatc ggcgtggagc gccgggcc 518 <210> 3 <211> 526 <212> DNA <213> Homo sapiens <400> 3 ttcaagatct aagtgagagg ccggtcagac agaggcaaga gctcagcgca ccgggatgga 60 ccaggtcagg ccctgggcgg cagaactggg gtcgcgggga acccagtctg ccctgcacct 120 gtttcaggcc gctggctcgg gtcgtgggcg cgctcggcta gccggtgccc accgggggag 180 ggggctgaga cagcaagtaa ggcctttgca cgcatgcatg ggggcctaca ggccgccgcc 240 ctggtcccag cgcgtgcggt gcccgcagag gccagcgagt ggacgtcctg gttcaacgtg 300 gaccaccccg gaggcgacgg cgacttcgag agcctggctg ccatccgctt ctactacggg 360 ccagcgcgcg tgtgcccgcg accgctggcg ctggaagcgc gcaccacgga ctgggccctg 420 ccgtccgccg tcggcgagcg cgtgcacttg aaccccacgc gcggcttctg gtgcctcaac 480 cgcgagcaac cgcgtggccg ccgctgctcc aactaccacg tgcgct 526

Claims

1. The use of a reagent or apparatus for determining the methylation level of a DNA sequence in a sample of a subject, or the methylation state or level of one or more CpG dinucleotides in said DNA sequence or a fragment thereof, in the preparation of a kit for the identification of pancreatic cancer and pancreatitis. in, The DNA sequence contains SEQ ID NO:3 or its complementary sequence.

2. The use as described in claim 1, characterized in that, The kit also includes: a processed nucleic acid molecule of the DNA sequence, wherein the processing converts unmethylated cytosine into bases that have a lower binding affinity to guanine than cytosine.

3. The use as described in claim 1, characterized in that, The DNA sequence may also include any of the following: (1) SEQ ID NO:1, (2) SEQ ID NO:2, (3) SEQ ID NO:1, SEQ ID NO:

2.

4. The use as described in any one of claims 1-3, characterized in that, The reagent contains primer molecules that hybridize with the DNA sequence or a fragment thereof, and / or The reagent contains probe molecules that hybridize with the DNA sequence or fragments thereof, and / or The reagent contains a medium containing a DNA sequence or fragment thereof as shown in SEQ ID NO:3 and / or its methylation information.

5. The use as described in claim 4, characterized in that, The samples are derived from mammalian tissues, cells, or body fluids.

6. The use as described in claim 4, characterized in that, The samples were derived from pancreatic tissue or blood.

7. The use as described in claim 4, characterized in that, The sample includes genomic DNA or cfDNA.

8. The use as described in claim 4, characterized in that, The DNA sequence was transformed, in which unmethylated cytosine was converted into bases that have a lower binding affinity to guanine than cytosine.

9. The use as described in claim 4, characterized in that, The DNA sequence was treated with a methylation-sensitive restriction endonuclease.

10. The use as described in any one of claims 1-3, characterized in that, The identification includes comparing with a control sample or calculating a score, and differentiating between pancreatic cancer and pancreatitis based on the score.

11. The use as described in claim 10, characterized in that, The calculations are performed by constructing a support vector machine model.

12. An apparatus for differentiating between pancreatic cancer and pancreatitis, the apparatus comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it performs the following steps: (1) Obtain the methylation level of a DNA sequence in a sample of the object, or the methylation state or level of one or more CpG dinucleotides in the DNA sequence, wherein the DNA sequence contains SEQ ID NO:3 or its complementary sequence. (2) Compare with the control sample, or calculate the score, and (3) Differentiate between pancreatic cancer and pancreatitis based on the scoring.

13. The apparatus as claimed in claim 12, characterized in that, The DNA sequence may also include any of the following: (1) SEQ ID NO:1, (2) SEQ ID NO:2, (3) SEQ ID NO:1, SEQ ID NO:

2.

14. The apparatus as claimed in claim 12, characterized in that, The sample includes genomic DNA or cfDNA.

15. The apparatus as claimed in claim 12, characterized in that, The sequence is transformed, wherein unmethylated cytosine is converted into bases that have a lower binding affinity to guanine than cytosine.

16. The apparatus as claimed in claim 12, characterized in that, The DNA sequence was treated with a methylation-sensitive restriction endonuclease.

17. The apparatus as claimed in claim 12, characterized in that, The score in step (2) is calculated by constructing a support vector machine model.