Diagnosis-related DNA methylation markers for pancreatic cancer and use thereof
Patent Information
- Application Number
- CN202110679281.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-18
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2041-06-18
AI Technical Summary
然而,大多数这些研究未取得有效的结果
[0157] Based on the methylated nucleic acid fragment biomarkers of this invention, pancreatic cancer can be effectively identified. This invention provides a diagnostic model of the relationship between cfDNA methylation biomarkers and pancreatic cancer based on high-throughput methylation sequencing of plasma cfDNA. This model has the advantages of non-invasive detection, safe and convenient detection, high throughput, and high detection specificity. Based on the optimal sequencing volume obtained by this invention, detection costs can be effectively controlled while achieving good detection performance.
Smart Images

Figure BDA0003122215630000331 
Figure BDA0003122215630000341 
Figure BDA0003122215630000351
Abstract
Description
Technical Field
[0001] This invention belongs to the field of molecular biomedical technology, specifically relating to a pancreatic cancer-related methylation marker and its application for early identification of pancreatic cancer. Background Technology
[0002] Pancreatic cancer (e.g., pancreatic ductal adenocarcinoma (PDAC)) is one of the deadliest diseases in the world. The 5-year relative survival rate is 9%, which drops further to only 3% for patients with distant metastases. A major reason for the high mortality rate is the limited availability of methods for early detection of PDAC, which is crucial for PDAC patients to undergo surgical resection. Currently, carbohydrate antigen 19-9 (CA19-9) is the most commonly used clinical serum biomarker for the auxiliary detection of PDAC, achieving a sensitivity of 79-90% and a specificity of 75-90% in symptomatic patients before resection. However, several large population studies have demonstrated that CA19-9 is ineffective in detecting PDAC in asymptomatic individuals due to its low positive predictive value, essentially ruling it out for early screening of PDAC (Kim et al., 2004) (Chang et al., 2006; Homma & Tsuchiya, 1991; Kim et al., 2004; Satake, Takeuchi, Homma, & Ozaki, 1994). Endoscopic ultrasound-guided fine-needle aspiration (EUS-FNA) is another commonly used method to obtain pathological diagnosis without open surgery, but it is invasive and requires clear imaging evidence, which usually indicates that PDAC has progressed. During tumorigenesis and development, the DNA methylation patterns and levels of malignant cell genomic DNA undergo profound changes. Some tumor-specific DNA methylations have been shown to occur early in tumorigenesis and may become "driving factors" in tumorigenesis.
[0003] Circulating tumor DNA (ctDNA) molecules originate from apoptotic or necrotic tumor cells and carry tumor-specific DNA methylation markers from early-stage malignancies. In recent years, they have been investigated as promising new targets for developing non-invasive early screening tools for various cancers. However, most of these studies have not yielded effective results. Existing research indicates that the proportion of ctDNA in the plasma DNA of patients with early-stage tumors is very small (Abbosh et al., 2017), making the identification of stable and consistent pancreatic cancer tumor-specific markers from plasma DNA a significant challenge. Summary of the Invention
[0004] This invention provides a method for detecting the methylation levels of multiple genes in a sample, and using the differential gene methylation levels in the detection results to distinguish pancreatic cancer, thereby achieving the goal of non-invasive and precise diagnosis of pancreatic cancer with higher accuracy and lower cost.
[0005] Specifically, a first aspect of the present invention provides an isolated nucleic acid molecule from a mammal, said nucleic acid molecule being a methylation marker of a gene associated with pancreatic cancer, the sequence of said nucleic acid molecule comprising (1) a sequence selected from one or more of the following sequences or variants having at least 70% identity with them: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:2 ... NO:29, SEQ ID NO:30, SEQ ID NO:31, SEQ ID NO:32, SEQ ID NO:33, SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:46, SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50, SEQ ID NO:51, SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, the methylation site in the variant is not mutated, the complementary sequence of (2)(1), the treated sequence of (3)(1) or (2), the treatment converting unmethylated cytosine into a base with a lower binding affinity to guanine than cytosine.
[0006] In one or more embodiments, the methylation sites are continuous CpG.
[0007] In one or more embodiments, the methylation marker may be any one or more CpG sites in the sequence region.
[0008] In one or more embodiments, the nucleic acid molecule is used as an internal standard or control for detecting the DNA methylation level of a corresponding sequence in a sample.
[0009] In one or more embodiments, the pancreatic cancer is pancreatic ductal adenocarcinoma.
[0010] A second aspect of the present invention provides a reagent for detecting DNA methylation, the reagent comprising a reagent for the methylation level of a DNA sequence or fragment thereof in a sample of the test subject, or the methylation status or level of one or more CpG dinucleotides in the DNA sequence or fragment thereof, wherein the DNA sequence is selected from one or more (e.g., at least seven) or all of the following gene sequences, or sequences within 20 kb upstream or downstream thereof: DMRTA2, FOXD3, TBX15, BCAN, TRIM58, SIX3, VAX2, EMX1, LBX2, TLX2, POU3F3, TBR1, EVX2, HO XD12, HOXD8, HOXD4, TOPAZ1, SHOX2, DRD5, RPL9, HOPX, SFRP2, IRX4, TBX18, OLIG3, ULBP1, HOXA13, TBX20, IKZF1, INSIG1, SOX7, E BF2, MOS, MKX, KCNA6, SYT10, AGAP2, TBX3, CCNA1, ZIC2, CLEC14A, OTX2, C14orf39, BNC1, AHSP, ZFHX3, LHX1, TIMP2, ZNF750, SIM2.
[0011] In one or more embodiments, the DNA sequence includes a sense strand or an antisense strand.
[0012] In one or more embodiments, the fragment length is 1-1000bp, preferably 1-700bp.
[0013] In one or more embodiments, the fragment contains at least one CpG dinucleotide.
[0014] In one or more embodiments, the DNA sequence is selected from one or more (e.g., at least seven) or all of the following sequences or their complementary sequences: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:31, SEQ ID NO:32, SEQ ID NO:3 ... SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:46, SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50, SEQ ID NO:51, SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, or variants thereof having at least 70% identity, wherein the methylation sites in the variants are not mutated.
[0015] In one or more embodiments, the reagent is a primer molecule that hybridizes with the DNA sequence or a fragment thereof. The primer molecule amplifies the DNA sequence or a fragment thereof. In one or more embodiments, the primer sequence is methylation-specific or non-specific. The primer molecule is at least 9 bp.
[0016] In one or more embodiments, the reagent is a probe molecule that hybridizes with the DNA sequence or a fragment thereof. In one or more embodiments, the probe further contains a detectable. In one or more embodiments, the detectable is a 5' fluorescent reporter group and a 3' labeled quencher group. In one or more embodiments, the fluorescent reporter gene is selected from Cy5, FAM, and VIC. Preferably, the probe sequence contains MGB (Minor Groove Binder) or LNA (Locked Nucleic Acid). The probe molecule is at least 12 bp.
[0017] In one or more embodiments, the reagent comprises the nucleic acid molecules described in the first aspect of this document.
[0018] In one or more embodiments, the sample is derived from a mammal, preferably a human.
[0019] A third aspect of the present invention provides a medium containing a DNA sequence or a fragment thereof and / or its methylation information, wherein the DNA sequence is (i) selected from one or more (e.g., at least seven) or all of the following gene sequences, or sequences upstream or downstream of the following gene sequences within 20 kb: DMRTA2, FOXD3, TBX15, BCAN, TRIM58, SIX3, VAX2, EMX1, LBX2, TLX2, POU3F3, TBR1, EVX2, HOXD12, HOXD8, HOXD4, TOPAZ1, SHOX2, DRD5, RPL9, HOPX, SFRP2 IRX4, TBX18, OLIG3, ULBP1, HOXA13, TBX20, IKZF1, INSIG1, SOX7, EBF2, MOS, MKX, KCNA6, SYT10, AGAP2, TBX3, CCNA1, ZIC2, CLEC14A, OTX2, C14orf39, BNC1, AHSP, ZFHX3, LHX1, TIMP2, ZNF750, SIM2, or (ii)(i) the processed sequence, wherein the processing converts unmethylated cytosine into bases with a lower binding affinity to guanine than cytosine.
[0020] In one or more embodiments, the DNA sequence is (i) selected from any of the gene sequences shown in the following group, or sequences within 20 kb upstream or downstream of them: (1) LBX2, TBR1, EVX2, SFRP2, SYT10, CCNA1, ZFHX3; (2) TRIM58, HOXD4, INSIG1, SYT10, CCNA1, ZIC2, CLEC14A; (3) EMX1, POU3F3, TOPAZ1, ZIC2, OTX2, AHSP, TIMP2; (4) EMX1, EVX2, RPL9, SFRP2, HOXA13, SYT10, CLEC14A; (5) TBX15, EMX1, LBX2, OLIG3, SYT10, AGAP2, TBX3; 6) TRIM58, VAX2, EMX1, HOXD4, ZIC2, CLEC14A, LHX1; (7) POU3F3, HOXD8, RPL9, TBX18, SYT10, TBX3, CLEC14A; (8) TRIM58, EMX1, TLX2, EVX2, HOXD4, HOXD4, IRX4; (9) SIX3, POU3F3, TOPAZ1, RPL9, SFRP2, CLEC14A, BNC1; (10) DMRTA2, HOXD4, IRX4, INSIG1, MOS, CLEC14A, CLEC14A, or (ii)(i) a treated sequence, said treatment converting unmethylated cytosine into bases with a lower binding affinity to guanine than cytosine.
[0021] In one or more embodiments, the medium is used to compare with gene methylation sequencing data to determine the presence, content, and / or methylation level of nucleic acid molecules containing the sequence or fragment.
[0022] In one or more embodiments, the DNA sequence includes a sense strand or an antisense strand.
[0023] In one or more embodiments, the fragment length is 1-1000bp, preferably 1-700bp.
[0024] In one or more embodiments, the fragment contains at least one CpG dinucleotide.
[0025] In one or more embodiments, the DNA sequence is selected from one or more (e.g., at least seven) or all of the following sequences or their complementary sequences: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:31, SEQ ID NO:32, SEQ ID NO:3 ... SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:46, SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50, SEQ ID NO:51, SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, or variants thereof having at least 70% identity, wherein the methylation sites in the variants are not mutated.
[0026] In one or more embodiments, the methylation information includes information related to cytosine that may be methylated in the sequence of the nucleic acid molecule. Preferably, the cytosine that may be methylated is the C in CpG. In one or more embodiments, the methylation information is the location of a methylation site (such as a CpG dinucleotide) in the nucleic acid molecule.
[0027] In one or more embodiments, the medium is a carrier printed with the DNA sequence or fragments thereof and / or its methylation information, including cards, such as paper, plastic, metal, or glass cards.
[0028] In one or more embodiments, the medium is a computer-readable medium storing the sequence and / or its methylation information and a computer program that, when executed by a processor, performs the following steps: comparing the methylation sequencing data of a sample with the sequence to obtain the presence, abundance, and / or methylation level of nucleic acid molecules containing the sequence in the sample. The presence, abundance, and / or methylation level of nucleic acid molecules containing the sequence are used for the diagnosis of pancreatic cancer.
[0029] In another aspect, the present invention also provides (a) and / or (b) use in the preparation of a kit for diagnosing pancreatic cancer in a target.
[0030] (a) A reagent or apparatus for determining the methylation level of a DNA sequence or fragment thereof in a sample of an object, or the methylation state or level of one or more CpG dinucleotides in the DNA sequence or fragment thereof.
[0031] (b) A treated nucleic acid molecule of the DNA sequence or a fragment thereof, wherein the treatment converts unmethylated cytosine into bases that have a lower binding affinity to guanine than cytosine.
[0032] The DNA sequence is selected from one or more (e.g., at least seven) or all of the following gene sequences, or sequences within 20 kb upstream or downstream of them: DMRTA2, FOXD3, TBX15, BCAN, TRIM58, SIX3, VAX2, EMX1, LBX2, TLX2, POU3F3, TBR1, EVX2, HOXD12, HOXD8, HOXD4, TOPAZ1, SHOX2, DRD5, RPL9, HOP X, SFRP2, IRX4, TBX18, OLIG3, ULBP1, HOXA13, TBX20, IKZF1, INSIG1, SOX7, EBF2, MOS, MKX, KCNA6, SYT1 0. AGAP2, TBX3, CCNA1, ZIC2, CLEC14A, OTX2, C14orf39, BNC1, AHSP, ZFHX3, LHX1, TIMP2, ZNF750, SIM2.
[0033] In one or more embodiments, the DNA sequence comprises a gene sequence selected from any of the following groups: (1) LBX2, TBR1, EVX2, SFRP2, SYT10, CCNA1, ZFHX3; (2) TRIM58, HOXD4, INSIG1, SYT10, CCNA1, ZIC2, CLEC14A; (3) EMX1, POU3F3, TOPAZ1, ZIC2, OTX2, AHSP, TIMP2; (4) EMX1, EVX2, RPL9, SFRP2, HOXA13, SYT10, CLEC14A; (5) TBX15, EMX1, LBX2, OLIG3, SYT 10. AGAP2, TBX3; (6) TRIM58, VAX2, EMX1, HOXD4, ZIC2, CLEC14A, LHX1; (7) POU3F3, HOXD8, RPL9, TBX18, SYT10, TBX3, CLEC14A; (8) TRIM58, EMX1, T LX2, EVX2, HOXD4, HOXD4, IRX4; (9) SIX3, POU3F3, TOPAZ1, RPL9, SFRP2, CLEC14A, BNC1; (10) DMRTA2, HOXD4, IRX4, INSIG1, MOS, CLEC14A, CLEC14A.
[0034] In one or more embodiments, the DNA sequence includes a sense strand or an antisense strand.
[0035] In one or more embodiments, the fragment length is 1-1000bp, preferably 1-700bp.
[0036] In one or more embodiments, the fragment contains at least one CpG dinucleotide.
[0037] In one or more embodiments, the DNA sequence is selected from one or more (e.g., at least seven) or all of the following sequences or their complementary sequences: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:31, SEQ ID NO:32, SEQ ID NO:3 ... SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:46, SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50, SEQ ID NO:51, SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, or variants thereof having at least 70% identity, wherein the methylation sites in the variants are not mutated.
[0038] In one or more embodiments, the DNA sequence comprises a sequence selected from any of the following groups or the complementary sequence thereof: (1) SEQ ID NO: 9, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 26, SEQ ID NO: 40, SEQ ID NO: 43, SEQ ID NO: 52; (2) SEQ ID NO: 5, SEQ ID NO: 18, SEQ ID NO: 34, SEQ ID NO: 40, SEQ ID NO: 43, SEQ ID NO: 45, SEQ ID NO: 46; (3) SEQ ID NO: 8, SEQ ID NO: 11, SEQ ID NO: 20, SEQ ID NO: 44, SEQ ID NO: 48, SEQ ID NO: 51, SEQ ID NO: 54; (4) SEQ ID NO: 8, SEQ ID NO: 14, SEQ ID NO: 24, SEQ ID NO: 26, SEQ ID NO: 31, SEQ ID NO: 40, SEQ ID NO: 46; (5) SEQ ID NO: 3, SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 29, SEQ ID NO: 40, SEQ ID NO: 41, SEQ ID NO: 42; (6) SEQ ID NO: 5, SEQ ID NO: 7, SEQ ID NO: 8, SEQ ID NO: 19, SEQ ID NO: 44, SEQ ID NO: 47, SEQ ID NO: 53; (7) SEQ ID NO: 12, SEQ ID NO: 17, SEQ ID NO: 24, SEQ ID NO: 28, SEQ ID NO: 40, SEQ ID NO: 42, SEQ ID NO: 47; (8) SEQ ID NO: 5, SEQ ID NO: 8, SEQ ID NO: 10, SEQ ID NO: 14, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 27; (9) SEQ ID NO: 6, SEQ ID NO: 12, SEQ ID NO: 20, SEQ ID NO: 24, SEQ ID NO: 26, SEQ ID NO: 47, SEQ ID NO: 50; (10) SEQ ID NO: 1, SEQ ID NO: 19, SEQ ID NO: 27, SEQ ID NO: 34, SEQ ID NO: 37, SEQ ID NO: 46, SEQ ID NO: 47.
[0039] In one or more embodiments, the nucleic acid molecule is the nucleic acid molecule described in the first aspect of this document.
[0040] In one or more embodiments, the reagent comprises primer molecules and / or probe molecules.
[0041] In one or more embodiments, the reagent comprises primer molecules that hybridize with the DNA sequence or a fragment thereof. The primer molecules amplify the DNA sequence or a fragment thereof. In one or more embodiments, the primer sequence is methylation-specific or non-specific. The primer molecule is at least 9 bp.
[0042] In one or more embodiments, the reagent is a probe molecule that hybridizes with the DNA sequence or a fragment thereof. In one or more embodiments, the probe further contains a detectable. In one or more embodiments, the detectable is a 5' fluorescent reporter group and a 3' labeled quencher group. In one or more embodiments, the fluorescent reporter gene is selected from Cy5, FAM, and VIC. Preferably, the probe sequence contains MGB (Minor Groove Binder) or LNA (Locked Nucleic Acid). The probe molecule is at least 12 bp.
[0043] In one or more embodiments, the reagent comprises the medium described in any of the embodiments herein.
[0044] In one or more embodiments, the kit is a non-invasive diagnostic kit.
[0045] In one or more embodiments, the object is a mammal, preferably a human.
[0046] In one or more embodiments, the sample is derived from mammalian tissue, cells, or bodily fluids, such as pancreatic tissue or blood. In one or more embodiments, the sample is pancreatic cancer tissue, preferably a fine-needle aspiration biopsy. In one or more embodiments, the sample is plasma.
[0047] In one or more embodiments, the sample comprises genomic DNA or cfDNA.
[0048] In one or more embodiments, the DNA sequence is transformed, wherein unmethylated cytosine is converted into bases with a lower binding affinity to guanine than cytosine. The transformation is performed using an enzymatic method, preferably deaminase treatment, or the transformation is performed using a non-enzymatic method, preferably treatment with bisulfite, acid sulfite, or metabisulfite, or a combination thereof.
[0049] In one or more embodiments, the DNA sequence is treated with a methylation-sensitive restriction endonuclease.
[0050] In one or more embodiments, the kit further includes PCR reaction reagents. Preferably, the PCR reaction reagents include DNA polymerase, PCR buffer, dNTPs, and Mg2+.
[0051] In one or more embodiments, the kit further includes other reagents for detecting DNA methylation, said other reagents being selected from one or more of the following methods: bisulfite-based PCR (e.g., methylation-specific PCR), DNA sequencing (e.g., bisulfite sequencing, whole-genome methylation sequencing, simplified methylation sequencing), methylation-sensitive restriction endonuclease assays, quantitative fluorescence assays, methylation-sensitive high-resolution melting curve assays, chip-based methylation mapping, and mass spectrometry (e.g., mass spectrometry of flight). Preferably, said other reagents are selected from one or more of the following: bisulfite, bisulfite, acid sulfite, or metabisulfite or derivatives thereof, methylation-sensitive or insensitive restriction endonucleases, enzyme digestion buffers, fluorescent dyes, fluorescence quenchers, fluorescent reporters, exonucleases, alkaline phosphatases, internal standards, and controls.
[0052] In one or more embodiments, the PCR reaction solution comprises Taq DNA polymerase, PCR buffer, dNTPs, KCl, MgCl2, and (NH4)2SO4. Preferably, the Taq DNA polymerase is a hot-start Taq DNA polymerase. Preferably, MgCl2... 2+ The final concentration is 1.0-10.0 mM.
[0053] In one or more embodiments, the diagnosis includes: comparing with a control sample or calculating a score, and diagnosing pancreatic cancer based on the score. In one or more embodiments, the calculation is performed by constructing a support vector machine model.
[0054] Another aspect of the present invention provides a method for pancreatic cancer screening, comprising:
[0055] (1) The methylation level of the DNA sequence or fragment thereof in the sample of the test subject, or the methylation status or level of one or more CpG dinucleotides in the DNA sequence or fragment thereof, wherein the DNA sequence is selected from one or more of the following gene sequences: DMRTA2, FOXD3, TBX15, BCAN, TRIM58, SIX3, VAX2, EMX1, LBX2, TLX2, POU3F3, TBR1, EVX2, HOXD12, HOXD8, HOXD4, TOPAZ1, SH OX2, DRD5, RPL9, HOPX, SFRP2, IRX4, TBX18, OLIG3, ULBP1, HOXA13, TBX20, IKZF1, INSIG1, SOX7, EBF2, MOS, MKX, K CNA6, SYT10, AGAP2, TBX3, CCNA1, ZIC2, CLEC14A, OTX2, C14orf39, BNC1, AHSP, ZFHX3, LHX1, TIMP2, ZNF750, SIM2,
[0056] (2) Compare with the control sample, or calculate the score.
[0057] (3) Diagnose pancreatic cancer based on the scoring.
[0058] In one or more embodiments, the DNA sequence includes a sense strand or an antisense strand.
[0059] In one or more embodiments, the fragment length is 1-1000bp, preferably 1-700bp.
[0060] In one or more embodiments, the fragment contains at least one CpG dinucleotide.
[0061] In one or more embodiments, the method further includes DNA extraction and / or quality control prior to step (1).
[0062] In one or more embodiments, the DNA sequence is selected from one or more of the following sequences or their complementary sequences: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:31, SEQ ID NO:32, SEQ ID NO:3 ... SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:46, SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50, SEQ ID NO:51, SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, or variants thereof having at least 70% identity, wherein the methylation sites in the variants are not mutated.
[0063] In one or more embodiments, step (1) includes performing the detection using the nucleic acid molecules, primer molecules, probe molecules and / or media described herein.
[0064] In one or more embodiments, the detection includes, but is not limited to: bisulfite-based PCR (e.g., methylation-specific PCR), DNA sequencing (e.g., bisulfite sequencing, whole-genome methylation sequencing, simplified methylation sequencing), methylation-sensitive restriction endonuclease assays, quantitative fluorescence assays, methylation-sensitive high-resolution melting curve assays, chip-based methylation mapping analysis, and mass spectrometry (e.g., mass spectrometry of flight).
[0065] In one or more embodiments, the detection is DNA sequencing. In one or more embodiments, the sequencing depth of the DNA sequencing is greater than or equal to 5M, preferably 7M, 11M, 13M, or 15M.
[0066] In one or more embodiments, the sample is derived from the tissue, cells, or body fluids of a mammal, such as pancreatic tissue or blood. The mammal is preferably a human. In one or more embodiments, the sample is pancreatic tumor tissue, preferably a fine-needle aspiration biopsy. In one or more embodiments, the sample is plasma.
[0067] In one or more embodiments, the sample comprises genomic DNA or cfDNA.
[0068] In one or more embodiments, the DNA sequence is transformed, wherein unmethylated cytosine is converted into bases that do not bind to guanine. The transformation is performed using an enzymatic method, preferably deaminase treatment, or the transformation is performed using a non-enzymatic method, preferably treatment with bisulfite, acid sulfite, or metabisulfite, or a combination thereof.
[0069] In one or more embodiments, the DNA sequence is treated with a methylation-sensitive restriction endonuclease.
[0070] In one or more implementations, the score in step (2) is calculated by constructing a support vector machine model.
[0071] In one or more embodiments, step (3) includes: comparing the methylation level of the subject sample with that of a control sample, and identifying the subject as having pancreatic cancer when the methylation level meets a threshold.
[0072] In one or more implementations, step (3) includes identifying the subject as having pancreatic cancer when the score meets a threshold.
[0073] Another aspect of the present invention provides a kit for identifying pancreatic cancer, comprising:
[0074] (a) A reagent or apparatus for determining the methylation level of a DNA sequence or fragment thereof in a sample of an object, or the methylation state or level of one or more CpG dinucleotides in said DNA sequence or fragment thereof, and
[0075] Option (b) is a treated nucleic acid molecule of the DNA sequence or a fragment thereof, wherein the treatment converts unmethylated cytosine into bases that have a lower binding affinity to guanine than cytosine.
[0076] The DNA sequence is selected from one or more (e.g., at least seven) or all of the following gene sequences, or sequences within 20 kb upstream or downstream of them: DMRTA2, FOXD3, TBX15, BCAN, TRIM58, SIX3, VAX2, EMX1, LBX2, TLX2, POU3F3, TBR1, EVX2, HOXD12, HOXD8, HOXD4, TOPAZ1, SHOX2, DRD5, RPL9, HOP X, SFRP2, IRX4, TBX18, OLIG3, ULBP1, HOXA13, TBX20, IKZF1, INSIG1, SOX7, EBF2, MOS, MKX, KCNA6, SYT1 0. AGAP2, TBX3, CCNA1, ZIC2, CLEC14A, OTX2, C14orf39, BNC1, AHSP, ZFHX3, LHX1, TIMP2, ZNF750, SIM2.
[0077] In one or more embodiments, the DNA sequence is selected from one or more (e.g., at least seven) or all of the following sequences or their complementary sequences: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:31, SEQ ID NO:32, SEQ ID NO:3 ... SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:46, SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50, SEQ ID NO:51, SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, or variants thereof having at least 70% identity, wherein the methylation sites in the variants are not mutated.
[0078] In one or more embodiments, the kit is suitable for the use described in any of the embodiments herein.
[0079] In one or more embodiments, the nucleic acid molecule is the nucleic acid molecule described in the first aspect of this document.
[0080] In one or more embodiments, the reagent comprises primer molecules and / or probe molecules.
[0081] In one or more embodiments, the reagent comprises primer molecules that hybridize with the DNA sequence or a fragment thereof. The primer molecules amplify the DNA sequence or a fragment thereof. In one or more embodiments, the primer sequence is methylation-specific or non-specific. The primer molecule is at least 9 bp.
[0082] In one or more embodiments, the reagent is a probe molecule that hybridizes with the DNA sequence or a fragment thereof. In one or more embodiments, the probe further contains a detectable. In one or more embodiments, the detectable is a 5' fluorescent reporter group and a 3' labeled quencher group. In one or more embodiments, the fluorescent reporter gene is selected from Cy5, FAM, and VIC. Preferably, the probe sequence contains MGB (Minor Groove Binder) or LNA (Locked Nucleic Acid). The probe molecule is at least 12 bp.
[0083] In one or more embodiments, the reagent comprises the medium described in any of the embodiments herein.
[0084] In one or more embodiments, the kit is a non-invasive diagnostic kit.
[0085] In one or more embodiments, the object is a mammal, preferably a human.
[0086] In one or more embodiments, the sample is derived from mammalian tissue, cells, or bodily fluids, such as pancreatic tissue or blood. In one or more embodiments, the sample is pancreatic cancer tissue, preferably a fine-needle aspiration biopsy. In one or more embodiments, the sample is plasma.
[0087] In one or more embodiments, the sample comprises genomic DNA or cfDNA.
[0088] In one or more embodiments, the DNA sequence is transformed, wherein unmethylated cytosine is converted into bases with a lower binding affinity to guanine than cytosine. The transformation is performed using an enzymatic method, preferably deaminase treatment, or the transformation is performed using a non-enzymatic method, preferably treatment with bisulfite, acid sulfite, or metabisulfite, or a combination thereof.
[0089] In one or more embodiments, the DNA sequence is treated with a methylation-sensitive restriction endonuclease.
[0090] In one or more embodiments, the kit further includes PCR reaction reagents. Preferably, the PCR reaction reagents include DNA polymerase, PCR buffer, dNTPs, and Mg2+. 2+ .
[0091] In one or more embodiments, the kit further includes reagents for detecting DNA methylation, said reagents being selected from one or more of the following methods: bisulfite-based PCR (e.g., methylation-specific PCR), DNA sequencing (e.g., bisulfite sequencing, whole-genome methylation sequencing, simplified methylation sequencing), methylation-sensitive restriction endonuclease assays, quantitative fluorescence assays, methylation-sensitive high-resolution melting curve assays, chip-based methylation mapping, and mass spectrometry (e.g., mass spectrometry of flight). Preferably, the reagents are selected from one or more of the following: bisulfite and its derivatives, methylation-sensitive or insensitive restriction endonucleases, enzyme digestion buffers, fluorescent dyes, fluorescence quenchers, fluorescent reporter agents, exonucleases, alkaline phosphatases, internal standards, and controls.
[0092] In another aspect, the present invention provides an apparatus for diagnosing pancreatic cancer, the apparatus comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor, when executing the program, performs the following steps:
[0093] (1) Obtain the methylation level of the DNA sequence or fragment thereof in the sample of the subject, or the methylation status or level of one or more CpGs in the DNA sequence or fragment thereof, wherein the DNA sequence is selected from one or more of the following gene sequences: DMRTA2, FOXD3, TBX15, BCAN, TRIM58, SIX3, VAX2, EMX1, LBX2, TLX2, POU3F3, TBR1, EVX2, HOXD12, HOXD8, HOXD4, TOPAZ1, SHOX2 , DRD5, RPL9, HOPX, SFRP2, IRX4, TBX18, OLIG3, ULBP1, HOXA13, TBX20, IKZF1, INSIG1, SOX7, EBF2, MOS, MKX, KCN A6, SYT10, AGAP2, TBX3, CCNA1, ZIC2, CLEC14A, OTX2, C14orf39, BNC1, AHSP, ZFHX3, LHX1, TIMP2, ZNF750, SIM2,
[0094] (2) Compare with the control sample, or calculate the score, and
[0095] (3) Diagnose pancreatic cancer based on the scoring.
[0096] In one or more embodiments, step (1) further includes a step of obtaining DNA, such as DNA extraction and / or quality control.
[0097] In one or more embodiments, the DNA sequence is selected from one or more of the following sequences or their complementary sequences: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, SEQ ID NO:4, SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:30, SEQ ID NO:31, SEQ ID NO:32, SEQ ID NO:3 ... SEQ ID NO:34, SEQ ID NO:35, SEQ ID NO:36, SEQ ID NO:37, SEQ ID NO:38, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, SEQ ID NO:43, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:46, SEQ ID NO:47, SEQ ID NO:48, SEQ ID NO:49, SEQ ID NO:50, SEQ ID NO:51, SEQ ID NO:52, SEQ ID NO:53, SEQ ID NO:54, SEQ ID NO:55, SEQ ID NO:56, or variants thereof having at least 70% identity, wherein the methylation sites in the variants are not mutated.
[0098] In one or more embodiments, step (1) includes detecting the methylation level of the sequence in the sample using the nucleic acid molecules, primer molecules, probe molecules, and / or media described herein. In one or more embodiments, the detection includes, but is not limited to: bisulfite-based PCR (e.g., methylation-specific PCR), DNA sequencing (e.g., bisulfite sequencing, whole-genome methylation sequencing, simplified methylation sequencing), methylation-sensitive restriction endonuclease assays, quantitative fluorescence assays, methylation-sensitive high-resolution melting curve assays, chip-based methylation mapping, and mass spectrometry (e.g., mass spectrometry of flight). In one or more embodiments, the detection is DNA sequencing. Preferably, the sequencing depth of the DNA sequencing is greater than or equal to 5M, more preferably 7M, 11M, 13M, or 15M.
[0099] In one or more embodiments, the sample is derived from the tissue, cells, or body fluids of a mammal, such as pancreatic tissue or blood. The mammal is preferably a human. In one or more embodiments, the sample is pancreatic tumor tissue, preferably a fine-needle aspiration biopsy. In one or more embodiments, the sample is plasma.
[0100] In one or more embodiments, the sample comprises genomic DNA or cfDNA.
[0101] In one or more embodiments, the sequence is transformed, wherein unmethylated cytosine is converted into a base that does not bind to guanine. The transformation is performed using an enzymatic method, preferably deaminase treatment, or the transformation is performed using a non-enzymatic method, preferably treatment with bisulfite, acid sulfite, or metabisulfite, or a combination thereof.
[0102] In one or more embodiments, the DNA sequence is treated with a methylation-sensitive restriction endonuclease.
[0103] In one or more implementations, the score in step (2) is calculated by constructing a support vector machine model.
[0104] In one or more embodiments, step (3) includes: comparing the methylation level of the subject sample with that of a control sample, and identifying the subject as having pancreatic cancer when the methylation level meets a threshold.
[0105] In one or more implementations, step (3) includes identifying the subject as having pancreatic cancer when the score meets a threshold. Attached Figure Description
[0106] Figure 1 This is a flowchart of the technical solution of one embodiment of the present invention.
[0107] Figure 2This is the ROC curve of the pancreatic cancer prediction model Model CN in the test group for diagnosing pancreatic cancer.
[0108] Figure 3 This is the distribution of prediction scores for the pancreatic cancer prediction model Model CN across different groups.
[0109] Figure 4 It represents the methylation levels of 56 sequences SEQ ID NO:1-56 in the training group.
[0110] Figure 5 It represents the methylation levels of 56 sequences SEQ ID NO:1-56 in the test group.
[0111] Figure 6 The classification ROC curves are obtained by using CA19-9 alone, using Example 2 alone to build the SVM model Model CN, and using the model built in Example 2 combined with CA19-9.
[0112] Figure 7 The distribution of classification prediction scores is obtained by using CA19-9 alone, using the SVM model Model CN built alone using Example 2, and combining the model built in Example 2 with CA19-9.
[0113] Figure 8 ROC curve of the SVM model Model CN constructed in Example 2 in samples that are negative for the tumor marker CA19-9 (CA19-9 measurement value is less than 37).
[0114] Figure 9 ROC curves of the combined model of seven biomarkers SEQ ID NO:9,14,13,26,40,43,52
[0115] Figure 10 ROC curves of the combined model of seven markers SEQ ID NO:5,18,34,40,43,45,46
[0116] Figure 11 ROC curves of the combined model of seven biomarkers SEQ ID NO:11,8,20,44,48,51,54
[0117] Figure 12 ROC curves of the combined model of seven markers SEQ ID NO:14,8,26,24,31,40,46
[0118] Figure 13 ROC curves of the combined model of seven biomarkers SEQ ID NO:3,9,8,29,42,40,41
[0119] Figure 14ROC curves of the combined model of seven biomarkers SEQ ID NO:5,8,19,7,44,47,53
[0120] Figure 15 ROC curves of the combined model of seven biomarkers SEQ ID NO:12,17,24,28,40,42,47
[0121] Figure 16 ROC curves of the combined model of seven biomarkers SEQ ID NO:5,18,14,10,8,19,27
[0122] Figure 17 ROC curves of the combined model of seven biomarkers SEQ ID NO:6,12,20,26,24,47,50
[0123] Figure 18 ROC curves of the combined model of seven biomarkers SEQ ID NO:1,19,27,34,37,46,47 Detailed Implementation
[0124] This invention explores the relationship between methylation and pancreatic cancer. The aim is to improve the accuracy of non-invasive diagnosis of pancreatic cancer by utilizing the methylation levels of relevant genes as differential markers for pancreatic cancer through a non-invasive method.
[0125] The inventors discovered that the nature of pancreatic cancer is related to the methylation level of sequences within 20 kb upstream or downstream of the following genes: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, and 50: DMRTA2, FOXD3, TBX15, BCAN, TRIM58, SIX3, VAX2, EMX1, and LBX2. TLX2, POU3F3, TBR1, EVX2, HOXD12, HOXD8, HOXD4, TOPAZ1, SHOX2, DRD5, RPL9, HOPX, SFRP2, IRX4, TBX18, OLIG3, ULBP1, HOXA13, TBX20, IKZF1, I NSIG1, SOX7, EBF2, MOS, MKX, KCNA6, SYT10, AGAP2, TBX3, CCNA1, ZIC2, CLEC14A, OTX2, C14orf39, BNC1, AHSP, ZFHX3, LHX1, TIMP2, ZNF750, SIM2. In one or more embodiments, the nature of pancreatic cancer is associated with the methylation level of genes selected from any of the following groups: (1) LBX2, TBR1, EVX2, SFRP2, SYT10, CCNA1, ZFHX3; (2) TRIM58, HOXD4, INSIG1, SYT10, CCNA1, ZIC2, CLEC14A; (3) EMX1, POU3F3, TOPAZ1, ZIC2, OTX2, AHSP, TIMP2; (4) EMX1, EVX2, RPL9, SFRP2, HOXA13, SYT10, CLEC14A; (5) TBX15, EMX1, LBX2, OLIG3, S YT10, AGAP2, TBX3; (6) TRIM58, VAX2, EMX1, HOXD4, ZIC2, CLEC14A, LHX1; (7) POU3F3, HOXD8, RPL9, TBX18, SYT10, TBX3, CLEC14A; (8) TRIM58, EMX1, TLX2, EVX2, HOXD4, HOXD4, IRX4; (9) SIX3, POU3F3, TOPAZ1, RPL9, SFRP2, CLEC14A, BNC1; (10) DMRTA2, HOXD4, IRX4, INSIG1, MOS, CLEC14A, CLEC14A. This invention provides nucleic acid molecules containing one or more CpGs of the above-mentioned genes or fragments thereof.
[0126] In this article, the term "gene" includes both coding and non-coding sequences of the gene in question on the genome. Non-coding sequences include introns, promoters, and regulatory elements or sequences.
[0127] Furthermore, the nature of pancreatic cancer is associated with the methylation levels of any one of the following segments or random segments 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55 or all 56 segments: SEQ ID NO:1 of the DMRTA2 gene region, SEQ ID NO:2 of the FOXD3 gene region, SEQ ID NO:3 of the TBX15 gene region, SEQ ID NO:4 of the BCAN gene region, and SEQ ID NO:4 of the TRIM58 gene region. SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:15, SEQ ID NO:16, SEQ ID NO:17, SEQ ID NO:18, SEQ ID NO:19, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:28, SEQ ID NO:29, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO:2 ...29, SEQ ID NO:20, SEQ ID NO:20, SEQ ID NO:21, SEQ ID NO:22, SEQ ID NO:23, SEQ ID NO:24, SEQ ID NO:25, SEQ ID NO:26, SEQ ID NO:27, SEQ ID NO SEQ ID NO:27, TBX18 gene region; SEQ ID NO:28, OLIG3 gene region; SEQ ID NO:29, ULBP1 gene region; SEQ ID NO:30, HOXA13 gene region; SEQ ID NO:31, TBX20 gene region; SEQ ID NO:32, IKZF1 gene region; SEQ ID NO:33, INSIG1 gene region; SEQ ID NO:34, SOX7 gene region; SEQ ID NO:35, EBF2 gene region; SEQ ID NO:36, MOS gene region; SEQ ID NO:37, MKX gene region.SEQ ID NO:38, KCNA6 gene region SEQ ID NO:39, SYT10 gene region SEQ ID NO:40, AGAP2 gene region SEQ ID NO:41, TBX3 gene region SEQ ID NO:42, CCNA1 gene region SEQ ID NO:43, ZIC2 gene region SEQ ID NO:44, SEQ ID NO:45, CLEC14A gene region SEQ ID NO:46, SEQ ID NO:47, OTX2 gene region SEQ ID NO:48, C14orf39 gene region SEQ ID NO:49, BNC1 gene region SEQ ID NO:50, AHSP gene region SEQ ID NO:51, ZFHX3 gene region SEQ ID NO:52, LHX1 gene region SEQ ID NO:53, TIMP2 gene region SEQ ID NO:54, ZNF750 gene region SEQ ID NO:55, SIM2 gene region SEQ ID NO:56.
[0128] In some implementations, the nature of pancreatic cancer is associated with the methylation level of sequences selected from any of the following groups or their complementary sequences: (1) SEQ ID NO:9, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:26, SEQ ID NO:40, SEQ ID NO:43, SEQ ID NO:52, (2) SEQ ID NO:5, SEQ ID NO:18, SEQ ID NO:34, SEQ ID NO:40, SEQ ID NO:43, SEQ ID NO:45, SEQ ID NO:46, (3) SEQ ID NO:8, SEQ ID NO:11, SEQ ID NO:20, SEQ ID NO:44, SEQ ID NO:48, SEQ ID NO:51, SEQ ID NO:54, (4) SEQ ID NO:8, SEQ ID NO:14, SEQ ID NO:24, SEQ ID NO:26, SEQ ID NO:31, SEQ ID NO:40, SEQ ID NO:46, (5) SEQ ID NO:3, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:13, SEQ ID NO:1 ...6) SEQ ID NO:3, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:26, SEQ ID NO:31, SEQ ID NO:40, SEQ ID NO:46, (7) SEQ ID NO:3, SEQ ID NO:8, SEQ ID NO:46, SEQ ID NO NO:9, SEQ ID NO:29, SEQ ID NO:40, SEQ ID NO:41, SEQ ID NO:42, (6) SEQ ID NO:5, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:19, SEQ ID NO:44, SEQ ID NO:47, SEQ ID NO:53, (7) SEQ ID NO:12, SEQ ID NO:17, SEQ ID NO: 24, SEQ ID NO: 28, SEQ ID NO: 40, SEQ ID NO: 42, SEQ ID NO: 47, (8) SEQ ID NO: 5, SEQ ID NO: 8, SEQ ID NO: 10, SEQ ID NO: 14, SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 27, (9) SEQ ID NO: 6, SEQ ID NO: 12, SEQ ID NO: 20, SEQ ID NO: 24, SEQ ID NO: 26, SEQ ID NO: 47, SEQ ID NO: 50, (10) SEQ ID NO: 1, SEQ ID NO: 19, SEQ ID NO: 27, SEQ ID NO: 34, SEQ ID NO: 37, SEQ ID NO: 46, SEQ ID NO: 47.
[0129] The “pancreatic cancer-related sequences” mentioned in this article include the above 50 genes, sequences within 20kb upstream or downstream of them, the above 56 sequences (SEQ ID NO:1-56) or their complementary sequences.
[0130] The locations of the above 56 sequences in human chromosomes are as follows: SEQ ID NO:1: chr1 50884507-50885207bps, SEQ ID NO:2: chr1 63788611-63789152bps, SEQ ID NO:3: chr1 119522143-119522719bps, SEQ ID NO:4: chr1 156611710-156612211bps, SEQ ID NO:5: chr1 248020391-248020979bps, SEQ ID NO:6: chr2 45028796-45029378bps, SEQ ID NO:7: chr2 71115731-71116272bps, SEQ ID NO: ...5028796-45029378bps, SEQ ID NO:75028796-71116272bps, SEQ ID NO:45028796-45029378bps, SEQ ID NO:75028796-71116272bps, SEQ ID NO:45028796-45029378bps, SEQ ID NO:75028796-71116272bps, SEQ ID NO:45028796-45029378bps, SEQ ID NO:75 NO:8:73147334-73147835bps of chr2,SEQ ID NO:9:74726401-74726922bps of chr2,SEQ ID NO:10:74742861-74743362bps of chr2,SEQ ID NO:11:105480130-105480830bps of chr2,SEQ ID NO:12:105480157-105480659bps of chr2,SEQ ID NO:13:162280233-162280736bps of chr2,SEQ ID NO:14:176945095-176945601bps of chr2,SEQ ID SEQ ID NO:15: chr2 176945320-176945821bps, SEQ ID NO:16: chr2 176964629-176965209bps, SEQ ID NO:17: chr2 176994514-176995015bps, SEQ ID NO:18: chr2 177016987-177017501bps, SEQ ID NO:19: chr2 177024355-177024866bps, SEQ ID NO:20: chr3 44063336-44063893bps, SEQ ID NO:21: chr3 157812057-157812604bps, SEQ ID NO:22: chr4 9783025-9783527bps, SEQ ID NO:23: chr4 39448278-39448779bps, SEQ ID NO:24: chr4 39448327-39448879bpsSEQ ID NO:25: chr4 57521127-57521736bps, SEQ ID NO:26: chr4 154709362-154709867bps, SEQ ID NO:27: chr5 1876136-1876645bps, SEQ ID NO:28: chr6 85476916-85477417bps, SEQ ID NO:29: chr6 137814499-137815053bps, SEQ ID NO:30: chr6 150285594-150286095bps, SEQ ID NO:31: chr7 27244522-27245037bps, SEQ ID SEQ ID NO:32: chr7 35293435-35293950bps, SEQ ID NO:33: chr7 50343543-50344243bps, SEQ ID NO:34: chr7 155167312-155167828bps, SEQ ID NO:35: chr8 10588692-10589253bps, SEQ ID NO:36: chr8 25907648-25908150bps, SEQ ID NO:37: chr8 57069450-57070150bps, SEQ ID NO:38: chr10 28034404-28034908bps, SEQ ID SEQ ID NO:39: chr12 4918941-4919489bps, SEQ ID NO:40: chr12 33592612-33593117bps, SEQ ID NO:41: chr12 58131095-58131654bps, SEQ ID NO:42: chr12 115124763-115125348bps, SEQ ID NO:43: chr13 37005444-37005945bps, SEQ ID NO:44: chr13 100649468-100649995bps, SEQ ID NO:45: chr13 100649513-100650027bps, SEQ ID NO:46: 38724419-38724935bps of chr14, SEQ ID NO: 47: 38724602-38725108bps of chr14, SEQ ID NO: 48: 57275646-57276162bps of chr14, SEQ ID NO:49: chr14 60952384-60952933bps,SEQ ID NO:50:chr15, 83952059-83952595bps; SEQ ID NO:51:chr16, 31579970-31580561bps; SEQ ID NO:52:chr16, 73096773-73097473bps; SEQ ID NO:53:chr17, 35299694-35300224bps; SEQ ID NO:54:chr17, 76929623-76930176bps; SEQ ID NO:55:chr17, 80846617-80847210bps; SEQ ID NO:56:chr21, 38081247-38081752bps. In this paper, the base numbers of each sequence and methylation site correspond to the reference genome HG19. ,
[0131] In one or more embodiments, the nucleic acid molecules described herein are selected from DMRTA2, FOXD3, TBX15, BCAN, TRIM58, SIX3, VAX2, EMX1, LBX2, TLX2, POU3F3, TBR1, EVX2, HOXD12, HOXD8, HOXD4, TOPAZ1, SHOX2, DRD5, RPL9, HOPX, SFRP2, IRX4, TBX18, OLIG3, ULBP1, HOXA13, TBX20, and IKZF1. A fragment of one or more genes selected from INSIG1, SOX7, EBF2, MOS, MKX, KCNA6, SYT10, AGAP2, TBX3, CCNA1, ZIC2, CLEC14A, OTX2, C14orf39, BNC1, AHSP, ZFHX3, LHX1, TIMP2, ZNF750, and SIM2; the fragment is 1 bp to 1 kb in length, preferably 1 bp to 700 bp; the fragment contains one or more methylation sites in the chromosomal region of the corresponding gene. The methylation sites in the genes or their fragments mentioned in this article include, but are not limited to: 50884514, 50884531, 50884533, 50884541, 50884544, 50884547, 50884550, 50884552, 50884566, 50884582, 50884586, 50884589, 50884591, and 50884 on the chr1 chromosome. 598, 50884606, 50884610, 50884612, 50884615, 50884621, 50884633, 50884646, 50884649, 50884658, 50884662, 50884673, 50884682, 50884691, 50884699, 50884702, 50884724, 508847 32,50884735,50884742,50884751,50884754,50884774,50884777,50884780,50884783,50884786,50884789,50884792,50884795,50884798,50884801,50884804,50884807,5088480 9,50884820,50884822,50884825,50884849,50884852,50884868,50884871,50884885,50884889,50884902,50884924,50884939,50884942,50884945,50884948,50884975,50884980,50884983,50884999,50885001,63788628,63788660,63788672,63788685,63788689,63788703,63788706,63788709,63788721,63788741,63788744,63788747,63788753,63788759,63788768,63788776,63788785,63788789,63788795,63788804,63788816,63788822,63788825,63788828,63788849,63788852,63788861,63788870,63788872,63788878,63788881,63788889,63788897,63788902,63788906,63788917,63788920,63788933,63788947,63788983,63788987,63788993,63788999,63789004,63789011,63789014,63789020,63789022,63789025,63789031,63789035,63789047,63789056,63789059,63789068,63789071,63789073,63789077,63789080,63789083,63789092,63789094,63789101,63789106,63789109,63789124,119522172,119522188,119522190,119522233,119522239,119522313,119522368,119522386,119522393,119522409,119522425,119522427,119522436,119522440,119522444,119522446,119522449,119522451,119522456,119522459,119522464,119522469,119522474,119522486,119522488,119522500,119522502,119522516,119522529,119522537,119522548,119522550,119522559,119522563,119522566,119522571,119522577,119522579,119522582,119522594,119522599,119522607,119522615,119522621,119522629,119522631,119522637,119522665,119522673,156611713,156611720,156611733,156611737,156611749,156611752,156611761,156611767,156611784,156611791,156611797,156611802,156611811,156611813,156611819,156611830,156611836,156611842,156611851,156611862,156611890,156611893,156611902,156611905,156611915,156611926,156611945,156611949,156611951,156611960,156611963,156611994,156612002,156612015,156612024,156612034,156612042,156612044,156612079,156612087,156612090,156612094,156612097,156612105,156612140,156612147,156612166,156612188,156612191,156612204,156612209,248020399,248020410,248020436,248020447,248020450,248020453,248020470,248020495,248020497,248020507,248020512,248020516,248020520,248020526,248020536,248020543,248020559,248020562,248020566,248020573,248020579,248020581,248020589,248020591,248020598,248020625,248020632,248020641,248020671,248020680,248020688,248020692,248020695,248020697,248020704,248020707,248020713,248020721,248020729, 248020741, 248020748, 248020756, 248020765, 248020775, 248020791, 248020795, 248020798, 248020812, 248020814, 248020821, 2480 20826,248020828,248020831,248020836,248020838,248020840,248020845,248020848,248020861,248020869,248020878,248020883,248020886 ,248020902,248020905,248020908,248020914,248020925,248020930,2 48020934,248020937,248020940,248020953,248020956,248020975; chr2 Chromosomes 45028802, 45028816, 45028832, 45028839, 45028956, 45028961, 45028965, 45028973, 45029004, 45029017, 45029035, 45029046, 45029057, 4502 9060,45029063,45029065,45029071,45029106,45029112,45029117,45029128,45029146,45029176,45029179,45029184,45029189,45029192,450 29195,45029218,45029226,45029228,45029231,45029235,45029263,45029273,45029285,45029288,45029295,45029307,45029317,45029353,45 029357,71115760,71115787,71115789,71115837,71115928,71115936,71115948,71115962,71115968,71115978,71115981,71115983,71115985,7 1115987,71115994,71116000,71116022,71116024,71116030,71116036,71116047,71116054,71116067,71116096,71116101,71116103,71116107,71116117,71116119,71116130,71116137,71116141,71116152,71116154,71116158,71116174,71116188,71116190,71116194,71116203,71116215,71116226,71116233,71116242,71116257,71116259,71116261,71116268,71116271,73147340,73147350,73147364,73147369,73147382,73147405,73147408,73147432,73147438,73147444,73147481,73147491,73147493,73147523,73147529,73147537,73147559,73147571,73147582,73147584,73147592,73147595,73147598,73147607,73147613,73147620,73147623,73147631,73147644,73147668,73147673,73147678,73147687,73147690,73147693,73147695,73147710,73147720,73147738,73147755,73147767,73147771,73147789,73147798,73147803,73147811,73147814,73147816,73147822,73147825,73147827,73147829,74726438,74726440,74726449,74726478,74726480,74726482,74726484,74726493,74726495,74726524,74726526,74726533,74726536,74726539,74726548,74726554,74726569,74726572,74726585,74726597,74726599,74726616,74726633,74726642,74726649,74726651,74726656,74726668,74726672,74726682,74726687,74726695,74726700,74726710,74726716,74726734,74726746,74726760,74726766,74726772,74726784,74726791,74726809,74726828,74726833,74726835,74726861,74726892,74726894,74726908,74742879,74742882,74742891,74742913,74742922,74742925,74742942,74742950,74742953,74742967,74742981,74742984,74742996,74743004,74743006,74743009,74743011,74743015,74743021,74743035,74743056,74743059,74743061,74743064,74743068,74743073,74743082,74743084,74743101,74743108,74743111,74743119,74743121,74743127,74743131,74743137,74743139,74743141,74743146,74743172,74743174,74743182,74743186,74743191,74743195,74743198,74743207,74743231,74743234,74743241,74743243,74743268,74743295,74743301,74743306,74743318,74743321,74743325,74743329,74743333,74743336,74743343,74743346,74743352,74743357,105480130,105480161,105480179,105480198,105480207,105480210,105480212,105480226,105480254,105480258,105480272,105480291,105480337,105480360,105480377,105480383,105480387,105480390,105480407,105480409,105480412,105480424,105480426,105480429,105480433,105480438,105480461,105480464,105480475,105480481,105480488,105480490,105480503,105480546,105480556,105480571,105480577,105480581,105480604,105480621,105480623,105480630,105480634,105480637,162280237,162280239,162280242,162280245,162280249,162280257,162280263,162280289,162280293,162280297,162280306,162280309,162280314,162280317,162280327,162280331,162280341,162280351,162280362,162280368,162280393,162280396,162280398,162280402,162280405,162280407,162280409,162280417,162280420,162280438,162280447,162280459,162280462,162280466,162280470,162280473,162280479,162280483,162280486,162280489,162280492,162280498,162280519,162280534,162280539,162280548,162280561,162280570,162280575,162280585,162280598,162280604,162280611,162280614,162280618,162280623,162280627,162280633,162280641,162280647,162280657,162280673,162280681,162280693,162280708,162280728,176945102,176945119,176945122,176945132,176945134,176945137,176945141,176945144,176945147,176945150,176945159,176945165,176945170,176945177,176945179,176945186,176945188,176945198,176945200,176945213,176945215,176945218,176945222,176945224,176945250,176945270,176945274,176945288,176945296,176945298,176945316,176945329,176945336,176945339,176945345,176945347,176945351,176945354,176945356,176945372,176945374,176945378,176945381,176945384,176945387,176945392,176945398,176945402,176945417,176945422,176945426,176945452,176945458,176945462,176945464,176945468,176945497,176945507,176945526,176945532,176945547,176945550,176945570,176945580,176945582,176945585,176945604,176945609,176945647,176945679,176945695,176945732,176945747,176945750,176945761,176945770,176945789,176945791,176945795,176964640,176964642,176964663,176964665,176964667,176964670,176964672,176964685,176964690,176964694,176964703,176964709,176964711,176964720,176964724,176964736,176964739,176964747,176964769,176964778,176964805,176964811,176964834,176964838,176964843,176964847,176964863,176964865,176964869,176964875,176964879,176964886,176964892,176964930,176964946,176964959,176964966,176964969,176964978,176965003,176965021,176965035,176965062,176965065,176965069,176965085,176965099,176965102,176965109,176965125,176965130,176965140,176965186,176965196,176994516,176994525,176994528,176994531,176994537,176994546,176994557,176994559,176994568,176994570,176994583,176994586,176994623,176994637,176994654,176994661,176994665,176994682,176994688,176994728,176994738,176994747,176994750,176994753,176994764,176994768,176994773,176994778,176994780,176994783,176994793,176994801,176994804,176994807,176994809,176994811,176994822,176994830,176994832,176994837,176994839,176994848,176994851,176994853,176994859,176994864,176994867,176994871,176994880,176994890,176994905,176994909,176994911,176994931,176994934,176994936,176994938,176994942,176994944,176994948,176994952,176994961,176994964,176994971,176994974,176994980,176994983,176994986,176994996,176995011,176995013,177017050,177017079,177017124,177017173,177017179,177017182,177017193,177017211,177017223,177017225,177017227,177017237,177017239,177017246,177017251,177017253, 177017267, 177017270, 177017276, 177017296, 177017300, 177017331, 177017352, 177017368, 177017374, 177017378, 177017389, 1770 17446, 177017449, 177017452, 177017463, 177017483, 177017488, 177024359, 177024367, 177024415, 177024502, 177024514, 177024528, 177024531 ,177024540,177024548,177024550,177024558,177024582,177024605,177024616,177024619,177024634,177024642,177024655,177024698,1770 24709, 177024714, 177024723, 177024725, 177024748, 177024756, 177024769, 177024771, 177024776, 177024783, 177024800, 177024836, 177024838 ,177024856,177024861; chr3 chromosomes 44063356,44063391,44063404,44063411,44063417,44063423,44063450,44063516,44063541,44063544,440635 59,44063565,44063567,44063574,44063586,44063593,44063602,44063606,44063620,44063633,44063638,44063643,44063649,44063657,44063 660,44063662,44063682,44063686,44063719,44063745,44063756,44063768,44063779,44063807,44063821,44063832,44063836,44063858,4406 3877,157812071,157812085,157812092,157812117,157812131,157812152,157812170,157812173,157812175,157812184,157812206,157812212,157812226, 157812256, 157812259, 157812275, 157812277, 157812287, 157812294, 157812296, 157812302, 157812305, 157812307, 157812312, 1578 12319, 157812321, 157812329, 157812331, 157812334, 157812354, 157812358, 157812369, 157812380, 157812383, 157812385, 157812404, 157812411 ,157812414,157812420,157812437,157812442,157812457,157812468,157812470,157812475,157812498,157812542,157812548; 978303 of chromosome chr4 6,9783050,9783059,9783075,9783080,9783097,9783105,9783112,9783120,9783126,9783142,9783144,9783153,9783160,9783166,9783185,978 3192,9783196,9783198,9783206,9783213,9783218,9783220,9783233,9783244,9783246,9783252,9783271,9783275,9783277,9783304,9783322, 9783327,9783342,9783348,9783354,9783358,9783361,9783363,9783376,9783398,9783409,9783425,9783427,9783442,9783449,9783467,97834 92,9783494,9783496,9783501,9783508,9783511,39448284,39448302,39448320,39448323,39448340,39448343,39448347,39448365,39448422,3 9448432,39448453,39448464,39448473,39448478,39448481,39448503,39448516,39448524,39448528,39448549,39448551,39448557,39448562,39448568, 39448575, 39448577, 39448586, 39448593, 39448613, 39448625, 39448629, 39448633, 39448647, 39448653, 39448662, 39448665, 39448670 ,39448683,39448695,39448697,39448729,39448732,39448748,39448757,39448759,39448767,39448773,39448796,39448800,39448809,3944881 1,39448836,39448845,39448857,39448864,39448869,39448874,57521138,57521209,57521237,57521297,57521304,57521310,57521336,575213 48,57521377,57521397,57521411,57521419,57521426,57521442,57521449,57521486,57521506,57521518,57521537,57521545,57521581,57521 603,57521622,57521631,57521652,57521657,57521665,57521680,57521687,57521701,57521716,57521725,57521733,154709378,154709414,15 4709425, 154709441, 154709492, 154709513, 154709522, 154709540, 154709557, 154709561, 154709576, 154709591, 154709597, 154709607, 1547096 12,154709617,154709633,154709640,154709663,154709675,154709684,154709690,154709697,154709721,154709745,154709756,154709759,15 4709789, 154709812, 154709828, 154709834; chr5 chromosomes: 1876139, 1876168, 1876200, 1876208, 1876213, 1876215, 1876286, 1876290, 1876298, 1876308.1876311, 1876337, 1876339, 1876347, 1876354, 1876368, 1876372, 1876374, 1876386, 1876395, 1876397, 1876399, 1876403, 1876420, 1876424, 1876 432,1876436,1876449,1876456,1876459,1876463,1876483,1876498,1876525,1876527,1876557,1876563,1876570,1876576,1876605,1876630,1 876634, 1876638; Chr6 chromosome 85476921, 85476930, 85476974, 85477014, 85477032, 85477035, 85477070, 85477083, 85477106, 85477124, 85477151, 85 477153,85477166,85477175,85477186,85477217,85477228,85477230,85477236,85477245,85477249,85477251,85477253,85477261,85477283,1 37814512, 137814516, 137814523, 137814548, 137814558, 137814561, 137814564, 137814567, 137814620, 137814636, 137814638, 137814642, 13781 4645,137814654,137814666,137814679,137814689,137814695,137814707,137814710,137814717,137814723,137814728,137814744,137814746, 137814749, 137814768, 137814776, 137814786, 137814788, 137814792, 137814794, 137814803, 137814807, 137814818, 137814824, 137814837, 13781 4860, 137814920, 137814935, 137814952, 137814957, 137814960, 137814969, 137814971, 137814986, 137814988, 137814995, 137815016, 137815024,137815030, 137815034, 137815036, 137815040, 150285620, 150285634, 150285641, 150285652, 150285659, 150285661, 150285670, 150285677, 15028 5688,150285695,150285697,150285706,150285713,150285715,150285724,150285731,150285733,150285742,150285760,150285767,150285769, 150285775, 150285778, 150285788, 150285813, 150285815, 150285826, 150285829, 150285844, 150285860, 150285887, 150285890, 150285892, 15028 5901,150285908,150285910,150285926,150285928,150285937,150285944,150285956,150285963,150285966,150285974,150285981,150285983, 150285992, 150285999, 150286001, 150286010, 150286017, 150286019, 150286028, 150286035, 150286038, 150286046, 150286055, 150286063, 15028 6073, 150286082, 150286089, 150286091; chr7 chromosome 27244531, 27244533, 27244537, 27244555, 27244564, 27244578, 27244603, 27244609, 27244612, 2 7244619,27244621,27244627,27244631,27244657,27244673,27244702,27244704,27244714,27244723,27244755,27244772,27244780,27244787, 27244789,27244798,27244800,27244810,27244833,27244856,27244869,27244874,27244881,27244885,27244887,27244892,27244897,27244907,27244911,27244917,27244920,27244931,27244948,27244951,27244980,27244982,27244986,27245014,27245018,35293441,35293451,35293470,35293479,35293482,35293488,35293492,35293497,35293502,35293506,35293514,35293531,35293537,35293543,35293588,35293590,35293621,35293652,35293656,35293658,35293670,35293676,35293685,35293687,35293690,35293692,35293700,35293717,35293721,35293731,35293747,35293750,35293753,35293759,35293767,35293780,35293783,35293790,35293796,35293809,35293812,35293815,35293821,35293827,35293829,35293834,35293838,35293840,35293847,35293849,35293860,35293863,35293867,35293869,35293879,35293884,35293892,35293940,50343545,50343548,50343552,50343555,50343562,50343566,50343572,50343574,50343577,50343579,50343587,50343603,50343605,50343608,50343611,50343624,50343628,50343630,50343635,50343637,50343639,50343648,50343651,50343654,50343656,50343659,50343663,50343669,50343672,50343674,50343678,50343682,50343693,50343696,50343699,50343702,50343714,50343719,50343725,50343728,50343731,50343736,50343739,50343758,50343765,50343768,50343770,50343785,50343789,50343791,50343805,50343813,50343822,50343824,50343826,50343829,50343831,50343833,50343838,50343847,50343850,50343853,50343858,50343864,50343869,50343872,50343883,50343890,50343897,50343907,50343909,50343914,50343926,50343934,50343939,50343946,50343950,50343959,50343961,50343963,50343969,50343974,50343980,50343990,50344001,50344007,50344011,50344028,50344041,155167320,155167333,155167340,155167343,155167345,155167347,155167350,155167357,155167379,155167382,155167394,155167401,155167423,155167430,155167467,155167478,155167480,155167486,155167499,155167505,155167507,155167511,155167513,155167516,155167518,155167528,155167543,155167552,155167555,155167560,155167562,155167568,155167570,155167578,155167602,155167608,155167611,155167617,155167662,155167702,155167707,155167716,155167718,155167739,155167750,155167753,155167757,155167759,155167771,155167773,155167791,155167801,155167803,155167805,155167813,155167819,155167821,155167827; Chr8 chromosome numbers 10588729, 10588742, 10588820, 10588833, 10588841, 10588851, 10588857, 10588865, 10588867, 10588883, 10588888, 1058889 5,10588938,10588942,10588946,10588948,10588951,10588959,10588992,10589003,10589007,10589009,10589016,10589034,10589060,105890 62,10589076,10589079,10589093,10589152,10589193,10589206,10589241,25907660,25907702,25907709,25907724,25907747,25907752,25907 754,25907757,25907769,25907796,25907800,25907814,25907818,25907821,25907824,25907838,25907848,25907866,25907874,25907880,2590 7884,25907893,25907898,25907900,25907902,25907906,25907918,25907947,25907976,25908055,25908057,25908064,25908071,25908098,259 08101,57069480,57069544,57069569,57069606,57069631,57069648,57069688,57069698,57069709,57069712,57069722,57069735,57069739,57 069755, 57069764, 57069773, 57069775, 57069784, 57069786, 57069791, 57069793, 57069800, 57069812, 57069816, 57069823, 57069825, 57069827, 5 7069839,57069842,57069847,57069851,57069853,57069884,57069889,57069894,57069907,57069914,57069919,57069931,57069940,57069948,57069958,57069968,57069973,57069978,57070013,57070035,57070038,57070042,57070046,57070066,57070079,57070087,57070091,5707012 6,57070143; chr10 chromosome 28034412,28034415,28034418,28034442,28034444,28034467,28034469,28034494,28034501,28034505,28034545,280345 56,28034559,28034568,28034582,28034591,28034596,28034599,28034605,28034616,28034619,28034622,28034624,28034645,28034651,2803 4654,28034658,28034669,28034682,28034687,28034697,28034711,28034714,28034727,28034729,28034739,28034741,28034751,28034757,280 34760,28034763,28034768,28034787,28034790,28034792,28034794,28034797,28034801,28034816,28034843,28034853,28034856,28034867,2 8034871, 28034873, 28034882, 28034888, 28034892, 28034907; 4918962, 4918966, 4918968, 4918975, 4918982, 4919001, 4919056, 4919065 of chromosome chr12 ,4919079,4919081,4919086,4919095,4919097,4919118,4919124,4919138,4919145,4919147,4919164,4919170,4919173,4919184,4919191,491 9199,4919215,4919230,4919236,4919239,4919242,4919253,4919260,4919281,4919293,4919300,4919303,4919309,4919327,4919331,4919351,4919358,4919376,4919386,4919395,4919401,4919408,4919421,4919424,4919430,4919438,4919453,4919465,4919469,4919475,4919486,33592615,33592629,33592635,33592642,33592659,33592661,33592663,33592674,33592681,33592683,33592692,33592704,33592707,33592709,33592711,33592715,33592720,33592725,33592727,33592744,33592774,33592798,33592803,33592811,33592831,33592848,33592859,33592862,33592865,33592867,33592875,33592882,33592885,33592887,33592891,33592905,33592908,33592913,33592915,33592923,33592931,33592933,33592953,33592955,33592977,33592981,33592986,33592989,33592998,33593004,33593017,33593035,33593049,33593090,33593093,58131100,58131102,58131111,58131133,58131154,58131168,58131175,58131181,58131224,58131242,58131261,58131277,58131300,58131303,58131306,58131309,58131312,58131318,58131321,58131331,58131345,58131348,58131384,58131390,58131404,58131412,58131414,58131426,58131429,58131445,58131453,58131475,58131478,58131487,58131503,58131510,58131523,58131546,58131549,58131553,58131557,58131564,58131571,58131576,58131586,58131605,58131608,58131624,58131642,115124768,115124773,115124782,115124811,115124838,115124853,11 5124871,115124874,115124894,115124904,115124924,115124930,115124933,115124935,115124946,115124970,115124973,115124981,1151249 99,115125013,115125034,115125053,115125060,115125098,115125107,115125114,115125121,115125131,115125141,115125151,115125177,11 5125192, 115125225, 115125305, 115125335; chr13 chromosome 37005452, 37005489, 37005501, 37005520, 37005551, 37005553, 37005557, 37005562, 370055 66,37005570,37005582,37005596,37005608,37005629,37005633,37005635,37005673,37005678,37005686,37005694,37005704,37005706,37005 721,37005732,37005738,37005741,37005745,37005773,37005778,37005794,37005801,37005805,37005814,37005816,37005821,37005833,3700 5835,37005844,37005855,37005857,37005878,37005881,37005883,37005892,37005899,37005909,37005924,37005929,37005934,37005939,370 05941, 100649486, 100649489, 100649519, 100649538, 100649567, 100649569, 100649577, 100649584, 100649601, 100649603, 100649605, 100649623,100649625, 100649628, 100649648, 100649671, 100649673, 100649686, 100649689, 100649691, 100649701, 100649705, 100649715, 100649718, 10064 9721,100649725,100649731,100649734,100649738,100649740,100649745,100649763,100649769,100649777,100649785,100649792,100649800, 100649847,100649886,100649912,100649915,100649917,100649941,100649945,100649949,100649965,100649975,100649982,100650005; chr14 Chromosomes 38724435, 38724459, 38724473, 38724486, 38724507, 38724511, 38724527, 38724531, 38724534, 38724540, 38724544, 38724546, 38724565, 3872 4578,38724586,38724597,38724624,38724627,38724646,38724648,38724650,38724669,38724675,38724680,38724682,38724685,38724726,387 24732,38724734,38724746,38724765,38724771,38724780,38724796,38724798,38724806,38724808,38724810,38724821,38724847,38724852,38 724858,38724864,38724867,38724873,38724896,38724906,38724929,38724935,38724945,38724978,38724995,38725003,38725005,38725014,3 8725016,38725023,38725026,38725030,38725034,38725038,38725048,38725058,38725077,38725081,38725088,38725101,57275669,57275674,57275677,57275681,57275683,57275687,57275690,57275706,57275725,57275749,57275752,57275761,57275768,57275772,57275778,5727578 5,57275821,57275823,57275827,57275829,57275831,57275835,57275852,57275874,57275876,57275885,57275896,57275908,57275912,572759 14,57275924,57275956,57275967,57275969,57275971,57275981,57275988,57275993,57275995,57276000,57276031,57276035,57276039,57276 057,57276066,57276073,57276090,60952394,60952398,60952405,60952418,60952421,60952425,60952464,60952468,60952482,60952500,6095 2503, 60952505, 60952517, 60952522, 60952544, 60952550, 60952554, 60952593, 60952599, 60952615, 60952618, 60952634, 60952658, 60952683, 609 52687,60952730,60952738,60952755,60952762,60952781,60952791,60952799,60952827,60952829,60952836,60952839,60952841,60952848,60 952855, 60952857, 60952870, 60952876, 60952878, 60952887, 60952896, 60952898, 60952908, 60952919, 60952921, 60952931; 83952068, 8 3952081, 83952084, 83952087, 83952095, 83952105, 83952108, 83952114, 83952125, 83952135, 83952140, 83952156, 83952160, 83952162, 83952175,83952178, 83952181, 83952184, 83952188, 83952200, 83952206, 83952209, 83952214, 83952220, 83952225, 83952229, 83952236, 83952238, 8395224 2,83952266,83952285,83952291,83952298,83952309,83952314,83952317,83952345,83952352,83952358,83952360,83952367,83952406,839524 11,83952414,83952418,83952420,83952425,83952430,83952453,83952464,83952472,83952486,83952496,83952498,83952500,83952506,83952 508, 83952527, 83952553, 83952559, 83952566, 83952570, 83952582, 83952592; chr16 chromosome 31579976, 31580071, 31580078, 31580081, 31580089, 3158 0100,31580110,31580117,31580138,31580150,31580153,31580159,31580165,31580220,31580246,31580254,31580269,31580287,31580296,315 80299,31580309,31580311,31580316,31580343,31580424,31580496,31580524,31580560,73096786,73096842,73096889,73096894,73096903,73 096914,73096923,73096929,73096934,73096943,73096948,73096966,73096970,73096979,73097000,73097015,73097017,73097019,73097028,7 3097037,73097045,73097057,73097060,73097066,73097069,73097078,73097080,73097082,73097084,73097108,73097114,73097142,73097156,73097183,73097260,73097267,73097284,73097296,73097301,73097329,73097357,73097364,73097377,73097381,73097387,73097470; chr17 staining The numbers 35299698, 35299703, 35299710, 35299719, 35299729, 35299731, 35299741, 35299746, 35299776, 35299813, 35299816, 35299822, 35299837, and 352998 are likely related to a specific character or index. 50,35299877,35299885,35299913,35299915,35299926,35299928,35299933,35299935,35299944,35299946,35299963,35299966,35299972,35299 974,35299990,35299996,35299999,35300006,35300010,35300020,35300027,35300036,35300039,35300044,35300059,35300068,35300074,3530 0086, 35300097, 35300109, 35300115, 35300146, 35300151, 35300163, 35300167, 35300172, 35300196, 35300202, 35300214, 35300217, 35300221, 769 29645,76929709,76929713,76929742,76929769,76929829,76929873,76929926,76929982,76930043,76930095,76930148,76930169,80846623,80 846652, 80846683, 80846709, 80846717, 80846730, 80846745, 80846763, 80846794, 80846860, 80846867, 80846886, 80846960, 80846965, 80847079, 8 0847092, 80847115, 80847128, 80847137, 80847153, 80847158, 80847209; chr21 chromosome numbers 38081248, 38081253, 38081300, 38081303, 38081306, 38081321.38081327, 38081333, 38081341, 38081344, 38081352, 38081354, 38081356, 38081363, 38081394, 38081396, 38081407, 38081421, 38081430, 38081443, 38081454, 38081461, 38081478, 38081480, 38081492, 38081497, 3808 The base numbers of the above methylation sites correspond to HG19 in the reference genome. (List of methylation sites follows, but is not translated as it's not part of the main text.)
[0132] In one or more embodiments, the nucleic acid molecule length is 1bp-1000bp, 1bp-900bp, 1bp-800bp, or 1bp-700bp. The nucleic acid molecule length can be any of the above-mentioned ranges.
[0133] In this document, methods for detecting DNA methylation are well known in the art, such as bisulfite-conversion-based PCR (e.g., methylation-specific PCR, MSP), DNA sequencing, whole-genome methylation sequencing, simplified methylation sequencing, methylation-sensitive restriction endonuclease assays, quantitative fluorescence assays, methylation-sensitive high-resolution melting curve assays, chip-based methylation mapping, and mass spectrometry. In one or more embodiments, the detection includes detecting any strand at a gene or site.
[0134] Therefore, this invention relates to reagents for detecting DNA methylation. Reagents used in the above-described methods for detecting DNA methylation are well known in the art. In detection methods involving DNA amplification, reagents for detecting DNA methylation include primers. The primer sequences may be methylation-specific or non-specific. The primer sequences may include non-methylation-specific blocking sequences. Blocking sequences can improve the specificity of methylation detection. Reagents for detecting DNA methylation may also include probes. Typically, the 5' end of the probe sequence is labeled with a fluorescent reporter group, and the 3' end is labeled with a quencher group. Exemplarily, the probe sequence contains an MGB (Minor Groove Binder) or an LNA (Locked Nucleic Acid). MGBs and LNAs are used to increase the Tm value, increase the specificity of the analysis, and improve the flexibility of probe design.
[0135] The term "primer" as used in this article refers to a nucleic acid molecule with a specific nucleotide sequence that guides the synthesis of nucleotides at the initiation of nucleotide polymerization. Primers are typically two artificially synthesized oligonucleotide sequences. One primer is complementary to one DNA template strand at one end of the target region, and the other primer is complementary to the other DNA template strand at the other end of the target region. Their function is to serve as the initiation point for nucleotide polymerization. Primers are usually at least 9 bp. Artificially designed primers are widely used in polymerase chain reaction (PCR), qPCR, sequencing, and probe synthesis. Typically, primers are designed to amplify products with lengths of 1-2000 bp, 10-1000 bp, 30-900 bp, 40-800 bp, 50-700 bp, or at least 150 bp, at least 140 bp, at least 130 bp, or at least 120 bp.
[0136] The term "variant" or "mutant" in this document refers to a polynucleotide whose nucleic acid sequence is altered compared to a reference sequence by the insertion, deletion, or substitution of one or more nucleotides while retaining its ability to hybridize with other nucleic acids. A mutant described in any embodiment of this document comprises a nucleotide sequence having at least 70%, preferably at least 80%, preferably at least 85%, preferably at least 90%, preferably at least 95%, preferably at least 97% sequence identity with a reference sequence and retaining the biological activity of the reference sequence. Sequence identity between two aligned sequences can be calculated using, for example, NCBI's BLASTn. A mutant also includes a nucleotide sequence having one or more mutations (insertions, deletions, or substitutions) in the reference sequence and its nucleotide sequence while still retaining the biological activity of the reference sequence. The multiple mutations typically refer to 1-10, for example, 1-8, 1-5, or 1-3. Substitution can be between purine nucleotides and pyrimidine nucleotides, or between purine nucleotides or pyrimidine nucleotides. Substitution is preferably conserved. For example, in the art, conserved substitution with nucleotides of similar or comparable properties generally does not alter the stability and function of the polynucleotide. Conservative substitutions include, for example, the interchange of (A and G) between purine nucleotides and the interchange of (T or U and C) between pyrimidine nucleotides. Therefore, replacing one or more sites with residues from the same source in the polynucleotides of this invention will not substantially affect their activity. Furthermore, methylation sites (e.g., consecutive CG) in the variants of this invention are not mutated. That is, the method of this invention detects the methylation status of methylable sites in the corresponding sequence; mutations can occur at bases in non-methylable sites. Typically, methylation sites are consecutive CpG dinucleotides.
[0137] As described herein, base conversions can occur between DNA or RNA bases. The terms "conversion," "cytosine conversion," or "CT conversion" used herein refer to the process of treating DNA using non-enzymatic or enzymatic methods to convert unmodified cytosine bases (C) into bases with a lower binding affinity to guanine (e.g., uracil bases (U)). Non-enzymatic or enzymatic methods for performing cytosine conversions are well known in the art. Exemplarily, non-enzymatic methods include treatment with conversion reagents such as bisulfites, acid sulfites, or metabisulfites, such as calcium bisulfite, sodium bisulfite, potassium bisulfite, ammonium bisulfite, sodium disulfite, potassium disulfite, and ammonium disulfite. Exemplarily, enzymatic methods include deaminase treatment. The converted DNA may optionally be purified. DNA purification methods suitable for use herein are well known in the art.
[0138] This invention also provides a methylation detection kit for diagnosing pancreatic cancer, the kit comprising the primers and / or probes described herein for detecting methylation levels of pancreatic cancer-related sequences discovered by the inventors. The kit may also contain nucleic acid molecules described herein, particularly those described in the first aspect, as internal standards or positive controls.
[0139] The term "hybridization" as used in this article primarily refers to nucleic acid sequence pairing under stringent conditions. An exemplary stringent condition is hybridization followed by membrane washing in a solution of 0.1×SSPE (or 0.1×SSC) and 0.1% SDS at 65°C.
[0140] In addition to the primers, probes, and nucleic acid molecules described above, the kit also contains other reagents required for the detection of DNA methylation. Exemplarily, these other reagents for detecting DNA methylation may include one or more of the following: bisulfite and its derivatives, PCR buffer, polymerase, dNTPs, primers, probes, methylation-sensitive or non-methylation-sensitive restriction endonucleases, enzyme digestion buffers, fluorescent dyes, fluorescence quenchers, fluorescent reporter agents, exonucleases, alkaline phosphatase, internal standards, and controls.
[0141] The kit may further include a converted positive standard, wherein unmethylated cytosine is converted to a base that does not bind to guanine. The positive standard may be fully methylated. The kit may also include PCR reaction reagents. Preferably, the PCR reaction reagents include Taq DNA polymerase, PCR buffer, dNTPs, and Mg2+. 2+ .
[0142] The present invention also provides a method for pancreatic cancer screening, comprising: (1) detecting the methylation level of pancreatic cancer-related sequences described herein in a sample of the subject; (2) comparing with a control sample, or calculating a score; and (3) identifying pancreatic cancer in the subject based on the score. Typically, the method further comprises, prior to step (1), extraction of sample DNA, quality control, and / or conversion of unmethylated cytosine on the DNA into bases that do not bind to guanine.
[0143] In a specific implementation, step (1) includes: treating genomic DNA or cfDNA with a transformation reagent to convert unmethylated cytosine into a base (e.g., uracil) that has a lower binding affinity to guanine; performing PCR amplification using primers suitable for amplifying the transformed sequence of the pancreatic cancer-related sequence described herein; and determining the methylation status or level of at least one CpG by the presence or absence of the amplification product or by sequence identification (e.g., probe-based PCR detection or DNA sequencing).
[0144] Alternatively, step (1) may also include: treating genomic DNA or cfDNA with a methylation-sensitive restriction endonuclease; performing PCR amplification using primers suitable for amplifying sequences having at least one CpG in the pancreatic cancer-related sequences described herein; and determining the methylation status or level of at least one CpG by the presence or absence of the amplification product.
[0145] The term "methylation level" as used herein refers to the relationship between the methylation status of any number and any position of CpGs in the sequence in question. This relationship can be the result of addition or subtraction of methylation status parameters (e.g., 0 or 1) or calculations using mathematical algorithms (e.g., mean, percentage, number, proportion, degree, or calculations using mathematical models), including but not limited to methylation level measures, methylation haplotype ratios, or methylation haplotype loadings. The term "methylation status" indicates the methylation of a specific CpG site, typically including methylated or unmethylated sites (e.g., methylation status parameter 0 or 1).
[0146] In one or more embodiments, the methylation level of the target sample increases or decreases when compared with a control sample. When the methylation marker level meets a certain threshold, pancreatic cancer is identified. Alternatively, mathematical analysis can be performed on the methylation level of the tested gene to obtain a score. For the tested sample, if the score is greater than the threshold, the result is considered positive, i.e., pancreatic cancer; otherwise, it is considered negative, i.e., no pancreatic cancer plasma. Conventional mathematical analysis methods and procedures for determining thresholds are known in the art. Exemplary methods include mathematical models, for example, for differential methylation markers, constructing a support vector machine (SVM) model for two sets of samples, using the model to statistically analyze the accuracy, sensitivity, and specificity of the detection results, as well as the area under the predictive value characteristic curve (ROC) (AUC), and statistically analyzing the predicted scores for the test set samples.
[0147] In a preferred embodiment, the model training process is as follows: First, differentially methylated regions are obtained based on the methylation level of each site, and a differentially methylated region matrix is constructed. For example, a methylation data matrix can be constructed from the methylation level data of a single CpG dinucleotide position in the HG19 genome using software such as samtools; then, SVM model training is performed.
[0148] An example SVM model training process is as follows:
[0149] a) Construct the training model mode. Use the sklearn package (0.23.1) of Python software (v3.6.9) to construct the training mode for cross-validation training of the model. Command line: model=SVR().
[0150] b) Using the sklearn package (0.23.1), input the data matrix and build an SVM model, model.fit(x_train, y_train), where x_train represents the training set data matrix and y_train represents the phenotypic information of the training set.
[0151] Typically, during model construction, pancreatic cancer is encoded as 1, and the absence of pancreatic cancer is encoded as 0. In this invention, the threshold is set to 0.895 using Python software (v3.6.9) and the sklearn package (0.23.1). The constructed model ultimately uses 0.895 to distinguish between samples with and without pancreatic cancer.
[0152] In this document, the samples are derived from mammals, preferably humans. Samples can be derived from any organ (e.g., pancreas), tissue (e.g., epithelial tissue, connective tissue, muscle tissue, and nerve tissue), cell (e.g., pancreatic cancer biopsy), or bodily fluid (e.g., blood, plasma, serum, tissue fluid, urine). Generally, the sample is acceptable as long as it contains genomic DNA or cfDNA (circulating free DNA or cell free DNA). cfDNA, also known as circulating cell-free DNA or cell-free DNA, is a fragment of degraded DNA released into the plasma. Exemplarily, the sample is a pancreatic cancer biopsy, preferably a fine-needle aspiration biopsy. Alternatively, the sample may be plasma or cfDNA.
[0153] This article also describes a method for obtaining methylation haplotype ratios associated with pancreatic cancer. Taking methylation data obtained from methylation-targeted sequencing (MethylTitan) as an example, the process of screening and testing biomarker sites is as follows: raw paired-end sequencing reads – readings are merged to obtain merged single-end reads – adapters are removed to obtain adapter-removed reads – Bismark is aligned to the human DNA genome to form a BAM file – samtools extracts the methylation level of CpG sites for each read to form a haplotype file – the proportion of methylated haplotypes at C sites is statistically analyzed to form a meth file – the MHF (Methylated Haplotype Fraction) methylated haplotype ratio is calculated – Coverage200 filters sites to form a meth.matrix matrix file – filtering is performed according to NA values greater than 0.1 – the samples are pre-divided into training and test sets – for each haplotype in the training set, a logistic regression model is constructed for the phenotype, and the regression P-value of each methylated haplotype ratio is selected – the methylated haplotype with the most significant P-value in each MethylTitan amplification region is selected to represent the methylation level of that region and modeled using a support vector machine – the results of the training set (ROC plot) are generated and the model is used to predict the test set for validation. Specifically, the method for obtaining methylated haplotypes associated with pancreatic cancer includes the following steps: (1) obtaining plasma samples from patients with or without pancreatic cancer, extracting cfDNA, and performing library construction and sequencing using the MethylTitan method to obtain sequencing reads; (2) preprocessing the sequencing data, including adapter removal and splicing of the sequencing data generated by the sequencer; (3) aligning the preprocessed sequencing data to the HG19 reference genome sequence of the human genome to determine the position of each fragment. The data in step (2) can be obtained from paired-end 150bp sequencing on the Illumina sequencing platform. Adapter removal in step (2) involves removing the sequencing adapters at the 5' and 3' ends of the two paired-end sequencing data, as well as removing low-quality bases after adapter removal. Splicing in step (2) involves merging the paired-end sequencing data to restore the original library fragment. This allows for better alignment and accurate location of sequencing fragments. For example, the length of the sequencing library is around 180bp, and paired-end 150bp can completely cover the entire library fragment. Step (3) includes: (a) converting the HG19 reference genome data to CT and GA respectively, constructing two sets of converted reference genomes, and constructing alignment indexes for the converted reference genomes respectively; (b) converting the merged sequencing sequence data to CT and GA respectively; (c) aligning the converted reference genome sequences respectively, and finally summarizing the alignment results to determine the position of the sequencing data in the reference genome.
[0154] Furthermore, the method for obtaining methylation values associated with pancreatic cancer also includes (4) calculating MHF; (5) constructing a methylation haplotype MHF data matrix; and (6) constructing a logistic regression model for each methylation haplotype based on sample grouping. Step (4) involves obtaining the methylation haplotype status and sequencing depth information at the location of the HG19 reference genome based on the alignment results obtained in step (3). Step (5) involves merging the methylation haplotype status and sequencing depth information data into a data matrix. Among them, each data point with a depth less than 200 is treated as a missing value, and the missing values are filled using the K nearest neighbor (KNN) method. Step (6) involves statistically modeling each location in the above matrix using logistic regression and screening for haplotypes with significant regression coefficients between the two groups.
[0155] The term "multiple" as used herein refers to any integer. Preferably, "multiple" in "one or more" can be any integer greater than or equal to 2, including 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60 or more.
[0156] The beneficial effects of this invention are:
[0157] Based on the methylated nucleic acid fragment biomarkers of this invention, pancreatic cancer can be effectively identified. This invention provides a diagnostic model of the relationship between cfDNA methylation biomarkers and pancreatic cancer based on high-throughput methylation sequencing of plasma cfDNA. This model has the advantages of non-invasive detection, safe and convenient detection, high throughput, and high detection specificity. Based on the optimal sequencing volume obtained by this invention, detection costs can be effectively controlled while achieving good detection performance.
[0158] Example
[0159] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments. In the following embodiments, experimental methods without specific conditions are generally performed according to the methods described under conventional conditions.
[0160] Example 1: Methylation-targeted sequencing to screen differentially methylation sites in pancreatic cancer
[0161] The inventors collected a total of 94 blood samples from patients with pancreatic cancer and 80 blood samples from patients without pancreatic cancer. All enrolled patients signed informed consent forms. Sample information is shown in the table below.
[0162]
[0163] Methylation sequencing data of plasma DNA was obtained using the MethylTitan method, and methylation taxonomic markers were identified. The process is as follows:
[0164] 1. Extraction of plasma cfDNA samples
[0165] A 2ml whole blood sample was collected from the patient using a streck blood collection tube. Plasma was separated by centrifugation within 3 days and then transferred to the laboratory. cfDNA was extracted using the QIAGEN QIAamp Circulating Nucleic Acid Kit according to the instructions.
[0166] 2. Sequencing and Data Preprocessing
[0167] 1) The library was sequenced using an Illumina Nextseq 500 sequencer with paired ends.
[0168] 2) The Pear (v0.6.0) software merges the paired-end sequencing data of the same fragment of 150bp from the Illumina Hiseq X10 / Nextseq 500 / Nova seq sequencer into a single sequence with a minimum overlap length of 20bp and a minimum length of 30bp after merging.
[0169] 3) Trim_galore v 0.6.0 and cutadapt v1.8.1 software were used to remove adapters from the merged sequencing data. The adapter sequence “AGATCGGAAGAGCAC” was removed from the 5' end of the sequence, and bases with sequencing quality values lower than 20 at both ends were removed.
[0170] 3. Sequencing data alignment
[0171] The reference genome data used in this article came from the UCSC database (UCSC:HG19, http: / / hgdownload.soe.ucsc.edu / goldenPath / hg19 / bigZips / hg19.fa.gz).
[0172] 1) First, HG19 was transformed into cytosine to thymine (CT) and adenine to guanine (GA) using Bismark software, and the transformed genomes were indexed using Bowtie2 software.
[0173] 2) Perform CT and GA conversion on the preprocessed data as well.
[0174] 3) Use Bowtie2 software to align the transformed sequences to the transformed HG19 reference genome. The minimum seed sequence length is 20, and mismatches in the seed sequence are not allowed.
[0175] 4. Calculation of MHF
[0176] For each CpG site of HG19 in the target region, the methylation level corresponding to each site is obtained based on the alignment results above. In this paper, the nucleotide number of the site corresponds to the nucleotide position number of HG19. A target methylation region may have multiple methylation haplotypes; this value needs to be calculated for each methylation haplotype within the target region. An example of the MHF calculation formula is as follows:
[0177]
[0178] Where i represents the target methylation region, h represents the target methylation haplotype, and N i N represents the number of reads located in the target methylation region. i,h This indicates the number of reads containing the target methylation haplotype.
[0179] 5. Methylation data matrix
[0180] 1) Merge the methylation sequencing data of each sample in the training set and the test set into a data matrix, and perform missing value processing on each site with a depth of less than 200.
[0181] 2) Remove sites with a missing value ratio higher than 10%.
[0182] 3) For missing values in the data matrix, the KNN algorithm is used to impute the missing data.
[0183] 6. Identify characteristic methylation regions by grouping training set samples.
[0184] 1) For each methylation segment, a logistic regression model is constructed for the phenotype. For each amplified target region, the methylation segment with the most significant regression coefficient is selected to form a candidate methylation segment.
[0185] 2) Randomly divide the training set into ten parts and perform incremental feature selection with tenfold cross-validation.
[0186] 3) The candidate methylation segments in each region are sorted from largest to smallest according to the significance of the regression coefficient. One methylation segment is added at a time to predict the test data.
[0187] 4) In step 3), use the 10 datasets generated in step 2) to calculate the AUC 10 times each time, and take the average of the 10 calculations. If the AUC of the training data increases, retain the candidate methylation region as a feature methylation region; otherwise, discard it.
[0188] 5) Take the feature combination corresponding to the median of the average AUC under different feature counts in the training set as the final determined feature methylation segment combination.
[0189] The distribution of the selected characteristic methylated nucleic acid sequences is as follows: SEQ ID NO:1 of the DMRTA2 gene region, SEQ ID NO:2 of the FOXD3 gene region, SEQ ID NO:3 of the TBX15 gene region, SEQ ID NO:4 of the BCAN gene region, SEQ ID NO:5 of the TRIM58 gene region, SEQ ID NO:6 of the SIX3 gene region, SEQ ID NO:7 of the VAX2 gene region, SEQ ID NO:8 of the EMX1 gene region, SEQ ID NO:9 of the LBX2 gene region, SEQ ID NO:10 of the TLX2 gene region, SEQ ID NO:11 and SEQ ID NO:12 of the POU3F3 gene region, SEQ ID NO:13 of the TBR1 gene region, SEQ ID NO:14 and SEQ ID NO:15 of the EVX2 gene region, SEQ ID NO:16 of the HOXD12 gene region, SEQ ID NO:17 of the HOXD8 gene region, SEQ ID NO:18 and SEQ ID NO:19 of the HOXD4 gene region, SEQ ID NO:20 of the TOPAZ1 gene region, and SEQ ID NO:20 of the SHOX2 gene region. SEQ ID NO:21, DRD5 gene region; SEQ ID NO:22, RPL9 gene region; SEQ ID NO:23, SEQ ID NO:24, HOPX gene region; SEQ ID NO:25, SFRP2 gene region; SEQ ID NO:26, IRX4 gene region; SEQ ID NO:27, TBX18 gene region; SEQ ID NO:28, OLIG3 gene region; SEQ ID NO:29, ULBP1 gene region; SEQ ID NO:30, HOXA13 gene region; SEQ ID NO:31, TBX20 gene region; SEQ ID NO:32, IKZF1 gene region; SEQ ID NO:33, INSIG1 gene region; SEQ ID NO:34, SOX7 gene region; SEQ ID NO:35, EBF2 gene region; SEQ ID NO:36, MOS gene region; SEQ ID NO:37, MKX gene region; SEQ ID NO:38, KCNA6 gene region; SEQ ID NO:39, SYT10 gene region; SEQ ID NO:40, AGAP2 gene region; SEQ ID NO:39, SYT10 gene region; SEQ ID NO:40, AGAP2 gene region; SEQ ID NO:41, TBX3 gene region; SEQ ID NO:42, CCNA1 gene region; SEQ ID NO:43, ZIC2 gene region; SEQ ID NO:44, SEQ ID NO:45The methylation levels of the above methylation biomarkers were found to be elevated or decreased in the cfDNA of pancreatic cancer patients (Table 1). The sequences of the 56 biomarker regions are shown in SEQ ID NO:1-56. The methylation levels of all CpG sites in each biomarker region were obtained using MethylTitan sequencing. The mean methylation level of all CpG sites in each region, as well as the methylation level of a single CpG site, can be used as biomarkers for diagnosing pancreatic cancer.
[0190] Table 1: Average level of methylation biomarkers in the training set
[0191]
[0192]
[0193]
[0194] Table 2 shows the methylation levels of methylation markers in the pancreatic cancer and non-pancreatic cancer populations tested. As can be seen from the table, the selected methylation markers showed significantly different distributions between the two populations, demonstrating good discriminatory power.
[0195] Table 2: Methylation levels of methylation biomarkers in the test set
[0196]
[0197]
[0198] Table 3 lists the correlation (Pearson correlation coefficient) between the methylation levels of 10 random CpG sites or combinations within each selected biomarker and the overall biomarker methylation level, along with the corresponding significance p-values. It can be seen that the methylation levels of individual CpG sites or combinations of multiple CpG sites within a biomarker are significantly correlated with the overall methylation level of the region (p<0.05), and the correlation coefficients are all above 0.8, indicating strong or very strong correlation. This suggests that individual CpG sites or combinations of multiple CpG sites within a biomarker, like the entire biomarker, also have good distinguishing effects.
[0199] Table 3: Correlation between methylation levels of random CpG sites or combinations of sites in 56 biomarkers and overall biomarker methylation levels.
[0200]
[0201]
[0202]
[0203]
[0204]
[0205]
[0206]
[0207]
[0208]
[0209]
[0210]
[0211]
[0212]
[0213] Example 2: Predictive performance of a single methylation marker
[0214] To verify the ability of a single methylation biomarker to differentiate between patients with and without pancreatic cancer, the methylation level of the single biomarker was used to validate its predictive performance.
[0215] First, the methylation levels of 56 methylation markers were used individually in the training set samples to determine the threshold for distinguishing between pancreatic cancer and its sensitivity and specificity. Then, the threshold was used to statistically analyze the sensitivity and specificity of the test set samples. The results are shown in Table 4 below. It can be seen that a single marker can also achieve good distinguishing performance.
[0216] Table 4: Predictive performance of 56 methylation markers
[0217]
[0218]
[0219]
[0220]
[0221] Example 3: Predictive Model for All Biomarker Combinations
[0222] To verify the potential of using methylated nucleic acid fragment biomarkers for pancreatic cancer differentiation, a support vector machine disease classification model was constructed based on 56 methylated nucleic acid fragment biomarkers in the training group. The classification prediction performance of this set of methylated biomarkers was then validated in the test group. The training and test groups were divided proportionally, with 117 cases in the training group (samples 1-117) and 57 cases in the test group (samples 118-174).
[0223] Support vector machine models were built on the training set using the discovered methylation markers from both groups of samples.
[0224] 1) Divide the samples into two parts in advance, one part for training the model and the other part for testing the model.
[0225] 2) The SVM model was trained using the methylation marker levels in the training set. The specific training process is as follows:
[0226] a) Use the sklearn package (0.23.1) of Python software (v3.6.9) to build the training mode of the cross-validation training model. Command line: model=SVR().
[0227] b) Using the sklearn package (0.23.1), input the methylation numerical matrix and build an SVM model, model.fit(x_train,y_train), where x_train represents the methylation numerical matrix of the training set and y_train represents the phenotypic information of the training set.
[0228] During model construction, pancreatic cancer samples were coded as 1, and samples without pancreatic cancer were coded as 0. The sklearn package (0.23.1) defaulted to setting the threshold to 0.895. The final model also used a scoring threshold of 0.895 to distinguish between samples with and without pancreatic cancer. The prediction scores of the two models for the training set samples are shown in Table 5.
[0229] Table 5: Model prediction scores on the training set
[0230]
[0231]
[0232]
[0233] Based on the methylated nucleic acid fragment biomarkers of this invention, predictions were made on the test set using the SVM model established in this embodiment. A prediction function was used to predict the test set, and the output was the prediction result (disease probability: the default scoring threshold is 0.895; a score greater than 0.895 indicates the subject is considered malignant). The test group consisted of 57 cases (samples 118-174), and the calculation process is as follows:
[0234] Command line:
[0235] test_pred=model.predict(test_df)
[0236] Where test_pred represents the prediction score obtained by the SVM prediction model constructed in this embodiment for the test set samples, model represents the SVM prediction model constructed in this embodiment, and test_df represents the test set data.
[0237] The predicted scores for the test group are shown in Table 6, and the ROC curves are as follows: Figure 2 As shown, the distribution of predicted scores is as follows: Figure 3 As shown, the area under the overall AUC of the test group is 0.911. In the training set, the model achieves a sensitivity of 71.4% with a specificity of 90.7%; in the test set, the model achieves a sensitivity of 83.9% with a specificity of 88.5%. This indicates that the SVM models built from the selected variables all exhibit good discriminative power.
[0238] Figure 4 and Figure 5 The distribution of the 56 methylated nucleic acid fragment biomarkers in the training and testing groups was shown separately. It was found that the differences of the methylation biomarkers in the plasma of subjects without pancreatic cancer and patients with pancreatic cancer were relatively stable.
[0239] Table 6: Model prediction scores for test set samples
[0240]
[0241]
[0242] Example 4: Comparison of Tumor Marker Predictions
[0243] Based on the methylation biomarker group of this invention, predictions were made in the test set according to the model established by SVM in Example 3. Pancreatic cancer prediction was performed using the CA19-9 biomarker. A sample of 130 cases was used (Table 7), and the calculation process is as follows:
[0244] Command line:
[0245] Combine_scalar=RobustScaler().fit(combine_train_df)
[0246] scaled_combine_train_df=combine_scalar.transform(combine_train_df)
[0247] scaled_combine_test_df=combine_scalar.transform(combine_test_df)
[0248] combine_model=LogisticRegression().fit(scaled_combine_train_df,train_ca19_pheno)
[0249] Where `combine_train_df` represents the training set data matrix combining the predicted scores of the test set samples obtained by the SVM prediction model constructed in Example 3 with CA19-9, `scaled_combine_train_df` represents the standardized training set data matrix, `scaled_combine_test_df` represents the standardized test set data matrix, and `combine_model` represents the logistic regression model fitted using the standardized training set data matrix.
[0250] The predicted scores for the samples are shown in Table 7, and the ROC curves are shown in... Figure 6 As shown, the distribution of predicted scores is as follows: Figure 7 As shown, the AUC of the test group in the overall dataset was 0.935. The figure also shows that the established logistic regression models exhibited good discrimination.
[0251] Figure 7The distribution of classification prediction scores using CA19-9 alone, the SVM model built using Example 3 alone, and the model built using Example 3 combined with CA19-9 are shown respectively. It can be found that this method performs more stably in pancreatic cancer identification.
[0252] Table 7: CA19-9 Predicted Scores and Model-Merged CA19-9 Predicted Scores
[0253]
[0254]
[0255]
[0256]
[0257] Example 5: Performance of the classification prediction model in negative samples with traditional biomarkers
[0258] Based on the methylation biomarker group of the present invention, the samples that were negative for the conventional tumor biomarker CA19-9 (CA19-9 measurement value <37) were tested according to the model established by SVM in Example 3.
[0259] The CA19-9 measurements and model predictions for the relevant samples are shown in Table 8, and the ROC curves are shown in Table 8. Figure 8 Using 0.895 as the scoring threshold, the AUC value reached 0.885 in the test set, indicating that the SVM model constructed in Example 3 can still achieve relatively good results for patients who cannot be identified using CA19-9.
[0260] Table 8: CA19-9 Measurement Values and SVM Model Predicted Scores
[0261]
[0262]
[0263] Example 6: Model Construction and Performance Evaluation of 7 Biomarker Combinations SEQ ID NO:9,14,13,26,40,43,52
[0264] To verify the predictive performance of different biomarker combinations, based on the 56 methylation biomarker groups of this invention, 7 biomarkers SEQ ID NO: 9, 14, 13, 26, 40, 43, 52 were selected for model construction and performance testing. A training group and a test group were established, with 117 cases in the training group (samples 1-117) and 57 cases in the test group (samples 118-174).
[0265] Using these 7 methylation markers, support vector machine models were constructed on the training set for both groups of samples:
[0266] 1. Divide the samples into two parts in advance, one part for training the model and the other part for testing the model.
[0267] 2. The SVM model was trained using the methylation marker levels in the training set. The specific training process is as follows:
[0268] a) Use the sklearn package (0.23.1) of Python software (v3.6.9) to build the training mode of the cross-validation training model. Command line: model=SVR().
[0269] b) Using the sklearn package (0.23.1), input the methylation numerical matrix and build an SVM model, model.fit(x_train,y_train), where x_train represents the methylation numerical matrix of the training set and y_train represents the phenotypic information of the training set.
[0270] 3. Test using test set data: Test the above model using the test set data. Command line: test_pred = model.predict(test_df), where test_pred represents the prediction score obtained by the SVM prediction model constructed in this embodiment for the test set samples, model represents the SVM prediction model constructed in this embodiment, and test_df represents the test set data.
[0271] The ROC curve of this 7-marker combination model is as follows: Figure 9 As shown, the AUC of the established model is 0.881. In the test set, when the specificity is 0.846, the sensitivity can reach 0.774 (Table 9), which can achieve good differentiation performance between pancreatic cancer patients and healthy people.
[0272] Table 9: Performance of the 7-marker combination model
[0273] training set 0.8586 0.7302 0.8519 0.5786 test set 0.8809 0.7742 0.8462 0.5786
[0274] Example 7: Model Construction and Performance Evaluation of 7 Biomarker Combinations SEQ ID NO:5,18,34,40,43,45,46
[0275] To verify the predictive performance of different biomarker combinations, based on the 56 methylation biomarker groups of this invention, 7 biomarkers SEQ ID NO:5,18,34,40,43,45,46 were selected for model construction and performance testing. A training group and a test group were established, with 117 cases in the training group (samples 1-117) and 57 cases in the test group (samples 118-174).
[0276] Using these 7 methylation markers, support vector machine models were constructed on the training set for both groups of samples:
[0277] 1. Divide the samples into two parts in advance, one part for training the model and the other part for testing the model.
[0278] 2. The SVM model was trained using the methylation marker levels in the training set. The specific training process is as follows:
[0279] a) Use the sklearn package (0.23.1) of Python software (v3.6.9) to build the training mode of the cross-validation training model. Command line: model=SVR().
[0280] b) Using the sklearn package (0.23.1), input the methylation numerical matrix and build an SVM model, model.fit(x_train,y_train), where x_train represents the methylation numerical matrix of the training set and y_train represents the phenotypic information of the training set.
[0281] 3. Test using test set data: Test the above model using the test set data. Command line: test_pred = model.predict(test_df), where test_pred represents the prediction score obtained by the SVM prediction model constructed in this embodiment for the test set samples, model represents the SVM prediction model constructed in this embodiment, and test_df represents the test set data.
[0282] The ROC curve of this 7-marker combination model is as follows: Figure 10 As shown, the AUC of the established model is 0.881. In the test set, when the specificity is 0.692, the sensitivity can reach 0.839 (Table 10), which can achieve good differentiation performance between pancreatic cancer patients and healthy people.
[0283] Table 10: Performance of the model combining these 7 markers
[0284] training set 0.8898 0.8095 0.8519 0.4179 test set 0.8809 0.8387 0.6923 0.4179
[0285] Example 8: Model Construction and Performance Evaluation of 7 Biomarker Combinations SEQ ID NO:8,11,20,44,48,51,54
[0286] To verify the predictive performance of different biomarker combinations, based on the 56 methylation biomarker groups of this invention, 7 biomarkers SEQ ID NO:8,11,20,44,48,51,54 were selected for model construction and performance testing. A training group and a test group were established, with 117 cases in the training group (samples 1-117) and 57 cases in the test group (samples 118-174).
[0287] Using these 7 methylation markers, support vector machine models were constructed on the training set for both groups of samples:
[0288] 1. Divide the samples into two parts in advance, one part for training the model and the other part for testing the model.
[0289] 2. The SVM model was trained using the methylation marker levels in the training set. The specific training process is as follows:
[0290] a) Use the sklearn package (0.23.1) of Python software (v3.6.9) to build the training mode of the cross-validation training model. Command line: model=SVR().
[0291] b) Using the sklearn package (0.23.1), input the methylation numerical matrix and build an SVM model, model.fit(x_train,y_train), where x_train represents the methylation numerical matrix of the training set and y_train represents the phenotypic information of the training set.
[0292] 3. Test using test set data: Test the above model using the test set data. Command line: test_pred = model.predict(test_df), where test_pred represents the prediction score obtained by the SVM prediction model constructed in this embodiment for the test set samples, model represents the SVM prediction model constructed in this embodiment, and test_df represents the test set data.
[0293] The ROC curve of this 7-marker combination model is as follows: Figure 11 As shown, the AUC of the established model is 0.880. In the test set, when the specificity is 0.769, the sensitivity can reach 0.839 (Table 11), which can achieve good differentiation performance between pancreatic cancer patients and healthy people.
[0294] Table 11: Performance of the model combining the 7 markers
[0295] training set 0.8812 0.7143 0.8519 0.4434 test set 0.8797 0.8387 0.7692 0.4434
[0296] Example 9: Model Construction and Performance Evaluation of 7 Biomarker Combinations SEQ ID NO:8,14,26,24,31,40,46
[0297] To verify the predictive performance of different biomarker combinations, based on the 56 methylation biomarker groups of this invention, 7 biomarkers SEQ ID NO:8,14,26,24,31,40,46 were selected for model construction and performance testing. A training group and a test group were established, with 117 cases in the training group (samples 1-117) and 57 cases in the test group (samples 118-174).
[0298] Using these 7 methylation markers, support vector machine models were constructed on the training set for both groups of samples:
[0299] 1. Divide the samples into two parts in advance, one part for training the model and the other part for testing the model.
[0300] 2. The SVM model was trained using the methylation marker levels in the training set. The specific training process is as follows:
[0301] a) Use the sklearn package (0.23.1) of Python software (v3.6.9) to build the training mode of the cross-validation training model. Command line: model=SVR().
[0302] b) Using the sklearn package (0.23.1), input the methylation numerical matrix and build an SVM model, model.fit(x_train,y_train), where x_train represents the methylation numerical matrix of the training set and y_train represents the phenotypic information of the training set.
[0303] 3. Test using test set data: Test the above model using the test set data. Command line: test_pred = model.predict(test_df), where test_pred represents the prediction score obtained by the SVM prediction model constructed in this embodiment for the test set samples, model represents the SVM prediction model constructed in this embodiment, and test_df represents the test set data.
[0304] The ROC curve of this 7-marker combination model is as follows: Figure 12 As shown, the AUC of the established model is 0.871. In the test set, when the specificity is 0.885, the sensitivity can reach 0.710 (Table 12), which can achieve good differentiation performance between pancreatic cancer patients and healthy people.
[0305] Table 12: Performance of the model combining the 7 markers
[0306] training set 0.8745 0.6984 0.8519 0.5380 test set 0.8710 0.7097 0.8846 0.5380
[0307] Example 10: Model Construction and Performance Evaluation of 7 Biomarker Combinations SEQ ID NO:3,9,8,29,42,40,41
[0308] To verify the predictive performance of different biomarker combinations, based on the 56 methylation biomarker groups of this invention, 7 biomarkers SEQ ID NO:3,9,8,29,42,40,41 were selected for model construction and performance testing. A training group and a test group were established, with 117 cases in the training group (samples 1-117) and 57 cases in the test group (samples 118-174).
[0309] Using these 7 methylation markers, support vector machine models were constructed on the training set for both groups of samples:
[0310] 1. Divide the samples into two parts in advance, one part for training the model and the other part for testing the model.
[0311] 2. The SVM model was trained using the methylation marker levels in the training set. The specific training process is as follows:
[0312] a) Use the sklearn package (0.23.1) of Python software (v3.6.9) to build the training mode of the cross-validation training model. Command line: model=SVR().
[0313] b) Using the sklearn package (0.23.1), input the methylation numerical matrix and build an SVM model, model.fit(x_train,y_train), where x_train represents the methylation numerical matrix of the training set and y_train represents the phenotypic information of the training set.
[0314] 3. Test using test set data: Test the above model using the test set data. Command line: test_pred = model.predict(test_df), where test_pred represents the prediction score obtained by the SVM prediction model constructed in this embodiment for the test set samples, model represents the SVM prediction model constructed in this embodiment, and test_df represents the test set data.
[0315] The ROC curve of this 7-marker combination model is as follows: Figure 13As shown, the AUC of the established model is 0.866, and the sensitivity can reach 0.903 when the specificity is 0.538 in the test set (Table 13), which can achieve good differentiation performance between pancreatic cancer patients and healthy people.
[0316] Table 13: Performance of the model combining the 7 markers
[0317] training set 0.8930 0.8413 0.8519 0.4014 test set 0.8660 0.9032 0.5385 0.4014
[0318] Example 11: Model Construction and Performance Evaluation of 7 Biomarker Combinations SEQ ID NO:5,8,19,7,44,47,53
[0319] To verify the predictive performance of different biomarker combinations, based on the 56 methylation biomarker groups of this invention, 7 biomarkers SEQ ID NO: 5, 8, 19, 7, 44, 47, 53 were selected for model construction and performance testing. A training group and a test group were established, with 117 cases in the training group (samples 1-117) and 57 cases in the test group (samples 118-174).
[0320] Using these 7 methylation markers, support vector machine models were constructed on the training set for both groups of samples:
[0321] 1. Divide the samples into two parts in advance, one part for training the model and the other part for testing the model.
[0322] 2. The SVM model was trained using the methylation marker levels in the training set. The specific training process is as follows:
[0323] a) Use the sklearn package (0.23.1) of Python software (v3.6.9) to build the training mode of the cross-validation training model. Command line: model=SVR().
[0324] b) Using the sklearn package (0.23.1), input the methylation numerical matrix and build an SVM model, model.fit(x_train,y_train), where x_train represents the methylation numerical matrix of the training set and y_train represents the phenotypic information of the training set.
[0325] 3. Test using test set data: Test the above model using the test set data. Command line: test_pred = model.predict(test_df), where test_pred represents the prediction score obtained by the SVM prediction model constructed in this embodiment for the test set samples, model represents the SVM prediction model constructed in this embodiment, and test_df represents the test set data.
[0326] The ROC curve of this 7-marker combination model is as follows: Figure 14 As shown, the AUC of the established model is 0.864. In the test set, when the specificity is 0.577, the sensitivity can reach 0.774 (Table 14), which can achieve good differentiation performance between pancreatic cancer patients and healthy people.
[0327] Table 14: Performance of the model combining the 7 markers
[0328] training set 0.8704 0.6984 0.8519 0.4803 test set 0.8635 0.7742 0.5769 0.4803
[0329] Example 12: Model Construction and Performance Evaluation of 7 Biomarker Combinations SEQ ID NO:12,17,24,28,40,42,47
[0330] To verify the predictive performance of different biomarker combinations, based on the 56 methylation biomarker groups of this invention, 7 biomarkers SEQ ID NO:12,17,24,28,40,42,47 were selected for model construction and performance testing. A training group and a test group were established, with 117 cases in the training group (samples 1-117) and 57 cases in the test group (samples 118-174).
[0331] Using these 7 methylation markers, support vector machine models were constructed on the training set for both groups of samples:
[0332] 1. Divide the samples into two parts in advance, one part for training the model and the other part for testing the model.
[0333] 2. The SVM model was trained using the methylation marker levels in the training set. The specific training process is as follows:
[0334] a) Use the sklearn package (0.23.1) of Python software (v3.6.9) to build the training mode of the cross-validation training model. Command line: model=SVR().
[0335] b) Using the sklearn package (0.23.1), input the methylation numerical matrix and build an SVM model, model.fit(x_train,y_train), where x_train represents the methylation numerical matrix of the training set and y_train represents the phenotypic information of the training set.
[0336] 3. Test using test set data: Test the above model using the test set data. Command line: test_pred = model.predict(test_df), where test_pred represents the prediction score obtained by the SVM prediction model constructed in this embodiment for the test set samples, model represents the SVM prediction model constructed in this embodiment, and test_df represents the test set data.
[0337] The ROC curve of this 7-marker combination model is as follows: Figure 15 As shown, the AUC of the established model is 0.862. In the test set, when the specificity is 0.731, the sensitivity can reach 0.871 (Table 15), which can achieve good differentiation performance between pancreatic cancer patients and healthy people.
[0338] Table 15: Performance of the model combining these 7 markers
[0339] training set 0.8859 0.8571 0.8519 0.4514 test set 0.8623 0.8710 0.7308 0.4514
[0340] Example 13: Model Construction and Performance Evaluation of 7 Biomarker Combinations SEQ ID NO:5,18,14,10,8,19,27
[0341] To verify the predictive performance of different biomarker combinations, based on the 56 methylation biomarker groups of this invention, 7 biomarkers SEQ ID NO:5,18,14,10,8,19,27 were selected for model construction and performance testing. A training group and a test group were established, with 117 cases in the training group (samples 1-117) and 57 cases in the test group (samples 118-174).
[0342] Using these 7 methylation markers, support vector machine models were constructed on the training set for both groups of samples:
[0343] 1. Divide the samples into two parts in advance, one part for training the model and the other part for testing the model.
[0344] 2. The SVM model was trained using the methylation marker levels in the training set. The specific training process is as follows:
[0345] a) Use the sklearn package (0.23.1) of Python software (v3.6.9) to build the training mode of the cross-validation training model. Command line: model=SVR().
[0346] b) Using the sklearn package (0.23.1), input the methylation numerical matrix and build an SVM model, model.fit(x_train,y_train), where x_train represents the methylation numerical matrix of the training set and y_train represents the phenotypic information of the training set.
[0347] 3. Test using test set data: Test the above model using the test set data. Command line: test_pred = model.predict(test_df), where test_pred represents the prediction score obtained by the SVM prediction model constructed in this embodiment for the test set samples, model represents the SVM prediction model constructed in this embodiment, and test_df represents the test set data.
[0348] The ROC curve of this 7-marker combination model is as follows: Figure 16 As shown, the AUC of the established model is 0.859. In the test set, when the specificity is 0.615, the sensitivity can reach 0.839 (Table 16), which can achieve good differentiation performance between pancreatic cancer patients and healthy people.
[0349] Table 16: Performance of the model combining these 7 markers
[0350] training set 0.8510 0.6667 0.8519 0.4124 test set 0.8586 0.8387 0.6154 0.4124
[0351] Example 14: Model Construction and Performance Evaluation of 7 Biomarker Combinations SEQ ID NO:6,12,20,26,24,47,50
[0352] To verify the predictive performance of different biomarker combinations, based on the 56 methylation biomarker groups of this invention, 7 biomarkers SEQ ID NO:6,12,20,26,24,47,50 were selected for model construction and performance testing. A training group and a test group were established, with 117 cases in the training group (samples 1-117) and 57 cases in the test group (samples 118-174).
[0353] Using these 7 methylation markers, support vector machine models were constructed on the training set for both groups of samples:
[0354] 1. Divide the samples into two parts in advance, one part for training the model and the other part for testing the model.
[0355] 2. The SVM model was trained using the methylation marker levels in the training set. The specific training process is as follows:
[0356] a) Use the sklearn package (0.23.1) of Python software (v3.6.9) to build the training mode of the cross-validation training model. Command line: model=SVR().
[0357] b) Using the sklearn package (0.23.1), input the methylation numerical matrix and build an SVM model, model.fit(x_train,y_train), where x_train represents the methylation numerical matrix of the training set and y_train represents the phenotypic information of the training set.
[0358] 3. Test using test set data: Test the above model using the test set data. Command line: test_pred = model.predict(test_df), where test_pred represents the prediction score obtained by the SVM prediction model constructed in this embodiment for the test set samples, model represents the SVM prediction model constructed in this embodiment, and test_df represents the test set data.
[0359] The ROC curve of this 7-marker combination model is as follows: Figure 17 As shown, the AUC of the established model is 0.857. In the test set, when the specificity is 0.846, the sensitivity can reach 0.774 (Table 17), which can achieve good differentiation performance between pancreatic cancer patients and healthy people.
[0360] Table 17: Performance of the model combining these 7 markers
[0361] training set 0.8695 0.6984 0.8519 0.5177 test set 0.8573 0.7742 0.8462 0.5177
[0362] Example 15: Model Construction and Performance Evaluation of 7 Biomarker Combinations SEQ ID NO:1,19,27,34,37,46,47
[0363] To verify the predictive performance of different biomarker combinations, based on the 56 methylation biomarker groups of this invention, 7 biomarkers SEQ ID NO:1,19,27,34,37,46,47 were selected for model construction and performance testing. A training group and a test group were established, with 117 cases in the training group (samples 1-117) and 57 cases in the test group (samples 118-174).
[0364] Using these 7 methylation markers, support vector machine models were constructed on the training set for both groups of samples:
[0365] 1. Divide the samples into two parts in advance, one part for training the model and the other part for testing the model.
[0366] 2. The SVM model was trained using the methylation marker levels in the training set. The specific training process is as follows:
[0367] a) Use the sklearn package (0.23.1) of Python software (v3.6.9) to build the training mode of the cross-validation training model. Command line: model=SVR().
[0368] b) Using the sklearn package (0.23.1), input the methylation numerical matrix and build an SVM model, model.fit(x_train,y_train), where x_train represents the methylation numerical matrix of the training set and y_train represents the phenotypic information of the training set.
[0369] 3. Test using test set data: Test the above model using the test set data. Command line: test_pred = model.predict(test_df), where test_pred represents the prediction score obtained by the SVM prediction model constructed in this embodiment for the test set samples, model represents the SVM prediction model constructed in this embodiment, and test_df represents the test set data.
[0370] The ROC curve of this 7-marker combination model is as follows: Figure 18 As shown, the AUC of the established model is 0.856, and the sensitivity can reach 0.742 when the specificity is 0.808 in the test set (Table 18), which can achieve good differentiation performance between pancreatic cancer patients and healthy people.
[0371] Table 18: Performance of the 7-marker combination model
[0372] training set 0.8492 0.6508 0.8519 0.5503 test set 0.8561 0.7419 0.8077 0.5503
[0373] This study investigated the differences between the plasma of individuals without pancreatic cancer and those with pancreatic cancer by analyzing the methylation levels of relevant genes in plasma cell-mediated DNA (cfDNA). Fifty-six significantly different methylated nucleic acid fragments were identified. Based on this group of methylated nucleic acid fragment biomarkers, a pancreatic cancer risk prediction model was established using support vector machines. This model effectively identifies pancreatic cancer with high sensitivity and specificity, making it suitable for pancreatic cancer screening and diagnosis. sequence list <110> Jiangsu Kunyuan Biotechnology Co., Ltd. <120> DNA methylation biomarkers for pancreatic cancer diagnosis and their applications <130> 207241 <160> 56 <170> SIPOSequenceListing 1.0 <210> 1 <211> 501 <212> DNA <213> Homo sapiens <400> 1 agtagggcgc catgaaggcc agaccgcggc tgtgcgccgc cgccgcggag taggccaggc 60 gcagggggct gaggccgagc ggcgcgccca gcgggtaggc gcccgcgtcg gcaccgaagt 120 gactggcgtt gggctgcagc ggcgagaagg ccgagcggct gctcagcgag cccagcgcc 180 caggcgccat ggcgccggcc agcaagggtc tgtggtgcgg aggtgcggcg ggccccgcct 240 gcagcggcgc aggcagccca ggccccccgg cggcggcggc ggcggcggcg gcggcgtcga 300 cgcggctggg ccacgcgtcg tctgcagctg ctgcagcacc ccggcggcc ttatctgggg 360 gcgccgcagg gcccaggccg gccgccaggc ccccacggtg gtggttcagc acctgctcga 420 tgcctgcac cacgtcgccg ccgcagccct gcaacaccag ctccaggacg cctcgcggt 480 ggcctgggaa cacgcgtgtc a 501 <210> 2 <211> 542 <212> DNA <213> Homo sapiens <400> 2 60. cctgccccc catctttcgg gggcactcaa accctcttcc cctgagctcc gtggcagccc ccgaacaccc tcatcgccccg ctgccccctc cccgccgccg ctaccaaccc cgaggaggga tgaccctctc cggcggcggc agcgccagcg acatgtccgg ccagacggtg ctgacggccg 180 aggacgtgga catcgatgtg gtgggcgagg gcgacgacgg gctggagag aaggacagcg 240 acgcaggttg cgatagcccc gcggggccgc cggagctgcg cctggacgag gcggacgagg 300 tgcccccggc ggcaccccat cacggacagc ctcagccgcc ccaccagcag cccctgacat 360 tgcccaagga ggcggccgga gccggggccg gaccgggggg cgacgtgggc gcgccggagg 420 cggacggctg caagggcggt gttggcggcg aggaggcgg cgcgagcggc ggcggggcctg 480 gcgcgggcag cggttcggcg ggaggcctgg ccccgagcaa gcccaagaac agcctagtga 540 ag 542 <210> 3 <211> 577 <212> DNA <213> Homo sapiens <400> 3 atttgttctg cctgatgaaa gcaaaagctc gaactcccct cagggcgcga ggtgtgagac ccttgggttc catttgcatt tctggtttgt cgttggcggg ttcctgattt gtttttgttt 120 tgttttggtc tgttctgttt tttggggggt gtctttcacc agggccttcc cggttagccc 180 agggtcccca catttctcca ggatgtaatt agagctaaga acagccgcca tccctcaggg 240 ttccgggtcc cgggtttcca gggtcccggg tttccaaggc cccgcgataa ccccgggcgc 300 acgcggcgcg atgcggcgag gcgaggcgag gcggtggggc cagcgcggag ccccaggcgc 360 gagaacagga actcgggctg gcacaccgag gcctcgcagc caagccgcgc ctgacccgtt 420 cgccgttccg gccccgcggc gcctccaagg ccgggccgag gggccgaggg gccgagggcg 480 ggcagacgcg gccacggcct aattctgact tctgaaggtc accgaaactg cgctgttttt 540 ccagagatgg gttgaagaga agagatgcaa tcccagt 577 <210> 4 <211> 501 <212> DNA <213> Homo sapiens <400> 4 atccgctgaa cgatgtccta cttcgctcgt ccttgctctc gccgctgctg ccggagccga 60 agcagagaag gcagcgggtc ccgtgaccgt cccgagagcc ccgcgctccc gaccaggggg 120 cgggggcggc cccggggagg gcggggcagg ggcgggggga agaaaggggg ttttgtgctg 180 cgccgggagg gccggcgccc tcttccgaat gtcctgcggc cccagcctct cctcacgctc 240 gcgcagtctc cgccgcagtc tcagctgcag ctgcaggact gagccgtgca cccggaggag 300 acccccggag gaggcgacaa acttcgcagt gccgcgaccc aaccccagcc ctgggtaggt 360 gagtgcctcc gcagccccgc cgcccgccgt ggggtcgggg acagggagaa gggagtgcct 420 gcctggtctg cgccccccgc ctgtcagccc ttgcctcgag gctctggggc acccaactcg 480 tcgactcctg acaccgcagc g 501 <210> 5 <211> 589 <212> DNA <213> Homo sapiens <400> 5 cttttcaccg ggtgtggctc gtctgagctc ttgaactgaa gccagcggac accacccgtc 60 ggcgcctgct ttcctggggc gtgggctcct ccccctgtgc agaccgcgag gggagacggt 120 gcgggcggcc gggagcgcag ccctccggga ggcgggtcat ggcctgggcg ccgcccgggg 180 agcggctgcg cgaggatgcg cggtgcccgg tgtgcctgga tttcctgcag gagccggtca 240 gcgtggactg cggccacagc ttctgcctca ggtgcatctc cgagttctgc gagaagtcgg 300 acggcgcgca gggcggcgtc tacgcctgtc cgcagtgccg gggccccttc cggccctcgg 360 gctttcgccc caaccggcag ctggcgggcc tggtggagag cgtgcggcgg ctggggttgg 420 gcgcggggcc cggggcgcgg cgatgcgcgc ggcacggcga ggacctgagc cgcttctgcg 480 aggaggacga ggcggcgctg tgctgggtgt gcgacgccgg ccccgagcac aggacgcacc 540 gcacggcgcc gctgcaggag gccgccggca gctaccaggt gaggcgccc 589 <210> 6 <211> 583 <212> DNA <213> Homo sapiens <400> 6 atccaccgtc acactctctc cgagcagcca gctccccgct taacgggga attgaagcag acagcctttg tctaaacact tcttttgccc agaattctt aattttccta tttgaatgtt taataaggtt tggggtgcag cagcttcctt ttaattgtga cggtgcggcc gcttgggcgt gatcccttgg ctggggctgc agggggcccg tcctccaggg gcgcagaggg aaggaccagc 240 gtttccaagc cgggctctgg ccgccggcgc gagagcgagg ccaaggtctg ggggcagttc agggggaccc cgaagtcggg acggcccaga aacgctttgc ccacagccac cgccctttcc 360 tttgtgagtt tccccaaagc cgtcggtgcg acccggcgcc gactctcctc ctcttctccc 420 tgcgagggcc cgcgccgcc gggcccagtc ctgggggata gatccctcgg ggcccaacgg 480 ctgggccacc gccggtctcc ggccactgct gcgaggacag gcgctgccta actaatttct 540 cctctaaggg ggctgtgcgt gcgtctcctt cccaactgat gtc 583 <210> 7 <211> 542 <212> DNA <213> Homo sapiens <400> 7 cagaaggtgt cacactctgg gatctgttcc gcagggaggc cacaggtgcc aggagacgcg 60 ggagagactg gctcctgcca agaacatttc tttgcattgt tcagtgcggt ttttatttt 120 tatttttgac tgtttgtttg cctaagagat gacttccctt gccagaaaaa aaaaagtgtt 180 gtaaaaataa aaggaaacgg gattacgatg taaaagacga atagataaac ccggttccgc 240 agatctgcgg cgcgcgcgc tggcgacctc ggatacattc attgaagttg ccgcgcactc 300 gtacccgggt tcacctcgcc cctcgctcat tcctcccgac caaggcccat ggtcagaggt 360 gtcctcgccc cgcggccgtc agagggcgcg gcctacactc gaatgccggc cgagccctcc 420 acgcgctcgg aacttgggct tcccggtgca gcctccccgc gatcgcaatg cccgctgcct 480 ttcccgagcc cagtccggaa cccgcctctc tcggggacct tgacctcgcg cggacctcgt 540 cg 542 <210> 8 <211> 501 <212> DNA <213> Homo sapiens <400> 8 cagctccggc tctgagcgtc tccagtcagg cgaggcggat aaatccttcg caaaaccctc 60 ttggaaattg ccgccgcttc ctgagccatc agtcccagcg ggtacgttat cgagtagcac 120 aaacagttgg atttttccct caagaaccga gtctggacgc ggagatggag ccaagtgtgg 180 ctgcattttc ggacccggaa atccgttggg cactgaagga cttttcgaac cctgtagcgc 240 tgttgcttcg cggtccatcg tcgccgctgc agacggatgc gctccccggc ggctctacgc 300 cctccagtcc cggccaggcc tctgggctgg gagccgagcc gtctcgggcc ctccggcgcc 360 gcgttttcta gagaaccggg tctcagcgat gctcatttca gccccgtctt aatgcaacaa 420 acgaacccc acacgaacga aaaggaacat gtctgcgctc tctgcgcagc gcttgggcgg 480 cgcggtcccg gcgcgcgggg a 501 <210> 9 <211> 522 <212> DNA <213> Homo sapiens <400> 9 aggaggagag aggtgagga aaggctaagt cagagtccgc gaccttgccg gctctatacc ttcagagggc tgcagagcgc gcgcgtcaag tccgcggaaa gttttactag tcagctcctc 120 cagcgcgcac agcggcgacg ttggacccgg acccgactct ggaagctgcg gcgcagaggg 180 tgctcggggg accatgcgcg gggctaggat gtctgcgatg cttaagagtg tccggggtgt 240 tcggggctcg cgtcccgagt tcatggtcgg ccggggctggg gcggtccggc tgtccgttgc 300 gctaggctcc gcaaacgcct gggccccagt gctcggctcc caatccggggc ccccagcctc 360 ggacccgccc ccggctctgg gcccgagtcc cgtgtgcccc tcctcctgcg cccccacctc 420 tccaccccgg gccgcggtgg atctggagct cctagatgtc cggggagggt atttctacag 480 gctggggcag gcgcggagag cagaagccga ccaaactacc ca <210> 10 <211> 501 <212> DNA <213> Homo sapiens <400> 10 caggtgctgg agttggagcg gcgcttcctg cgccagaagt acctggcctc tgcggagagg 60 gcggcgctgg ccaaggcctt gcgcatgacc gacgcacagg tcaaaacgtg gttccagaac 120 cgacgcacca agtggcggtg aggcgcggcg cgggcgaggg cggactgggg ttcccgagca 180 gggcctggtg agaagcgacg cggcgggcgc cccgctgacc ccgcgtctcc ctcccttagg 240 cgccagacgg cggaggagcg cgaggccgag cggcaccgcg cgggccggct gctcctgcat 300 ctgcagcagg acgcgttgcc acggccgctg cggccgccgc tgcccccgga ccctctctgc 360 ctgcacaact cgtcgctctt cgcgctgcag aacctgcagc cctgggccga ggacaacaaa 420 gtggcttcag tgtccgggct cgcctcggtg gtgtgagcga cgcccgtccg atcggcgtgg 480 agcgccgggc ccggagcggt g 501 <210> 11 <211> 501 <212> DNA <213> Homo sapiens <400> 11 cggactgtgg cccagcccac agaccagggc ccgaaattga ggtggggggc gtactctgtt 60 tgtcttcccg aaggatgcgg cgcgtggaag gagatgcgct gacttgttcc aacccataac 120 cttcgctcg ggtccccatg tgcgggcaga agaagtcaga gcggaacagc ctagtgcact 180 ggcagggctc attgtctggg aagacaccga ggtctaggca gctgggactg cggagtggag 240 gcaaggccgg aggcggccgg cggctttgtg gaagttttcgc gccgccaggc cctgcgcgcc 300 gcacggggcg gtggagttct tgggcagccc ccggcgcttg gcccacgcct ccgcttcccg 360 cgtgtgggaa actcgagcac cctacaggca ccagggtaaa ctgcctgtgc ctggcccggt 420 gagggtcgct cccccaggcc ccgtctccgc ccgaggactg caggcctagg cctgcgggga 480 gatcctgaga ccgcggtgtg c 501 <210> 12 <211> 503 <212> DNA <213> Homo sapiens <400> 12 ggcccgaaat tgaggtgggg ggcgtactct gtttgtcttc ccgaaggatg cggcgcgtgg 60 aaggagatgc gctgacttgt tccaacccat aacctttcgc tcgggtcccc atgtgcggc 120 agaagaagtc agagcggaac agcctagtgc actggcaggg ctcattgtct gggaagacac 180 cgaggtctag gcagctggga ctgcggagtg gaggcaaggc cggaggcggc cggcggcttt 240 gtggaagttt cgcgccgcca ggccctgcgc gccgcacggg gcggtggagt tcttgggcag 300 cccccggcgc ttggcccacg cctccgcttc ccgcgtgtgg gaaactcgag caccctacag 360 gcaccagggt aaactgcctg tgcctggccc ggtgagggtc gctcccccag gccccgtctc 420 cgcccgagga ctgcaggcct aggcctgcgg ggagatcctg agaccgcggt gtgcgggcgc 480 cggcagcagg gcaaggcagg gac 503 <210> 13 <211> 504 <212> DNA <213> Homo sapiens <400> 13 cttacgcggc ggcgggcgtg aaggcgctgc cgctgcaggc tgcaggctgc actggccgcc 60 cgctcggcta ctacgccgac ccgtcgggct ggggcgcccg cagtcccccg cagtactgcg 120 gcaccaagtc gggctcggtg ctgccctgct ggcccaacag cgccgcggcc gccgcgcgca 180 tggccggcgc caatccctac ctgggcgagg aggccgaggg cctggccgcc gagcgctcgc 240 cgctgccgcc cggcgccgcc gaggacgcca agcccaagga cctgtccgat tccagctgga 300 tcgagacgcc ctcctcgatc aagtccatcg actccagcga ctcggggatt tacgagcagg 360 ccaagcggag gcggatctg ccggccgaca cgcccgtgtc cgagagttcg tcccgctca 420 agagcgaggt gctggcccag cgggactgcg agaagaactg cgccaaggac attagcggct 480 actatggctt ctactcgcac agct 504 <210> 14 <211> 507 <212> DNA <213> Homo sapiens <400> 14 tgaggcacga gcagggtgca gagccgccgc tggggggcgc gccggccgcc gccgccgagg 60 aggccgcagc cgctgcggct gccgcggctg ccgcggcaga ggccgcgctg ttgagccccg 120 cggcggccgc gggagcctgg tagagaccag ggtggcggaa gctacacagc agctccggcc 180 gagagtaggg gtgcgagagg gcgcggaagg tgtccagtgg ccggatggaa gtagcgaagg 240 gcgacgaagc cgcggccgcc gcgcctgagg ctgcagccgc ggcgcgcc gcgtgacgc 300 ccacgtgcgg gtagtagtgc agcggcacgt gcgagtggaa ggggtagggc aggcttccgg 360 tggcggccgc gtgcgtcatc atgtaggtgt agaagctggg gtcggctggg tgcggccagg 420 acatggccag gcgctgccgc ttgtccttca tgcgccggtt ctggaaccac acctgcgggg 480 agagacgcgc cgcagcctgg gttaggg 507 <210> 15 <211> 501 <212> DNA <213> Homo sapiens <400> 15 tggaagtagc gaagggcgac gaagccgcgg ccgccgcgcc tgaggctgca gccgcggccg 60 ccgccgccgt gacgcccacg tgcgggtagt agtgcagcgg cacgtgcgag tggaaggggt 120 agggcaggct tccggtggcg gccgcgtgcg tcatcatgta ggtgtagaag ctggggtcgg 180 ctgggtgcgg ccaggacatg gccaggcgct gccgcttgtc cttcatgcgc cggttctgga 240 accacacctg cggggagaga cgcgccgcag cctgggttag ggagcgcccc gtgttcccag 300 ctcctgtccc aggacctctg ccccttccgg acctctgaat ggcttggtct acttctctcc 360 gaccaagccc aaccccgagt accctgtggt ctcccagctg ggaaagtgtg gacggcagtg 420 tgtggaccgc cgtgggcaca ccgtcctcaa cgaagagggt cctctccccc gcgtccggct 480 gctgctgctc ctcaggcttt t 501 <210> 16 <211> 581 <212> DNA <213> Homo sapiens <400> 16 ggccagttgg ccgcgcttcc ccctatctcc tacccgcgcg gcgcgctgcc ctgggccgcc 60 acgcccgcct cctgcgcccc cgcgcagcct gcgggcgcca ctgccttcgg cggcttctcg 120 cagccctacc tggctggctc cgggcctctc ggcctgcagc ccccaacagc caaagacgga 180 cccgaagagc aggctaagtt ctatgcgccc gaagcggccg ctgggccaga ggagcgcggt 240 cgtacccggc cgtccttcgc ccccgagtct agcctggctc ctgcagtggc tgctctcaaa 300 gcggccaagt atgactacgc tggtgtgggt cgtgccacgc cgggctccac gaccctgctc 360 cagggggctc cctgcgcccc tggcttcaag gacgacacca agggcccgct caacttgaac 420 atgacagtgc aggcggcggg cgttgcctct tgcctgcgac cttcactgcc cgacggtaaa 480 cggtgcccat gctccccggg ccggtttggg ccgggatggg aggtggggtt caagggagag 540 tgtaagggga ggtgaaccgc ctgggggcgg gcaatagaca g 581 <210> 17 <211> 501 <212> DNA <213> Homo sapiens <400> 17 gtcggcagcc tcggcggcgg gggcgagatt ggcgggaggg gggcgcgggg ggggcgcggt 60 aagaggtggc ggcgggcaga gggtgtttt ttcttttcc ctccagagcc ggggtttgta 120 aaccgaggcc agagtgtccc cgtgggccga gcgcactttt ttcttgtccg ggtgcgctca 180 gtcactggtg cctgagga aacagtggag gcagcggggc aggtcgcctg gggcgtcggc 240 gattatattg cggccgagcc ggggcgcgcc gggaaaggcc gggagggcgg cggcgcgcgg 300 gggctgggcg aggccccgcg acccgcgagg gaggcggcgc gaagccgagg cggcgggcgc 360 aagagccggg catgagcgcc footgctga gcgcccgcgg ctgcctggcc tcagaagcga 420 cgcgcgagcg cgggcgggcg gcagcagcga cgtagcccgg cggtcccggc ggcgagagca 480 gccgccccac aggccccccgc g 501 <210> 18 <211> 515 <212> DNA <213> Homo sapiens <400> 18 gggtggggat gggggggtgg gggaggactc cattttcaga gcagggggaa ggctgtggag 60 gagcggggga tttccaaaat gcttgagggt tccggacctg gtggtgggcc cagaagaagg 120 agcacatttg gggatcccgc aagcctgggg tatgtgggtg tgtttgagga ggtgggtggg 180 agtgagcgtg tgcgccgggg agagggcggg agggag caagcgagct tgggagcgcg 240 cggggagggc cgcggggcctc ggggcgcgcc aggaagtgag cggcggaggc gaggggccta 300. actagtggcc gggcgctgac ctgcctgtcc tgtctgtttt gtctcgcagt gaaccccaac 360 tacaccggtg gggaccca gcggtcccga acggcctaca cccggcagca agtcctagaa 420 480. ctggaaaaag aatttcattt taacaggtat ctgacaaggc gccgtcggat tgaaatcgct cacaccctgt gtctgtcgga gcgccagatc aagat 515 <210> 19 <211> 512 <212> DNA <213> Homo sapiens <400> 19 ctggcgctgg cacgcttaat tctttttttcc cacattgcag aatcattccc accagccact cggagagtgg tgggaatctg tcttggttta atattctaa aatataagtt tcattgtccc ccaggttagc ccagccagga ctcattgcgc agtcctcctc gccttcctgg aggcgccgca 180 ggaagcggga agtcgcggct tggcggttgc tggggcctgtg ggatctgcgg gtcctgccca 240 gacctggagt cgcacagatc acggcgggca gtggctcagc gcctaggcgg ctccaggcct 300 cgaaggacca ggttggggtg ctcagggatc agagagggga ggtcgctctg ggtccgggtc 360 gcctgctacg cgccttttct gtctcagaag tggcggtgac tcggctgctg agtccgcgga 420 acgagccacg gaatggtggt ggtggcgggg ttttctgagg tgactggcca gagctgagag 480 tcgcggcttc cacctttggg ccggagcggg tc 512 <210> 20 <211> 558 <212> DNA <213> Homo sapiens <400> 20 tgggattgat tttggcccc cgctgcagca agttgggggc tggtgaggag tgtagcggtg 60 actgggggcg gagtgcggac tcgcatccgc tgtaccagga gcccactgcc acctcgggat 120 tttttttta acttggaatt tccatatgac aaaaaagaaa gaggttctc ctcaatctaa 180 cggagccatt aacatctatt aataacgccg acagggtaag taacggagcc gcgctcctcg 240 gggtggtcac cgggctgcgt ggtcctcggc cggcctcctg catccgctgc cctgtgcgc 300 tccgggccgg atgcgcaagg gcggcgcggg gaccaagcct ggctgccggc cgcctactcc 360 tccccttccc taaggtaagg ggtcgttttc acactcacca gagctcctgc gggctgagct 420 cgccccctcc cccgacttct ttgcggggca ttttctcttg ctggtgtatt acgtgtcatt 480 tctcacgggg cattgccggc cgcttttctg caactgtcct ttcggatttg gtgatctggt 540 ccggcacaga ggctctcc 558 <210> 21 <211> 548 <212> DNA <213> Homo sapiens <400> 21 gagagcaggc cttgcgggag tctggacccg aagggcgaga ctccacaggg ccaaggaaag 60 cggcctctgt cctccgttag tcttggggga gcagacgcaa gaggaggcaa gggcgccgcg 120 agctccccgg atgcactggt cccacaggcc gtgcccgagt ggagcactgc gaatggggcc 180 aagaaatttt ggcctttctc gccggacctg gctgcctccg cgggcctctc cgcctaccgc 240 gctcccgccg cggcccgact cccgcgggtc tccgcgccga acccacctgg ctcctatcgc 300 acgggacatt cccgacccac ccacgccgcg tcactgagcc tctgtaccga tacccggcgc 360 ctccgccagc agggcctgga cgcaccgcct cctttgacct cgggcttccc ccgcgctccg 420 ctgcttgggg cagactggcc ccgagaggga gccaccatct cccctgctcc agggtctcca 480 gggtccgaac ccgtgttggg atctgggtta ggattagggt ttggagcttg gagcctgcct 540 gttaggac 548 <210> 22 <211> 503 <212> DNA <213> Homo sapiens <400> 22 ctccagggat gcgccaagca cccttcggtt ttcccgggga gaattttccc cggcccgggg 60 actagggtct ggcgctgggg cgcccctcgg acctgcggga tcgcccctac actctggcgc 120 gctgagggcg gtgagcgagg gcgccaaggc acaggtgggg cgggagtcga gcgcggaggc 180 tcggggggcg ggacgcgggg cctgggagcg gccagggacc gcggcagcgc ctcagtgcca 240 gcctggcgcc cgcgactgcc tgccccagcc cctcagtggc ggcttgctct cttctctcgc 300 tccgaaccag acacagccgc tgccgctgcc gtccggcgcg ctacagactc ccgagaacag 360 ccctggctgt cagcgagcac cagccgcttc ctgtccccat cgcggagact ggaggggcgc 420 accacggcca tggagccaga ggcgcttcag gaggcaagag aagtccccgc gcgctccgca 480 gcccggcgca gctcatggtg agc 503 <210> 23 <211> 501 <212> DNA <213> Homo sapiens <400> 23 gacccacgcc cacctaggcc tccccgagcc tctgttgcat gccgacgggt ggctgaaccc 60 atcgacggcc gaggccttcc aggcctacgc tgggctgtgc ttccaggagc tgggggacct 120 ggtgaagctc tggatcacca tcaacgagcc taaccggcta agtgacatct acaaccgctc 180 tggcaacgac acctacgggg cggcgcacaa cctgctggtg gcccacgccc tggcctggcg 240 cctctacgac cggcagttca ggccctcaca gcgcggggcc gtgtcgctgt cgctgcacgc 300 ggactgggcg gaacccgcca acccctatgc tgactcgcac tggagggcgg ccgagcgctt 360 cctgcagttc gagatcgcct ggttcgccga gccgctcttc aagaccgggg actaccccgc 420 ggccatgagg gaatacattg cctccaagca ccgacggggg ctttccagct cggccctgcc 480 gcgcctcacc gaggccgaaa g 501 <210> 24 <211> 553 <212> DNA <213> Homo sapiens <400> 24 tggctgaacc catcgacggc cgaggccttc caggcctacg ctgggctgtg cttccaggag 60 ctgggggacc tggtgaagct ctggatcacc atcaacgagc ctaaccggct aagtgacatc 180. tacaaccgct ctggcaacga cacctacggg gcggcgcaca acctgctggt ggcccacgcc ctggcctggc gcctctacga ccggcagttc aggccctcac agcgcggggc cgtgtcgctg 240 tcgctgcacg cggactgggc ggaacccgcc aacccctatg ctgactcgca ctggaggggcg 300 gccgagcgct tcctgcagtt cgagatcgcc tggttcgccg agccgctctt caagaccggg 360 420. gactaccccg cggccatgag ggatacatt gcctccaagc accgacgggg gctttccagc tcggccctgc cgcgcctcac cgaggccga aggaggctgc tcaagggcac ggtcgacttc 480 tgcgcgctca accacttcac cactaggttc gtgatgcacg agcagctggc cggcagccgc 540 tacgactcgg here <210> 25 <211> 610 <212> DNA <213> Homo sapiens <400> 25 60. aaaagagaag tcggagttta gacagggttt taaaagtcag ctaaaggctc ccacattgca cctgtggtta acaaccacag gccgtgttgc attctttacc tggcactttt cgggatata caggagcatt taaaaaatag ataagtcaat gatgcactt aggggacat cggctgccgc 180 tgccgtcagc tgaatgtta gctatctacc gtcttataaa acgccaggaa aaacctctaa 240 accttagagc cggggaattt ttaaaaat cggaaccaaa tctccgtggc tcgtgcagc 300 gtgagttctg cagctcgggg gacgctgcag tgtgatgtgg tggagagc atgcttcacc 360 gctcctgcca tcctgacagc gccctccctc ccggcctcag cctcctggtt cgccaaccg 420 gaggactgaa ttatgcta gctggtctct ggggcgcctt ccagctctga cattcccgcc 480 tagatagat ctcccgaag gtttcgcaga cagaccagag gggaccgagc cgggaaggcg 540 agacagggac aggcgagaga cgctgctccc aactcgcaga gggagaaagc gtgtatcccg 600 ggctgccggg 610 <210> 26 <211> 506 <212> DNA <213> Homo sapiens <400> 26 gaacttctgc ccttcccgct actgcaccc caagcaggga tgcactggga tgcgtggcag 60 gggcgggatc tcctgggagc gtctcagccc agcagggagt ggggaagcaa gagggaggc 120 ttaccttcct cggtggctgg caggaggtgg tcgctgctag cgagggggat gcaaaggtcg 180 ttgtcctggg ggaaacggtc gcactcaagc atgtcgggcc aggggaagcc gaaggcggac 240 atgaccgggg cgcagcggtc cttcacctgc acgcagagcg agtggcatgg ctggatggtc 300 tcgtctaggt catcgaggca gacgggggcg aagagcgagc acaggaactt cttggtgtcc 360 gggtggcact gcttcatgac cagcgggatc caagcgccgg cctgctccag cacctccttc 420 atggtctcgt ggcccagcag gttgggcagc cgcatgttct ggtattcgat gccgtggcac 480 agctgcaggt tggcagggat gggctt 506 <210> 27 <211> 510 <212> DNA <213> Homo sapiens <400> 27 attcgagtc tttgccctt ttcagtctaa gacgtgggct ttctgcaaag cctccccctg 60 ccagcgagct ctcggagcgc ggagccttta gaaattgagg ggtttactgt caaaatgaaa 120 atttcacttc aaattacctt ggctgatgct cgctcgccag gccgggggct cccgccgcag 180 ccttttgaca ggcacatgag ccgcgagctt ccgaacctcg ataatatcat ctcgagcgcg 240 aaagtcaata cggtgacagc gcgcggccgg attacaatcca attacgctcg gctgcccggg 300 cgctcctggg gctcggggtc cggcggccga gggtccccct cagggcccgg tccaggccct 360 gtcgccaggg ttcagggcag gccccaccac gcgggggact ttggtggccc aggggtcccc 420 acgaggccgc agtccgggtc cgcccagccc caggctccta gaggaaagcc gagcctagtg 480 agtccctcca aggccgcccg cccgcaagac 510 <210> 28 <211> 501 <212> DNA <213> Homo sapiens <400> 28 acttgcgtta agttcggctc aggctactgg attgggcagg accagctaac ccaggtcccg 60 aggggcagtg tgtcacagac tgcagcccac tccaacctcg gctcctggag aaggggcgtc 120 gaatctctct tgggcatggg agggaaagac attccgagtt ggctgggcgg agtggcagcc 180 ttgagagtga cgagtgacag caaagcctcg tcctagcaag gccttttacc aacagcgcgg 240 catgcccttt cgaggagagc gccaggccct cgcactttgc aagtcaagag agcaaagaaa 300 gcggggacag ggcgcgtaat cgcaatgtcc ggtcgcgcgt gtgcacgtgt ctgtgtttgc 360 atgtgtgcgt gagcatgtgc acctgctcaa gtgtaaatgt gtctgttggc agttggggtc 420 taagtacctg agaatgtgtg tcttctgttg cttaggaga ttaaaatgtc ttttcccagt 480 attgagctac attgaggaaa c 501 <210> 29 <211> 555 <212> DNA <213> Homo sapiens <400> 29 aagtccttgg actcggccga cagccgggcc atgttggctg tggagagagc ggacaggtgc 60 ggcggcggcg gcatctggca gatggtgcag gggcagggca gaccagccca gtgctggaag 120 ccgctgccca gctgcagcgc gggcggcgtg gagggcgcct tgagtagcga gtggggaggc 180 cggatggtgc cgatggcggg aagtgaggcg gcggacagcg gtgacgaggc gttgccagat 240 gagagcgcgc cgcccaagat ggggtgcacc gggtgcacgg agttggccgc gtgcgcgggg 300 tggccggccg agtggcccac ggtcccgcag tgaaaggccg agtggtggcc cccatagatc 360 tcgccaacca gcctcttcat ctcctccagg gagctggtga gcatgaggat gtagttctg 420 gcgagcagga gtgtggcgat cttggagagc ttgcgcaccg acggcccatg cgcgtagggc 480 atgacttcgc gcagcccgtc catggctagg ttcaggtcgt gcatccgctt gcgttcgcgt 540 ccgttgatct tcagc 555 <210> 30 <211> 501 <212> DNA <213> Homo sapiens <400> 30 tcaggccagg aggtttctgg aaggaccggt gctgtctccc cgaacatcgt ggtctccccg 60 aacatcgcgg cctctccgaa catcgccctc tctccgagca acgcgatctc cccgaacatc 120 gcggtctccc cgaaaatcgc gatctccccg aacattgcca tctcaccgaa catcgcgatc 180 tcgccgaaca tgcccggctg aaggcactca gttcccctcc gcggctcctt tccgccgggt 240 ctgattcctg cggctgctgc ttgccccgca ggccaggagg cttctggtag caccggcgcg 300 atgcccccga acatcgcgtt ctaccccaac atcgcgatcc ctccgaacat cgtgatcccc 360 cccgaacatc gccgtccccc cgagtaacgc ggtctccccg aacatcgcgg tccccccgaa 420 catcgcggta cccccgaaca tcgccgtctc cccgtacatt gcgatccccc gaaacattgc 480 gatctccccg aacatcgcga t 501 <210> 31 <211> 516 <212> DNA <213> Homo sapiens <400> 31 ttggccagcc gcgcccggac tcctcagagc tggcgcaaac tccgtcctcc aaaactcggc 60 tctgggaggc ctaagtgact ccgaagccgg cggcagccgc ggcagcggcc gtggtggtgg 120 aagagctctt ttccccgaca gtgccactga tcgctcttca ctggagctgg aaacagcctt 180 cgcggaaagg accggagcat gcgttagaag cagagggagc ttggtgaagg gctcggctgg 240 aaggaggaaa cgccttctg cagtgcgcgg ccagcccgcg ggggacaccg gcttgctgga 300 ctgcaggggc ccgtgccacc caggaagtga cctgcgggtc actcagccgg ggcgctgggc 360 gagcgcggga cggcccggag aattccgtgc ggctgcgacg ggaaaaggac gaggggtctc 420 tgtacccgac gctgccactg gcccaaagga atttacccg cgagcgccca ccccacccta 480 gcttgatgct tacgcccgca acaaaacagg aaacca 516 <210> 32 <211> 516 <212> DNA <213> Homo sapiens <400> 32 agacttcgaa ggcagccgga gaggagaggg cccaccgagc actacggcgg gtgcgcacgc 60 cccggggcgc tcggcaggac gacagtctgc acagcccgaa ggcggaaacg agcatcaact 120 gcacaaagtc ctggggtcct ggagcatccc ctccgcgtcc ttcctccctc tggggctggg 180 gacagccggg atgtcccagg ctgaggtggc caccagccga gcgcggctgc taggacgctg 240 gcgtggggag cgcggcgcgg aactacggac agtgagccct ggcgctcgct gccctgcgcc 300 ttaatttgct ggcggcggcg atcccggagg cccgcagcca gtcagcgccg tctcacgtca 360 ccgcttcctg attccgccgc cgggggcggg gccgcgggcc gggcgcggag ggcgcgccca 420 gggtgcggcg cccgcgtggc ctgtcgcccc ggctgttcgg taccccagca caggttcagg 480 gaaaagggtg ccaccactag gctgacgcag cagcca 516 <210> 33 <211> 501 <212> DNA <213> Homo sapiens <400> 33 agcgccggcc gccgcatccc gtgcggggcc gcggcgcgat gctgcgctgg aatgaggaag 60 cgcggcggcg aggggagggc ccgggcgcgg tgcgcgcggg ggtggcggcg gcgcgccgag 120 cgggcccggc gcgggcgagc gggctgcagc cggcggcggc gccagcaggt acggcccgca 180 cccgccgccg ccccggcggc ctttgggggc tgagccggag cccggcgcga ttgcaaagtt 240 ttcgtgcgcg gcccctctgg cccggagttg cggctgagac gcgcgccgcg cgagccgggg 300 gactcggcga cggggcgggg acgggacgac gcaccctctc cgtgtcccgc tctgcgccct 360 tctgcgcgcc ccgctccctg taccggagca gcgatccggg aggcggccga gaggtgcgcg 420 cggggccgag ccggctgcgg ggcaggtcga gcagggaccg ccagcgtgcg tcaccccaaa 480 gtttgcgggg tggcagggcg c 501 <210> 34 <211> 517 <212> DNA <213> Homo sapiens <400> 34 agagctcccg gagggcttgg ccggccaccg ccgcgcggcg ctgctcgggg actgctactt 60 tgcaaggcgg cggctgcccc tgcggggttc gggttgcagg gtcaagtgtc acgtcctccg 120 caatctccaa tattcctgta atgtatttaa atggacgaat tcattacgcg gggccgtgtg 180 aatggggcga ggccgcgagc gcggcgcgat cagtagcgcc cactaacagt tcgttctgca 240 cggcggagcg cgagaccgcg gacccacgga agccccctca atggtgtttg cgtcctcgcc 300 gccaccggct tggtagggtc ctttagggaa ggaggaagag ttcaggcacc cggacagatc 360 ctaatggtct ttctgatttt tctttccctt cggtccgctt tccccgcgac ctcctccacc 420 ctcagtccgc ctttcaaacg tcgtccgcgg ggatggctgc gcgatggaga aattggtctc 480 gtccagagac gcgcgcacag ccgtccccgc gcacacg 517 <210> 35 <211> 562 <212> DNA <213> Homo sapiens <400> 35 tctcagccac ctgattgatt tctctctca ctccacccgc acccagtctc cgggtccagg 60 cctccagctc cctcacttct ggctctctc acctgaatt ttctccttat atttttctt 120 tctctccg attggcagtc ccgcttctcc gagtggagtc gctcccgcc tctcgcgtcc 180 cccctggct gcgctgcgac ctgcgaactc ccccagtttc cctcatctgc acaccctggt 240 gtagaccgac cgtgcgcgcc gggcccacgt gcagcctggg gactgcaggc tgggagctca 300 cggccatctc tcggccgcc tcaccgcagc tcccctgtca cccggcccc tgtgaggagc 360 tctgttcccg cgctctcata taagcgccgg cacacagtag gcgctcaagg cctgcagaat 420 480. gagtgagcaa atatagctca gacacctact gaatgaagt cggcaggttt gactagatcc tggaatttaa aatttactga gcgccaccca tgtgcggggc tccacagagg tgatcctgga 540 aggaggcagc gttgtggggg tg 562 <210> 36 <211> 503 <212> DNA <213> Homo sapiens <400> 36 60. ctcagtgata accgaagagc tactctgaaa tgcccccctt ttcctggtgg tgcccgccag ccggcagggg aaagcccgag ggacctccca gctccttccc ggatcgcggc ggaggtgtga 120 gcgatgtgtt gattattcat atttttaccg agcgcatact ctgctgcggc cggcgccgcc 180 240. acatttcaca cgtacactga cgtacccaca tgcacaagcg ctcactcggc cccgcacgca agcagcgccc cgcgcgcccg gggccctcct cggataaggg aggggtgaca aaagtctccc 300 360. gctcactgct gcctacccac ccccaacccg gctgcctttt cctccaggcc cccacaaaca cccttggctt tcagatccaa ctttcttcct catatatac tagtcaccgc gactcccgcc 420 tcccggattt gaggatgggg gagactttgg cggcggggggt cagctgcaaa tatggcacca 480 catcattt agc 503 <210> 37 <211> 701 <212> DNA <213> Homo sapiens <400> 37 agtattagca tagagaatcc agtaatgtgt cgacacaag cagatagttc ccaaaatgcc 60 aaacctgtttc aaaaagatg aaaacaccaa taaacgaaaa gtagaaaaac ctagtggac 120 gcatcacag atgctgaaaa ggcattcct agaagtcggc agccaactt ggtattctt 180 gcgtgtgata aaggcagccg tctgttctgc tcagaagggg tttcctaca ggaggggccg 240 aatgcaggcg tcacacc gccgccccag gtcgtacacc taggccgtcc gggctgtccc 300 agagccgcag gccccgcatc atccgcgtcc ttagcgcggg gcgcggagcc cgcagccagg 360 tgcggccgag acccgcgcgc cagggaagc ggcgcagcgg acggcgggaga aggctggtgg 420 gtacaggttg cctccgggcc ggagcgccca tgcagggcga gctgcgctcc gcacaaaatt 480 gcggtggggg cgccagaccg ccttgctccg cccctgagcg gggcgccccg gcccacccc 540 ctgagggagg agggtccagg tgccgcagac tcttagcccc tggcccggcg tccccccggc 600 aggttctggc actcctcgtt ggtaagcccc gttatttcgt gcgcagtgtt tacagaatat 660 aaagttcttc aggaaacgat gttataggag aaacgcctgg a 701 <210> 38 <211> 505 <212> DNA <213> Homo sapiens <400> 38 tcccccaacg ccggcgaata attttaaagc aaaggaggcg cggccaggtg ggctcccaag 60 ctccgcgcag acccttgggc cagccttggc cgctacccga gcgcctctcc accagacctt 120 ggagggaagt tgggggaagg gcgggagagc accggcgccc agggcgcagg ggccagagcg 180 agcctggcgt tccgccgcag ccggctgaga ctcggcgacg cgggggctgt acctgtggct 240 gcggggccga cggccggctg cagggcggct ggctctcccg cctcgagact aggcgcactc 300 ccatccccgc cgcatgttct ccacgcgggc tccagcgcgc tcaccaccgc caccgccgtc 360 gtctcggctt tatttaccca gcccggcgcg cgccgcccgg gaacaggaat agcgaggcct 420 tctcatgttt cctgactgcc ggtcccagcc ggcgaacatc ctgcgggcgc ggtatccacg 480 ttcccgggcg ggtggagagg aagcg 505 <210> 39 <211> 549 <212> DNA <213> Homo sapiens <400> 39 atcccagtaa gctctagcac ccggggcgcgg gtaacggga gcgcagaacc aaatccccag cgcccaggtc acctccccag acccagcctt gcagggacca gggctttagg gctcacggac 120 ccaacggcca ggtcagaccg cgaaccgggga ggagcgcggg ccccacccta aagaggggcgc agccgggagc tggggagcgg gtgccgcgct ccagagattg tgtcgtgggc gccgtcctag 240 tggcggggag cgcacctccg agggggcatg agatcggaga aatcccttac gctggcggcg 300 ccggggggagg tccgtgggcc ggaggggagg caacaggatg cggggagactt cccggaggcc 360 ggcggggggcg ggggctgctg tagtagcgag cggctggtga tcaatatctc cgggctgcgc 420 480. tttgagacac aattgcgcac cctgtcgctg tttccggaca cgctgctcgg agaccctggc cggcgagtcc gcttcttcga ccccctgagg aacgagtact tcttcgaccg caaccggccc 540 agcttcgac 549 <210> 40 <211> 506 <212> DNA <213> Homo sapiens <400> 40 ccacgttggc cccatggcgg gagcggaggg cgtaggggaa ggagaggcgc gcgaggaggc 60 tgcggctgcc gcgaggttg cgccaactct cccgccgcgc gagcgagccg aggcgcgctg 120 gaactagaga cccggcatgg agtgctgagg ggagggggga gccgtaaaaa agccaaagca 180 agccctcgac tcgcaagcac gcccccctcc tctccccagc gcactggtgt ttctggcggg 240 tgcctggcgg cgacgcgtcc aatcgcagcc cggcgcgggc gctaggtgac aggcggcgga 300 gcgcgcagac ccggctcccc gcgtcctctg aagaagggac tcgcgaggga gggagggagg 360 gaggcgggc ggcccggcgc cctgccgag gccggggatg ctcatcgttg cccagagttg 420 gcccgaggag ccctctccgt tttcccaata cttttccctg catcagtgca gccatccccg 480 ccgccttgtgt ctctccaact tttcca 506 <210> 41 <211> 560 <212> DNA <213> Homo sapiens <400> 41 ggctgcgcgg aagcagcggt gacagcagtg gctggactcg gagttggtgg gagggttagc 60 ggaggag agccggcagg cggtcccgga tgcaagtcac tgttgtccaa ggtcttactc 120 ttgcctttcc gaggggacaa cttccctcgg gctccagccc cagccccgac cccaccagag 180 gtcgaagctg tagagccccc tcccccggcg gcggcggcgg tggcggcggc agagaccgaa 240 gctccagtcc cggcgctgct ctttgacccc ttgaccctgg gcttgccctc gctttcgggc 300 catgacaggc ggctacccgc gcccttgccc ccgccggctt tggctccact cgtggtcacg 360 gtcttgcaag gcttgggagc cggcggagga ggcgccacct tgagcctccg gctgccggtg 420 ccagggtgcg gagaggatga gccagggatg ccgccgcccg cccggccttc gggctccggg 480 ccgccccagc tcgggctgct gagcaggggg cgccgggagg aggtgggggc gcccccaggc 540 ttggggtcgg ggctcagtcc 560 <210> 42 <211> 586 <212> DNA <213> Homo sapiens <400> 42 gggttcgaat cgaaaatgtc gacatcttgc taatggtctg caaacttccg ccaattatga 60 ctgacctccc agactcggcc ccaggaggct cgtattaggc agggaggccg ccgtaattct 120 gggatcaaaa gcgggaaggt gcgaactcct ctttgtctct gcgtgcccgg cgcgcccccc 180 tcccggtggg tgataaaccc actctggcgc cggccatgcg ctgggtgat aatttgcgaa 240 caacaaag cggcctggtg gccactgcat tcgggttaaa cattggccag cgtgttccga 300 aggcttgtgc tgggctggc ctccaggaga acccacgagg ccagcgctcc ccggaccccg 360 gcattaggcg ccagctgccg gctatctgcg gtctttttct ctctgcagac ccctcgatcc 420 tctttccttc gtctcacac tcaaaag acagactaga gacgttgaaa gagcctgccc 480 ttcacagag tcccagaaaa gggtgactta aggggaggag aagggagaa gaggcagat 540 tccgggtcag aaagaccca gataatttct gggcgtctg aaatat 586 <210> 43 <211> 501 <212> DNA <213> Homo sapiens <400> 43 aaagttcg attatttcac ctggcttgtc agtcacctat gcaggcgtct gagcccccgg 60 gtttccagga gccccgta taaggacccc agggactcct ctcccacgc ggccggccg 120 cccgcccggc cccccccg gagagctgcc accgacccc tcacgccc aagccccagc 180 tctgtcgcc gggttccttc ctctcctg gccacaatct tggcttccc gggccggctt 240 cacgcagttg cgcaggagcc cgcgggggaa gacctctcgt ggggacctcg agcacgacgt 300 gcgaccctaa atccccacat ctcctctgcc gcctcgcagg ccacatgcac cgggagccgg 360 gcggggcagg cgcggcccgc aaggaccccc gcgatggaga cgcaacactg ccgcgactgc 420 acttggggca gccccgccgc gtcccagccg cctcccggca ggaagcgtag gtgtgtgagc 480 cgacccggag cgagccgcgc c 501 <210> 44 <211> 528 <212> DNA <213> Homo sapiens <400> 44 agactccctc cttgggaacg tcgaactctc tctgccttgg ggagtggggc tcgataaagg 60 gtacctaggt cgcaccctgg caggggagca ctagagggcc gcgaggtccc gggtttcgcc 120 atcctgagac ccccgcgcgg atggcccagg aggggcgcgg cggccctgag tcaaggtggg 180 cgggggcagg tgcttccctc caccgcgttg tcctatgccg gcgcggtccc caccgcccga 240 cctagcccgg cgccggccga gcacggcggc cgcgcttcgc actccttcct cccaccgggt 300 ccgcaggccc ggcttcacga ttcccgggcc ctcgggcatg tgagggactt gagtgaatgc 360 agctccctca actcactccc gcaaaaccac agccaagagg gccttaagtc agagaacccg 420 gcctaggagc ctcccctaga gcctcggcgc gggccccttc cccttcccca catcggtcgg 480 ccgagggagc ctagagccgg tgggagacgg gcagcggcct ctcctgat 528 <210> 45 <211> 515 <212> DNA <213> Homo sapiens <400> 45 ggggctcgat aaagggtacc taggtcgcac cctggcaggg gagcactaga gggccgcgag 60 gtcccgggtt tcgccatcct gagacccccg cgcggatggc ccaggagggg cgcggcggcc 120 ctgagtcaag gtgggcgggg gcaggtgctt ccctccaccg cgttgtccta tgccggcgcg 180 gtcccccaccg cccgacctag cccggcgccg gccgagcacg gcggccgcgc ttcgcactcc 240 ttcctcccac cgggtccgca ggcccggctt cacgattccc gggccctcgg gcatgtgagg 300 gacttgagtg aatgcagctc cctcaactca ctcccgcaaa accacagcca agaggccctt 360 aagtcagaga acccggccta ggagcctccc ctagagcctc ggcgcgggcc ccttcccctt 420 ccccacatcg gtcggccgag ggagcctaga gccggtggga gacgggcagc ggcctctcct 480 gatcctttcc tgcggtcata caagttccta gggtg 515 <210> 46 <211> 517 <212> DNA <213> Homo sapiens <400> 46 gcaaagcctc ccaagtcgtc taggcagtta gggagctctg cgcatttgcc agcacggagg 60 tacctcccgg ggcagggaca caacacatcg cccgagagtt tgtcccagcg agcgccgatt 120 tcgtccgcga tgcaagtaac tgagatcggg agctgtcccc ggcagagcgc actcacctcg 180 gtcccaggtg gactgaagtc cagagcggcg ctgtgcagct ggaagggcgc gcgatagctc 240 aagttagagg cggccccggg gcgcggcgca ggacacaaga cctcaaactg gtacttgcac 300 aggtagccgt tggcgcgcag gtggcatcgc atctccttcc agcctgcggg ctcgacccca 360 ccggtggcct ggagtaccgc gcatctccgc gcggtgcagg agcgttgggg ctcctccacc 420 cactgcagcg tgtcgctttc gagaccgccg gggtcggagg acagccagga gaaaccccgc 480 aaaggctcgt tctccagggt gcagtgggaa cgcctgc 517 <210> 47 <211> 507 <212> DNA <213> Homo sapiens <400> 47 ccaggtggac tgaagtccag agcggcgctg tgcagctgga agggcgcgcg atagctcaag 60 ttagaggcgg ccccggggcg cggcgcagga cacaagacct caaactggta cttgcacagg 120 tagccgttgg cgcgcaggtg gcatcgcatc tccttccagc ctgcgggctc gaccccaccg 180 gtggcctgga gtaccgcgca tctccgcgcg gtgcaggagc gttggggctc ctccacccac 240 tgcagcgtgt cgctttcgag accgccgggg tcggaggaca gccaggagaa accccgcaaa 300 ggctcgttct ccagggtgca gtgggaacgc ctgcgctcca gtgcgaccca gaacagcagg 360 tctttggagc cccctccggg ccctgggcct gcccgcagga gcgcgagcac agcgcgcagc 420 tcggcgcccg cacgcacggt gctgagcgcc ccacctcgca ggatgcaggc ctcctcggcc 480 gcctgccgct tcatggtagc gtggtgc 507 <210> 48 <211> 517 <212> DNA <213> Homo sapiens <400> 48 agacttcttg ggagtttgca gagcgacccg tcgcccgcgc ccggcgctgg cagggacctt 60 cggatggttc ttactgggcc gatccatggc acaggctggg cctcggcgaa cccctcggcc 120 cccgcccggc cccgagccac gacacctcat tgtcctggag cctgggaagg gggtgcgcga 180 gcgcgcgggc gagccctgcc tctccccgcc agagaacagc tgaggggccg cggtcccagc 240 gggaggattc cggtccctgg cccggccgcg gccttgggcg gagcaggggc cactagctgc 300 cacttctgcc cgccccaggt gcgcgcggag ggctacgtgg ggcgggccgc gacccggcaa 360 agtcatgttg aaaaaacact cttcacgttc gctcggcctg gtgaccaggg tcgggggacca 420 cgacaaccgg gggttgggag gctgcgtaat tacaacccag ggtggtttgg attttgggg 480 gtggtggata tttaaaaaaca aaaaggagat ctggaag 517 <210> 49 <211> 550 <212> DNA <213> Homo sapiens <400> 49 aaccacagcc cgtgcgctc ccgcagtggg agttcgccgg ccgactccca ccctcacagc 60 ctcctgtcct ggcttcccct cgccccgaggc tgcaacaccg catcccccc atccccccgcc 120 gcgccctcag cctcgggccg caccaaccca ggggataagg cgactccggt cgctctgagg 180 ggcagggcca gccagccccc tcccacccac gcacacgctc cccctcagag ccgccggccc 240 agagaaaaac cgccacatgc agctccctc cacacgcacc taacagctc ctctggaccc 300 gaacgcccac accctccctc cctggggtcc caactccac tcaggacgcc acagcggatc 360 ctaactacaa acggtccccg gagccctggg ctggactcgc tcagccccgc ccccacgccc 420 ctggtaccag ccctgagaga ccccgcggag cacgccgcgg gagccgcaga tcgcgctgaa 480 gagcagcgag atcgcgctct ggacgagacc tgcgcggctg caacgctcc ttctcgcgg 540 gtggaagccgc 550 <210> 50 <211> 537 <212> DNA <213> Homo sapiens <400> 50 gaggaactcc ggcaagcca ggcggcggg gggctccggg tctggcggc ggctccggg 60 gagcagcggg agaccccgca gcggctcct ccttctccgc ccgcggcccc cagcctcgcc 120 gccgccgccc ggctcccagc acggaccga cggggcgctc ccgagacggg cgagccacgc 180 gctcgcaggt cccaaggcca ggctggggcgg gactgttaag ggagctcga gtcggggggcc 240 gggggcttcc cgtcccggcg cttcccatgc aaacccctga aggaagcggc agggcgccc 300 gcgggctccg cagcccaggc ccacttcctg tcactccagg aaaacctcgg agcggcggac 360 gcggctcggc ccggcttcca gcccagagcc caagcgcctt agccccgtcc cagcgctttc 420 tgaaagacgg gccacctcgc gcggagccgc gacaaggact ccagggtccg cagtgaagct 480 ggtcaaatct gccccgcaca cggtcaacgc tcggtctgtg tcccggaagc tttcgga 537 <210> 51 <211> 592 <212> DNA <213> Homo sapiens <400> 51 ggatttcgat gaaatggtcc ctgaagttgt gctccttctg ggactccatc cttcagctct 60 ccaatctcaa cagcctgtat tctgttggga tggggtaaga ccggtgagcg acggtcaaac 120 gtctgtccca cgtggtaagg cgggaaccgc tgctgcctgt actgggggcg atactggggg 180 cggcgcagcc gattccgggc cccagagaac tgcctatcag tggtaggggg gtcaaatcct 240 tcactgctgc cgctcccttc ttcctcctcc tcccagcgta atcccgggga gggccatggc 300 gccttccata gtagccacgt ctgtaacggc gccaatcagc gcgtaacgac tgccctccac 360 aggaactcca ccccggccag tcacactggc tgcttctgca cccttctctc cttaaaccac 420 atcaaactct acagtttctc catctcctac actgcgcaga tatttctgtg ggttattctt 480 cttgatggca gtctgatgta taaatagatc ttctttggtg tcatttcgat ttataaatcc 540 atatccattt ctgacgttga accatttgac agtgccaagg actttggtgg cg 592 <210> 52 <211> 701 <212> DNA <213> Homo sapiens <400> 52 tctattgtat gtacgtgttg cagtcctttc atttgccaca acatatggat tccataaatg 60 cagacatgcc gaagtgcatc tgtctgggta gttaacatga tctaaacatc cctcttcgtt 120 ccgctaactc cggctcttct tcgggctcct cggcagcgct cggggccagc cggcccgtgc 180 cccaggcttg cagcgcccgg cagcctcgtc cttttgtggt ctctgcacgg gatccaaggt 240 gccgcgcgga ggaggcgggc tgctcgcagt gccggggtca gaggcgccgc caccggcggc 300 ctctgcgcgc gcggggagga aagggttaag ctgcccgagc ccggggaagg ggctgctctc 360 atcctggagc gaggtgcagc caccggcagc tgtgatttag gggtcaagtc cgagatcacc 420 tttctcctgc ctctggaaat ggcagaagat gagataggga gggagaaact agagagtggc 480 agccaggcgc agcacgtggg ctccatccat ccgacacccc catcgccccg gtccactccc 540 tgacccccag acaaatcgga cagttccctt ttctggtaga gatgcggggt gcgcttcttc 600 tgagcgtccg gaatcgctcc atccaaggct ctgccctaag gttaagccac tgtgccctga 660 gcctcaacca ccagatctca aaagtttgct ctcaatgcgc c 701 <210> 53 <211> 531 <212> DNA <213> Homo sapiens <400> 53 aggtcggggc gggcttcgtt ggaagcgggt ggcagcgcgg gggggcacgc ctcgctctct 60 gtaagccact ggagagttgg ggcgagtagg gagaaggctg ggagtaaatc aaggggaggc 120 ggcgagaccg aggacccaat tcacggccct gaataacggg ggtagctggt aaggggcagc 180 tcccgggct gcgcccagcc tcctccctgc acccaggccc gcgagggctc cccgcgatcc 240 gcgagttccc cgcgtggcct tctcagccc gcgaggtcg cgtcttccct cccttcggt 300 cccgccggcc cccggccggg ccctgacgtc ctgcgcctc ccccgctc cgcagattac 360 cagagcgagt actacgggcc cgggggcaac tacgacttct tcccgcaagg ccccccgtcc 420 tcgcaggccc agacaccagt ggacctaccc ttcgtgccgt catctgggcc gtccgggacg 480 cccctgggtg gcctggagca cccgctgccg ggccaccacc cgtcgagcga g 531 <210> 54 <211> 554 <212> DNA <213> Homo sapiens <400> 54 tcactggctc tacagactgc cacgggtaaa cagcttagac cagatgactc aggctgaaag 60 catatgaacc cttcctgcag ggaggccgcc cggatgcaac agtggtttca cccctgagcc 120 gggcagcctt ggcagacctt gcttcacgtg ggctgaaatg gcagtgtctc tcctctttgt 180 ggccaggttt tgcctcctct ttgactcgga agcatctcct atcctgcaag gacagtttga 240 gcagggcccc cgggccctcc ttccaagagg cttctgcagc tgtggacccc caagagttta 300 tgccgctgag ctctgctgtc tctccccacc tgctccccac ctgtctgccc ccacacctgc 360 gactctggct ctcctggatc ctcctgtaga ctggttccta taagcacaag gaggaacatg 420 cgagatgctg ggattggatg ctctgggcct ggggctggtg tttcctcatg cccgggctat 480 ttccttttgg ccctgggcat gcagtcatgt gcttcctttc atgggcgggt tggggaccag 540 ggccagcgag caga 554 <210> 55 <211> 594 <212> DNA <213> Homo sapiens <400> 55 ggtgcccgtc tgtgtgtgct cctcccagca gccatcgctc aaccttgctc tcaggaagcc 60 cccaggcgag tgttggcagg aatcctgcca ggcgggaggt cgctcctcca gagcgtggtc 120 cctgaagccg ccagcctccc tggcctcgcc ccttgctggt ggtgtgtgtg gtgtggccgt 180 gggtgcactt tgctgggtct tcctgggaca ctgaagtctc ctgtgtctcc agccctgaga 240 actcggagcc cgggtgcttt tgggaaggac ggggcaccag ctggtgacac atgggaaggg 300 aggtgtggtt gtcaccttgc ccaggtaacc tgctctgcct ggtcggtgcg cctaaggggg 360 gcagggtgtt tggggaggac atgagaggcc tcctggaagc acttcatcct gttgaagttc 420 acattttgac cttttcagca gcccttgctc tgggcctgtg cccggccctg ggactcggcc 480 tggagagcct attgacaccg tgccatgggt gcgggcaggg cgccctccct ggagggcggc 540 acgtggtgcc agttggtgac catgagctgc ctcactcctg aggaagagtg ttcg 594 <210> 56 <211> 506 <212> DNA <213> Homo sapiens <400> 56 ccgctgcggg gatttctccc ccagcctttt ctttttaaca gagggcaaag gggcgacggc 60 gagagcacag atggcggctg cggagccggg gaggcggcgg ggagacgcgc gggactcgtg 120 gggagggctg gcagggtgca ggggttccgc gtgacctgcc cggctcccag gcatcgggct 180 gggcgctgca gtttaccgat ttgctttcgt ccctcgtcca ggtttaggag acgcgtgggg 240 acagccgagc cgcgccgggc ccctggacgg cgtcgccaag gagctgggat cgcacttgct 300 gcaggtagag cggcctcgcc gggggaggag cgcagccgcc gcaggctccc ttcccacccc 360 gccaccccag cctccaggcg tcccttcccc aggagcgcca ggcagatcca gaggctgccg 420 ggggctgggg atggggtggt ccccactgcg gagggatgga cgcttagcat gtcggatgcg 480 gcctgcggcc aaccctaccc taaccc 506
Claims
1. Reagents or apparatus for determining the methylation level of DNA sequences in a sample of a subject, for use in the preparation of a kit for diagnosing pancreatic cancer in a subject. in, The DNA sequence comprises the following sequences or their complementary sequences: combinations of SEQ ID NO:8, SEQ ID NO:11, SEQ ID NO:20, SEQ ID NO:44, SEQ ID NO:48, SEQ ID NO:51, and SEQ ID NO:
54.
2. The use as described in claim 1, characterized in that, The kit also includes: a processed nucleic acid molecule of the DNA sequence, wherein the processing converts unmethylated cytosine into bases that have a lower binding affinity to guanine than cytosine.
3. The use as described in claim 1 or 2, characterized in that, The reagent contains primer molecules that hybridize with the DNA sequence or a fragment thereof, and / or The reagent contains probe molecules that hybridize with the DNA sequence or fragments thereof, and / or The reagent comprises a medium containing a combination of DNA sequences and / or their methylation information as shown in SEQ ID NO:8, SEQ ID NO:11, SEQ ID NO:20, SEQ ID NO:44, SEQ ID NO:48, SEQ ID NO:51, and SEQ ID NO:
54.
4. The use as described in claim 3, characterized in that, The samples are derived from mammalian tissues, cells, or body fluids.
5. The use as described in claim 3, characterized in that, The samples were derived from pancreatic tissue or blood.
6. The use as described in claim 3, characterized in that, The sample includes genomic DNA or cfDNA.
7. The use as described in claim 3, characterized in that, The DNA sequence was treated with a methylation-sensitive restriction endonuclease.
8. The use as described in claim 1 or 2, characterized in that, The diagnosis includes: comparing with a control sample or calculating a score, and diagnosing pancreatic cancer based on the score.
9. The use as described in claim 8, characterized in that, The calculations are performed by constructing a support vector machine model.
10. An apparatus for diagnosing pancreatic cancer, the apparatus comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it performs the following steps: (1) Obtain the methylation level of the DNA sequence in the sample of the object, wherein the DNA sequence comprises the following sequences or their complementary sequences: a combination of SEQ ID NO:8, SEQ ID NO:11, SEQ ID NO:20, SEQ ID NO:44, SEQ ID NO:48, SEQ ID NO:51, and SEQ ID NO:
54. (2) Compare with the control sample, or calculate the score, and (3) Diagnose pancreatic cancer based on the scoring.
11. The apparatus as claimed in claim 10, characterized in that, The sample includes genomic DNA or cfDNA.
12. The apparatus as claimed in claim 10, characterized in that, The sequence is transformed, wherein unmethylated cytosine is converted into bases that have a lower binding affinity to guanine than cytosine.
13. The apparatus as claimed in claim 10, characterized in that, The DNA sequence was treated with a methylation-sensitive restriction endonuclease.
14. The apparatus as claimed in claim 10, characterized in that, The score in step (2) is calculated by constructing a support vector machine model.
Citation Information
Patent Citations
Composition for distinguishing pancreatic cancer and chronic pancreatitis, and uses thereof
CN108277274A
Detecting pancreatic high-grade dysplasia
CN109153993A