A machine learning-based pancreatic cancer diagnosis method

By detecting the DNA methylation level and CA19-9 level in patient plasma samples, a machine learning model was constructed to address the shortcomings of early pancreatic cancer detection, achieving a more accurate non-invasive diagnosis, especially in the differential diagnosis of pancreatic cancer in patients with chronic pancreatitis.

CN115985486BActive Publication Date: 2026-05-01SINGLERA GENOMICS (SHANGHAI) LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SINGLERA GENOMICS (SHANGHAI) LTD
Filing Date
2021-10-13
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies have limited methods for early detection of pancreatic cancer, especially in asymptomatic individuals where detection is ineffective. Furthermore, the accuracy of differentiating between chronic pancreatitis and pancreatic cancer is low, necessitating more stable and specific biomarkers.

Method used

By detecting DNA methylation levels and CA19-9 levels in patient plasma samples, a machine learning model was constructed. A pancreatic cancer diagnostic model was built by combining methylation scores and CA19-9 levels, and mathematical models such as support vector machines or logistic regression were used for diagnosis.

Benefits of technology

It improves the accuracy of non-invasive diagnosis of pancreatic cancer, enabling earlier and more accurate identification of pancreatic cancer, especially in the differential diagnosis of pancreatic cancer in patients with chronic pancreatitis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115985486B_ABST
    Figure CN115985486B_ABST
Patent Text Reader

Abstract

The present application relates to a machine learning-based pancreatic cancer diagnosis method, and specifically provides a method for constructing a pancreatic cancer diagnosis model, comprising: (1) obtaining the methylation level of a DNA sequence or a fragment thereof, or the methylation state or level of one or more CpG dinucleotides in the DNA sequence or the fragment thereof, and the CA19-9 level of the subject, (2) using a mathematical model to calculate the methylation score using the methylation state or level, (3) combining the methylation score and the CA19-9 level into a data matrix, and (4) constructing a pancreatic cancer diagnosis model based on the data matrix. The method can significantly improve the accuracy of non-invasive diagnosis of pancreatic cancer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of pancreatic disease diagnosis, specifically relating to a machine learning-based method for diagnosing pancreatic cancer. Background Technology

[0002] The 5-year relative survival rate for pancreatic cancer is 9%, decreasing to only 3% for patients with distant metastases. A major reason for the high mortality rate is the limited availability of methods for early detection of pancreatic cancer, which is crucial for patients undergoing surgical resection. Currently, carbohydrate antigen 19-9 (CA19-9) is the most commonly used clinical serum biomarker for the auxiliary detection of pancreatic cancer, achieving a sensitivity of 79-90% and a specificity of 75-90% in symptomatic patients before resection. However, several large population studies have demonstrated that CA19-9 is ineffective in detecting pancreatic cancer in asymptomatic individuals due to its low positive predictive value, essentially ruling it out for early screening of pancreatic cancer (Kim et al., 2004).

[0003] Typical early symptoms of pancreatic cancer, including abdominal and back pain, diarrhea, weight loss, and jaundice, are not specific and may be associated with other gastrointestinal disorders. These complications are particularly common in the diagnosis of chronic pancreatitis, especially since patients with chronic pancreatitis have a significantly higher long-term risk of developing pancreatic cancer. Therefore, accurate differential diagnosis between pancreatic cancer and chronic pancreatitis is crucial for screening patients with chronic pancreatitis for pancreatic cancer. However, the current accuracy rate for differentiating between chronic pancreatitis and pancreatic cancer is 65% or lower, leaving much room for improvement. Therefore, stable and consistent specific markers for differentiating between chronic pancreatitis and pancreatic cancer are needed. Summary of the Invention

[0004] This invention provides a method for detecting DNA methylation in patient plasma samples and constructing a machine learning model to diagnose pancreatic cancer based on the methylation level data of target methylation markers and the detection results of CA19-9, in order to achieve a non-invasive and precise diagnosis of pancreatic cancer with higher accuracy and lower cost.

[0005] The first aspect of this invention provides a method for diagnosing pancreatic cancer or constructing a diagnostic model for pancreatic cancer, comprising:

[0006] (1) Obtain the methylation level of the DNA sequence or fragment thereof in the target sample, or the methylation status or level of one or more CpG dinucleotides in the DNA sequence or fragment thereof, and the CA19-9 level of the target.

[0007] (2) Methylation scores are obtained by using mathematical models to calculate methylation status or level.

[0008] (3) Combine the methylation score and CA19-9 level data matrix.

[0009] (4) Construct a pancreatic cancer diagnostic model based on the data matrix.

[0010] (5) Optional: Obtain a pancreatic cancer score; diagnose pancreatic cancer based on the pancreatic cancer score.

[0011] In one or more embodiments, the DNA sequence is selected from one or more (e.g., at least two) or all of the following gene sequences, or sequences within 20 kb upstream or downstream of them: SIX3, TLX2, CILP2. Preferably, the DNA sequence comprises a gene sequence selected from any of the following groups: (1) SIX3, TLX2; (2) SIX3, CILP2; (3) TLX2, CILP2; (4) SIX3, TLX2, CILP2.

[0012] In one or more embodiments, the fragment length is 1-1000 bp, preferably 1-700 bp. The fragment contains at least one CpG dinucleotide.

[0013] In one or more embodiments, the DNA sequence is selected from one or more (e.g., at least two) or all of the following sequences or their complementary sequences: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, or variants having at least 70% identity with them, wherein the methylation sites in the variants are not mutated. Preferably, the DNA sequence comprises a sequence selected from any of the following groups or their complementary sequences: (1) SEQ ID NO:1, SEQ ID NO:2, (2) SEQ ID NO:1, SEQ ID NO:3, (3) SEQ ID NO:2, SEQ ID NO:3, (4) SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3.

[0014] In one or more embodiments, step (1) includes detecting the methylation level of a DNA sequence or fragment thereof in a sample of the subject, or the methylation status or level of one or more CpG dinucleotides in the DNA sequence or fragment thereof.

[0015] In one or more embodiments, the method further includes DNA extraction and / or quality control prior to step (1).

[0016] In one or more embodiments, step (1) includes detecting methylation status or level using primer molecules and / or probe molecules.

[0017] In one or more embodiments, the primer molecule comprises a primer molecule that hybridizes to the DNA sequence or a fragment thereof. The primer molecule is capable of amplifying the DNA sequence or a fragment thereof. In one or more embodiments, the primer sequence is methylation-specific or non-specific. The primer molecule is at least 9 bp.

[0018] In one or more embodiments, the probe molecule comprises a probe molecule that hybridizes to the DNA sequence or a fragment thereof. In one or more embodiments, the probe further comprises a detectable. In one or more embodiments, the detectable is a 5' fluorescent reporter group and a 3' labeled quencher group. In one or more embodiments, the fluorescent reporter gene is selected from Cy5, FAM, and VIC. Preferably, the sequence of the probe comprises MGB (Minor Groovebinder) or LNA (Locked Nucleic Acid). The probe molecule is at least 12 bp.

[0019] In one or more embodiments, the detection includes, but is not limited to: PCR based on bisulfite conversion, DNA sequencing, methylation-sensitive restriction endonuclease analysis, quantitative fluorescence method, methylation-sensitive high-resolution melting curve method, chip-based methylation mapping analysis, and mass spectrometry.

[0020] In one or more embodiments, the detection is DNA sequencing. In one or more embodiments, the sequencing depth of the DNA sequencing is greater than or equal to 5M, preferably at least 7M, 11M, 13M, or 15M.

[0021] In one or more embodiments, the detection is performed using MethylTitan sequencing.

[0022] In one or more embodiments, the sample is derived from mammalian tissue, cells, or body fluids, such as pancreatic tissue or blood. The mammal is preferably human. In one or more embodiments, the sample is a fine-needle aspiration biopsy. In one or more embodiments, the sample is plasma.

[0023] In one or more embodiments, the sample comprises genomic DNA or cfDNA.

[0024] In one or more embodiments, the DNA sequence is transformed, wherein unmethylated cytosine is converted into bases that do not bind to guanine. The transformation is performed using an enzymatic method, preferably deaminase treatment, or the transformation is performed using a non-enzymatic method, preferably treatment with bisulfite, acid sulfite, or metabisulfite, or a combination thereof.

[0025] In one or more embodiments, the DNA sequence is treated with a methylation-sensitive restriction endonuclease.

[0026] In one or more embodiments, the CA19-9 level is the blood or plasma CA19-9 level.

[0027] In one or more implementations, the mathematical model described in step (2) is a support vector machine model.

[0028] In one or more implementations, the pancreatic cancer diagnostic model described in step (4) is a logistic regression model.

[0029] In one or more embodiments, step (5) includes diagnosing pancreatic cancer based on whether a pancreatic cancer score reaches a threshold.

[0030] In one or more embodiments, the diagnosis of pancreatic cancer is to differentiate between pancreatic cancer and pancreatitis.

[0031] A second aspect of the present invention also provides a method for diagnosing pancreatic cancer, comprising:

[0032] (1) Obtain the methylation level of the DNA sequence or fragment thereof in the target sample, or the methylation status or level of one or more CpG dinucleotides in the DNA sequence or fragment thereof, and the CA19-9 level of the target.

[0033] (2) Methylation scores are obtained by using mathematical models to calculate methylation status or level.

[0034] (3) Obtain the pancreatic cancer score according to the model shown below, and diagnose pancreatic cancer based on the pancreatic cancer score:

[0035]

[0036] Where M is the methylation score of the sample calculated in step (2), and C is the CA19-9 level of the sample.

[0037] In one or more embodiments, the DNA sequence is selected from one or more (e.g., at least two) or all of the following gene sequences, or sequences within 20 kb upstream or downstream of them: SIX3, TLX2, CILP2. Preferably, the DNA sequence comprises a gene sequence selected from any of the following groups: (1) SIX3, TLX2; (2) SIX3, CILP2; (3) TLX2, CILP2; (4) SIX3, TLX2, CILP2.

[0038] In one or more embodiments, the fragment length is 1-1000 bp, preferably 1-700 bp. The fragment contains at least one CpG dinucleotide.

[0039] In one or more embodiments, the DNA sequence is selected from one or more (e.g., at least two) or all of the following sequences or their complementary sequences: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, or variants having at least 70% identity with them, wherein the methylation sites in the variants are not mutated. Preferably, the DNA sequence comprises a sequence selected from any of the following groups or their complementary sequences: (1) SEQ ID NO:1, SEQ ID NO:2, (2) SEQ ID NO:1, SEQ ID NO:3, (3) SEQ ID NO:2, SEQ ID NO:3, (4) SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3.

[0040] In one or more embodiments, step (1) includes detecting the methylation level of a DNA sequence or fragment thereof in a sample of the subject, or the methylation status or level of one or more CpG dinucleotides in the DNA sequence or fragment thereof.

[0041] In one or more embodiments, the method further includes DNA extraction and / or quality control prior to step (1).

[0042] In one or more embodiments, step (1) includes detecting methylation status or level using primer molecules and / or probe molecules.

[0043] In one or more embodiments, the primer molecule comprises a primer molecule that hybridizes to the DNA sequence or a fragment thereof. The primer molecule is capable of amplifying the DNA sequence or a fragment thereof. In one or more embodiments, the primer sequence is methylation-specific or non-specific. The primer molecule is at least 9 bp.

[0044] In one or more embodiments, the probe molecule comprises a probe molecule that hybridizes to the DNA sequence or a fragment thereof. In one or more embodiments, the probe further comprises a detectable. In one or more embodiments, the detectable is a 5' fluorescent reporter group and a 3' labeled quencher group. In one or more embodiments, the fluorescent reporter gene is selected from Cy5, FAM, and VIC. Preferably, the sequence of the probe comprises MGB (Minor Groovebinder) or LNA (Locked Nucleic Acid). The probe molecule is at least 12 bp.

[0045] In one or more embodiments, the detection includes, but is not limited to: PCR based on bisulfite conversion, DNA sequencing, methylation-sensitive restriction endonuclease analysis, quantitative fluorescence method, methylation-sensitive high-resolution melting curve method, chip-based methylation mapping analysis, and mass spectrometry.

[0046] In one or more embodiments, the detection is DNA sequencing. In one or more embodiments, the sequencing depth of the DNA sequencing is greater than or equal to 5M, preferably at least 7M, 11M, 13M, or 15M.

[0047] In one or more embodiments, the detection is performed using MethylTitan sequencing.

[0048] In one or more embodiments, the sample is derived from mammalian tissue, cells, or body fluids, such as pancreatic tissue or blood. The mammal is preferably human. In one or more embodiments, the sample is a fine-needle aspiration biopsy. In one or more embodiments, the sample is plasma.

[0049] In one or more embodiments, the sample comprises genomic DNA or cfDNA.

[0050] In one or more embodiments, the DNA sequence is transformed, wherein unmethylated cytosine is converted into bases that do not bind to guanine. The transformation is performed using an enzymatic method, preferably deaminase treatment, or the transformation is performed using a non-enzymatic method, preferably treatment with bisulfite, acid sulfite, or metabisulfite, or a combination thereof.

[0051] In one or more embodiments, the DNA sequence is treated with a methylation-sensitive restriction endonuclease.

[0052] In one or more embodiments, the CA19-9 level is the blood or plasma CA19-9 level.

[0053] In one or more implementations, the mathematical model described in step (2) is a support vector machine model.

[0054] In one or more embodiments, step (3) includes diagnosing pancreatic cancer based on whether a pancreatic cancer score reaches a threshold. Preferably, the threshold of the model is approximately 0.885.

[0055] In one or more embodiments, the diagnosis of pancreatic cancer is to differentiate between pancreatic cancer and pancreatitis.

[0056] A third aspect of the present invention provides a method for constructing a diagnostic model for pancreatic cancer, comprising:

[0057] (1) Obtain the methylation haplotype ratio and sequencing depth of the target genomic DNA segment.

[0058] Optionally (2) preprocesses the methylation haplotype ratios and sequencing depth data.

[0059] (3) Perform cross-validation incremental feature screening to obtain the characteristic methylation region.

[0060] (4) A mathematical model is constructed based on the methylation detection results of the characteristic methylated regions to obtain a methylation score.

[0061] (5) A pancreatic cancer diagnostic model was constructed based on the methylation score and the corresponding CA19-9 level.

[0062] In one or more embodiments, step (1) includes:

[0063] 1.1) DNA methylation detection was performed on the target sample to obtain sequencing read data.

[0064] 1.2) Optional preprocessing of sequencing data, such as adapter removal and / or splicing,

[0065] 1.3) Align the sequencing data to the reference genome to obtain the location and sequencing depth information of methylated regions.

[0066] 1.4) Calculate the methylation haplotype ratio (MHF) of the segment according to the following formula:

[0067]

[0068] Where i represents the target methylation region, h represents the target methylation haplotype, and N i N represents the number of reads located in the target methylation region. i,h This indicates the number of reads containing the target methylated haplotype.

[0069] In one or more embodiments, the methylation detection is performed by MethylTitan sequencing.

[0070] In one or more embodiments, the sample is cfDNA.

[0071] In one or more implementations, a methylation haplotype ratio is calculated for each methylation haplotype within the target region.

[0072] In one or more embodiments, step (2) includes: 2.1) merging methylation haplotype ratio status and sequencing depth information data into a data matrix.

[0073] In one or more embodiments, step (2) further includes: 2.2) removing sites in the data matrix with a missing value ratio higher than 5-15% (e.g., 10%).

[0074] In one or more embodiments, step (2) further includes: 2.3) treating each data point with a depth less than 300 (e.g., less than 200) as a missing value and filling the missing values ​​(e.g., using the K nearest neighbor method).

[0075] In one or more embodiments, step (3) includes: performing cross-validation incremental feature screening on the training data using a mathematical model, wherein the DNA segment that increases the AUC of the mathematical model is a feature methylation segment. In one or more embodiments, the mathematical model is a support vector machine (SVM) model.

[0076] In one or more embodiments, step (3) includes:

[0077] (3.1) Based on the methylation haplotype ratio and sequencing depth, DNA segments are ranked according to their relevance to obtain candidate methylated segments with high relevance.

[0078] (3.2) Perform cross-validation incremental feature screening, in which candidate methylation segments are sorted according to their correlation (e.g., regression coefficients from large to small), and one or more candidate methylation segment data are added each time to predict the test data. Among them, the candidate methylation segments whose mean cross-validation AUC increases are the characteristic methylation segments.

[0079] In one or more implementations, step (3.1) is: (3.1) construct a logistic regression model based on the methylation haplotype ratio of DNA segments and the sequencing depth relative to the object phenotype, and screen out DNA segments with large regression coefficients to form candidate methylation segments.

[0080] In one or more implementations, the prediction in step (3.2) is made by constructing a model (e.g., a support vector machine model).

[0081] In one or more implementations, the mathematical model in step (4) is a vector machine (SVM) model.

[0082] In one or more embodiments, the methylation detection result in step (4) is a combined matrix of methylation haplotype ratio and sequencing depth.

[0083] In one or more implementations, step (5) includes merging a data matrix of methylation scores and CA19-9 levels, and constructing a pancreatic cancer diagnostic model based on the data matrix.

[0084] In one or more implementations, the pancreatic cancer diagnostic model in step (5) is a logistic regression model.

[0085] The present invention also provides (a) the methylation level of a DNA sequence or a fragment thereof, or the methylation state or level of one or more CpG dinucleotides in the DNA sequence or a fragment thereof, and (b) the application of CA19-9 level in constructing a diagnostic model for pancreatic cancer.

[0086] In one or more embodiments, the pancreatic cancer diagnostic model is a logistic regression model.

[0087] In one or more embodiments, the pancreatic cancer diagnostic model is as described in any of the embodiments of the first, second, and third aspects herein.

[0088] In one or more embodiments, the method for constructing a pancreatic cancer diagnostic model is as described in any of the embodiments of the first, second, and third aspects herein.

[0089] In one or more embodiments, the DNA sequence is selected from one or more (e.g., at least two) or all of the following gene sequences, or sequences within 20 kb upstream or downstream of them: SIX3, TLX2, CILP2. Preferably, the DNA sequence comprises a gene sequence selected from any of the following groups: (1) SIX3, TLX2; (2) SIX3, CILP2; (3) TLX2, CILP2; (4) SIX3, TLX2, CILP2.

[0090] In one or more embodiments, the fragment length is 1-1000 bp, preferably 1-700 bp. The fragment contains at least one CpG dinucleotide.

[0091] In one or more embodiments, the DNA sequence is selected from one or more (e.g., at least two) or all of the following sequences or their complementary sequences: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, or variants having at least 70% identity with them, wherein the methylation sites in the variants are not mutated. Preferably, the DNA sequence comprises a sequence selected from any of the following groups or their complementary sequences: (1) SEQ ID NO:1, SEQ ID NO:2, (2) SEQ ID NO:1, SEQ ID NO:3, (3) SEQ ID NO:2, SEQ ID NO:3, (4) SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3.

[0092] In one or more embodiments, the CA19-9 level is the blood or plasma CA19-9 level.

[0093] Another aspect of the present invention provides the use of reagents or devices for detecting DNA methylation and reagents or devices for detecting CA19-9 levels in the preparation of kits for diagnosing pancreatic cancer, wherein the reagents or devices for detecting DNA methylation are used to determine the methylation level of a DNA sequence or fragment thereof in a sample of a subject, or the methylation status or level of one or more CpG dinucleotides in the DNA sequence or fragment thereof.

[0094] In one or more embodiments, the DNA sequence is selected from one or more (e.g., at least two) or all of the following gene sequences, or sequences within 20 kb upstream or downstream of them: SIX3, TLX2, CILP2.

[0095] In one or more embodiments, the DNA sequence comprises a gene sequence selected from any of the following groups: (1) SIX3, TLX2; (2) SIX3, CILP2; (3) TLX2, CILP2; (4) SIX3, TLX2, CILP2.

[0096] In one or more embodiments, the DNA sequence includes a sense strand or an antisense strand.

[0097] In one or more embodiments, the fragment length is 1-1000bp, preferably 1-700bp.

[0098] In one or more embodiments, the fragment contains at least one CpG dinucleotide.

[0099] In one or more embodiments, the DNA sequence is selected from one or more (e.g., at least two) or all of the following sequences or their complementary sequences: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, or variants thereof having at least 70% identity, wherein the methylation sites in the variants are not mutated.

[0100] In one or more embodiments, the DNA sequence comprises a sequence selected from any of the following groups or a complementary sequence thereof: (1) SEQ ID NO:1, SEQ ID NO:2, (2) SEQ ID NO:1, SEQ ID NO:3, (3) SEQ ID NO:2, SEQ ID NO:3, (4) SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3.

[0101] In one or more embodiments, the reagent for detecting DNA methylation comprises primer molecules and / or probe molecules.

[0102] In one or more embodiments, the reagent for detecting DNA methylation comprises primer molecules that hybridize with the DNA sequence or a fragment thereof. The primer molecules amplify the DNA sequence or a fragment thereof. In one or more embodiments, the primer sequence is methylation-specific or non-specific. The primer molecule is at least 9 bp.

[0103] In one or more embodiments, the reagent for detecting DNA methylation comprises a probe molecule that hybridizes to the DNA sequence or a fragment thereof. In one or more embodiments, the probe further comprises a detectable. In one or more embodiments, the detectable is a 5' fluorescent reporter group and a 3' labeled quencher group. In one or more embodiments, the fluorescent reporter gene is selected from Cy5, FAM, and VIC. Preferably, the sequence of the probe comprises MGB (Minorgroove binder) or LNA (Locked nucleic acid). The probe molecule is at least 12 bp.

[0104] In one or more embodiments, the CA19-9 level is a blood or plasma level.

[0105] In one or more embodiments, the reagent for detecting CA19-9 levels is an immunoreaction-based detection reagent; including: an antibody against CA19-9, and optionally buffer solutions, washing solutions, etc.

[0106] In one or more embodiments, the kit is a non-invasive diagnostic kit.

[0107] In one or more embodiments, the kit is an auxiliary diagnostic kit.

[0108] In one or more embodiments, the object is a mammal, preferably a human.

[0109] In one or more embodiments, the object is an object diagnosed with pancreatitis (e.g., chronic pancreatitis).

[0110] In one or more embodiments, the sample is derived from mammalian tissue, cells, or body fluids, such as pancreatic tissue or blood, preferably a fine-needle aspiration biopsy or plasma.

[0111] In one or more embodiments, the sample comprises genomic DNA or cfDNA.

[0112] In one or more embodiments, the DNA sequence is transformed, wherein unmethylated cytosine is converted into bases with a lower binding affinity to guanine than cytosine. The transformation is performed using an enzymatic method, preferably deaminase treatment, or the transformation is performed using a non-enzymatic method, preferably treatment with bisulfite, acid sulfite, or metabisulfite, or a combination thereof.

[0113] In one or more embodiments, the DNA sequence is treated with a methylation-sensitive restriction endonuclease.

[0114] In one or more embodiments, the kit further includes PCR reaction reagents. Preferably, the PCR reaction reagents include DNA polymerase, PCR buffer, dNTPs, and Mg2+.

[0115] In one or more embodiments, the kit further includes other reagents for detecting DNA methylation, said other reagents being selected from one or more of the following methods: PCR based on bisulfite conversion, DNA sequencing, methylation-sensitive restriction endonuclease analysis, quantitative fluorescence assay, methylation-sensitive high-resolution melting curve assay, chip-based methylation mapping analysis, and mass spectrometry. Preferably, said other reagents are selected from one or more of the following: bisulfite, bisulfite, acid sulfite, or metabisulfite or derivatives thereof, methylation-sensitive or insensitive restriction endonucleases, enzyme digestion buffers, fluorescent dyes, fluorescence quenchers, fluorescent reporter agents, exonucleases, alkaline phosphatases, internal standards, and controls.

[0116] In one or more embodiments, the PCR reaction solution comprises Taq DNA polymerase, PCR buffer, dNTPs, KCl, MgCl2, and (NH4)2SO4. Preferably, the Taq DNA polymerase is a hot-start Taq DNA polymerase. Preferably, the final concentration of Mg2+ is 1.0-10.0 mM.

[0117] In one or more embodiments, the diagnosis includes: calculating by constructing a pancreatic cancer diagnostic model as described in any of the embodiments herein, and diagnosing pancreatic cancer based on a score.

[0118] Another aspect of the present invention provides a kit for diagnosing pancreatic cancer, comprising:

[0119] (a) A reagent or apparatus for detecting DNA methylation, used to determine the methylation level of a DNA sequence or fragment thereof in a sample of a subject, or the methylation state or level of one or more CpG dinucleotides in said DNA sequence or fragment thereof, and

[0120] (b) Reagents or apparatus for detecting CA19-9 levels.

[0121] In one or more embodiments, the DNA sequence is selected from one or more (e.g., at least two) or all of the following gene sequences, or sequences within 20 kb upstream or downstream of them: SIX3, TLX2, CILP2.

[0122] In one or more embodiments, the DNA sequence includes a sense strand or an antisense strand.

[0123] In one or more embodiments, the fragment length is 1-1000bp, preferably 1-700bp.

[0124] In one or more embodiments, the fragment contains at least one CpG dinucleotide.

[0125] In one or more embodiments, the DNA sequence is selected from one or more (e.g., at least two) or all of the following sequences or their complementary sequences: SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, or variants thereof having at least 70% identity, wherein the methylation sites in the variants are not mutated.

[0126] In one or more embodiments, the CA19-9 level is a blood or plasma level.

[0127] In one or more embodiments, the reagent for detecting CA19-9 levels is an immunoreaction-based detection reagent; including: an antibody against CA19-9, and optionally buffer solutions, washing solutions, etc.

[0128] In one or more embodiments, the kit is suitable for the use described in any of the embodiments herein.

[0129] In one or more embodiments, the reagent for detecting DNA methylation comprises primer molecules and / or probe molecules.

[0130] In one or more embodiments, the reagent for detecting DNA methylation comprises primer molecules that hybridize with the DNA sequence or a fragment thereof. The primer molecules amplify the DNA sequence or a fragment thereof. In one or more embodiments, the primer sequence is methylation-specific or non-specific. The primer molecule is at least 9 bp.

[0131] In one or more embodiments, the reagent for detecting DNA methylation comprises a probe molecule that hybridizes to the DNA sequence or a fragment thereof. In one or more embodiments, the probe further comprises a detectable. In one or more embodiments, the detectable is a 5' fluorescent reporter group and a 3' labeled quencher group. In one or more embodiments, the fluorescent reporter gene is selected from Cy5, FAM, and VIC. Preferably, the sequence of the probe comprises MGB (Minorgroove binder) or LNA (Locked nucleic acid). The probe molecule is at least 12 bp.

[0132] In one or more embodiments, the kit is a non-invasive diagnostic kit.

[0133] In one or more embodiments, the object is a mammal, preferably a human.

[0134] In one or more embodiments, the sample is derived from mammalian tissue, cells, or body fluids, such as pancreatic tissue or blood. In one or more embodiments, the sample is a fine-needle aspiration biopsy. In one or more embodiments, the sample is plasma.

[0135] In one or more embodiments, the sample comprises genomic DNA or cfDNA.

[0136] In one or more embodiments, the DNA sequence is transformed, wherein unmethylated cytosine is converted into bases with a lower binding affinity to guanine than cytosine. The transformation is performed using an enzymatic method, preferably deaminase treatment, or the transformation is performed using a non-enzymatic method, preferably treatment with bisulfite, acid sulfite, or metabisulfite, or a combination thereof.

[0137] In one or more embodiments, the DNA sequence is treated with a methylation-sensitive restriction endonuclease.

[0138] In one or more embodiments, the kit further includes PCR reaction reagents. Preferably, the PCR reaction reagents include DNA polymerase, PCR buffer, dNTPs, and Mg2+.

[0139] In one or more embodiments, the kit further includes reagents for detecting DNA methylation, said reagents being selected from one or more of the following methods: PCR based on bisulfite conversion, DNA sequencing, methylation-sensitive restriction endonuclease analysis, quantitative fluorescence assay, methylation-sensitive high-resolution melting curve assay, chip-based methylation mapping analysis, and mass spectrometry. Preferably, the reagents are selected from one or more of the following: bisulfite and its derivatives, methylation-sensitive or insensitive restriction endonucleases, enzyme digestion buffers, fluorescent dyes, fluorescence quenchers, fluorescent reporter agents, exonucleases, alkaline phosphatases, internal standards, and controls.

[0140] In one or more embodiments, the diagnosis of pancreatic cancer is to differentiate between pancreatic cancer and pancreatitis.

[0141] In another aspect, the present invention provides an apparatus for diagnosing pancreatic cancer or constructing a diagnostic model for pancreatic cancer, the apparatus comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor, when executing the program, performs the following steps:

[0142] (1) Obtain the methylation level of the DNA sequence or fragment thereof in the target sample, or the methylation status or level of one or more CpG dinucleotides in the DNA sequence or fragment thereof, and the CA19-9 level of the target.

[0143] (2) Methylation scores are obtained by using mathematical models to calculate methylation status or level.

[0144] (3) Combine the methylation score and CA19-9 level data matrix.

[0145] (4) Construct a pancreatic cancer diagnostic model based on the data matrix.

[0146] (5) Optional: Obtain a pancreatic cancer score; diagnose pancreatic cancer based on the pancreatic cancer score.

[0147] Other features of the device are as described in the first aspect of this document.

[0148] In another aspect, the present invention provides an apparatus for diagnosing pancreatic cancer or constructing a diagnostic model for pancreatic cancer, the apparatus comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor, when executing the program, performs the following steps:

[0149] (1) Obtain the methylation level of the DNA sequence or fragment thereof in the target sample, or the methylation status or level of one or more CpG dinucleotides in the DNA sequence or fragment thereof, and the CA19-9 level of the target.

[0150] (2) Methylation scores are obtained by using mathematical models to calculate methylation status or level.

[0151] (3) Obtain the pancreatic cancer score according to the model shown below, and diagnose pancreatic cancer based on the pancreatic cancer score:

[0152]

[0153] Where M is the methylation score of the sample calculated in step (2), and C is the CA19-9 level of the sample.

[0154] Other features of the device are as described in the second aspect of this document.

[0155] In another aspect, the present invention provides an apparatus for constructing a diagnostic model for pancreatic cancer, the apparatus comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor, when executing the program, performs the following steps:

[0156] (1) Obtain the methylation haplotype ratio and sequencing depth of the target genomic DNA segment.

[0157] (2) A logistic regression model was constructed based on the methylation haplotype ratio of DNA segments and the sequencing depth relative to the subject phenotype.

[0158] (3) Perform cross-validation incremental feature screening, where the DNA segments with increased AUC are characteristic methylation segments.

[0159] (4) A mathematical model is constructed based on the methylation detection results of the characteristic methylated regions to obtain a methylation score.

[0160] (5) A pancreatic cancer diagnostic model was constructed based on the methylation score and the corresponding CA19-9 level.

[0161] Other features of the device are as described in the third aspect of this document. Attached Figure Description

[0162] Figure 1 This is a flowchart of the technical solution of one embodiment of the present invention.

[0163] Figure 2 This shows the methylation level distribution of the three methylation markers in the training group.

[0164] Figure 3 This shows the methylation level distribution of the three methylation markers in the test group.

[0165] Figure 4 The ROC curves of the CA19-9 pancreatic cancer and pancreatitis differentiation prediction models pp_model and cpp_model on the test set are shown.

[0166] Figure 5 This is the distribution of prediction scores for the CA19-9 pancreatic cancer and pancreatitis differentiation prediction models pp_model and cpp_model in the test set samples (the values ​​have been normalized by maximizing and minima). Detailed Implementation

[0167] This invention explores the relationship between DNA methylation and CA19-9 levels and pancreatic cancer and pancreatitis. The aim is to improve the accuracy of non-invasive diagnosis of pancreatic cancer by utilizing the biomarkers DNA methylation and CA19-9 levels as differential markers between pancreatic cancer and chronic pancreatitis using a non-invasive method.

[0168] The inventors discovered that combining CA19-9 levels in pancreatic cancer marker screening and diagnosis can significantly improve diagnostic accuracy.

[0169] The present invention first provides a method for screening pancreatic cancer methylation biomarkers, comprising: (1) obtaining the methylation haplotype ratio and sequencing depth of DNA segments of the target genome (e.g., cfDNA); optionally (2) preprocessing the methylation haplotype ratio and sequencing depth data; and (3) performing cross-validation incremental feature screening to obtain characteristic methylation segments.

[0170] The acquisition of step (1) can be data analysis after methylation detection or direct reading from a file. In the implementation scheme for methylation detection, step (1) includes: 1.1) performing DNA methylation detection on the target sample to obtain sequencing read data; 1.3) aligning the sequencing data to a reference genome to obtain the location and sequencing depth information of the methylated region; 1.4) calculating the methylation haplotype ratio (MHF) of the region according to the following formula:

[0171]

[0172] Where i represents the target methylation region, h represents the target methylation haplotype, and N i N represents the number of reads located in the target methylation region. i,h This indicates the number of reads containing the target methylated haplotype. Typically, the methylation haplotype ratio needs to be calculated for each methylated haplotype within the target region. This step may also include steps 1.2) for preprocessing the sequencing data, such as adapter removal and / or splicing.

[0173] Step (2) includes merging methylation haplotype ratio status and sequencing depth information data into a data matrix. Furthermore, to make the results more accurate, step (2) also includes removing sites with a missing value ratio higher than 5-15% (e.g., 10%) from the data matrix, treating each data point with a depth less than 300 (e.g., less than 200) as a missing value, and filling the missing values ​​using the K-nearest neighbor method.

[0174] In one or more implementations, step (3) includes: using a mathematical model to perform cross-validation incremental feature screening in the training data, wherein the DNA segments that increase the AUC of the mathematical model are characteristic methylated segments. The mathematical model can be a support vector machine (SVM) model or a random forest model. Preferably, step (3) includes: (3.1) ranking the DNA segments according to their methylation haplotype ratio and sequencing depth to obtain highly correlated candidate methylated segments, and (3.2) performing cross-validation incremental feature screening, wherein the candidate methylated segments are ranked according to their correlation (e.g., regression coefficients from largest to smallest), and one or more candidate methylated segment data are added each time to predict the test data, wherein the candidate methylated segments whose mean cross-validation AUC increases are characteristic methylated segments. Specifically, step (3.1) can be: constructing a logistic regression model based on the methylation haplotype ratio of the DNA segments and sequencing depth relative to the object phenotype, and screening out DNA segments with large regression coefficients to form candidate methylated segments. The predictions in step (3.2) can be made by building a model (e.g., a support vector machine model or a random forest model).

[0175] After obtaining the characteristic methylated regions, they can be combined with CA19-9 levels to construct a more accurate pancreatic cancer diagnostic model. Therefore, in the method for constructing a pancreatic cancer diagnostic model, in addition to the above steps (1)-(3), it also includes (4) constructing a mathematical model for the data of the characteristic methylated regions to obtain a methylation score, and (5) merging the methylation score and CA19-9 levels into a data matrix, and constructing a pancreatic cancer diagnostic model based on the data matrix. The "data" in step (4) is the methylation detection result of the characteristic methylated regions, preferably a merged matrix of methylation haplotype ratio and sequencing depth.

[0176] The mathematical model in step (4) can be any mathematical model commonly used for diagnostic data analysis, such as a support vector machine (SVM) model, random forest, regression model, etc. In this paper, the exemplary mathematical model is the vector machine (SVM) model.

[0177] The pancreatic cancer diagnostic model in step (5) can be any mathematical model used for diagnostic data analysis, such as a support vector machine (SVM) model, random forest, regression model, etc. In this paper, an exemplary pancreatic cancer diagnostic model is the logistic regression pancreatic cancer model shown below:

[0178]

[0179] Where M is the methylation score of the sample, and C is the CA19-9 level of the sample. In one or more embodiments, the model threshold is 0.885; values ​​above this value are considered pancreatic cancer, while values ​​below or equal to this value are considered non-pancreatic cancer.

[0180] In specific implementation plans, machine learning-based methods for differentiating between pancreatitis and pancreatic cancer include:

[0181] (1) Collect blood from patients with pancreatic cancer or pancreatitis to be tested, and collect information such as patient age, gender, and CA19-9 test value; (2) Obtain plasma samples from patients with pancreatic cancer or pancreatitis to be tested, extract cfDNA, and use the MethylTitan method for library construction and sequencing to obtain sequencing reads; (3) Preprocess the sequencing data, including adapter removal and splicing of the sequencing data generated by the sequencer; (4) Align the preprocessed sequencing data to the reference genome sequence to determine the position of each fragment; (5) Calculation of MHF (Methylated Haplotype Fraction) methylation numerical matrix: A target methylation region may have multiple methylation haplotypes. For each methylation haplotype in the target region, the value needs to be calculated. The calculation formula of MHF is as follows:

[0182]

[0183] Where i represents the target methylation region, h represents the target methylation haplotype, Ni represents the number of reads located in the target methylation region, and Ni,h represents the number of reads containing the target methylation haplotype; (6) For the location of the reference genome, obtain the methylation haplotype ratio status and sequencing depth information at that location, and merge the methylation haplotype ratio status and sequencing depth information data into a data matrix. Remove sites with a missing value ratio higher than 10%, and treat each data point with a depth less than 200 as a missing value, and use the K nearest neighbor (KNN) method to fill the missing values; (7) Divide all samples into two parts, one as the training set and the other as the test set; (8) Discover characteristic methylation segments based on the grouping of training set samples: construct a logistic regression model for each methylation segment for the phenotype, and select the methylation segment with the most significant regression coefficient for each amplified target region to form candidate methylation segments. Randomly divide the training set into ten parts for tenfold cross-validation incremental feature screening. Candidate methylation segments in each region are sorted from largest to smallest according to the significance of the regression coefficient. One methylation segment is added at a time, and the test data is predicted (using a Support Vector Machine (SVM) model). The discrimination index is the mean of the AUC of 10 cross-validations. If the AUC of the training data increases, the candidate methylation segment is retained as a feature methylation segment; otherwise, it is discarded. (9) The feature methylation segments selected in step (8) are input into the support vector machine (SVM) model, and the performance of the model is verified in the test set. (10) The combined data matrix of the predicted score of the training set SVM model in step (9) and the CA19-9 measurement value corresponding to the training set sample is input into the logistic regression model, and the performance of the model after merging CA19-9 is verified in the test set.

[0184] Using the method described above, the inventors screened three genes associated with pancreatic cancer diagnosis (especially the differentiation between pancreatic cancer and pancreatitis): SIX3, TLX2, and CILP2. In this document, the term "gene" includes both the coding and non-coding sequences of the gene in question on the genome. Non-coding sequences include introns, promoters, and regulatory elements or sequences.

[0185] Furthermore, the differentiation between pancreatic cancer and pancreatitis is associated with the methylation level of any one of the following segments or two or all three segments randomly selected from: SEQ ID NO:1 in the SIX3 gene region, SEQ ID NO:2 in the TLX2 gene region, and SEQ ID NO:3 in the CILP2 gene region.

[0186] The “sequences relevant to the differentiation of pancreatic cancer and pancreatitis” mentioned in this article include the above three genes, sequences within 20kb upstream or downstream of them, the above three sequences (SEQ ID NO:1-3) or their complementary sequences.

[0187] The locations of the above three sequences in human chromosomes are as follows: SEQ ID NO:1:chr2, 45028785-45029307; SEQ ID NO:2:chr2, 74742834-74743351; SEQ ID NO:3:chr19, 19650745-19651270. In this document, the base numbers of each sequence and methylation site correspond to the reference genome HG19. Methylatable sites in the genes or fragments described herein can be obtained through routine detection (e.g., sequencing) or by searching on NCBI, see patent application CN202110680924.

[0188] In this document, methods for detecting DNA methylation are well known in the art, such as bisulfite-based PCR (e.g., methylation-specific PCR), DNA sequencing (e.g., bisulfite sequencing, whole-genome methylation sequencing, simplified methylation sequencing), methylation-sensitive restriction endonuclease assays, quantitative fluorescence assays, methylation-sensitive high-resolution melting curve assays, chip-based methylation mapping, and mass spectrometry (e.g., mass spectrometry of flight). In one or more embodiments, the detection includes detecting any strand at a gene or site. The sequencing depth of the DNA sequencing is typically greater than or equal to 5M, preferably at least 7M, 11M, 13M, or 15M. In a specific implementation plan, methylation detection is performed using MethylTitan sequencing, which includes the following steps: First, the extracted DNA is converted to bisulfite and dephosphorylated, then a universal adapter with a UMI (unique molecular identifier) ​​is ligated, and after purification, a second strand is synthesized. Then, multiplex PCR is performed to amplify the target methylated region. A portion of the PCR product is ligated with barcodes and adapters to construct a sequencing library, and the library is sequenced at PE150 (paired-end 150bp) on an Illumina NextSeq or NovaSeq sequencer.

[0189] Therefore, this invention relates to reagents for detecting DNA methylation. Reagents used in the above-described methods for detecting DNA methylation are well known in the art. In detection methods involving DNA amplification, reagents for detecting DNA methylation include primers. The primer sequences may be methylation-specific or non-specific. Preferably, the primer sequences may include non-methylation-specific blocking sequences. Reagents for detecting DNA methylation may also include probes. Typically, the probe sequence is labeled with a fluorescent reporter group at the 5' end and a quencher group at the 3' end. Exemplarily, the probe sequence contains MGB or LNA.

[0190] The term "primer" as used in this article refers to a nucleic acid molecule with a specific nucleotide sequence that guides the synthesis of nucleotides at the initiation of nucleotide polymerization. Primers are typically two artificially synthesized oligonucleotide sequences. One primer is complementary to one DNA template strand at one end of the target region, and the other primer is complementary to the other DNA template strand at the other end of the target region. Their function is to serve as the initiation point for nucleotide polymerization. Primers are usually at least 9 bp. Artificially designed primers are widely used in polymerase chain reaction (PCR), qPCR, sequencing, and probe synthesis. Typically, primers are designed to amplify products with lengths of 1-2000 bp, 10-1000 bp, 30-900 bp, 40-800 bp, 50-700 bp, or at least 150 bp, at least 140 bp, at least 130 bp, or at least 120 bp.

[0191] As described herein, base conversions can occur between DNA or RNA bases. The terms "conversion," "cytosine conversion," or "CT conversion" used herein refer to the process of treating DNA using non-enzymatic or enzymatic methods to convert unmodified cytosine bases (C) into bases with a lower binding affinity to guanine (e.g., uracil bases (U)). Non-enzymatic or enzymatic methods for performing cytosine conversions are well known in the art. Exemplarily, non-enzymatic methods include treatment with conversion reagents such as bisulfites, acid sulfites, or metabisulfites, such as calcium bisulfite, sodium bisulfite, potassium bisulfite, ammonium bisulfite, sodium disulfite, potassium disulfite, and ammonium disulfite. Exemplarily, enzymatic methods include deaminase treatment. The converted DNA may optionally be purified. DNA purification methods suitable for use herein are well known in the art.

[0192] The present invention also provides a kit for diagnosing pancreatic cancer, the kit comprising reagents or devices for detecting DNA methylation and reagents or devices for detecting CA19-9 levels.

[0193] Reagents for detecting DNA methylation are used to determine the methylation level of a DNA sequence or fragment thereof in a sample of a subject, or the methylation status or level of one or more CpG dinucleotides in said DNA sequence or fragment thereof. Exemplary reagents for detecting DNA methylation include the primers and / or probes described herein for detecting the methylation level of sequences related to the identification of pancreatic cancer and pancreatitis discovered by the inventors.

[0194] The term "hybridization" as used in this article primarily refers to nucleic acid sequence pairing under stringent conditions. An exemplary stringent condition is hybridization followed by membrane washing in a solution of 0.1×SSPE (or 0.1×SSC) and 0.1% SDS at 65°C.

[0195] The kit may also include a converted positive standard, wherein unmethylated cytosine is converted into a base that does not bind to guanine. The positive standard may be fully methylated. In addition, the kit contains other reagents required for the detection of DNA methylation. Exemplarily, other reagents for detecting DNA methylation may include one or more of the following: bisulfite and its derivatives, PCR reaction reagents, methylation-sensitive or non-methylation-sensitive restriction endonucleases, enzyme digestion buffers, fluorescent dyes, fluorescence quenchers, fluorescent reporter agents, exonucleases, alkaline phosphatases, internal standards, and controls. The PCR reaction reagents include polymerase, PCR buffer, dNTPs, and Mg2+. 2+ .

[0196] The CA19-9 level described herein primarily refers to the CA19-9 level in body fluids (e.g., blood or plasma). Reagents for detecting CA19-9 levels can be any reagent known in the art that can be used in CA19-9 detection methods, such as immunoreaction-based detection reagents, including but not limited to: CA19-9 antibodies, and optionally buffer solutions, washing solutions, etc. The exemplary detection method used in this invention detects CA19-9 content using chemiluminescent immunoassay. The specific steps are as follows: First, CA19-9 antibodies are labeled with a chemiluminescent marker (acridinium ester). The labeled antibody and CA19-9 antigen undergo an immunoreaction to form a CA19-9 antigen-acridinium ester-labeled antibody complex. Then, an oxidant (H2O2) and NaOH are added to create an alkaline environment. At this point, the acridine ester decomposes and emits light without a catalyst. The photon energy generated per unit time is received and recorded by a photoconductor and a photomultiplier tube (chemiluminescence detector). The integral of this light is proportional to the amount of CA19-9 antigen, and the CA19-9 content can be calculated based on a standard curve.

[0197] The present invention also includes a method for diagnosing pancreatic cancer, comprising: (1) obtaining the methylation level of a DNA sequence or fragment thereof in a sample, or the methylation status or level of one or more CpG dinucleotides in the DNA sequence or fragment thereof, and the CA19-9 level of the sample; (2) calculating a methylation score using the methylation status or level using a mathematical model (e.g., a support vector machine model or a random forest model); (3) merging the methylation score and the CA19-9 level into a data matrix; (4) constructing a pancreatic cancer diagnostic model (e.g., a logistic regression model) based on the data matrix; and optionally (5) obtaining a pancreatic cancer score; and diagnosing pancreatic cancer based on whether the pancreatic cancer score reaches a threshold. The method may further include DNA extraction and / or quality control prior to step (1). The present invention is particularly suitable for differentiating pancreatic cancer from patients with pancreatitis, i.e., differentiating between pancreatic cancer and pancreatitis.

[0198] In this document, the samples are derived from mammalian subjects, preferably humans. Samples can be derived from any organ (e.g., pancreas), tissue (e.g., epithelial tissue, connective tissue, muscle tissue, and nerve tissue), cell, or body fluid (e.g., blood, plasma, serum, tissue fluid, urine). Generally, the sample is acceptable as long as it contains genomic DNA or cfDNA (circulating free DNA or cell-free DNA). cfDNA, also known as circulating cell-free DNA or cell-free DNA, is a fragment of degraded DNA released into the plasma. Exemplarily, the sample is a pancreatic cancer biopsy, preferably a fine-needle aspiration biopsy. Alternatively, the sample may be plasma or cfDNA.

[0199] The subjects are, for example, patients diagnosed with or previously diagnosed with pancreatitis. That is, in one or more embodiments, the method identifies pancreatic cancer in patients diagnosed with chronic pancreatitis (including those previously diagnosed). Of course, the method of the present invention is not limited to the above-described subjects and can also be used to directly diagnose or differentiate pancreatitis or pancreatic cancer in undiagnosed subjects.

[0200] In a specific implementation, step (1) includes detecting the methylation level of a DNA sequence or fragment thereof in a sample of the target, or the methylation state or level of one or more CpG dinucleotides in the DNA sequence or fragment thereof, for example, by detecting the methylation state or level using primer molecules and / or probe molecules described herein.

[0201] Methods for detecting methylation status or levels, and for detecting CA19-9 levels, are described elsewhere in this document. One specific method for detecting methylation status or levels includes: treating genomic DNA or cfDNA with a transformation reagent to convert unmethylated cytosine into bases with a lower binding affinity for guanine (e.g., uracil); performing PCR amplification using primers suitable for amplifying transformed sequences relevant to the identification of pancreatic cancer and pancreatitis described herein; and determining the methylation level of at least one CpG by the presence or absence of the amplification product, or by sequence identification (e.g., probe-based PCR detection or DNA sequencing).

[0202] Alternatively, step (1) may also include: treating genomic DNA or cfDNA with a methylation-sensitive restriction endonuclease; performing PCR amplification using primers suitable for amplifying sequences having at least one CpG in the pancreatic cancer and pancreatitis identification-related sequences described herein; and determining the methylation level of at least one CpG by the presence or absence of the amplification product.

[0203] The term "methylation level" as used herein refers to the relationship between the methylation levels of any number and any position of CpGs in the sequence in question. This relationship can be the result of addition or subtraction of methylation level parameters (e.g., 0 or 1) or calculations using mathematical algorithms (e.g., mean, percentage, fraction, proportion, degree, or calculations using mathematical models), including but not limited to methylation level measures, methylation haplotype ratios, or methylation haplotype loadings. For example, the methylation level may represent the average methylation level of multiple CpG sites within a region. The methylation level of each CpG site here refers to the percentage of C methylated at that site. Therefore, an increase or decrease in the methylation level of a region does not necessarily indicate an increase or decrease in the methylation level of all CpG sites within the region. The process of converting the results of methods for detecting DNA methylation (e.g., simplified methylation sequencing) into methylation levels is known in the art. The term "methylation status" indicates the methylation of a specific CpG site, typically including methylated or unmethylated (e.g., methylation status parameter 0 or 1).

[0204] The terms "methylation score" or "methylation-based disease predictive value" have the same meaning, referring to a disease predictive value obtained by calculating the methylation status or level using a mathematical model. Conventional mathematical model analysis methods are known in the art; an exemplary method is the Support Vector Machine (SVM) mathematical model. For example, for differential methylation biomarkers, an SVM is constructed on the training set samples, and the accuracy, sensitivity, and specificity of the detection results, as well as the area under the predictive value characteristic curve (ROC) (AUC), are used to calculate the predicted score for the test set samples.

[0205] In a preferred embodiment, the model training process is as follows: First, differentially methylated regions are obtained based on the methylation level of each site, and a differentially methylated region matrix is ​​constructed. For example, a methylation data matrix can be constructed from the methylation level data of a single CpG dinucleotide position in the HG19 genome using software such as samtools; then, SVM model training is performed.

[0206] An exemplary SVM model training process is as follows:

[0207] a) Use the sklearn package (v0.23.1) of Python software (v3.6.9) to build the training mode of the cross-validation training model. Command line: model=SVR().

[0208] b) Using the sklearn package (v0.23.1), input the data matrix and build an SVM model, model.fit(x_train,y_train), where x_train represents the training set data matrix and y_train represents the phenotypic information of the training set.

[0209] This article also relates to methods for obtaining methylation haplotype ratios associated with pancreatic cancer and pancreatitis. Taking methylation data obtained from methylation-targeted sequencing (MethylTitan) as an example, the process of screening and testing biomarker sites is as follows: raw paired-end sequencing reads – readings are merged to obtain merged single-end reads – adapters are removed to obtain adapter-removed reads – Bismark is aligned to the human DNA genome to form a BAM file – samtools extracts the CpG site methylation level of each read to form a haplotype file – the proportion of C site methylation haplotype ratios is statistically analyzed to form a meth file – MHF (Methylated Haplotype Fraction) methylation values ​​are calculated – Coverage 200 filter sites are used to form a meth.matrix matrix file – filter sites according to NA values ​​greater than 0.1 – samples are pre-divided into training and test sets – for each haplotype in the training set, a logistic regression model is constructed for the phenotype, and the regression P-value of each methylation haplotype ratio is selected – the methylation haplotype ratio with the most significant P-value in each MethylTitan amplification region is selected to represent the methylation haplotype ratio level of that region and modeled using a support vector machine – the results of the training set (ROC plot) are generated and the model is used to predict the test set for validation. Specifically, the method for obtaining the methylation haplotype ratio associated with pancreatic cancer includes the following steps: (1) obtaining plasma samples from patients with pancreatic cancer or pancreatitis to be tested, extracting cfDNA, and performing library construction and sequencing using the MethylTitan method to obtain sequencing reads; (2) preprocessing the sequencing data, including adapter removal and splicing of the sequencing data generated by the sequencer; (3) aligning the preprocessed sequencing data to the HG19 reference genome sequence of the human genome to determine the position of each fragment. The data in step (2) can be obtained from paired-end 150bp sequencing on the Illumina sequencing platform. Adapter removal in step (2) involves removing the sequencing adapters at the 5' and 3' ends of the two paired-end sequencing data respectively, as well as low-quality base removal after adapter removal. Splicing in step (2) involves merging the paired-end sequencing data to restore the original library fragment. This allows for better alignment and accurate location of sequencing fragments. For example, the length of the sequencing library is about 180bp, and paired-end 150bp can completely cover the entire library fragment. Step (3) includes: (a) converting the HG19 reference genome data to CT and GA respectively, constructing two sets of converted reference genomes, and constructing alignment indexes for the converted reference genomes respectively; (b) converting the merged sequencing sequence data to CT and GA respectively; (c) aligning the converted reference genome sequences respectively, and finally summarizing the alignment results to determine the position of the sequencing data in the reference genome.

[0210] Furthermore, the method for obtaining methylation haplotype ratios also includes (4) calculating MHF; (5) constructing a methylation haplotype ratio MHF data matrix; and (6) constructing a logistic regression model for each methylation haplotype ratio based on sample grouping. Step (4) involves obtaining the methylation haplotype ratio status and sequencing depth information at the location of the HG19 reference genome based on the alignment results obtained in step (3). Step (5) involves merging the methylation haplotype ratio status and sequencing depth information data into a data matrix. Among them, each data point with a depth less than 200 is treated as a missing value, and the missing values ​​are filled using the K nearest neighbor (KNN) method. Step (6) involves statistically modeling each location in the above matrix using logistic regression and screening for haplotypes with significant regression coefficients between the two groups.

[0211] According to the inventors, combining methylation scores with CA19-9 levels can significantly improve diagnostic accuracy. Specifically, methylation scores and CA19-9 levels are combined into a data matrix, and then a pancreatic cancer diagnostic model (e.g., a logistic regression model) is constructed based on the data matrix to obtain a pancreatic cancer score.

[0212] The data matrix of methylation score and CA19-9 level is arbitrarily standardized. Standardization can be performed using conventional standardization methods in the art. In this embodiment of the invention, the RobustScaler standardization method is used, and the standardization formula is as follows:

[0213]

[0214] Where x and x' are the sample data before and after normalization, respectively, median is the median of the sample, and IQR is the interquartile range of the sample.

[0215] Similar to methylation scoring, methods using conventional mathematical models and processes for determining thresholds through data matrices are known in the art, such as support vector machine (SVM) models, random forest models, or logistic regression models. An exemplary approach is the logistic regression model. For example, for differential methylation biomarkers, a logistic regression model is used on the training set samples, and the accuracy, sensitivity, and specificity of the detection results, as well as the area under the predictive curve (ROC) (AUC), are used to calculate the predicted scores for the test set samples. When the pancreatic cancer score combining methylation level and CA19-9 level meets a certain threshold, pancreatic cancer is identified; otherwise, chronic pancreatitis is identified.

[0216] The term "multiple" as used herein refers to any integer. Preferably, "multiple" in "one or more" can be any integer greater than or equal to 2, including 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60 or more.

[0217] Example

[0218] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments. In the following embodiments, experimental methods without specific conditions are generally performed according to the methods described under conventional conditions.

[0219] Example 1: Methylation-targeted sequencing to screen for characteristic methylation sites

[0220] The inventors collected blood samples from a total of 94 pancreatic cancer patients and 25 patients with chronic pancreatitis. All enrolled patients signed informed consent forms. The pancreatic cancer patients had a history of pancreatitis diagnosis. Sample information is shown in the table below.

[0221]

[0222] Plasma DNA methylation sequencing data were obtained using the MethylTitan method to identify DNA methylation classification markers. The workflow is as follows: Figure 1 The specific process is as follows:

[0223] 1. Extraction of plasma cfDNA samples

[0224] A 2ml whole blood sample was collected from the patient using a streck blood collection tube. Plasma was separated by centrifugation within 3 days and then transferred to the laboratory. cfDNA was extracted using the QIAGEN QIAamp Circulating Nucleic Acid Kit according to the instructions.

[0225] 2. Sequencing and Data Preprocessing

[0226] 1) The library was sequenced using an Illumina Nextseq 500 sequencer for paired-end sequencing.

[0227] 2) The Pear (v0.6.0) software merges the paired-end sequencing data of the same fragment of 150bp from the Illumina Hiseq X10 / Nextseq 500 / Nova seq sequencer into a single sequence with a minimum overlap length of 20bp and a minimum length of 30bp after merging.

[0228] 3) Trim_galore v0.6.0 and cutadapt v1.8.1 software were used to remove adapters from the merged sequencing data. The adapter sequence “AGATCGGAAGAGCAC” was removed from the 5' end of the sequence, and bases with sequencing quality values ​​below 20 at both ends were removed.

[0229] 3. Sequencing data alignment

[0230] The reference genome data used in this article came from the UCSC database (UCSC:HG19, http: / / hgdownload.soe.ucsc.edu / goldenPath / hg19 / bigZips / hg19.fa.gz).

[0231] 1) First, HG19 was transformed into cytosine to thymine (CT) and adenine to guanine (GA) using Bismark software, and the transformed genomes were indexed using Bowtie2 software.

[0232] 2) Perform CT and GA conversion on the preprocessed data as well.

[0233] 3) Use Bowtie2 software to align the transformed sequences to the transformed HG19 reference genome. The minimum seed sequence length is 20, and mismatches in the seed sequence are not allowed.

[0234] 4. Calculation of MHF

[0235] For each target region HG19 CpG site, the methylation status corresponding to each site is obtained based on the alignment results above. In this paper, the nucleotide number of the site corresponds to the nucleotide position number of HG19. A target methylation region may have multiple methylation haplotypes. For each methylation haplotype within the target region, this value needs to be calculated. An example of the MHF calculation formula is as follows:

[0236]

[0237] Where i represents the target methylation region, h represents the target methylation haplotype, and N i N represents the number of reads located in the target methylation region. i,h This indicates the number of reads containing the target methylated haplotype.

[0238] 5. Methylation data matrix

[0239] 1) Merge the methylation sequencing data of each sample in the training set and the test set into a data matrix, and perform missing value processing on each site with a depth of less than 200.

[0240] 2) Remove sites with a missing value ratio higher than 10%.

[0241] 3) For missing values ​​in the data matrix, the KNN algorithm is used to impute the missing data.

[0242] 6. Identify characteristic methylation regions by grouping training set samples.

[0243] 1) For each methylation segment, a logistic regression model is constructed for the phenotype. For each amplified target region, the methylation segment with the most significant regression coefficient is selected to form a candidate methylation segment.

[0244] 2) Randomly divide the training set into ten parts and perform incremental feature selection with tenfold cross-validation.

[0245] 3) Candidate methylated segments in each region are sorted from largest to smallest according to the significance of the regression coefficient. One methylated segment is added at a time, and the test data is predicted (Support Vector Machine (SVM) model).

[0246] 4) In step 3), use the 10 datasets generated in step 2) to calculate the AUC 10 times each time, and take the average of the 10 calculations. If the AUC of the training data increases, retain the candidate methylation region as a feature methylation region; otherwise, discard it.

[0247] The distribution of the selected characteristic methylation markers in HG19 is as follows: SEQ ID NO:1 in the SIX3 gene region, SEQ ID NO:2 in the TLX2 gene region, and SEQ ID NO:3 in the CILP2 gene region. The levels of these methylation markers increased or decreased in cfDNA of pancreatic cancer patients (Table 1). The sequences of the above three marker regions are shown in SEQ ID NO:1-3.

[0248] The mean methylation levels of methylation markers in pancreatic cancer and chronic pancreatitis patients in the training and test sets are shown in Tables 1 and 2, respectively. The distributions of the methylation levels of the three methylation markers in pancreatic cancer and chronic pancreatitis patients in the training and test sets are shown in Tables 1 and 2, respectively. Figure 2 and Figure 3 As shown in the figure, the methylation levels of the three methylation markers differed significantly between pancreatic cancer and chronic pancreatitis patients, demonstrating good discriminatory power.

[0249] Table 1: Methylation levels of DNA methylation biomarkers in the training set

[0250] sequence markers pancreatic cancer Chronic pancreatitis SEQ ID NO:1 chr2:45028785-45029307 0.843731054 0.909570522 SEQ ID NO:2 chr2:74742834-74743351 0.953274962 0.978544302 SEQ ID NO:3 chr19:19650745-19651270 0.408843665 0.514101315

[0251] Table 2: Methylation levels of DNA methylation biomarkers in the test set

[0252] sequence markers pancreatic cancer Chronic pancreatitis SEQ ID NO:1 chr2:45028785-45029307 0.843896661 0.86791556 SEQ ID NO:2 chr2:74742834-74743351 0.926459851 0.954493044 SEQ ID NO:3 chr19:19650745-19651270 0.399831579 0.44918572

[0253] Example 2: Constructing a machine learning-based classification prediction model

[0254] To validate the potential of a pancreatic cancer-chronic pancreatitis patient classifier using DNA methylation markers (such as methylation haplotype ratios), a support vector machine (SVM) disease classification model (pp_model) based on a combination of three DNA methylation markers was constructed in the training group. Simultaneously, a logistic regression disease classification model (cpp_model) based on a combined data matrix of SVM model predicted scores and CA19-9 measurements was constructed. The classification prediction performance of both models was validated in the test group. The training and test groups were divided proportionally, with 80 cases (samples 1-80) in the training group and 39 cases (samples 80-119) in the test group.

[0255] A support vector machine model was built on the training set using the discovered DNA methylation markers.

[0256] 1) Divide the samples into two parts in advance, one part for training the model and the other part for testing the model.

[0257] 2) To explore the potential of using methylation biomarkers for pancreatic cancer identification, a disease classification system based on gene biomarkers was developed. An SVM model was trained using the methylation biomarker levels in the training set. The specific training process is as follows:

[0258] a) Use the sklearn package (v0.23.1) of Python software (v3.6.9) to build the training model. Command line: pp_model = SVR().

[0259] b) Using the sklearn package (v0.23.1), input the methylation numerical matrix and construct an SVM model: pp_model.fit(train_df,train_pheno), where train_df represents the methylation numerical matrix of the training set, train_pheno represents the phenotypic information of the training set, and pp_model represents the SVM model constructed using the numerical matrices of three methylation markers.

[0260] c) Input the training and test set data into the pp_model to obtain the predicted scores: train_pred

[0261] =pp_model.predict(train_df)

[0262] test_pred=pp_model.predict(test_df)

[0263] Where train_df and test_df are the methylation numerical matrices of the training and test sets, respectively, and train_pred,

[0264] test_pred represents the predicted scores of the pp_model model for the training and test sets, respectively.

[0265] 3) To improve the ability to differentiate between pancreatic cancer and pancreatitis, the CA19-9 detection value was included in the model. The specific process is as follows:

[0266] d) Combine the SVM model predictions and corresponding CA19-9 measurement data from the training set into a data matrix and then standardize it:

[0267] Combine_scalar_train=RobustScaler().fit(combine_train_df)

[0268] Combine_scalar_test=RobustScaler().fit(combine_test_df)

[0269] scaled_combine_train_df=Combine_scalar_train.transform

[0270] (combine_train_df)

[0271] scaled_combine_test_df=Combine_scalar_test.transform(combine_test_df)

[0272] Where combine_train_df and combine_test_df represent the test set and the data matrix of the training set samples obtained by the prediction scores of the pp_model prediction model constructed in this embodiment and the CA19-9 combined, respectively; scaled_combine_train_df and scaled_combine_test_df represent the training set and test set data matrices after standardization, respectively.

[0273] e) Construct a logistic regression model using the combined normalized data matrix of the model prediction scores and CA19-9 measurements from the training set pp_model, and then use this model to predict the model prediction scores and the combined normalized data matrix of CA19-9 measurements from the test set pp_model:

[0274] cpp_model=LogisticRegression().fit(scaled_combine_train_df,train_pheno)

[0275] combine_test_pred=cpp_model.predict(scaled_combine_test_df)

[0276] Where cpp_model represents the logistic regression model fitted using the training set data matrix after incorporating CA19-9 detection values ​​and standardization; combine_test_pred represents the predicted score of cpp_model on the test set.

[0277] In the process of building the model, pancreatic cancer type was coded as 1 and chronic pancreatitis type was coded as 0. According to the distribution of predicted scores by the model, the thresholds of pp_model and cpp_model were set to 0.892 and 0.885 respectively. Based on the two models, when the predicted score is higher than the threshold, the patient is identified as a pancreatic cancer patient, and otherwise as a pancreatitis patient.

[0278] The prediction scores of the two models for the training and test sets are shown in Tables 3 and 4, respectively. The distribution of the prediction scores is shown in Table 4. Figure 5 ROC curves for the two machine learning models and those using CA19-9 measurements alone are shown in [link to ROC curves]. Figure 4 Among them, the AUC value of CA19-9 alone was 0.84, the AUC value of pp_model was 0.88, and the AUC value of cpp_model was 0.90. The SVM model (pp_model) built using three methylation markers significantly outperformed CA19-9. The logistic regression model cpp_model, which was built by adding CA19-9 detection values ​​to the predicted values ​​of pp_model, outperformed pp_model.

[0279] Statistical analysis was performed on the test set using a defined threshold (CA19-9 used the accepted threshold of 37). Sensitivity and specificity are shown in Table 5. With 100% specificity on the test set, cpp_model achieved a sensitivity of 87% for pancreatic cancer patients, outperforming pp_model and CA19-9.

[0280] In addition, the performance of the two models in CA19-9 negative (<37) samples was analyzed. The results are shown in Table 6. It can be seen that cpp_model can still achieve a sensitivity of 63% and a specificity of 100% for CA19-9 negative pancreatic cancer patients in the test set.

[0281] Table 3: Prediction scores and discrimination results of the two models on the training set.

[0282]

[0283]

[0284]

[0285] Table 4: Prediction scores and discrimination results of the two models on the test set.

[0286]

[0287]

[0288] Table 5: Sensitivity and Specificity of CA19-9 and Two Machine Learning Models

[0289]

[0290] Table 6: Performance of the two machine learning models in distinguishing negative samples in CA19-9

[0291]

[0292] This study investigated the differences in plasma methylation levels of cfDNA markers between patients with chronic pancreatitis and those with pancreatic cancer by examining the methylation levels of cfDNA methylation markers in plasma. Three DNA methylation markers with significant differences were identified. Based on these DNA methylation markers and incorporating CA19-9 detection values, a risk prediction model for malignant pancreatic cancer was established using support vector machine and logistic regression. This model effectively distinguishes between pancreatic cancer and chronic pancreatitis in patients diagnosed with chronic pancreatitis, exhibiting high sensitivity and specificity, and is suitable for screening and diagnosing pancreatic cancer in patients with chronic pancreatitis. sequence list <110> Shanghai Kunyuan Biotechnology Co., Ltd. Jiangsu Kunyuan Biotechnology Co., Ltd. <120> A machine learning-based method for pancreatic cancer diagnosis <130> 215332 <160> 3 <170> SIPOSequenceListing 1.0 <210> 1 <211> 523 <212> DNA <213> Homo sapiens <400> 1 taatttatgg aatccaccgt cacactctct ccgagcagcc agctccccgc ttaacgggga 60 aattgaagca gacagccttt gtctaaacac ttctttttgcc cagaatatct taattttcct 120 atttgaatgt ttaataaggt ttggggtgca gcagcttcct tttaattgtg acggtgcggc 180 cgcttgggcg tgatcccttg gctggggctg cagggggccc gtcctccagg ggcgcagagg 240 gaaggaccag cgtttccaag ccgggctctg gccgccggcg cgagagcgag gccaaggtct 300 gggggcagtt cagggggacc ccgaagtcgg gacggcccag aaacgctttg cccacagcca 360 ccgccctttc ctttgtgagt ttccccaaag ccgtcggtgc gacccggcgc cgactctcct 420 cctctctcc ctgcgagggc cgcgcgcc cgggcccagt cctgggggat agatcctcg 480 gggcccaacg gctgggccac cgccggctc cggccactgc tgc 523 <210> 2 <211> 518 <212> DNA <213> Homo sapiens <400> 2 aagccgcgca cgtccttctc ccgctcacag gtgctggagt tggagcggcg cttcctgcgc 60 cagaagtacc tggcctctgc ggagagggcg gcgctggcca aggccttgcg catgaccgac 120 gcacaggtca aaacgtggtt ccagaaccga cgcaccaagt ggcggtgagg cgcggcgcgg 180 gcgagggcgg actggggttc ccgagcaggg cctggtgaga agcgacgcgg cgggcgcccc 240 gctgaccccg cgtctccctc ccttaggcgc cagacggcgg aggagcgcga ggccgagcgg 300 caccgcgcgg gccggctgct cctgcatctg cagcaggacg cgttgccacg gccgctgcgg 360 ccgccgctgc ccccggaccc tctctgcctg cacaactcgt cgctcttcgc gctgcagaac 420 ctgcagccct gggccgagga caacaaagtg gcttcagtgt ccgggctcgc ctcggtggtg 480 tgagcgacgc ccgtccgatc ggcgtggagc gccgggcc 518 <210> 3 <211> 526 <212> DNA <213> Homo sapiens <400> 3 ttcaagatct aagtgagagg ccggtcagac agaggcaaga gctcagcgca ccgggatgga 60 ccaggtcagg ccctgggcgg cagaactggg gtcgcgggga acccagtctg ccctgcacct 120 gtttcaggcc gctggctcgg gtcgtgggcg cgctcggcta gccggtgccc accgggggag 180 ggggctgaga cagcaagtaa ggcctttgca cgcatgcatg ggggcctaca ggccgccgcc 240 ctggtcccag cgcgtgcggt gcccgcagag gccagcgagt ggacgtcctg gttcaacgtg 300 gaccaccccg gaggcgacgg cgacttcgag agcctggctg ccatccgctt ctactacggg 360 ccagcgcgcg tgtgcccgcg accgctggcg ctggaagcgc gcaccacgga ctgggccctg 420 ccgtccgccg tcggcgagcg cgtgcacttg aaccccacgc gcggcttctg gtgcctcaac 480 cgcgagcaac cgcgtggccg ccgctgctcc aactaccacg tgcgct 526

Claims

1. A method for constructing a diagnostic model for pancreatic cancer, comprising: (1) Obtain the methylation level of a DNA sequence or fragment thereof in the subject sample, or the methylation status or level of one or more CpG dinucleotides in the DNA sequence or fragment thereof, and the CA19-9 level of the subject, wherein the DNA sequence is selected from the following gene sequences: (a) CILP2; (b) CILP2, SIX3; (c) CILP2, TLX2; or (d) CILP2, SIX3, TLX2, wherein the fragment contains at least one CpG dinucleotide. (2) A methylation score is calculated using a mathematical model based on the methylation state or level. (3) Combine the methylation score and CA19-9 level data matrix. (4) Construct a pancreatic cancer diagnostic model based on the data matrix.

2. The method as described in claim 1, characterized in that, The method also includes one or more features selected from the following: Step (1) includes detecting the methylation level of a DNA sequence or fragment thereof in the sample of the target, or the methylation status or level of one or more CpG dinucleotides in the DNA sequence or fragment thereof. The samples were derived from mammalian tissues, cells, or body fluids. CA19-9 level refers to the CA19-9 level in blood or plasma. The mathematical model described in step (2) is a support vector machine model. The pancreatic cancer diagnostic model described in step (4) is a logistic regression model.

3. The method as described in claim 1, characterized in that, The samples were derived from pancreatic tissue or blood.

4. A method for constructing a diagnostic model for pancreatic cancer, comprising: (1) Obtain the methylation haplotype ratio and sequencing depth of the target genomic DNA segment. Optionally (2) preprocesses the methylation haplotype ratios and sequencing depth data. (3) Perform cross-validation incremental feature screening to obtain characteristic methylation regions. (4) A mathematical model is constructed based on the methylation detection results of the characteristic methylated regions to obtain a methylation score. (5) Construct a pancreatic cancer diagnostic model based on methylation scores and corresponding CA19-9 levels; The DNA segment is selected from the regions corresponding to the following nucleotide sequences: (a) SEQ ID NO: 3; (b) SEQ ID NO: 3 and SEQ ID NO: 1; (c) SEQ ID NO: 3 and SEQ ID NO: 2; or (d) SEQ ID NO: 3, SEQ ID NO: 1 and SEQ ID NO:

2.

5. The method as described in claim 4, characterized in that, The method also includes one or more features selected from the following: Step (1) includes: 1.1) DNA methylation detection was performed on the target sample to obtain sequencing read data. 1.2) Optionally, preprocess the sequencing data. 1.3) Align the sequencing data to the reference genome to obtain the location and sequencing depth information of methylated regions. 1.4) Calculate the methylation haplotype ratio (MHF) of the segment according to the following formula: in i Indicates the target methylation region. h Indicates the target methylated haplotype. N i This indicates the number of reads located in the target methylation region. N i,h This indicates the number of reads containing the target methylated haplotype; Step (2) includes: (2.1) merging the methylation haplotype ratio status and sequencing depth information data into a data matrix, Step (3) includes: using a mathematical model to perform cross-validation incremental feature selection on the training data, wherein the DNA segment that increases the AUC of the mathematical model is the feature methylation segment. Step (5) includes: merging the methylation score and CA19-9 level into a data matrix, and constructing a pancreatic cancer diagnostic model based on the data matrix.

6. The method as described in claim 5, characterized in that, In step 1.2), the pretreatment refers to the removal of joints and / or splicing.

7. The method as described in claim 6, characterized in that, Step (2) further includes: 2.2) Remove sites in the data matrix with a missing value ratio higher than 5-15%, and / or 2.3) Treat each data point with a depth less than 300 as a missing value and fill in the missing values.

8. The method as described in claim 7, characterized in that, Step 2.2) removes sites in the data matrix with a missing value ratio higher than 10%.

9. The method as described in claim 7, characterized in that, In step 2.3), each data point with a depth less than 200 is treated as a missing value and the missing value is filled.

10. The method as described in claim 7, characterized in that, The method for filling missing values ​​in step 2.3) is the K-nearest neighbor method.

11. The method according to any one of claims 4-10, characterized in that, The method also includes one or more features selected from the following: The mathematical model in step (4) is a vector machine (SVM) model. The methylation detection result in step (4) is a combined matrix of methylation haplotype ratio and sequencing depth. The pancreatic cancer diagnostic model in step (5) is a logistic regression model.

12. Use of reagents or apparatus for detecting DNA methylation and reagents or apparatus for detecting CA19-9 levels in the preparation of a kit for diagnosing pancreatic cancer, wherein the reagents or apparatus for detecting DNA methylation are used to determine the methylation level of a DNA sequence or fragment thereof in a sample of a subject, or the methylation status or level of one or more CpG dinucleotides in the DNA sequence or fragment thereof; the fragment comprising at least one CpG dinucleotide, the DNA sequence being selected from the following gene sequences: (a) CILP2; (b) CILP2, SIX3; (c) CILP2, TLX2; or (d) CILP2, SIX3, TLX2.

13. The use as described in claim 12, characterized in that, The use also includes one or more of the following features: The reagent for detecting DNA methylation comprises primer molecules that hybridize with the DNA sequence or a fragment thereof, the primer molecules being capable of amplifying the DNA sequence or a fragment thereof. The reagent for detecting DNA methylation contains probe molecules that hybridize with the DNA sequence or fragments thereof. The reagents for detecting CA19-9 levels are based on immune response assays. The kit also includes PCR reaction reagents. The kit also includes other reagents for detecting DNA methylation, which are selected from one or more of the following methods: bisulfite-based PCR, DNA sequencing, methylation-sensitive restriction endonuclease analysis, quantitative PCR, methylation-sensitive high-resolution melting curve analysis, chip-based methylation mapping, and mass spectrometry. The diagnosis includes: calculating by constructing the pancreatic cancer diagnostic model as described in any one of claims 1-10, and diagnosing pancreatic cancer based on the score.

14. A kit for diagnosing pancreatic cancer, comprising: (1) A reagent or apparatus for detecting DNA methylation, used to determine the methylation level of a DNA sequence or fragment thereof in a sample of a subject, or the methylation state or level of one or more CpG dinucleotides in the DNA sequence or fragment thereof, and (2) Reagents or apparatus for detecting CA19-9 levels; The fragment contains at least one CpG dinucleotide. The DNA sequence is selected from the following gene sequences: (a) CILP2; (b) CILP2, SIX3; (c) CILP2, TLX2; or (d) CILP2, SIX3, TLX2.

15. The kit according to claim 14, characterized in that, The kit also includes one or more features selected from the following: The reagent for detecting DNA methylation comprises primer molecules that hybridize with the DNA sequence or a fragment thereof, the primer molecules being capable of amplifying the DNA sequence or a fragment thereof. The reagent for detecting DNA methylation contains probe molecules that hybridize with the DNA sequence or fragments thereof. The reagents for detecting CA19-9 levels are based on immune response assays. The kit also includes PCR reaction reagents. The kit also includes other reagents for detecting DNA methylation, which are selected from one or more of the following methods: PCR based on bisulfite conversion, DNA sequencing, methylation-sensitive restriction endonuclease analysis, quantitative fluorescence method, methylation-sensitive high-resolution melting curve method, chip-based methylation mapping analysis, and mass spectrometry.

16. An apparatus for diagnosing pancreatic cancer or constructing a diagnostic model of pancreatic cancer, the apparatus comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it performs the following steps: (1) Obtain the methylation level of the DNA sequence or fragment thereof in the target sample, or the methylation status or level of one or more CpG dinucleotides in the DNA sequence or fragment thereof, and the CA19-9 level of the target. (2) A methylation score is calculated using a mathematical model based on the methylation state or level. (3) Combine the methylation score and CA19-9 level data matrix. (4) Construct a pancreatic cancer diagnostic model based on the data matrix. Optional (5) Obtain a pancreatic cancer score; diagnose pancreatic cancer based on the pancreatic cancer score. or (1) Obtain the methylation level of the DNA sequence or fragment thereof in the target sample, or the methylation status or level of one or more CpG dinucleotides in the DNA sequence or fragment thereof, and the CA19-9 level of the target. (2) A methylation score is calculated using a mathematical model based on the methylation state or level. (3) Obtain the pancreatic cancer score according to the model shown below, and diagnose pancreatic cancer based on the pancreatic cancer score: Where M is the methylation score of the sample calculated in step (2), and C is the CA19-9 level of the sample. The fragment contains at least one CpG dinucleotide. The DNA sequence is selected from the following gene sequences: (a) CILP2; (b) CILP2, SIX3; (c) CILP2, TLX2; or (d) CILP2, SIX3, TLX2.

17. The apparatus as claimed in claim 16, characterized in that, The device further includes one or more features selected from the following: Step (1) includes detecting the methylation level of a DNA sequence or fragment thereof in the sample of the target, or the methylation status or level of one or more CpG dinucleotides in the DNA sequence or fragment thereof. The samples were derived from mammalian tissues, cells, or body fluids. CA19-9 level refers to the CA19-9 level in blood or plasma. The mathematical model described in step (2) is a support vector machine model. The pancreatic cancer diagnostic model described in step (4) is a logistic regression model.

18. The apparatus as claimed in claim 17, characterized in that, The samples were derived from pancreatic tissue or blood.

Citation Information

Patent Citations

  • Methylation marker for identifying pancreatitis and pancreatic cancer and application thereof

    CN115491411A