Cell-free DNA analysis in pancreatic cancer detection and monitoring using combinations of features
A non-invasive liquid biopsy using 5-hydroxymethylcytosine signatures and machine learning models effectively detects pancreatic cancer from blood samples, addressing the limitations of current methods by improving sensitivity and specificity for early detection.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-22
- Publication Date
- 2026-03-19
AI Technical Summary
Current methods for early detection of pancreatic cancer are invasive, costly, and lack sensitivity and specificity, leading to late-stage diagnoses and poor survival outcomes.
A non-invasive liquid biopsy method using a combination of epigenetic markers, including 5-hydroxymethylcytosine signatures and other features, analyzed through whole-genome sequencing and machine learning models to detect pancreatic cancer from blood samples.
The method achieves high accuracy, efficiency, and specificity in detecting pancreatic cancer at early stages, enabling timely intervention and improving survival rates.
Smart Images

Figure 2026509568000004 
Figure 2026509568000005 
Figure 2026509568000006
Abstract
Description
Technical Field
[0001] Submission as Table 1 of "20230322 3599-0016P MODEL" The large table entitled "20230322 3599-0016P MODEL" is included as an ASCII plain text file (in accordance with the provisions of 37 CFR§1.58(c)-(e)) in U.S. Provisional Patent Application No. 63 / 491,738, a priority document filed on March 22, 2023, and was electronically submitted to the United States Patent and Trademark Office together with the patent application. This file was created at 6:42 am on March 22, 2023, has a size of 4.2 MB, and is hereinafter referred to as "Table 1".
Background Art
[0002] Background of the Invention Cancer is the second leading cause of death worldwide. The mortality rate of cancer worsens due to being diagnosed at a late stage with a poor prognosis. Earlier cancer detection provides an opportunity to improve patient outcomes by identifying tumors at a point when treatment is more likely to be effective. Breast cancer, colorectal cancer, and lung cancer are some of the few cancers for which screening methods exist, but the screening tests currently used clinically are expensive, invasive, and may be limited to the detection of a single type of cancer; as a result, multiple tests are required, further increasing the cost for overall early cancer detection and potentially causing treatment delays. Multiple early cancer detection tests based on liquid biopsy aim to address these limitations and complement these screening approaches. Current non-invasive methods for early cancer detection rely on cell-free DNA (cfDNA) obtained from plasma or genetic, epigenetic, or proteomic changes in exosomes circulating in the blood. While these methods can achieve a certain level of performance in cancer detection and treatment response prediction, there is still a need to improve the performance of non-invasive tests in order to detect more cancers earlier (i.e., to increase sensitivity) and to differentiate non-cancer cases from being falsely identified as positive cases (i.e., to increase specificity).
[0003] Pancreatic cancer (PaC) is the fourth leading cause of cancer death in the United States (1). Currently, there are no available molecular early detection tools for PaC (2). The poor survival outcomes in PaC are primarily due to the late diagnosis of the disease in the majority of patients (87%), including distant metastases in 49% of cases (1). Late diagnosis deprives individuals of the opportunity for potentially curative therapeutic interventions such as surgery (3), negatively impacting survival rates, with the 5-year survival rate in localized PaC dropping to less than one-tenth of that in locally metastatic PaC, from 38% to 3% (1). Therefore, early diagnosis is clearly crucial for better survival outcomes in patients with PaC.
[0004] Epigenetic regulation of DNA state and chromatin regulation is known to reinforce cancer development and progression (4, 5), and there has been growing interest in assessing epigenetic profiles for tumor detection and characterization in recent years (6).
[0005] 5-hydroxymethylcytosine (5hmC) is a stable epigenetic marker that arises as the first step in the active demethylation of cytosine bases in DNA by 10-11 translocation enzymes (also known as TET enzymes), indicating a region of active transcription and gene regulation (7). 5hmC positively correlates with gene expression and regulation in multiple biological contexts (8-10). Therefore, 5hmC profiles provide a unique signature that allows for the definition of tissue identity and cellular state (9, 11), making 5hmC a valuable marker for identifying the tissue of origin. In particular, 5hmC profiling has been used to detect cancer in cfDNA in the plasma of individuals with cancer, including pancreatic cancer, lung cancer, hepatocellular carcinoma, colon cancer, and gastric cancer (12-14).
[0006] There are certain genetic syndromes (e.g., Peutz-Jeggers syndrome, or germline mutations in CDKN2A, BRCA2, or PALB2) for which surveillance is recommended in association with an increased risk of developing PaC (15, 16), but there are other well-known risk factors for PaC that are not included in surveillance and management guidelines. More than 50% of individuals diagnosed with PaC have a prior diagnosis of diabetes (17). Furthermore, individuals with diabetes have a 1.5 to 2 times higher risk of developing PaC compared to the general population (18), but the risk is 6 to 8 times higher in individuals aged 50 or older diagnosed with new-onset type 2 diabetes (NOD) (within the past 3 years) (19). In fact, nearly 25% of all new PaC diagnoses in the United States are identified in individuals with NOD (20). Therefore, surveillance of the NOD population for signs of PaC presents an opportunity to advance the diagnosis of PaC to an earlier stage of the disease, thereby improving outcomes through timely intervention. [Overview of the project] [Problems that the invention aims to solve]
[0007] In this field, there is still a need for the development and validation of novel, non-invasive, cfDNA-based methods for the detection, assessment, and monitoring of pancreatic cancer that can be used in clinical settings. More specifically, there is still a need to dramatically improve survival rates associated with pancreatic cancer by providing methods for reliably detecting pancreatic cancer at its early stages. [Means for solving the problem]
[0008] Summary of the Invention Therefore, the present invention provides a method for detecting, assessing, and monitoring pancreatic cancer without requiring surgical biopsy or other invasive means. The method is a “liquid biopsy”-based approach that relies on analysis involving the examination of multiple feature types, including, but not limited to, two or more of the following:
[0009] Epigenetic signatures, particularly 5-hydroxymethylcytosine signatures;
[0010] Count of 5hmC-containing fragments in annotated CpG island genomic regions;
[0011] Count of 5hmC-containing fragments in the CTCF binding region;
[0012] Count of 5hmC-containing fragments in annotated enhancer regions;
[0013] Count of 5hmC-containing fragments in annotated gene regions;
[0014] Count of 5hmC-containing fragments in annotated 3'-UTR genomic regions;
[0015] WGS fragment counts in each of the series of windows aligned with the genome and / or WGS fragment counts in each of the set of size range bins;
[0016] WGS fragment counting in a 100kb genomic region for copy number variation (CNV) determination;
[0017] Epigenetic features other than 5hmC, e.g., 5-methylcytosine (5mC)-related features such as the location, level, and / or fragment size of 5mC;
[0018] Plasma cfDNA concentration;
[0019] Plasma levels of CA 19-9;
[0020] Histone markers, e.g., H3K4me1, H3K4me3, H3K9me3, H3K27ac, H3K27me3, and H3K36me3;
[0021] Other biomarkers not explicitly included in the above categories; as well as
[0022] Any patient-specific clinical parameter that tends to correlate with the risk of pancreatic cancer.
[0023] With respect to the latter category, clinical features can also be examined in this method by one or both of the following methods. Firstly, one or more patient-specific clinical parameters can be used to exclude or include samples from specific patients (e.g., smokers, individuals under 35 years of age, presence of pancreatitis, etc.). Alternatively, such parameters may be incorporated into this analysis as feature types. Examples of patient-specific clinical features, not limited to but illustrative, include: lesion size; lesion grade; lesion stage; lesion location; presence or absence of pancreatitis; presence or absence of jaundice; presence or absence of diabetes, including type 1 and type 2 diabetes; presence or absence of other conditions or symptoms; levels of pro-inflammatory cytokines; patient age; patient weight; patient BMI; patient sex; patient ethnicity; family history; physical activity; diet; smoking status; and exposure to or absence of known carcinogens.
[0024] The method of the present invention finally includes the step of calculating a probability score indicating the likelihood that a patient has cancer. In contrast to previous attempts to detect the presence of pancreatic cancer from a blood sample, the method is highly accurate, efficient, shows excellent specificity and sensitivity, and can be implemented on very small samples, for example, cfDNA samples of 10 ng or less.
[0025] Briefly, the process includes the step of counting cfDNA fragments with specific characteristics, such as only 5hmC-containing fragments, 5hmC-containing fragments mapped to specific positions on the genome, 5hmC-containing fragments in a 5hmC assay, or DNA fragments generated using WGS, as described in detail elsewhere in this specification. The fragment counts are used in a count vector for a given set of features (e.g., promoters), which is normalized and subsequently scored using each of a plurality of base models. The base models are created as (penalized) logistic regression fits, which means that one score is generated for each base model by summing the product of the correlation coefficient × cpm value (or coefficient × cpm × log 1+ concentration).
[0026] Finally, it is these base model scores (one per base model per sample, calculated as described above) that are input into an ensemble model, e.g., a stacked ensemble model, to generate an ensemble score, rather than the individual count data. Table 1, which is incorporated herein by reference, shows some of the base models and ensemble models used in the method as follows:
[0027] The count of 5hmC fragments in the annotated CpG island genomic region of a decontaminated, CPM-normalized glmnet fit, titled "CpG-island_decontaminated_CPM_GLMNET" (the term "decontaminated" refers to a computational method used to remove noise in 5hmC data introduced by the presence of non-5hmC fragments in a 5hmC assay);
[0028] 5hmC fragment counts in annotated CTCF-binding genomic regions of decontaminated, CPM-normalized glmnet fits titled "CTCF_decontaminated_CPM_GLMNET";
[0029] 5hmC fragment counts in annotated enhancer genomic regions of decontaminated, glmnet fits with concentration interaction terms in cfDNA plasma converted and CPM-normalized titled "enhancer_plasmaconc_inter_CPM_GLMNET";
[0030] 5hmC fragment counts in annotated gene body genomic regions of decontaminated, glmnet fits with concentration interaction terms in cfDNA plasma converted and CPM-normalized titled "gene_plasmaconc_inter_CPM_GLMNET";
[0031] 5hmC fragment counts in annotated promoter genomic regions of decontaminated, CPM-normalized glmnet fits titled "promoter_plasmaconc_inter_CPM_GLMNET";
[0032] 5hmC fragment counts in annotated 3'UTR genomic regions of decontaminated, CPM-normalized glmnet fits titled "three_prime_UTR_decontaminated_CPM_GLMNET";
[0033] WGS fragment counts in 2MB genomic regions of glmnet fits divided by fragment length range, mean-centered, and scaled titled "Frag_Data_2MB_SML_scaled_GLMNET";
[0034] WGS fragment counts in 100kb genomic regions of CPM-normalized glmnet fits titled "WGS_CNV_100kb_CPM_GLMNET"; and
[0035] An ensemble combining the predictions of the above basic models and using glmnet fit.
[0036] In one embodiment, the present invention then provides a method for detecting the likelihood of pancreatic cancer in a patient from a blood sample obtained from the patient, the method comprising:
[0037] (a) A step of extracting cell-free DNA (cfDNA) from a blood sample;
[0038] (b) A step of counting 5-hydroxymethylcytosine (5hmC)-containing DNA fragments in a cfDNA sample to obtain a total 5hmC fragment count;
[0039] (c) A step of counting 5hmC-containing DNA fragments mapped to a first genomic location to give a first genomic location 5hmC fragment count;
[0040] (d) A step of counting 5hmC-containing DNA fragments mapped to a second genomic location to obtain a second genomic location 5hmC fragment count;
[0041] (e) If necessary, perform (d) and count 5hmC-containing DNA fragments that map to at least one additional genomic location to give at least one additional genomic 5hmC fragment count;
[0042] (f)(b)~(e) Normalize the 5hmC fragment count obtained in (f)(b)~(e) to obtain the normalized 5hmC fragment count;
[0043] (g) A step of scoring the normalized 5hmC fragment counts using each of a plurality of base models, each including a penalized logistic regression fit, wherein one score is generated for each base model by summing the products of the determined correlation coefficient and the normalized 5hmC fragment count, thereby providing a plurality of base model scores; and
[0044] (h) A step of inputting the plurality of basic model scores into an ensemble model to generate a final score in the form of a probability score p between 0 and 1 that indicates the likelihood of pancreatic cancer in the patient.
[0045] In another embodiment, the method further includes, in addition to 5hmC-containing DNA fragment counts, the incorporation of at least one additional feature type into the ensemble model of (h). Examples of additional feature types include: WGS fragment counts in each of a set of windows along the genome and / or WGS fragment counts in each of a set of size range bins; copy number variations; epigenetic feature information other than 5hmC-related information, e.g., 5-methylcytosine (5mC) levels and / or 5mC-containing cfDNA fragment counts; CA 19-9 levels; histone marker information; cfDNA concentration in blood samples; and any one of several patient-specific clinical parameters that correlate with the risk of pancreatic cancer, as discussed elsewhere in this specification.
[0046] In a further embodiment, the present invention provides a method for detecting the likelihood of pancreatic cancer in a patient from a blood sample obtained from the patient, the method comprising:
[0047] (a) A step of extracting a cell-free DNA (cfDNA) sample from the blood sample;
[0048] (b) A step of counting 5-hydroxymethylcytosine (5hmC)-containing DNA fragments in a cfDNA sample to obtain a total 5hmC fragment count;
[0049] (c) A step of determining the 5hmC-containing fragment count in the CTCF-binding region, annotated enhancer region, annotated gene body region, and annotated 3'-UTR genomic region to provide the 5hmC fragment count for each location;
[0050] (d) A step of determining the WGS fragment count in each of the set of genome windows along the genome and / or the WGS fragment count in each of the set of size range bins;
[0051] (e) A step to determine the WGS fragment count in a genomic region of a specific size for the determination of copy number variation (CNV); (f) A step to determine the cfDNA concentration in a blood sample;
[0052] (f)(b)~(e) A step to normalize the fragment count obtained in (f)(b)~(e) and give the normalized fragment count;
[0053] (g) A step of scoring normalized fragment counts using each of several basic models, each containing a penalized logistic regression fit, wherein one score is generated for each basic model by summing the products of the determined correlation coefficient and the normalized fragment count, thereby providing a set of basic model scores; and
[0054] (h) A step of inputting multiple base model scores into an ensemble model to generate a final score in the form of a probability score p between 0 and 1 that indicates the likelihood of pancreatic cancer in the patient.
[0055] In a related embodiment, the method described above further includes the step of incorporating at least one additional feature type into the ensemble model of (h). In one embodiment of this embodiment, the at least one additional feature type includes the cfDNA concentration in the blood sample.
[0056] In a further embodiment, the present invention provides a method for monitoring changes in the state of pancreatic cancer in a patient, comprising the steps of repeating any of the above processes at intervals over a long monitoring period, and determining the difference in probability scores over time. [Brief explanation of the drawing]
[0057] [Figure 1] Figure 1 schematically illustrates the 5hmC enrichment, WGS features, and machine learning workflow for developing and validating a predictive model for pancreatic cancer detection.
[0058] [Figure 2] Figures 2A and 2B clearly illustrate the characteristics of the training and validation datasets. Figure 2A shows the samples included in the training and validation datasets, and Figure 2B shows the distribution of pancreatic cancer risk factors in the training and validation datasets. Figure 2C shows the correlations found between 5hmC fold changes for the genetic features included in the model used in the examples, comparing non-diabetic PaC (Y-axis) and NOD PaC (X-axis) with the corresponding non-cancer samples. Linear regression equations and Pearson correlations with the relevant p-values are reported to the panel.
[0059] [Figure 3] Figure 3 presents the results of a clinical agreement analysis between CA19-9 and the PaC detection algorithm used herein.
[0060] [Figure 4]Figures 4A and 4B present the results of the analysis of hydroxymethylated genes in cfDNA. Figure 4A shows the mean CPM values in cfDNA for 366 low-hydroxymethylated genes, and Figure 4B shows the mean CPM values in cfDNA for 43 high-hydroxymethylated genes that are expressed with differences between pancreatic cancer tissue and normal pancreatic tissue. Each data point represents the subject and the corresponding mean CPM count for either 366 (Figure 4A) or 43 (Figure 4B) genes.
[0061] [Figure 5] Figures 5A, 5B, 5C, and 5D relate to model training. Figure 5A shows the ROC curves from predictive modeling using a regularized regression model (elastic net) on a training dataset containing 132 pancreatic cancer cfDNA samples and 528 non-cancer cfDNA samples. Figure 5B presents the ROC curves for each of the four disease stages. Figure 5C shows the pancreatic cancer prediction scores by disease stage, with the dashed line indicating the 98% specificity threshold. Figure 5D shows the mean sensitivity range (62.8–67.4%) and specificity range (97.5–98.1%) in 10-fold cross-validation with 10 iterations.
[0062] [Figure 6] Figures 6A and 6B show the 5hmC and WGS features included in the PaC prediction model. Figure 6A presents the total fragment count, and Figure 6B shows the relative proportion of features used to classify PaC and non-cancer cells.
[0063] [Figure 7] Figure 7 is a table that provides an overview of the demographic attributes and disease status of the subjects evaluated in the examples.
[0064] [Figure 8] Figure 8 is a table showing the sensitivity and specificity observed for pancreatic cancer and non-cancer cells in the training and validation datasets. [Modes for carrying out the invention]
[0065] Detailed description of the invention I. Definitions and Terms:
[0066] Unless otherwise specified, all technical and scientific terms used herein have meanings that are generally understood by those skilled in the art in the field to which this invention pertains. Certain terms that are particularly important for describing this invention are defined below. Other relevant terms are defined in the International Patent Application Publication WO2017 / 176630 to Quake et al., “Noninvasive Diagnostics by Sequencing 5-Hydroxymethylated Cell-Free DNA.” The aforementioned patent publications and all other patent documents and publications referenced herein are expressly incorporated by reference.
[0067] In this specification and the appended claims, the singular forms “a,” “an,” and “the” refer to multiple subjects unless the context clearly indicates otherwise.
[0068] Numerical ranges include the number that defines that range. Unless otherwise specified, nucleic acids are written from left to right in the 5' to 3' direction; amino acid sequences are written from left to right in the amino to carboxyl direction.
[0069] The abbreviations used herein are as follows: 5hmC, 5-hydroxymethylcytosine; auROC, area under the receiver operating characteristic curve; CA19-9, glycosylated antigen 19-9; cfDNA, cell-free DNA; CI, confidence interval; CNV, copy number variation; CPM, counts per million; gDNA, genomic DNA; IPMN, intraductal papillary mucinous neoplasm; IRB, institutional ethics board; N / A, not applicable; NOD, new-onset diabetes mellitus; NS, no significant difference; PaC, pancreatic cancer; T2D, type 2 diabetes mellitus; WGS, whole-genome sequencing.
[0070] The headings presented herein are not intended to limit the various aspects or embodiments of the invention. Therefore, the terms defined immediately thereafter are more fully defined by referring to this specification as a whole.
[0071] As used herein, the term “sample” refers to a substance or mixture of substances, typically in liquid form, containing one or more analytes of interest. The biological samples evaluated herein are blood samples obtained from patients.
[0072] Where the term is used herein, “nucleic acid sample” refers to a biological sample containing nucleic acids. A nucleic acid sample may be a genomic DNA sample, or it may consist of cfDNA that is substantially free of histones and other proteins, for example, after cfDNA purification.
[0073] A "sample fraction" refers to a subset of the original biological sample and may be a compositionally identical portion of the biological sample, such as when a blood sample is divided into multiple equal fractions. Alternatively, sample fractions may be compositionally different, for example, when certain components of the biological sample are removed, such as in the extraction of cell-free nucleic acids.
[0074] As used herein, the term “cell-free nucleic acid” encompasses both cfDNA and cfRNA, which may be present in cell-free fractions of biological samples, including bodily fluids. Bodily fluids may be blood, serum, or plasma, including peripheral blood. In most cases, the biological sample is a blood sample, and cell-free nucleic acid samples, such as cell-free DNA samples, are extracted therefrom using currently common methods known to those skilled in the art and / or described in relevant textbooks and literature; kits for performing cell-free nucleic acid extraction are commercially available (e.g., AllPrep® DNA / RNA Mini Kit and QIAmp DNA Blood Mini Kit (both available from Qiagen), or MagMAX Cell-Free Total Nucleic Acid Kit and MagMAX DNA Isolation Kit (available from ThermoFisher Scientific)). See also, e.g., Hui et al. and Fong et al. (2009) Clin. Chem. 55(3):587-598.
[0075] The term “adapter ligated” as used herein refers to a nucleic acid that has been ligated with an adapter. Adapters can be ligated to the 5' and / or 3' ends of a nucleic acid molecule. As used herein, the term “add adapter sequence” refers to the act of adding an adapter sequence to the ends of a fragment in a sample. This can be done by filling the ends of the fragment with polymerase, adding an A tail, and then ligating the fragment having the A tail with an adapter containing a T overhang. Adapters are typically ligated to a DNA double helix using ligase, but in the case of RNA, the adapter is covalently or otherwise attached to at least one end of a cDNA double helix, preferably in the absence of ligase.
[0076] As used herein, the term “amplify” refers to generating one or more copies or “amplicons” of a template nucleic acid, which can be carried out using any suitable nucleic acid amplification technique, such as PCR, NASBA, TMA, and SDA.
[0077] The terms "enrich" and "enrichment" refer to the partial purification of a template molecule having a specific characteristic (e.g., a nucleic acid containing 5-hydroxymethylcytosine) from an analyte lacking that characteristic (e.g., a nucleic acid not containing hydroxymethylcytosine). Enrichment typically increases the concentration of the characteristic analyte by at least 2 times, at least 5 times, or at least 10 times compared to the non-characteristic analyte. After enrichment, at least 10%, at least 20%, at least 50%, at least 80%, or at least 90% of the analyte in the sample may have the characteristic used for enrichment. For example, at least 10%, at least 20%, at least 50%, at least 80%, or at least 90% of the nucleic acid molecules in the enriched composition may contain chains having one or more hydroxymethylcytosines modified to contain a capture tag.
[0078] As used herein, the term "sequencing" refers to a method for obtaining a polynucleotide entity consisting of at least 10 consecutive nucleotides (for example, an entity consisting of at least 20, at least 50, at least 100, or at least 200, or more consecutive nucleotides).
[0079] The terms “next-generation sequencing” (NGS) or “high-throughput sequencing,” as used herein, refer to so-called parallelized synthesis sequencing or ligation sequencing platforms currently employed by Illumina, Life Technologies, Roche, and others. Next-generation sequencing methods also include nanopore sequencing, such as that commercialized by Oxford Nanopore Technologies; electronic detection methods, such as Ion Torrent technology commercialized by Life Technologies; and single-molecule fluorescence-based methods, such as that commercialized by Pacific Biosciences.
[0080] As used herein, the term “read” refers to the raw or processed output of a sequencing system, such as a large-scale parallel sequencing system. In some embodiments, the output of the methods described herein is reads. In some embodiments, these reads may need to be trimmed, filtered, and aligned to obtain raw reads, trimmed reads, and aligned reads.
[0081] A "unique molecular identifier" (UMI) refers to a relatively short nucleic acid sequence that is attached to every nucleic acid template molecule in a sample. This sequence is random, and as a result, if the UMI sequence is long enough, every nucleic acid template molecule is associated with a unique UMI sequence. The UMI sequence can be used to correct for errors in amplification and sequencing equipment, as is well known in the art, to allow users to track down duplicates and remove them from downstream analysis, and to enable molecular counting and, consequently, the determination of analyte concentrations. See, for example, Casbon et al. (2011) Nuc. Acids Res. 39(12):1-8. Here, "unique molecule" refers to the entity of the nucleic acid template molecule.
[0082] In some embodiments, the UMI may have a length in the range of 1 to about 35 nucleotides, for example, 3 to 30 nucleotides, 4 to 25 nucleotides, or 6 to 20 nucleotides. In certain cases, the UMI may also be error detection and / or error correction, meaning that the code can still be correctly interpreted even if an error is present. The use of error correction sequences is described in the literature (e.g., U.S. Patent Application Publication 2010 / 0323348 to Hamati et al. and U.S. Patent Application Publication 2009 / 0105959 to Braverman et al., both of which are incorporated herein by reference).
[0083] A "barcode," also a short nucleic acid sequence, can be attached to each DNA molecule in a sample, thereby helping to identify the origin of a sample after processing, amplification, and sequencing of a mixed sample group.
[0084] The term “detection” is used interchangeably with the terms “determining,” “measuring,” “evaluating,” “assessing,” and “analyzing,” and refers to any form of measurement, including determining whether or not an element is present. These terms include both quantitative and / or qualitative determinations. Assessment can be relative or absolute. Therefore, “assessing the presence of ~” includes determining the amount of a portion present, or determining whether or not it is present. Assessing the level at a hydroxymethylation biomarker locus refers to determining the degree of hydroxymethylation at that locus.
[0085] "Precision" refers to the degree of agreement between a measured or calculated quantity (the value reported by a test) and its accurate (or true) value. Clinical precision can be expressed as sensitivity, specificity, positive predictive value (PPV), or negative predictive value (NPV), or as likelihood or odds ratio, among several other measures, in relation to the proportion of true outcomes (true positive (TP) or true negative (TN)) to misclassified outcomes (false positive (FP) or false negative (FN)).
[0086] "Performance" is a term relating to the overall usefulness and quality of a diagnostic or prognostic test, including, in particular, clinical and analytical accuracy, other analytical and process characteristics such as usability (e.g., stability, ease of use), health economic value, and the relative cost of the test's components. Any of these factors can be a source of excellent performance, and therefore usefulness, and can be measured, where appropriate, by suitable "performance metrics" such as AUC, time to results, and shelf life.
[0087] "Clinical parameters" include, but are not limited to, all non-sample biomarkers of the subject's health status or other characteristics, including, for example, lesion size; lesion location; patient's age; patient's weight; patient's sex; patient's ethnicity; family history; genetic mutations; and PD-L1 tumor staining results currently used in the practice to determine whether anti-PD-1 therapy is appropriate.
[0088] A “formula,” “algorithm,” or “model” is any mathematical equation, algorithmic, analytical, or programmed process or statistical method that takes one or more sequential or categorical inputs and computes an output value, which may be referred to as a “probability score” or “index value.” Non-exclusive examples of a “formula” include sums, ratios, and regression operators, e.g., coefficients or exponents, transformations and normalizations of biomarker values (including, but not limited to, normalization schemes based on clinical parameters such as sex, age, or ethnicity), rules and guidelines, statistical classification models, and neural networks trained on historical populations.
[0089] Of particular interest in this specification are linear and nonlinear equations and statistical classification analyses for determining the correlation between hydroxymethylation levels at biomarker loci detected in patient samples and the likelihood that a patient has a particular type of cancer. Of particular interest in constructing panels and combinations are structural and syntactic statistical classification algorithms, as well as methods for constructing risk indicators that utilize pattern recognition and machine learning features, including established methods such as cross-correlation, principal component analysis (PCA), factor rotation, logistic regression (LogReg), linear discriminant analysis (LDA), intrinsic gene linear discriminant analysis (ELDA), support vector machines (SVM), random forests (RF), recursive split trees (RPART), and other related decision tree classification methods, shrinking centroid (SC), StepAIC, K-nearest neighbors, boosting, decision trees, neural networks, Bayesian networks, and hidden Markov models. Many such algorithmic methods, particularly ridge regression, lasso, and elastic nets, have been further implemented to perform both feature (locus) selection and regularization. Other methods, including the Cox model, Weibull model, Kaplan-Meier model, and Greenwood model, which are well known to those skilled in the art, may be used for survival and time to event hazard analysis. Many of these methods are either useful in combination with hydroxymethylated biomarker selection methods, such as forward selection, backward selection, or stepwise selection, complete enumeration of all potential biomarkers of a given size set or panel, or genetic algorithms, or they themselves may include biomarker selection methods. These methods may be combined with information criteria such as the Akaike Information Criterion (AIC) or Bayesian Information Criterion (BIC) to quantify the trade-off between additional biomarkers and model improvement, and to help minimize overfitting.The resulting predictive models may be validated in other trials using techniques such as Bootstrap, Leave-One-Out (LOO), and 10-fold cross-validation (10-Fold CV), or they may be cross-validated in the trials in which they were initially trained. At various steps, the false positive rate may be estimated by value substitution using techniques known in the art.
[0090] "Likelihood" is the probability, in the context of one embodiment of the present invention, that a patient has or does not have pancreatic cancer.
[0091] "Hydroxymethylation level" refers to the degree of hydroxymethylation within a hydroxymethylation biomarker locus. The degree of hydroxymethylation is usually measured as hydroxymethylation density, for example, the ratio of 5hmC residues to all cytosines, both modified and unmodified, within a nucleic acid region. Other measures of hydroxymethylation density, such as the ratio of 5hmC residues to all nucleotides in a nucleic acid region, are also possible.
[0092] A “hydroxymethylation profile” or “hydroxymethylation signature” refers to a dataset containing hydroxymethylation levels at each of several hydroxymethylation biomarker loci, pre-selected as being hydroxymethylated with respect to a specific disease phenotype, such as lung cancer, colorectal cancer, or breast cancer. A hydroxymethylation profile may also be a reference hydroxymethylation profile, which includes a composite hydroxymethylation profile for a population of individuals having at least one common feature, as described elsewhere in this specification. A hydroxymethylation profile may also be a patient's hydroxymethylation signature, constructed from measurements of hydroxymethylation levels at each of several hydroxymethylation biomarker sites.
[0093] As used throughout this application, the term “locus” refers to a site on a nucleic acid molecule, which may be single-stranded or double-stranded, and further, an individual locus (or a group of “locuses”) may be of any length, thus encompassing a single CpG site and a full-length gene, or may span larger features such as topologically related domains, including cases where several such loci collectively form a group of related sequence motifs, other homologous or functional features (regardless of their adjacency or topological relationships). Loci as used herein may be located inside a gene body; inside annotation features outside a gene body, such as promoters, enhancers, transcription start sites, transcription stop sites or DNA binding sites, or combinations thereof; or inside an untranslated region or “UTR” (including 3'UTR and 5'UTR).
[0094] It should be noted that some of the individual biomarkers disclosed herein, such as hydroxymethylated biomarkers, may not exhibit significant individual significance in specific assessments, but may be significant in the differentiation required by the methods of the present invention when used in combination with one or more other types of biomarkers and, as appropriate, clinical parameters that affect the detection and assessment of cancerous lesions.
[0095] As used herein, the term “correlated” in relation to two variables (e.g., two values, two sets of values, a value or set of values and a disease state, a value or set of values and the risk associated with a disease state) indicates a tendency for the two variables to change together. “Correlation” is a measure of the degree to which two or more variables change together. A positive correlation indicates the degree to which those variables increase or decrease in parallel. An example of a positive correlation is the relationship between, on the one hand, the level of hydroxymethylation at a hydroxymethylation biomarker locus and on the other hand, the likelihood of a patient having cancer or a particular type of cancer. Conversely, a negative correlation is thought to exist if the level of hydroxymethylation at a hydroxymethylation biomarker locus decreases as the likelihood of a subject having cancer or a particular type of cancer decreases.
[0096] II. Methods for identifying cancer in cfDNA samples:
[0097] The method of the present invention relies on combinations of “basic” models described in Table 1, which has been incorporated herein by prior reference. Table 1 provides a list of regions identified in sequenced cfDNA and information on how each region is used in each basic model, as described below.
[0098] Each feature name entry in the "Feature Name" column of Table 1 is a region in sequenced cfDNA identified by its corresponding coordinates in the human reference genome hg38, separated by a colon. More specifically, each feature name entry begins with a chromosome, followed by the position of a nucleotide terminal, with the start position identified first, then the end position, and finally a certain identifier, such as a gene name. For example, the feature name chr17:5111410:5112061:CpG-10567 refers to a region on chromosome 17 that has the identifier "CpG-10567" and begins at position 5111410 and ends at position 5112061. Each entry in the "hg38_genomic_coordinates" column is a gene region aligned with the corresponding identified region, with hg38 coordinates provided in the same format as above.
[0099] In the first column of Table 1, each entry under “Model” is a name chosen to correspond to a group of similar features used in one of the “Basic” models. For example, a basic model may include a region from a promoter. For instance, each row identified as “promoter_decontaminated_CPM_GLMNET” in the “Model” column lists the genomic regions corresponding to a promoter, where the number of fragments was obtained by counting promoters across our cancer and control samples and then using an elastic net to fit them to the model.
[0100] Certain feature types are used only to build models from 5hmC data obtained from 5hmC-containing cfDNA fragments using the 5hmC assay protocol described in the Examples, and are identified as such in the “Assay” column of Table 1. Other feature types are used only to build models from WGS data, and again they are designated as “WGS” in the “Assay” column. For example, only WGS fragments are considered in the base models “Frag_Data_2MB_scaled_GLMNET” and “WGS_CNV_100kb_GLMNET”.
[0101] Regarding fragment length, it should be noted that all sequenced fragments in Table 1, regardless of whether they were generated by the 5hmC assay or WGS, are filtered (i.e., for all samples and libraries) to exclude fragments outside the 50nt–1000nt size range. However, as mentioned earlier, fragment size is considered a separate feature for the basic model "Frag_Data_2MB_scaled_GLMNET," where fragments are counted and sorted into one of three size range bins (50–152nt, 153–240nt, or 241–1000nt); these are also identified in terms of their position within one of a series of 2Mbp windows along the genome. Thus, there are three rows for each 2Mbp window, insofar as fragments can exist within each of the three size ranges contained within the same 2Mbp window.
[0102] The column "cfdna_plasma_interaction_term" indicates whether the cfDNA concentration is used as a factor along with the count of fragments mapped to that region for that particular region in the genome. Specifically, if the column has a value of "yes", the cpm value is multiplied by 1 + the log value of the concentration. This is useful for performance, i.e., overall predictive power, because patients with cancer tend to have higher concentrations of cfDNA in their plasma.
[0103] Each value in the "Coefficient" column is used to multiply the cpm value from the sample in that region. The sum of coefficient × cpm value (or coefficient × cpm × log[1+conc]) gives the log odds or logit for the corresponding base model, which has a one-to-one correspondence with the p-value score.
[0104] The ensemble row shows how the basic model scores are combined to arrive at a final score, which is a probability score p between 0 and 1, representing the likelihood of pancreatic cancer.
[0105] Further information relating to various aspects of the present invention can be found in U.S. Patent Application No. 17 / 131,287, filed December 22, 2020, “Pancreatic Ductal Adenocarcinoma Evaluation Using Cell-Free DNA hmC Profile,” which includes methods for extending the invention to include patient monitoring and treatment assessment. The disclosures of that application are incorporated herein by reference in their entirety.
[0106] A typical embodiment of this process is described in detail in the following examples. [Examples]
[0107] Patient samples were collected under the clinical trial NCT03869814 case-control study, which aimed to optimize methods for detecting cancer in the blood of subjects with solid tumors using genomics, epigenomics, and proteomics. The study included subjects without cancer who were followed every six months for up to three years from blood collection.
[0108] A. Study design and methods:
[0109] (i) Clinical cohort and study design:
[0110] From June 2018 to May 2022, cancer and non-cancer subjects were recruited for a case-control trial (NCT03869814) from 146 sites across the United States. Written informed consent was obtained from all subjects, and the trial was approved by the Institutional Review Board (IRB) at each site. Submission of the trial protocol, IRB approval, and handling of specimens across all sites were managed by several contract research laboratories.
[0111] Study inclusion criteria:
[0112] (1) The subjects were between 45 and 75 years old at the time of registration;
[0113] (2) The patient gave full consent;
[0114] (3) A high probability of cancer diagnosis or clinical suspicion of cancer based on the standard of care (SOC) of the participating facilities and physicians; and
[0115] Applicants must have no history of cancer and must not have received any treatment at the time of registration.
[0116] Exclusion criteria for the exam:
[0117] (1) Under 45 years of age or over 75 years of age;
[0118] (2) A history of cancer diagnosis, whether or not treated (excluding non-melanoma skin cancer that was resolved / treated more than one year prior to registration);
[0119] (3) Receiving any cancer treatment, including chemotherapy, radiotherapy, palliative radiotherapy, hormone therapy, or natural therapy;
[0120] (4) Carcinoma in situ without invasive components;
[0121] (5) Having undergone surgery requiring general anesthesia within two months of collection (anesthesia used for procedures such as colonoscopy and EBUS is acceptable);
[0122] (6) Dental novocaine taken within one week prior to sample collection;
[0123] (7) Having received systemic immunomodulatory therapy within the past 12 months;
[0124] (8) Current pregnancy or pregnancy within the past 12 months;
[0125] (9) Organ transplantation;
[0126] (10) Receiving dialysis;
[0127] (11) Blood transfusion within the last month;
[0128] (12) Infection with HIV / AIDS, hepatitis A, hepatitis D or E, TB, any type of prion disease (e.g., CJD), or any other pathogen currently present or within the past five years.
[0129] The case-control study was divided into two datasets: a training set and a validation set. The validation set included a subset of individuals with NOD. Samples for the predictive logistic regression-based algorithm on the training set were obtained from subjects within the training cohort who were found to be free of intraductal papillary mucinous neoplasms (IPMNs), pancreatic cysts, pancreatitis, and NOD.
[0130] (ii) Sample collection and plasma preparation:
[0131] Plasma was isolated from whole blood samples obtained by standard venous phlebotomy at registration. Whole blood (2 × 10 mL) was collected in cell-free DNA BCT® tubes (Streck, La Vista, NE) according to the manufacturer's protocol, maintained at 15–25°C for transport to the laboratory, and processed within 72 hours of venous puncture. To separate the plasma, the tubes were centrifuged at 1600 × g for 10 minutes, the plasma was transferred to a new tube, and further centrifuged at 16,000 × g for 10 minutes. The final plasma was portioned and cryopreserved at -80°C before cfDNA isolation.
[0132] (iii) Isolation of cfDNA:
[0133] cfDNA isolation was performed using the MagMax® cfDNA Isolation Kit (Thermo Fisher Scientific US, Waltham, MA) on a liquid handler (Hamilton STAR Liquid Handling System, Reno, NV) according to the manufacturer's protocol. The isolated cfDNA was quantified using Quant-iT® PicoGreen® (Thermo Fisher Scientific US) and stored at -20°C until library preparation.
[0134] (iv) Handling of tissues and isolation of genomic DNA:
[0135] Forty-three PaC tissue samples (obtained surgically) and ten normal pancreatic tissue samples (two obtained as normal adjacent tumors; eight obtained from normal pancreases) were collected and stored in Hypo Thermosol(H) medium (Sigma, St. Louis, MO). All tissue samples were obtained surgically. Each sample was weighed and divided into approximately 35 mg sections, homogenized in 500 μl of RLT Buffer Plus using Tissue Lyser LT (QIAGEN, Germantown, MD), and stored at -80°C until DNA extraction. Genomic DNA was extracted using the DNeasy Blood & Tissue Kit (QIAGEN, Hilden, Germany) according to the manufacturer's instructions. Genomic DNA eluates were quantified using the Qubit dsDNA High Sensitivity Assay (Thermo Fisher Scientific US) and stored at -20°C until further processing. Prior to sequencing library construction, genomic DNA was fragmented to a mode of 150 base pairs using an ME220 Focused-ultrasonicator (Covaris, Woburn, MA); the mode of the fragmented DNA size was validated using a 2200 TapeStation dsDNA high-sensitivity assay (Agilent Technologies, Santa Clara, CA) and quantified as previously described before initiating library construction.
[0136] (v) 5hmC assay enrichment:
[0137] cfDNA isolated from a single BCT tube or DNA isolated from tissue was normalized to either 10 ng (cfDNA) or 20 ng (tissue DNA) in a 96-well plate using an automated liquid handler (Beckman Coulter Life Sciences, Biomek I7, Indianapolis, IN), end repaired, A-tails added, and ligated to a sequencing adapter to create a whole-genome sequencing (WGS) library. A portion of the WGS library for each sample was advanced to complete library preparation, and the remainder of the WGS library underwent further processing to enrich fragments containing 5hmC bases. 5hmC enrichment was carried out by a two-step chemistry biotinylation reaction as previously outlined (1), enriched by binding to streptavidin-coated Dynabeads M-270 (Thermo Fisher Scientific US). After enrichment with 5hmC fragments, both the 5hmC and WGS libraries were amplified by polymerase chain reaction and normalized to 1 ng / μL using an automated liquid handler (Hamilton Microlab STARLet, Reno, NV). Following normalization, the WGS and 5hmC libraries were pooled and sequenced on a NovaSeq6000 sequencer (Illumina, Inc., San Diego, CA).
[0138] B. Bioinformatics Analysis:
[0139] (i) Bioinformatics pipeline and quality control:
[0140] Sequencing data from 5hmC and WGS was generated using NovaSeq Control Software v1.7.0 (Illumina, Inc., San Diego, CA). Raw data processing and demultiplexing were performed using bcl2fastq Conversion Software (Illumina, Inc., San Diego, CA) to generate sample-specific FASTQ output. Sequencing reads were analyzed by a computational pipeline implemented as a Nextflow script, which aligned reads against the Human Genome Build 38 reference genome using the BWA-MEM2 algorithm. This pipeline divided the genome into functional regions that identified gene bodies, enhancers, CpG islands, CCCTC binding factor sites, promoters, and three prime untranslated regions from Gencode Human Annotation Version 31 (GRCh38.p12). The number of 5hmC library read pairs mapped to each region was then counted, and coverage variations were corrected using the count per million mapped reads. In addition, a feature set incorporating copy number variations across 100kb bins, confirmed by read coverage depth and fragment size variations across 2MB bins genomically, was created using WGS data. To assess the quality of the sequencing data, metrics were calculated by the pipeline via Picard. Samples that passed the quality control metrics were placed into two datasets for training and validation. The quality control failure rate for the validation sample set was 1.94%. Non-cancer samples in the training data were matched with various clinical features, including age, sex, body mass index, and smoking status.
[0141] (ii) Development of a pancreatic cancer algorithm:
[0142] The machine learning classification algorithm was trained as follows: Each sample in the training dataset was analyzed using the bioinformatics pipeline described in the previous section. An elastic net logistic regression algorithm was constructed for each feature set using the R package glmnet, with an optimized elastic net mixing ratio α and regularization parameter λ using 10-fold cross-validation. After removing highly variable features, the number of features was further reduced by regularization performed by the elastic net, and coefficients were assigned to each (Figures 6A and 6B). Combining all of the above, an equation is formed to calculate the log odds ratio of individuals with cancer from the coefficient β and features x from N different feature sets:
number
[0143] By solving the equation for p, we can obtain the probability of cancer ranging from 0 to 1.
[0144] The classification probability threshold was determined by setting a threshold that yielded a 98% specificity for non-cancer cells in the training dataset (Figure 5C). Subsequently, the final algorithm was integrated into the inventors' automated computation pipeline as a Docker container on Amazon Web Services (Amazon, Seattle, WA). Scoring and cancer classification were performed in a blinded manner, and the target label was revealed after scoring was complete.
[0145] (iii) Differential gene presentation analysis:
[0146] For differential 5hmC gene analysis, genes that were not mapped to autosomes were removed. Furthermore, weakly presented genes were removed by excluding those with counts exceeding 3 per million reads in at least 10 samples. This filter removed approximately 7.5% of all genes from the consensus code sequence database. Subsequently, the R package "edgeR" (2) was used to identify fold changes between PaC and non-cancer in both NOD and non-NOD (any subject not diagnosed with new-onset diabetes) cohorts.
[0147] (iv) CA 19-9 clinical concordance analysis:
[0148] 507 non-cancerous samples and 73 PaC samples were sent 500 μl of frozen plasma to a Clinical Laboratory Improvement Amendments (ARUP) accredited laboratory for evaluation of glycan antigen 19-9 (CA19-9). Samples with CA19-9 values above 37 were classified as cancerous according to standard clinical thresholds used for CA19-9 and PaC detection.
[0149] C. Result:
[0150] (i) Development of a pancreatic cancer detection algorithm:
[0151] A total of 660 plasma samples (training dataset) from 132 PaC subjects and 528 non-cancer subjects were used to develop a PaC detection algorithm capable of differentiating between PaC and non-cancer cases (Figure 2A). Key clinical features of the cohort, including body mass index, age, and other PaC risk factors such as smoking, diabetes status, family history, and genetic predisposition to PaC, are reported in Figure 2B and Table 2.
[0152] [Table 1]
[0153] Statistical analysis of clinical variables revealed that in the validation dataset, PaC subjects were older and had lower BMIs compared to non-cancer subjects (Student's t-test, p<.05). Furthermore, overall, the validation dataset had fewer men compared to the training dataset (Fisher's exact test, P<.05). In both sets, PaC subjects had a higher proportion of past smokers, while the proportion of current smokers was higher in the validation set (Fisher's exact test, P<.05).
[0154] The inventors first evaluated whether the 5hmC signal found in PaC tumor tissue could be detected in plasma cfDNA by identifying sets of genes with significant over- or under-presentation based on their 5hmC levels. Therefore, the inventors compared 42 PaC tumor tissue samples with 10 normal tissue samples and identified 366 genes with decreased hydroxymethylation and 43 genes with increased hydroxymethylation (FDR changes of 0.05 and 1.5 times). These same over-hydroxymethylated and under-hydroxymethylated genes were investigated in cfDNA from PaC and non-cancer subjects. A consistent trend in 5hmC presentation was found in the comparison of tumor tissue with normal pancreatic tissue for each set of genes with increased and decreased hydroxymethylation (Kruskal-Wallis P<3.1 × 10 sets (Reference 11) and P<2.2 × 10 sets (Reference 16)) (Figures 4A and 4B). This indicates that the cfDNA of subjects with PaC possesses a specific and differential 5hmC profile, and that it is possible to detect PaC using plasma-derived cfDNA by utilizing the 5hmC signal.
[0155] Next, the inventors developed a specific algorithm for cfDNA PaC and constructed a binomial predictive algorithm using elastic net logistic regression that combined predictors from both 5hmC and WGS features after matching for body mass index, age, and smoking status within the training cohort (Table 2). This yielded overall performance, measured by the area under the receiver operating characteristic (auROC) curve, as 0.93 overall and a stage-specific range of 0.84–0.95 (Figures 5A–5C). The feature sets that contributed most significantly to the algorithm were the 5hmC locus in the enhancer region, followed by the 5hmC locus in the gene body region (Figures 6A and 6B). To simulate how well the algorithm performs on new data, the algorithm was assessed using 10-fold cross-validation, which allowed 10% of the samples to be retained from the training set and used instead for validation. For each of the 10-fold divisions, the training data was divided into training and test sample portions in a 90:10 ratio. As shown in Figure 8, the overall test sensitivity of this 10-fold cross-validation analysis was 65.9% (95% confidence interval [CI], 57.2–73.9%), the sensitivity for the early stages (Stage I / II) was 57.1% (95% CI, 44.0–69.5%), and the specificity was 97.9% (95% CI, 96.3–99.0%). To evaluate the reproducibility of the 10-fold cross-validation, the process was repeated 10 times using different sample splitting assignments. This yielded consistent sensitivity (62.8–67.4%) and specificity (97.5–98.1%) (Figure 5D).
[0156] (ii) Verification of the pancreatic cancer detection algorithm:
[0157] Next, the inventors evaluated the performance of the PaC detection test in a separate set of samples processed independently and blindly from the training set. The validation dataset included 2,150 subjects, consisting of 102 PaC subjects and 2,048 non-cancer subjects. A sensitivity of 68.3% (95 CI, 51.9–81.9%) was observed in early disease (stage I / II samples combined), with an overall sensitivity of 66.7% (95 CI, 56.6–75.7%) and a specificity of 96.9% (95% CI, 96.0–97.6%) (Figure 8). Statistical testing revealed no significant differences in performance between the validation set and the training set, nor between NOD subjects and non-NOD subjects.
[0158] In addition to the validation dataset, a set of 74 non-cancerous samples with other pancreatic lesions (excluding PaC) was examined. These included 40 subjects with IPMN, 27 subjects with pancreatitis, and 7 subjects with both conditions. The algorithm predicted PaC in 26 out of 47 subjects (55.3%) with IPMN and in 14 out of 34 subjects (41.2%) with pancreatitis. Notably, the majority of samples from subjects with IPMN classified as PaC (18 / 26 [69.2%]) had moderate to severe dysplasia, and one case had PanIN-2 disease. In contrast, only 14.3% (3 / 21) of IPMN classified as “not detected” by our examination had moderate to severe dysplasia.
[0159] Furthermore, the inventors evaluated the algorithm on 1,524 subjects across 11 different diseases diagnosed as cancers other than PaC (Figure 2A and Table 3).
[0160] [Table 2]
[0161] Subjects with stage IV cancer were excluded to avoid detecting signals arising from occult metastases in the pancreas. Detection rates for non-pancreatic cancer samples were determined in independent validation sets for bladder cancer, breast cancer, colorectal cancer, esophageal cancer, gastric cancer, kidney cancer, liver cancer, lung cancer, ovarian cancer, prostate cancer, and uterine cancer. Liver cancer and ovarian cancer had very high detection rates of 55.2% and 58.3%, respectively. Notably, of the liver cancers detected by the algorithm, three had pancreatic and bile duct origins, four were intrahepatic cholangiocarcinomas, and one was cholangiocarcinoma, all of which shared pancreatic tissue characteristics. Gastrointestinal cancers (colorectal cancer, esophageal cancer, and gastric cancer) had moderate detection rates (20.2%, 25.9%, and 27.3%, respectively) compared to, for example, prostate cancer and breast cancer (4.6% and 14.2%, respectively).
[0162] (iii) Verification of pancreatic cancer detection in high-risk populations:
[0163] The validation dataset included 2,150 subjects, including individuals with high-risk clinical conditions for PaC, such as family history, genetic predisposition, long-term type 2 diabetes (more than 3 years since diabetes diagnosis), and NOD (type 2 diabetes diagnosed within 3 years of diagnosis) (Figure 2B and Table 2). As shown in Figure 2B, the training set did not include subjects with NOD.
[0164] The inventors conducted a separate performance evaluation in the larger group of subjects with NOD (Non-Onset Disease) within the validation set, which had a high risk of PaC (relative risk of 6 to 8 times). When comparing the performance of subjects with NOD and subjects without NOD in the validation set, no significant difference was found (Fisher's direct test, P>.05); the sensitivity was 57.5% (95% CI, 40.9–73.0%) and 72.6% (95% CI, 59.8–83.1%) for NOD and non-NOD, respectively. The specificity was also determined to be unsignificant, at 96.2% (95% CI, 92.4–98.5%) and 96.9% (95% CI, 96.1–97.7%) for NOD and non-NOD, respectively (Figure 8). Furthermore, when the inventors evaluated the performance of the pancreatic cancer detection algorithm in other high-risk groups, they observed a sensitivity of 88.9% (95% CI 51.8%, 99.7%) and a specificity of 94.2% (95% CI 90.1%, 97%) for long-term diabetic patients, and a sensitivity of 71.4% (95% CI 47.8%, 88.7%) and a specificity of 97.2% (95% CI 94.3%, 98.9%) for current smokers. Sensitivity was lower in subjects with a BMI above 32, at 41.2% (95% CI 18.4%, 67.1%), but specificity was maintained at 96% (95% CI 94.2%, 97.3%). Taken together, these data suggest that the algorithm detects PaC signals regardless of clinical risk status.
[0165] To verify these observations, gene-based differential changes in the 5hmC profile, which are found when comparing cancer and non-cancer individuals, were compared in a paired manner between subjects without diabetes and subjects with NOD. A high level of correlation was shown (r=0.96; P<2.2×10⁻⁶). -16 ), and the finding that the gene-based differential 5hmC profile is consistent between cancer and non-cancer individuals and is not affected by the status of type 2 diabetes was supported (Figure 2C).
[0166] (iv) Performance comparison with CA19-9 biomarker:
[0167] CA19-9 is the most commonly used biomarker for the diagnosis and management of PaC (25); therefore, we evaluated the agreement between CA19-9 and the PaC detection test presented in this study. CA19-9 analysis was performed on 73 PaC cases and 507 non-cancer cases and compared with predictions made using the PaC detection algorithm. PaC detection performance in early-stage cancer was higher with the PaC detection test than with CA19-9, with a sensitivity of 75.8% (95% CI, 57.7–88.9%) and a specificity of 97.4% (95% CI, 95.7–98.6%) compared to CA19-9 alone, which had a sensitivity of 57.6% (95% CI, 39.2–74.5%) and a specificity of 95.5% (95% CI, 93.3–97.1%) (Figure 3). In particular, the false positive rate for CA19-9 was almost twice that of the PaC detection test, at 4.5% (95% CI, 2.9–6.7%) and 2.6% (95% CI, 1.4–4.3%), respectively.
[0168] D. Significance of the results:
[0169] Data from our tests provide evidence that differential epigenomic signals identified by 5hmC measurement in pancreatic tumor tissue are also found in cfDNA from different patients. Using cfDNA-derived 5hmC and genomic signals enables the construction of a robust pancreatic cancer detector, as demonstrated by consistent performance across training (n=660) and validation (n=2,150) datasets. The large validation dataset shows a cancer incidence of approximately 5%, which is close to the clinical incidence of cancer in high-risk populations (e.g., NOD, family history, and mutation) of approximately 1%. PaC detection performance was maintained with 66.7% sensitivity and 96.9% specificity compared to the training dataset, and was also retained in samples from patients with early PaC (stage I or II; 68.3% sensitivity).
[0170] PaC signal detectors using epigenomic and genomic signals perform better than the CA19-9 scale for matched cohorts for analysis, particularly for early PaC.
[0171] CA19-9 is the only biomarker routinely used in the management of PaC; however, due to its low susceptibility in early disease, lack of expression in individuals with Lewis-negative genotypes, and its elevated levels in many other benign and malignant diseases, it is rarely used in early detection protocols (22). We compared the performance of CA19-9 and our PaC detection test in a subset of patients. Our PaC detection test showed a sensitivity of 75.8% (95% CI, 57.7–88.9%) and a specificity of 97.4% (95% CI, 95.7–98.6%) compared to CA19-9 alone, which showed a sensitivity of 57.6% (95% CI, 39.2–74.5%) and a specificity of 95.5% (95% CI, 93.3–97.1%), outperforming CA19-9 in its early stage performance. This supports the applicability of our test for early disease detection.
[0172] Targeted, conventional clinical assessments are recommended only for individuals with a lifetime risk of developing pancreatic ductal adenocarcinoma greater than 5%, including those with a family history and certain genetic syndromes, or those with mucinous cystic lesions of the pancreas (15). In addition, another high-risk group includes NOD patients, who have a 6- to 8-fold increased risk of developing the disease (17, 19). Previous studies have demonstrated a 3-year cumulative incidence of PaC in individuals with NOD of 0.85% (19). Other studies have reported Pac incidences ranging from 0.3% (23) to 3% (24) to 10% (25) in NOD subjects. The variability in reported incidences is likely influenced by the demographic features (e.g., age, sex, and ethnicity) or size of the studies.
[0173] In summary, in the United States, nearly 1.4 million individuals are diagnosed with NOD annually; approximately 900,000 are 50 years of age or older; up to 1% develop PaC within 3 years of type 2 diabetes diagnosis; and it is suggested that this NOD group would benefit from early PaC detection. In our validation cohort, we demonstrated PaC detection in a subpopulation of NOD subjects with the same level of performance as shown in the complete dataset including subjects with and without type 2 diabetes. Notably, the training set did not include subjects with NOD, which supports the evidence that the algorithm detects PaC independently of type 2 diabetes status. Furthermore, the close correlation between 5hmC signatures of PaC in NOD and PaC in non-NOD subjects (see Figure 2C) indicates that the biological phenomena of PaC detected by 5hmC are similar in NOD and non-NOD populations, which supports the use of this test in subjects with NOD.
[0174] Characterized by intraductal papillary proliferation of mucin-producing epithelial cells, IPMN is also known as a precursor lesion to PaC26, despite the low percentage of patients with IPMN progressing to PaC. There is evidence that epigenetic changes are present in IPMN and increase with progression to PaC (27). Notably, the PaC algorithm can detect subjects with IPMN showing moderate to high levels of dysplasia. Further development using the 5hmC signature in the relevant cohort suggests that we may be able to train the algorithm to detect which precursor IPMNs lead to PaC, and that these should be closely monitored (28).
[0175] In conclusion, a highly accurate, scalable, and efficient assay requiring only 10 ng of cfDNA is provided. The combination of the 5hmC assay and the associated PaC-specific detection algorithm enables effective measurement of cancer presence in individuals at high risk of PaC, thereby providing a molecular tool for earlier detection and timely intervention.
Claims
1. A method for detecting the likelihood of pancreatic cancer in a patient from a blood sample obtained from the patient, (a) A step of extracting a cell-free DNA (cfDNA) sample from the blood sample; (b) A step of counting the 5-hydroxymethylcytosine (5hmC)-containing DNA fragments in the cfDNA sample to obtain a total 5hmC fragment count; (c) A step of counting 5hmC-containing DNA fragments mapped to a first genomic location to give a first genomic location 5hmC fragment count; (d) A step of counting 5hmC-containing DNA fragments mapped to a second genomic location to give a second genomic location 5hmC fragment count; (e) If necessary, perform (d) and count 5hmC-containing DNA fragments that are mapped to at least one additional genomic location to give at least one additional genomic 5hmC fragment count; (f) A step of normalizing the 5hmC fragment count obtained in (b) to (e) to give the normalized 5hmC fragment count; (g) A step of scoring the normalized 5hmC fragment counts using each of a plurality of base models, each including a penalized logistic regression fit, wherein one score is generated for each base model by summing the products of the determined correlation coefficient and the normalized 5hmC fragment count, thereby providing a plurality of base model scores; and (h) A step of inputting the plurality of basic model scores into an ensemble model to generate a final score in the form of a probability score p between 0 and 1 that indicates the likelihood of pancreatic cancer in the patient. A method that includes this.
2. The method according to claim 1, further comprising incorporating (h) at least one additional feature type into the ensemble model, in addition to the 5hmC-containing DNA fragment count.
3. The method according to claim 2, wherein the at least one additional feature type comprises WGS fragment counts in each of a set of windows along the genome and / or WGS fragment counts in each of a set of size range bins.
4. The method according to claim 2, wherein the at least one additional feature type includes copy number variation.
5. The method according to claim 2, wherein the at least one additional feature type includes an epigenetic feature other than 5hmC.
6. The method according to claim 2, wherein the at least one additional feature type includes the CA 19-9 level.
7. The method according to claim 2, wherein the at least one additional feature type includes a histone marker.
8. The method according to claim 2, wherein the at least one additional feature type comprises the cfDNA concentration in the blood sample.
9. The method according to claim 2, wherein the at least one additional feature type includes a patient-specific clinical parameter that correlates with the risk of pancreatic cancer.
10. A method for detecting the likelihood of pancreatic cancer in a patient from a blood sample obtained from the patient, (a) A step of extracting a cell-free DNA (cfDNA) sample from the blood sample; (b) A step of counting the 5-hydroxymethylcytosine (5hmC)-containing DNA fragments in the cfDNA sample to obtain a total 5hmC fragment count; (c) A step of determining the 5hmC-containing fragment count in the CTCF-binding region, annotated enhancer region, annotated gene body region, and annotated 3'-UTR genomic region to provide the 5hmC fragment count for each location; (d) A step of determining the WGS fragment count in each of a set of genome windows along the genome and / or the WGS fragment count in each of the set of size range bins; (e) a step of determining the WGS fragment count in a genomic region of a specific size for the purpose of determining copy number variation (CNV); (f) a step of determining the cfDNA concentration in the blood sample; (f) The step of normalizing the fragment counts obtained in (b) to (e) to give the normalized fragment counts; (g) a step of scoring the normalized fragment counts using each of a plurality of base models, each including a penalized logistic regression fit, wherein one score is generated for each base model by summing the products of the determined correlation coefficient and the normalized fragment count, thereby providing a plurality of base model scores; and (h) A step of inputting the plurality of basic model scores into an ensemble model to generate a final score in the form of a probability score p between 0 and 1 that indicates the likelihood of pancreatic cancer in the patient. A method that includes this.
11. The method according to claim 10, further comprising the step of incorporating at least one additional feature type into the ensemble model of (h).
12. The method according to claim 11, wherein the at least one additional feature type comprises the cfDNA concentration in the blood sample.
13. A method for monitoring changes in the state of a patient's pancreatic cancer, comprising the steps of repeating the process described in claim 1 at intervals over a long monitoring period, and determining the difference in probability scores over time.