Methods for detecting breast cancer
A buccal sample-based DNA methylation assay provides a convenient and accurate method for breast cancer detection, addressing the limitations of mammography by offering high diagnostic performance and early identification.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SOLA DIAGNOSTICS GMBH
- Filing Date
- 2025-10-17
- Publication Date
- 2026-04-23
AI Technical Summary
Current breast cancer detection methods, particularly mammography, suffer from overdiagnosis and high false positive rates, necessitating the development of more accurate and convenient molecular biomarkers for early detection.
A non-invasive assay that analyzes DNA methylation profiles from buccal samples using a panel of CpG dinucleotides to assess breast cancer presence or absence, allowing for self-sampling and providing high diagnostic performance with a simplified test.
The assay offers convenient, cost-effective, and rapid breast cancer detection with high accuracy, enabling early identification and reducing the need for invasive procedures.
Smart Images

Figure EP2025079995_23042026_PF_FP_ABST
Abstract
Description
[0001] METHODS FOR DETECTING BREAST CANCER
[0002] FIELD OF THE INVENTION
[0003] The present invention relates to assays for assessing the presence or absence of breast cancer in an individual. The invention also relates to in vitro methods of assaying DNA and detecting methylation of the DNA therein, including amplification-based assay and detection methods such as PCR. The invention also relates to methods of treating breast cancer in an individual, and methods of monitoring the breast cancer status in an individual. The invention also relates to arrays and kits for performing the assay and detection methods.
[0004] BACKGROUND TO THE INVENTION
[0005] Breast Cancer (BC) is the most common and second most fatal cancer affecting women, emphasizing the demand for effective detection methods. Current screening or detection approaches predominantly revolve around imaging techniques, primarily mammography. While mammography screening has been proven to reduce breast cancer mortality it can lead to overdiagnosis and not insignificant false positive rates. Therefore, the combination of existing screening and early detection approaches with molecular biomarkers has been named a key priority in a recent consensus statement for breast cancer diagnosis and effective treatment.
[0006] SUMMARY OF THE INVENTION
[0007] The current inventors set out to understand whether DNAme (DNA methylation) profiles within non-invasive surrogate samples may be used to detect the presence or absence of breast cancer. The inventors also set out to understand whether said DNAme profiles within non-invasive surrogate samples may be associated with the breast cancer, and therefore whether such profiles may be capable of functioning as surrogate markers for diagnosing an individual with breast cancer.
[0008] In contrast to existing assays that may be used for assessing the presence or absence of breast cancer in an individual, the present invention provides hugely convenient, simple and accurate test for breast cancer diagnosis or determination of breast cancer risk.
[0009] The assay according to the present invention is more convenient because it allows the individual to obtain a sample from their own buccal area and for this sample to be assayed for DNAme. The assay according to the present invention is more convenient because it merely requires the assessment of CpG methylation at a single CpG dinucleotide, yet nevertheless maintains a high level of diagnostic performance, as indicated by a high AUC. It is not required for the individual to visit a clinic and / or for a clinician to obtain the sample from the individual. The convenience of the assay would increases the likelihood of a breast cancer being identified in individuals at an early stage.
[0010] Overall, such a simplified test provides not only cost benefits to individuals and public health organisations, it also allows assessments to be conducted more rapidly and more frequently, thereby allowing women to subject to the appropriate course of further cancer tests (i.e. such as obtaining tissue by means of invasive procedures to make a histological diagnosis), and potentially cancer treatment, without delay.
[0011] Accordingly, the invention provides an assay for assessing the presence or absence of breast cancer in an individual, the assay comprising: a. providing a sample which has been taken from the buccal area of the individual, the sample comprising a population of DNA molecules; b. determining in the population of DNA molecules in the sample the methylation status of a test panel of one or more CpGs selected from the 30,037 CpGs identified at nucleotide positions 15 to 16 in SEQ ID NOs 1 - 30,037; and c. assessing the presence or absence of breast cancer in the individual based on the methylation status of the CpGs in the test panel.
[0012] The invention further provides an in vitro method of assaying DNA and detecting methylation of the DNA therein, the method comprising measuring a methylation status of one or more CpGs, wherein selected from the 30,037 CpGs identified at nucleotide positions 15 to 16 in SEQ ID NOs 1 - 30,037.
[0013] The invention further provides a method of treating breast cancer in an individual, the method comprising: i. assessing the presence or absence of breast cancer in an individual according to the assay of the invention; and ii. administering one or more therapeutic treatments or measures to the individual based on the assessment.
[0014] The invention further provides a method of monitoring the breast cancer status in an individual, the method comprising: (a) assessing the presence or absence of breast cancer in an individual by performing the assay of the invention at a first time point; (b) assessing the presence or absence of breast cancer in the individual by performing the assay of the invention at one or more further time points; and (c) monitoring any change in the breast cancer status of the individual.
[0015] The invention also provides an array capable of discriminating between methylated and non-methylated forms of CpGs; the array comprising oligonucleotide probes specific for a methylated form of each CpG in a CpG panel and oligonucleotide probes specific for a non-methylated form of each CpG in the panel; wherein the panel consists of at least 500 CpGs selected from the CpGs identified at nucleotide position 15 to 16 in SEQ ID NOs 1 to 30,037.
[0016] BRIEF DESCRIPTION OF THE FIGURES
[0017] Figure 1. Study overview and epigenome-wide association study for breast cancer in buccal, cervical and blood samples a) Overview of the study and datasets. b) eFORGE results of significant enrichments of tissue or cell type-specific signals in the subset DMPs with hypomethylation in buccal samples and hypermethylation in cervical samples. c) Mean methylation of subset of DMPs with hypomethylation in buccal samples and hypermethylation in cervical samples in breast tissue (GSE225845). d) Mean methylation of subset of DMPs with hypermethylation in buccal samples and hypermethylation in cervical samples in breast tissue (GSE225845). e) Mean delta beta values between BC cases and controls for overlapping DMPs within the buccal and cervical discovery sets, prior to implementing multiple testing correction. Annotated is the corresponding odds ratio [95% confidence interval] . f) eFORGE results of significant enrichments of tissue or cell type-specific signals in the subset DMPs with hypomethylation in buccal samples and hypomethylation in cervical samples. g) Mean methylation of subset of DMPs with hypomethylation in buccal samples and hypomethylation in cervical samples in breast tissue (GSE225845). h) eFORGE results of significant enrichments of tissue or cell type-specific signals in the subset DMPs with hypermethylation in buccal samples and hypomethylation in cervical samples. i) Mean methylation of subset of DMPs with hypermethylation in buccal samples and hypomethylation in cervical samples in breast tissue (GSE225845). p values are derived from Wilcoxon tests (no multiple testing correction was applied). Box plots correspond to standard Tukey representation, with boxes indicating mean and interquartile range, and lines indicating smallest and largest values within 1.5 times of the 25th and 75th percentile, respectively. Individual data points are overlaid. No corrections for multiple testing were carried out.
[0018] Figure 2. Validation of the developed indices in independent datasets a) Distribution of the WID-buccal-BC index with respect to immune cell proportion in the External Validation Set. b) Distribution of the WID-buccal-BC index with respect to subject age in the External Validation Set. c) Evaluation of the WID-buccal-BC index, adjusted for age and immune cell proportion, in the External Validation Set. d) ROC curve of the WID-buccal-BC index, adjusted for age and immune cell proportion, stratified by median immune cell proportion in the External Validation Set. e) Distribution of the WID-cervical index with respect to immune cell proportion in the External Validation Set. f) Distribution of the WID-cervical -BC index with respect to subj ect age in the External Validation Set. g) Evaluation of the WID-cervical-BC index, adjusted for age and immune cell proportion, in the External Validation Set. h) ROC curve of the WID-cervical-BC index, adjusted for age and immune cell proportion, stratified by median immune cell proportion in the External Validation Set. i) Distribution of the WID-blood-BC index with respect to neutrophil proportion in GSE237036. j) Evaluation of the WID-blood-BC index, adjusted for neutrophil proportion, in GSE237036. k) ROC curve of the WID-blood-BC index, adjusted for neutrophil proportion, stratified by median neutrophil proportion in GSE237036. ****p < 0.0001, ***p < 0.001 in two-sided Wilcoxon test. Box plots correspond to standard Tukey representation, with boxes indicating mean and interquartile range, and lines indicating smallest and largest values within 1.5 times of the 25th and 75th percentile, respectively. Individual data points are overlaid. No corrections for multiple testing were carried out.
[0019] Figure 3. Association of the developed indices with clinical and epidemiological parameters and genetic risk factors. Presented index scores have been adjusted for age and either IC (WID-buccal / cervical-BC) or neutrophil proportion (WID-blood- BC). a) Assessment of the WID-buccal-BC index with cancer stage. b) Assessment of the WID-buccal-BC index with cancer grade. c) Assessment of the WID-buccal-BC index with ER status. d) Assessment of the WID-buccal-BC index with PR status. e) Assessment of the WID-buccal-BC index with HER2 status. f) Distribution of the WID-buccal-BC index with respect to polygenic risk score. g) The WID-buccal-BC index in BRCA1 and BRCA2 mutation carriers and controls. h) ROC curves of the WID-buccal-BC index in BRCA1- and BRCA2 mutation carriers versus controls. i) Assessment of the WID-cervicall index with cancer stage. j) Assessment of the WID-cervical-BC index with cancer grade. k) Assessment of the WID-cervical-BC index with ER status. l) Assessment of the WID-cervical-BC index with PR status. m) Assessment of the WID-cervical-BC index with HER2 status. n) Distribution of the WID-cervical-BC index with respect to polygenic risk score. o) The WID-cervical-BC index in BRCA1 and BRCA2 mutation carriers and controls. p) ROC curves of the WID-cervical-BC index in BRCA1- and BRCA2 mutation carriers versus controls. l) The WID-blood-BC index in BRCA1 and BRCA2 mutation carriers and controls. m) ROC curves of the WID-blood-BC index BRCA1- and BRCA2 mutation carriers versus controls.
[0020] ****p < 0.0001, ***p < 0.001, **p < 0.01 in two-sided Wilcoxon test. Box plots correspond to standard Tukey representation, with boxes indicating mean and interquartile range, and lines indicating smallest and largest values within 1.5 times of the 25th and 75th percentile, respectively. Individual data points are overlaid. No corrections for multiple testing were carried out.
[0021] Figure 4. The developed indices evaluated in breast tissue. Presented index scores have been adjusted for age and either IC (WID-buccal / cervical-BC) or neutrophil proportion (WID-blood-BC). a) Assessment of the WID indices in the tissue at risk set (**p < 0.01, ***p < 0.001 in two-sided Wilcoxon test). b) ROC curves of the WID indices in the tissue at risk set, normal-adjacent breast tissue (n=14) versus triple negative breast cancer tissue (n=14). c) Assessment of the WID indices in the TCGA set < 0.0001 in two-sided Wilcoxon test). d) ROC curves of the WID indices in the TCGA set (n = 97 controls, n = 792 cancer cases). e) Assessment of the WID indices in GSE225845 (**p < 0.01, ***p < 0.001 in two- sided Wilcoxon test). f) ROC curves of the WID indices in GSE225845, normal-adjacent breast tissue versus tumor tissue.
[0022] ****p < 0.0001, ***p < 0.001, **p < 0.01 in two-sided Wilcoxon test. Box plots correspond to standard Tukey representation, with boxes indicating mean and interquartile range, and lines indicating smallest and largest values within 1.5 times of the 25th and 75th percentile, respectively. Individual data points are overlaid. No corrections for multiple testing were carried out.
[0023] Figure 5. Discovery sets numbers and characteristics for buccal, cervical, and blood DNA methylation data used for epigenome-wide analysis and training the WID-indices Note that individuals in the buccal and cervical training sets are overlapping
[0024] Figure 6. Validation sets numbers and characteristics for buccal, cervical, and blood DNA methylation data used in validating the developed WID-indices
[0025] Note that individuals in the buccal and cervical validation sets are overlapping
[0026] Figure 7. BRCA1 / 2 mutation carrier dataset overview Figure 8. Breast tissue DNA methylation dataset numbers and characteristics
[0027] Figure 9. Principal Component Analysis (PCA) of the top 2010% variable CpGs from all methylation datasets utilized in the present study for both discovery and validation purposes. Principal components 1 and 2 are shown of: a) All samples, coloured by immune cell (IC) fraction. b) All samples, coloured by sample type. c) All samples, coloured by corresponding data set. d) All buccal samples, coloured by immune cell (IC) fraction. e) All buccal samples, coloured by sample type. f) All buccal samples, coloured by corresponding data set. g) All cervical samples, coloured by immune cell (IC) fraction. h) All cervical samples, coloured by sample type. i) All cervical samples, coloured by corresponding data set. j) All blood samples, coloured by neutrophil fraction. k) All blood samples, coloured by sample type. l) All blood samples, coloured by corresponding data set. m) All blood samples, coloured by immune cell (IC) fraction. n) All buccal samples, coloured by sample type. o) All buccal samples, coloured by corresponding data set.
[0028] Figure 10. Training of indices discriminating breast cancer cases from controls based on three tissue types. a) WID-buccal: area under the receiver operating characteristic curve (AUROC) in the internal validation set as a function of the number of CpGs used to train the classifier. The AUROC for the ridge model, at 30,000 CpGs based on ranked p-values, was 0.93. b) WID-buccal: area under the receiver operating characteristic curve (AUROC) of out- of-bag samples as a function of the number of CpGs used to train the classifier. The AUROC for the ridge model, at 30,000 CpGs based on ranked p-values, was 0.89. c) WID-cervical: area under the receiver operating characteristic curve (AUROC) in the internal validation set as a function of the number of CpGs used to train the classifier. The AUROC for the ridge model, at 30,000 CpGs based on ranked p- values, was 0.93. d) WID-cervical: area under the receiver operating characteristic curve (AUROC) of out-of-bag samples as a function of the number of CpGs used to train the classifier. The AUROC for the ridge model, at 30,000 CpGs based on ranked p-values, was 0.82. e) WID-blood: area under the receiver operating characteristic curve (AUROC) in the internal validation set as a function of the number of CpGs used to train the classifier. The AUROC for the ridge model, at 30,000 CpGs based on ranked p-values, was 0.98. f) WID-blood: area under the receiver operating characteristic curve (AUROC) of out- of-bag samples as a function of the number of CpGs used to train the classifier. The AUROC for the ridge model, at 30,000 CpGs based on ranked p-values, was 0.91.
[0029] Figure 11. Comparison of cell type proportion in breast cancer cases and controls in the Discovery and External Validation Sets a) Inferred immune cell fraction in breast cancer cases and controls in the buccal Discovery set versus the buccal Validation Set. b) Inferred immune cell fraction in breast cancer cases and controls in the cervical Discovery set versus the cervical Validation Set. c) Inferred neutrophil fraction in breast cancer cases and controls in the blood Discovery set versus GSE237036. d) Distribution of cell type proportions in buccal samples of the Discovery Set (left) and External Validation Set (right) inferred using the HEpiDISH algorithm. e) Distribution of cell type proportions in cervical samples of the Discovery Set (left) and External Validation Set (right) inferred using the HEpiDISH algorithm. f) Distribution of cell type proportions in blood samples of the Discovery Set (left) and GSE237036 (right) inferred using the HEpiDISH algorithm.
[0030] ***p < 0.001, *p < 0.05 in two-sided Wilcoxon test. Box plots correspond to standard Tukey representation, with boxes indicating mean and interquartile range, and lines indicating smallest and largest values within 1.5 times of the 25th and 75th percentile, respectively. Individual data points are overlaid. No corrections for multiple testing were carried out.
[0031] Figure 12. Assessing differentially methylated positions in the Discovery Sets of three tissue types. a) Histogram showing the distribution of P-values for methylation differences in CpG sites (after adjustment for immune cell fraction and age) between the breast cancer (BC) case and control groups within the buccal sample Discovery Set. b) Manhattan plot for buccal EWAS results. Significant p values after Benjamini- Hochberg correction are shown in orange. c) Histogram showing the distribution of P-values for methylation differences in CpG sites (after adjustment for immune cell fraction and age) between the breast cancer (BC) case and control groups within the cervical sample Discovery Set. d) Manhattan plot for cervical EWAS results. Significant p values after Benjamini- Hochberg correction are shown in orange. e) Histogram showing the distribution of P-values for methylation differences in CpG sites (after adjustment for immune cell fraction and age) between the breast cancer (BC) case and control groups within the blood sample Discovery Set. f) Manhattan plot for blood EWAS results. No results remained significant after Benjamini-Hochberg correction.
[0032] Note, significance thresholds for buccal, cervical and blood samples are different as Benjamini-Hochberg correction was applied, a method for controlling the false discovery rate using a ‘step-up’ method. Alternative approaches, such as correction generic thresholds such as p < 10'5have also been proposed and could alternatively be applied.
[0033] Figure 13. Relation to island and gene group of differentially methylated probes between breast cancer cases and controls a) Relation to island of differentially methylated probes between breast cancer cases and controls in buccal samples, that remain significant post Benjamini-Hochberg correction. b) Relation to island of differentially methylated probes between breast cancer cases and controls in cervical samples, that remain significant post Benjamini-Hochberg correction. c) Gene group of differentially methylated probes between breast cancer cases and controls in buccal samples, that remain significant post Benjamini-Hochberg correction. d) Gene group of differentially methylated probes between breast cancer cases and controls in cervical samples, that remain significant post Benjamini-Hochberg correction. Figure 14. Comparison of methylated signals in epithelial and immune fractions and differential methylated region analysis in buccal and cervical cells
[0034] Mean delta beta values between BC cases and controls in pure inferred epithelial cells. for overlapping DMPs within the buccal and cervical discovery sets, prior to implementing multiple testing correction. Annotated is the corresponding odds ratio [95% confidence interval], b) Mean delta beta values between BC cases and controls in pure inferred immune cells. for overlapping DMPs within the buccal and cervical discovery sets, prior to implementing multiple testing correction. Annotated is the corresponding odds ratio [95% confidence interval], c) Differentially methylated regions (DMRs) in buccal (blue) and cervical (red) samples within the discovery sets. d) Examples of DMRs overlapping between buccal (dashed lines) and cervical samples (solid lines). Mean methylation beta values in controls (blue) and BC cases (red) are visualized. NCK1 / RP11-85F14.1 (p=0.00094 for buccal samples; p=0.015 for cervical samples) and CCDC88C (p=0.0072 for buccal samples; p=0.012 for cervical samples) exhibit opposing differential methylation between cases and controls in buccal and cervical samples (i.e., hypermethylation in BC cases in cervical samples and hypomethylation in BC cases in buccal samples), whereas RP11-551L14.1 (p=0.011 for buccal samples; 0.0002 for cervical samples) and LTBP4 (p=0.002 for buccal samples. p=0.0045 for cervical samples) show the same directionality of differential methylation in the two samples.
[0035] Figure 15. Functional enrichment analysis of differentially methylated positions (adjusted p-value after BH correction: <0.05) identified between breast cancer cases and controls in the buccal Discovery Set (585 CpGs).
[0036] The figure illustrates the top 10 enriched biological processes. There were no significantly enriched molecular or cellular functions
[0037] Figure 16. Functional enrichment analysis of differentially methylated positions (adjusted p-value after BH correction: <0.05) identified between breast cancer cases and controls in the cervical Discovery Set (21,614 CpGs) a) Biological processes b) Molecular functions c) Cellular functions
[0038] Figure 17. ROC curves of the developed indices (unadjusted for age, IC or neutrophils) in the validation sets a) ROC curves of the WID-buccal-BC index in the external validation set (black) and stratified by median immune cell proportion (turquoise and blue). b) ROC curves of the WID-buccal-BC index in the external validation set (black) and stratified by menopausal status (turquoise and blue). c) ROC curves of the WID-buccal-BC index in the external validation set (black) and stratified by median age at menarche (turquoise and blue). d) ROC curves of the WID-cervical-BC index in the external validation set (black) and stratified by median immune cell proportion (turquoise and blue). e) ROC curves of the WID-cervical-BC index in the external validation set (black) and stratified by menopausal status (turquoise and blue). f) ROC curves of the WID-cervical-BC index in the external validation set (black) and stratified by median age at menarche (turquoise and blue). g) ROC curves of the WID-blood-BC index in GSE237036 (black) and stratified by median neutrophil proportion (turquoise and blue).
[0039] Figure 18. Assessment of the WID-buccal-BC and WID-cervical-BC index (adjusted for age and IC) with:
[0040] Age at menarche b) Family history c) Age at first live birth d) Menopausal status e) Hormone replacement therapy (post-menopausal women only) f) Nodal status h) Age at menarche i) Family history j) Age at first live birth h) Menopausal status k) Hormone replacement therapy (post-menopausal women only) l) Nodal status ****p < 0.0001, ***p < 0.001, **p < 0.01 in two-sided Wilcoxon test. Box plots correspond to standard Tukey representation, with boxes indicating mean and interquartile range, and lines indicating smallest and largest values within 1.5 times of the 25th and 75th percentile, respectively. Dots indicate outlier values. No corrections for multiple adjustment were carried out.
[0041] Figure 19. Assessing WID-buccal-BC and WID-cervical-BC in blood samples a) Distribution of the WID-buccal-BC index with respect to neutrophil proportion in the FORECEE (Blood training) dataset. b) ROC curve of the WID-buccal-BC index, adjusted for neutrophil proportion, in the FORECEE (Blood training) dataset c) Distribution of the WID-cervical-BC index with respect to neutrophil proportion in the FORECEE (Blood training) dataset. d) ROC curve of the WID-cervical-BC index, adjusted for neutrophil proportion, in the FORECEE (Blood training) dataset. e) Distribution of the WID-buccal-BC index with respect to neutrophil proportion in GSE237036. f) ROC curve of the WID-buccal-BC index, adjusted for neutrophil proportion, in GSE237036. g) Distribution of the WID-cervical-BC index with respect to neutrophil proportion in GSE237036. h) ROC curve of the WID-cervical-BC index, adjusted for neutrophil proportion, in GSE237036.
[0042] Figure 20. Exemplary diagnostic thresholds applied using assays of the invention for diagnosing breast cancer in an individual
[0043] Exemplary threshold for diagnosing an individual using an assay of the invention, wherein the assessment is based on the use of a WID-buccal-BC-index algorithm as described herein. AUC is indicated. Sensitivity is indicated in all instances whereby the specificity of each assay when applying the exemplary thresholds is 75%. The assay in each panel consists of the determination of the methylation status of (a) the 30,000 CpGs defined according to SEQ ID NOs 1 to 30,037, excluding 37 CpGs not present on EPIC version 2; (b) a random 25 CpGs among the 30,037 CpGs defined according to SEQ ID NOs 1-
[0044] 30,037, specifically the 25 CpGs defined according to SEQ ID NOs 1, 6, 25, 26, 48, 80, 99, 106, 154, 182, 265, 267, 348, 353, 372, 430, 436, 447, 460, 471, 489, 495, 506, 513, and 545; (c) a single CpG defined according to SEQ ID NO: 1.
[0045] Figure 21. Data demonstrating diagnostic capability of each individual CpG among the 30,037 CpGs of the invention described herein
[0046] The CpGs defined according to SEQ ID NOs 1-30,037 are identified in a table together with the Illumina manifest ‘eg’ identifier, chromosome number in which the CpG is located, and genomic position in that chromosome, SEQ ID NO, the methylation status (i.e. hyper = methylated; hypo = unmethylated) that is indicative of the presence of breast cancer, and AUC values yielded when the CpG is applied to the discovery or validation cohort of samples (see Examples).
[0047] Figure 22. Data demonstrating the diagnostic capability of each individual DMR among the 145 DMRs of the invention described herein
[0048] The DMRs defined according to SEQ ID NOs 30,038-30,182 are identified according to their SEQ ID NO a table which additionally identifies the chromosome number in which the DMR is located, and genomic position in that chromosome, DMR width, number of CpGs in the DMR, the methylation status (i.e. hyper = methylated; hypo = unmethylated) that is indicative of the presence of breast cancer, and AUC values yielded when the DMR is applied to the discovery or validation cohort of samples (see Examples).
[0049] SEQUENCES
[0050] The 30,037 CpGs described herein and identified at nucleotide positions 15 to 16 in SEQ ID NOs 1 - 30,037 are disclosed in Figure 21 together with the sequence listing that forms part of this application.
[0051] The 145 differentially methylated regions (DMRs) defined by SEQ ID NOs 30, OSS- SO, 182 are disclosed in the sequence listing that forms part of this application. The individual CpGs comprised within each of the 145 DMRs are identified according to their “eg” identifier as set out in Table 1. These eg identifiers may be used to correspond the CpGs comprised within the 145 DMR to the CpGs defined according to SEQ ID NOs 1- 30,037 (see Figure 21 which identifies each of the 30,037 CpGs according to its “eg” identifier and SEQ ID NO). Each eg identifier thereby corresponds to a CpG at nucleotide positions 15 to 16 of a particular sequence defined according to SEQ ID NOs 1-30,037. The DMR defined by SEQ ID NO 30,060 consists of the nucleotide sequence CGCGAGGC.
[0052] The DMR defined by SEQ ID NO 30,115 consists of the nucleotide sequence CGCAATC. All sequences are shown in the 5 ’ to 3 ’ direction.
[0053] able 1
[0054]
[0055]
[0056]
[0057]
[0058]
[0059]
[0060]
[0061] DETAILED DESCRIPTION OF THE INVENTION
[0062] The present invention will be described with respect to particular embodiments and with reference to certain drawings but the disclosure is not limited thereto. Any reference signs in the claims shall not be construed as limiting the scope.
[0063] It should be appreciated that “embodiments” of the disclosure can be specifically combined together unless the context indicates otherwise. The specific combinations of all disclosed embodiments (unless implied otherwise by the context) are further disclosed embodiments of the claimed invention.
[0064] In addition, as used in this specification and the appended claims, the singular forms “a”, “an”, and “the” include plural referents unless the content clearly dictates otherwise. Thus, for example, reference to “a polynucleotide” includes two or more polynucleotides, reference to “a protein” includes two or more proteins, and the like.
[0065] "About" as used herein when referring to a measurable value such as an amount, a temporal duration, and the like, is meant to encompass variations of ± 20 % or ± 10 %, more preferably ± 5 %, even more preferably ± 1 %, and still more preferably ± 0.1 % from the specified value, as such variations are appropriate to perform the disclosed methods.
[0066] “Polynucleotide”, “nucleotide sequence”, “DNA sequence”, or “nucleic acid molecule(s)” as used herein refers to a polymeric form of nucleotides of any length, which comprises ribonucleotides or deoxyribonucleotides. This term refers only to the primary structure of the molecule. Thus, this term includes double- and single-stranded DNA, and RNA. The term “polynucleotide ” as used herein, may be a single or double stranded covalently-linked sequence of nucleotides in which the 3' and 5' ends on each nucleotide are joined by phosphodiester bonds. The polynucleotide may be made up of deoxyribonucleotide bases or ribonucleotide bases. Polynucleotides may be manufactured synthetically in vitro or isolated from natural sources. Polynucleotides may further include modified DNA or RNA, for example DNA or RNA that has been methylated, or RNA that has been subject to post-translational modification, for example 5 ’-capping with 7- methylguanosine, 3 ’-processing such as cleavage and polyadenylation, and splicing. Polynucleotides may also include synthetic nucleic acids (XNA), such as hexitol nucleic acid (HNA), cyclohexene nucleic acid (CeNA), threose nucleic acid (TNA), glycerol nucleic acid (GNA), locked nucleic acid (LNA) and peptide nucleic acid (PNA). Chromosomal coordinates with reference to a human polynucleotide sequence are defined with respect to hgl9.
[0067] A Differentially Methylated Region (DMR) refers to a double-stranded polynucleotide sequence, formed of complementary polynucleotide sequences, whereby differential CpG methylation of one or more CpGs may be observed. The DMR may be an isolated double-stranded polynucleotide sequence. The DMR may be a region within a double stranded sequence optionally comprised within a population of DNA molecules. The DMR may be a region within a genome. The skilled person is aware that CpG methylation observed in a first strand of a DMR effectively acts as a surrogate indicator for the presence of CpG methylation at the complementary CpG dinucleotide in the complementary second strand of said DMR. For the purposes of determining the methylation status of CpG in a DMR, the skilled person would therefore be aware that determining CpG methylation in one strand of a DMR is sufficient to define a CpG within a DMR as a whole (i.e. on either strand) as methylated. Described particularly herein are methods for determining methylation status of CpGs in a test panel whereby the methods rely on PCR directed towards one strand of a DMR, in order to thereby determine the methylation status of CpGs within the DMR as a whole. The skilled person would be able to devise a method likewise relying on PCR but directed to the complementary strand in order to determine the methylation status of CpGs within the DMR as a whole.
[0068] As used herein, a percentage of “sequence identity” will be understood to arise from a comparison of two sequences in which they are aligned together to give a maximum correlation between the sequences. This may include inserting “gaps” in either one or both sequences to enhance the degree of alignment. The percentage of sequence identity may then be determined over the length of each of the sequences being compared. For example, a nucleotide sequence (“subject sequence”) having at least 95% “sequence identity” with another nucleotide sequence (“query sequence”) is intended to mean that the subject sequence is identical to the query sequence except that the subject sequence may include up to five nucleotide alterations per 100 nucleotides of the query sequence. In other words, to obtain a nucleotide sequence of at least 95% sequence identity to a query sequence, up to 5% (i.e. 5 in 100) of the nucleotides in the subject sequence may be inserted or substituted with another nucleotide or deleted. Percentage identity is also used equivalently in relation to protein sequences but in terms of comparing the corresponding amino acids. Assay
[0069] A CpG as defined herein refers to the CG dinucleotide motif identified in relation to each SEQ ID NO. nucleotide sequence (SEQ ID NOs 1 to 30,037), and identified in the nucleotide sequences defined by SEQ ID NOs 30,038-30,099, wherein the cytosine base of the CG dinucleotide may be modified. Thus, by determining the methylation status of any panel of one or more CpGs defined by or identified in a given SEQ ID NO, it is meant that a determination is made as to the methylation status of the cytosine of the CG dinucleotide motif, accepting that variations in the sequence upstream and downstream of any given CpG may exist due to sequencing errors or variation between individuals. The CpGs selected from SEQ ID NOs 1-30,037 are identified at nucleotide positions 15 to 16 in the sequences defined by SEQ ID NOs 1-30,037. The CpGs selected from the CpGs comprised in the differentially methylated regions (DMRs) defined by SEQ ID NOs 30,0371-30,182 are identified by a CG dinucleotide.
[0070] Any of the assays described herein for assessing the presence or absence of breast cancer in an individual are capable of being utilised for assessing the presence or absence of breast cancer. Most preferably, the assay is a diagnostic assay for determining the presence or absence of breast cancer in the individual. The present inventors determined CpG methylation levels in a sample which has been taken from the buccal area using an assay according to the present invention and compared the ability of the assay to accurately diagnose individuals with breast cancer based on their DNA methylation status with a conventional assessment pathway.
[0071] In any of the assays described herein, in the sample which has been taken from the buccal area of the individual, the sample comprises a population of DNA molecules. The individual may be asymptomatic or symptomatic of a breast cancer. Symptoms of breast cancer will be well known to a person of skill in the art
[0072] The sample may have been self-collected or collected by a medical professional, such as a clinician. The sample may have been obtained using any suitable means, such as a buccal swab. Any suitable device for sample collection from the buccal area may be used. Preferably the sample for example has been obtained by contacting the buccal area with a brush device which collects epithelial cells from the buccal area, e.g. using a Copan Medical Diagnostics buccal swab.
[0073] Alternative to the providing step of the assays of the invention comprising the use of a sample which has been taken from the buccal area of the individual, the providing step of the assays of the invention may comprise obtaining a sample from the buccal area of the individual, the sample comprising a population of DNA molecules.
[0074] The invention provides for assessing the presence or absence of breast cancer in an individual, the assay comprising: a. providing a sample which has been taken from the buccal area of the individual, the sample comprising a population of DNA molecules; b. determining in the population of DNA molecules in the sample the methylation status of a test panel of one or more CpGs selected from the 30,037 CpGs identified at nucleotide positions 15 to 16 in SEQ ID NOs 1 - 30,037; and c. assessing the presence or absence of breast cancer in the individual based on the methylation status of the CpGs in the test panel.
[0075] The assay of the invention may be further defined wherein the test panel of one or more CpGs is selected from the 585 CpGs identified at nucleotide positions 15 to 16 in SEQ ID NOs 1-585, preferably selected from the 222 CpGs identified at nucleotide positions 15 to 16 in SEQ ID NOs 1 - 222, and more preferably selected from the 53 CpGs identified at nucleotide positions 15 to 16 in SEQ ID NOs 1 - 53.
[0076] The assay may be further defined wherein the test panel of one or more CpGs comprises two or more CpGs selected from the CpGs identified in the differentially methylated regions (DMRs) defined by SEQ ID NOs 30,038 - 30,182, preferably identified in a single DMR defined by SEQ ID NOs 30,038 - 30,099, and more preferably identified in a single DMR defined by SEQ ID NOs 30,038 - 30,051.
[0077] The assay of the invention may be further defined wherein the test panel of one or more CpGs comprises at least three CpGs, at least three CpGs, at least four CpGs, at least five CpGs, or preferably all of the CpGs identified in a single DMR defined by SEQ ID NOs 30,038 - 30,182, preferably identified in a single DMR defined by SEQ ID NOs 30,038 - 30,099, and more preferably identified in a single DMR defined by SEQ ID NOs 30,038 - 30,051.
[0078] The assay of the invention may be further defined wherein the test panel of one or more CpGs comprises 100, 500, 1,000, 2,000, 5,000, 10,000, 20,000, or all 30,037 CpGs selected from the 30,037 CpGs identified at nucleotide positions 15 to 16 in SEQ ID NOs 1 - 30,037.
[0079] The assay comprises assessing the presence or absence of breast cancer in the individual based on the methylation status of the CpGs in the test panel. Accordingly, the individual may be assessed as having breast cancer by the methylation status of the test panel of the one or more CpGs. The CpGs identified at nucleotide positions 15 to 16 in SEQ ID NOs 1-30,037 may indicate the presence of presence of breast cancer when determined to be methylated or unmethylated, depending on the specific CpG comprised in the test panel. The methylation status of each CpG indicative of the presence of breast cancer individual is shown in Figures 21 and 22.
[0080] Breast cancer may comprise any type of breast cancer or breast cancer metastasis known in the art, for example ductal carcinoma in situ; an invasive ductal carcinoma such as tubular type invasive ductal carcinoma (IDC), medullary type IDC, mucinous type IDC, papillary type IDC, cribriform type IDC; invasive lobular carcinoma, inflammatory breast cancer, lobular carcinoma in situ, male breast cancer, luminal A breast cancer, luminal B breast cancer, triple-negative / basal-like breast cancer, HER2-enriched breast cancer, normal-like breast cancer, Paget’s Disease of the nipple, Phyllodes tumours of the breast, or metastatic breast cancer.
[0081] Described herein are assays that utilise a statistically robust test panel of one or more CpGs whose methylation status can be determined to provide a reliable prediction of the presence or absence of breast cancer in an individual. By determining the methylation status of each CpG within the panel of one or more CpGs, the individual’s cancer status can be determined with statistically robust sensitivity and specificity. The skilled person would understand that the methylation status of each CpG within a panel of one or more CpGs can be determined by any suitable means. Any one method, or a combination of methods, may be used to determine the methylation status of each CpG within a panel of one or more CpGs.
[0082] Various exemplary methods for determining the methylation status of each CpG within a panel of one or more CpGs are described herein. For example, in one method a percent methylated reference (PMR) value of a CpG may be determined. In another method the methylation P-values of a CpG may be determined. In another method, methylation-sensitive enzymes can be employed which digest or cut only in the presence of methylated DNA. Preferably such methods comprising the use of methylation-sensitive enzymes would not comprise subjecting the one or more CpGs in the test panel to bisulfite treatment, and would preferably further comprise digital PCR. Different mechanisms may be employed to determine specific values depending on the circumstances, such as PCR- based mechanisms or array-based mechanisms, and these different methods are discussed further herein in ‘Assessment of methylation status of CpGs’. The step of determining the methylation status of the CpGs in the test panel may particularly comprise: i. bisulfite treatment of the test panel of one or more CpGs; and / or ii. an amplification-based technique, preferably wherein the amplificationbased technique is PCR, more preferably further comprising quantification and / or the use of a probe comprising a detection moiety; and / or iii. sequencing the test panel of one or more CpGs.
[0083] As described herein, the methods of determining methylation status when applied to the assay of the invention can establish absence of breast cancer in the individual.
[0084] A skilled person will readily appreciate that methylation status of the test panel of one or more CpGs can indicate a “likelihood” or “risk” or “prediction” of any of the assays of the invention correctly assessing the presence or absence of breast cancer in an individual. This is because the assessment is based upon a correlation between DNA methylation profiles of tissue samples and individual disease status. Nevertheless, as demonstrated by data set out in the Examples and elsewhere herein, the assays of the invention provide such correlations with high statistical accuracy, thus providing the skilled person with a high degree of confidence that the methylation status of the CpGs of the test panel for any individual whose breast cancer status is required to be tested will provide an accurate correlation with actual disease status for the individual.
[0085] In the context of the present invention, “likelihood”, “risk” and “prediction” may be used synonymously with each other.
[0086] Any references herein to sequences, genomic sequences and / or genomic coordinates are derived based upon Homo sapiens (human) genome assembly GRCh37 (hgl9). The skilled person would understand variations in the nucleotide sequences of any given sequence, particularly the sequences of SEQ ID NOs 1 to 30,182, may exist due to sequencing errors and / or variation between individuals.
[0087] The assay of the invention represents a ‘prediction’ because the determination of any test panel of one or more CpGs in accordance with the invention is unlikely to be capable of diagnosing every individual as having or not having cancer with 100% specificity and 100% sensitivity. Rather, depending on the specific set of CpGs, the false positive and false negative rate will vary. In other words, the inventors have discovered that an assay comprising the determination of a test panel of CpGs according to the invention can achieve variable levels of sensitivity and specificity for predicting the presence or absence of breast cancer, as defined by receiver operating characteristics. Such sensitivity and specificity can be seen from the data disclosed herein to be achievable at high proportions, demonstrating accurate and statistically-significant discriminatory capability.
[0088] In any of the assays described herein, the step of assessing the presence or absence of breast cancer in an individual may involve the application of a threshold value. This value may typically depend on the method used for determining the methylation status of the CpGs in the test panel of one or more CpGs. Typically, the threshold value will relate to the proportion methylated DNA molecules in the sample having being methylated or unmethylated (depending on the methylation directionality indicative of breast cancer at specific CpGs as described herein) at one or more CpGs of the test panel. Determining the methylation status can be performed by any of the suitable methods described herein. In any of such assays of the invention, the assay may be characterised as having a ROC AUC of 0.55 or more, 0.56 or more, 0.57 or more, 0.575 or more, 0.58 or more, 0.59 or more, 0.60 or more, 0.61 or more, 0.62 or more, 0.625 or more, 0.63 or more, 0.64 or more, 0.65 or more, 0.66 or more, 0.67 or more, 0.675 or more, 0.68 or more, 0.69 or more, 0.70 or more, 0.71 or more, 0.72 or more, 0.73 or more, 0.74 or more, or 0.75 or more.
[0089] In any of the assays described herein, an index value may be calculated based on the methylation status of the one or more CpGs in the test panel which can be used to assess an individual as having breast cancer. Any suitable mathematical model may be used to determine the index value, such as an algorithm or formula. Preferably, the cancer index value is termed Women’s risk Identification for Breast Cancer Index using a buccal sample (WID-buccal-BC-index) and wherein the mathematical model which is applied to the methylation 0-value data set to generate the cancer index is calculated by an algorithm according to the following formula: wherein: Pnare methylation beta-values (between 0 and 1); t ... , w30 037are real valued coefficients, although i may preferably be ual to / ; c. |i and o are real valued parameters used to scale the index; and d. n refers to the number of CpGs in the set of test CpGs; preferably wherein the cancer is ovarian cancer.
[0090] In any of the assays described herein, the WID-buccal-BC-index algorithm applies real value coefficients inferred by initially training on a dataset (this dataset in the exemplary embodiments of the invention described in the Examples consisted of 133 breast cancer cases and 130 controls) to fit a ridge classifier using the R package glmnet with a mixing parameter value of alpha = 0 (ridge penalty) and binomial response type. Ten-fold cross-validation was used internally by the cv. glmnet function in order to determine the optimal value of the regularisation parameter lambda. The beta values from are used as inputs to the ridge classifier. The coefficients wlt... , wnare obtained from the fitted model. The following quantity was computed for each individual v in the training set: n ^ WiPi i = l
[0091] Any suitable real valued coefficients may be applied to the WID-buccal-BC-Index in any of the assays described herein.
[0092] The value of the parameters / z and cr are given by the mean and standard deviation of xvin the training dataset respectively.
[0093] Thus, any suitable / z and cr real valued parameters may be applied to the WID-BC- index in any of the assays described herein. Any suitable training data set may be applied to the assays described herein in order to infer real value parameters and coefficients that can subsequently be applied to the WID-buccal-BC-index formula according to the present invention.
[0094] Preferably, the step of determining the methylation status of the CpGs in the test panel may comprise an amplification-based technique. Amplification may be performed by any suitable method, such as polymerase chain reaction (PCR), polymerase spiral reaction (PSR), loop mediated isothermal amplification (LAMP), nucleic acid sequence based amplification (NASBA), self-sustained sequence replication (3SR), rolling circle amplification (RCA), strand displacement amplification (SDA), multiple displacement amplification (MDA), ligase chain reaction (LCR), helicase dependant amplification (HD A), ramification amplification method (RAM), recombinase polymerase amplification (RPA) etc. Preferably, amplification is performed by polymerase chain reaction (PCR). Optionally, wherein the PCR comprises quantification, further optionally wherein the PCR comprises real time quantitative PCR and / or digital PCR. In assays of the invention wherein the step of determining the methylation status comprising performing PCR, the PCR may comprise the use of any primer pair suitable for amplifying a polynucleotide, e.g. genomic, region which contains the any one or more CpGs in the test panel. The assay may comprise the use of one primer pair to amplify a polynucleotide comprising one CpG identified in SEQ ID NOs 1-30,037 in the test panel. The assay may comprise the use of one primer pair to amplify a polynucleotide region comprising two or more CpGs identified in SEQ ID NOs 1-30,037 in the test pane. The assay may comprise the use of one primer pair to amplify a polynucleotide region comprising two or more, or preferably all, CpGs contained within the a DMR defined according to SEQ ID NOs 30,038-30,182. The assay may comprise the use of one primer pair to amplify a DMR defined according to SEQ ID NOs 30,038-30,182. The protocol for designing suitable PCR primer pairs is well known to skilled person.
[0095] Prior to PCR, the population of DNA molecules in the sample may be subject to bisulfite treatment, and preferably the one or more CpGs in the test panel may be subject to bisulfite treatment. Bisulfite treatment of nucleic acid molecules is a well-known procedure to the skilled person. In particular the skilled person would be aware of how to treat nucleic acid molecules with bisulfite in order to ensure that all unmethylated cytosines in the nucleic acid molecule are converted to uracil.
[0096] Preferred methods for determining methylation status in the context of PCR may comprise amplification using forward and reverse primers together with a detection system. The detection system may comprise a probe for annealing to any one of the sequences comprising a CpG defined by SEQ ID NOs 1-30,037 according to the invention, or any one of the DMRs defined by SEQ ID NOs 30,038-30,182 of the invention. The probe may comprise a detection moiety, or the probe may comprise a portion of sequence capable of hybridising to a component of the system comprising a detection moiety. The system may for example be a MethyLight system or a QuARTS (Quantitative allelespecific real-time target and signal amplification) system. Detection systems, including MethyLight and QuARTS, for use as part of a PCR assay for determining methylation status are well-known in the art.
[0097] The PCR may further comprise a negative control. The negative control may be any suitable negative control known in the art. Preferably, the negative control is a PCR reaction absent any template nucleic acid sequence and optionally absent any PCR primer or probes. The negative control may be water. The assay may comprise assessing of the presence or absence of breast cancer in the individual based on the methylation status of the one or more CpGs in the test panel comprises determination of mean percent of fully methylated reference (PMR) values for each polynucleotide region targeted by a primer pair as described herein, or any DMR, based on the methylation status of the CpGs comprised therein and comprised in the test panel. The sum of the PMR values for each DMR (SPMR) may be used to assess the presence or absence of breast cancer in the individual.
[0098] Whilst the preferred exemplary PMR-based assessment of methylation status of the CpGs in the test panel, and the subsequent assessment of risk by SPMR and specified thresholds can be used in accordance with the invention to assess the presence or absence of breast cancer in an individual, the skilled person would appreciate that alternative methods for determining methylation status can be utilised as described herein and the presence or absence of breast cancer in the individual may be determined accordingly.
[0099] Any of the assays described herein may further comprise determining in a sample which has been taken from the individual the immune cell proportion. The immune cell proportion may in a sample may be performed by any suitable means. For example, the immune cell proportion may be inferred by using the epithelial, fibroblast and immune cell reference dataset in the computer program, EpiDISH.
[0100] In the assay of the invention, preferably at least 25% of the DNA molecules in the sample are derived from immune cells. Preferably at least 30%, at least 35%, at least 40%, at least 45%, or at least 50% of the DNA molecules in the sample are derived from immune cells.
[0101] In any of the assays described herein the individual may be a woman had not been diagnosed with breast cancer previously but had breast cancer when the sample was taken from her buccal area. The individual may be a woman that did not have breast cancer when the sample was taken from her buccal area.
[0102] The invention also provides a variety of assays, each comprising any 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more (or any range derivable therein) of a variety of steps and in no particular order, including methods of the following: measuring in a sample; analyzing a sample; assessing a sample; evaluating a sample; measuring nucleic acids in a sample; assessing nucleic acids in a sample; detecting nucleic acids in a sample; measuring methylation in nucleic acids in a sample; analyzing nucleic acids in a sample; assessing nucleic acids in a sample; measuring methylation at one or more CpG dinucleotides in a sample; detecting methylation at one or more CpG dinucleotides in a sample; assaying methylation at one or more CpG dinucleotides in a sample; assessing methylation at one or more CpG dinucleotides in a sample; measuring a methylation status in a sample; assaying a methylation status in a sample; detecting methylation status in a sample; determining methylation status in a sample; identifying methylation status in a sample; measuring one or more DNA methylation markers in a sample; assessing one or more DNA methylation markers in a sample; detecting one or more DNA methylation markers in a sample; measuring the presence of methylation at one or more markers in a sample; detecting the presence of methylation at one or more markers in a sample; assessing the presence of methylation at one or more markers in a sample; assaying the presence of one of more markers in a sample; measuring one or more DNA methylation markers in a sample but excluding the measuring of one or more other DNA methylation markers in the sample; assessing one or more DNA methylation markers in a sample but excluding the assessing of one or more other DNA methylation markers in the sample; analyzing one or more DNA methylation markers in a sample but excluding the analyzing of one or more other DNA methylation markers in the sample; detecting one or more DNA methylation markers in a sample but excluding the detecting of one or more other DNA methylation markers in the sample; measuring methylation status in nucleic acids from a sample from tissue from tissue of an individual suspected of, or at risk for, being cancerous; detecting methylation status in nucleic acids from a sample from tissue from an individual suspected of, or at risk for, being cancerous; analyzing methylation status in nucleic acids from a sample from tissue from an individual from the individual suspected of, or at risk for, being cancerous; assessing methylation status in nucleic acids from a sample from tissue from an individual suspected of, or at risk for, being cancerous; measuring methylation at one or more CpG dinucleotides in a sample but excluding the measuring of methylation at one or more CpG dinucleotides in the sample; assessing methylation at one or more CpG dinucleotides in a sample but excluding the assessing of methylation at one or more CpG dinucleotides in the sample; analyzing methylation at one or more CpG dinucleotides in a sample but excluding the analyzing of methylation at one or more CpG dinucleotides in the sample; detecting methylation at one or more CpG dinucleotides in a sample but excluding the detecting of methylation at one or more CpG dinucleotides in the sample; measuring methylation at one or more CpG dinucleotides in nucleic acids from a sample from tissue from an individual suspected of, or at risk for, being cancerous; detecting methylation at one or more CpG dinucleotides in nucleic acids from a sample from tissue from an individual suspected of, or at risk for, being cancerous; analyzing methylation at one or more CpG dinucleotides in nucleic acids from a sample from tissue from an individual suspected of, or at risk for, being cancerous; assessing methylation at one or more CpG dinucleotides in nucleic acids from a sample from tissue from an individual suspected of, or at risk for, being cancerous; treating an individual for cancer when the individual has been determined to have a methylation status at one or more methylation markers; treating an individual for cancer when the individual has been determined to have methylation at one or more CpG dinucleotides; wherein any of the aforementioned methods, or any other methods encompassed by the disclosure, may comprise any one or more of the following method steps: measuring methylation status, wherein the measuring identifies the methylation status of three or markers from nucleic acids in a sample; measuring methylation status, wherein the measuring identifies the presence of one or more markers from nucleic acids in a sample; measuring the presence of one or more methylation markers from a sample; providing DNA from a sample; providing nucleic acids from a sample; determining whether one or more methylation markers from nucleic acids from a sample are methylated; measuring whether one or more methylation markers from nucleic acids from a sample are methylated; performing a sequencing step on nucleic acids from a sample; determining a sequence of nucleic acids from a sample; performing bisulphite conversion on one or more markers; performing bisulphite conversion on one or more CpG dinucleotides; hybridizing DNA to an array comprising probes capable of determining methylated versus non-methylated markers; hybridizing DNA to an array comprising probes capable of determining methylated versus non-methylated CpG dinucleotides; hybridizing DNA to an array comprising probes capable of discriminating between methylated and non-methylated markers; hybridizing DNA to an array comprising probes capable of discriminating between methylated and non-methylated CpG dinucleotides; performing an amplification step on sequence from nucleic acids from a sample; performing an amplification step on sequence from nucleic acids using methylation-specific primers; performing amplification of sequence comprising one or more regions suspected of having methylation or in need of determination of a methylation status; performing PCR on sequence comprising one or more regions suspected of having methylation or in need of determination of a methylation status; performing a capturing step; performing a binding step; performing a purification step; performing a capturing step comprising binding of polynucleotides comprising one or more methylation markers to binding molecules specific to the one or more methylation markers and collecting complexes thereof; stratifying the grade of a cancer; determining the risk of cancer; determining the risk of recurrence of cancer; obtaining a sample from an individual; obtaining DNA from a sample from an individual; administering a treatment to an individual; providing DNA from a sample; determining whether one or more methylation markers from a panel of methylation markers comprises a specific sequence; and / or obtaining data that identifies whether each one of a group of methylation markers from a panel comprises a specific sequence.
[0103] Moreover, in some aspects of the invention, an individual who is administered a therapy or treatment has been subjected to any of the methods and steps described herein.
[0104] Assessment of methylation status of CpGs
[0105] Methylation of DNA is a recognised form of epigenetic modification which has the capability of altering the expression of genes and other elements such as microRNAs. In cancer development and progression, methylation may have the effect of e.g. silencing tumor suppressor genes and / or increasing the expression of oncogenes. Other forms of dysregulation may occur as a result of methylation. Methylation of DNA occurs at discrete loci which are predominately dinucleotides consisting of a CpG motif, but may also occur at CHH motifs (where H is A, C, or T). During methylation, a methyl group is added to the fifth carbon of cytosine bases to create methylcytosine.
[0106] Methylation can occur throughout the genome and is not limited to regions with respect to an expressed sequence such as a gene. Methylation typically, but not always, occurs in a promoter or other regulatory region of an expressed sequence such as enhancer elements. Most typically, the methylation status of CpGs is clustered in CpG islands, for example CpG islands present in the regulatory regions of genes, especially in their promoter regions.
[0107] Typically, an assessment of DNA methylation status involves analysing the presence or absence of methyl groups in DNA, for example methyl groups on the 5 position of one or more cytosine nucleotides. Preferably, the methylation status of one or more cytosine nucleotides present as a CpG dinucleotide (where C stands for Cytosine, G for Guanine and p for the phosphate group linking the two) is assessed.
[0108] A variety of techniques are available for the identification and assessment of CpG methylation status, as will be outlined briefly below. The assays described herein encompass any suitable technique for the determination of CpG methylation status.
[0109] Methyl groups are lost from a starting DNA molecule during conventional in vitro handling steps such as PCR. To avoid this, techniques for the detection of methyl groups commonly involve the preliminary treatment of DNA prior to subsequent processing, in a way that preserves the methylation status information of the original DNA molecule. Such preliminary techniques involve three main categories of processing, i.e. bisulphite modification, restriction enzyme digestion and affinity-based analysis. A bisulfite-free preliminary technique may comprise TET-assisted pyridine borane sequencing (TAPS) combined with ten-eleven translocation (TET) oxidation of 5mC and 5hmC to 5- carboxylcytosine (5caC), and subsequent pyridine borane reduction of 5caC to dihydrouracil (DHU) for the detection of DNA methylation. Products of these techniques can then be coupled with sequencing or array-based platforms for subsequent identification or qualitative assessment of CpG methylation status.
[0110] Techniques involving bisulphite modification of DNA have become the most common assays for detection and assessment of methylation status of CpG dinucleotides. Treatment of DNA with bisulphite, e.g. sodium bisulphite, converts cytosine bases to uracil bases, but has no effect on 5 -methylcytosines. Thus, the presence of a cytosine in bisulphite-treated DNA is indicative of the presence of a cytosine base which was previously methylated in the starting DNA molecule. Such cytosine bases can be detected by a variety of techniques. For example, primers specific for unmethylated versus methylated DNA can be generated and used for PCR-based identification of methylated CpG dinucleotides. DNA may be amplified, either before or after bisulphite conversion. A separation / capture step may be performed, e.g. using binding molecules such as complementary oligonucleotide sequences. Standard and next-generation DNA sequencing protocols can also be used. Accordingly, in any of the methods described and defined herein, steps of determining in the population of DNA molecules in the sample the methylation status of CpGs in a test panel may be performed by methods which comprise subjecting CpGs to bisulfite treatment, e.g. subjecting the population of DNA molecules in the sample to bisulfite treatment. Any such method may further comprise an amplification step, preferably a PCR amplification step, such as any PCR amplification step described and defined herein.
[0111] Quantitative allele-specific real-time target and signal amplification (QuARTS) may also be applied to the invention described herein in order to determine in the population of DNA molecules in the sample the methylation status of a test panel of one or more CpGs. QuARTS is known in the art. QuARTS combines a polymerase-based target amplification with an invasive cleavage -based signal amplification. The fluorescence signal is detected in a fashion similar to real-time PCR. Accordingly, in any of the methods described and defined herein, steps of determining in the population of DNA molecules in the sample the methylation status of CpGs in a test panel may be performed by methods which comprise subjecting CpGs to bisulfite treatment, e.g. subjecting the population of DNA molecules in the sample to bisulfite treatment. Any such method may further comprise a PCR amplification step, preferably wherein the PCR amplification step comprises the use of a forward primer, a reverse primer and a detection system, wherein the detection system may comprise a probe for annealing to any one of the DMRs or sequences assayed according to the invention, wherein the probe may comprise a portion of sequence capable of hybridising to a component of the system comprising a detection moiety.
[0112] In other approaches, methylation-sensitive enzymes can be employed which digest or cut only in the presence of methylated DNA. Such enzyme -based approaches optionally do not require bisulfite treatment of the DNA being assayed. Analysis of resulting fragments is commonly carried out using microarrays.
[0113] Affinity-based techniques exploit binding interactions to capture fragments of methylated DNA for the purposes of enrichment. Binding molecules such as anti-5- methylcytosine antibodies are commonly employed prior to subsequent processing steps such as PCR and sequencing.
[0114] Olkhov-Mitsel and Bapat (2012) provide a comprehensive review of techniques available for the identification and assessment of biomarkers involving methylcytosine.
[0115] For the purposes of assessing the methylation status of the CpG-based biomarkers characterised and described herein, any suitable assay can be employed.
[0116] Assays described herein may comprise determining methylation status of CpGs by bisulphite converting the DNA. Preferred assays involve bisulphite treatment of DNA, including amplification of the identified CpG loci for methylation specific PCR and / or sequencing and / or assessment of the methylation status of target loci using methylation- discriminatory microarrays.
[0117] Amplification of CpG loci can be achieved by a variety of approaches. Preferably, CpG loci are amplified using PCR. A variety of PCR-based approaches may be used. For example, methylation-specific primers may be hybridized to DNA containing the CpG sequence of interest. Such primers may be designed to anneal to a sequence derived from either a methylated or non-methylated CpG locus. Following annealing, a PCR reaction is performed and the presence of a subsequent PCR product indicates the presence of an annealed CpG of identifiable sequence. In such assays, DNA is bisulphite converted prior to amplification. Such techniques are commonly referred to as methylation specific PCR
[0118] In other techniques, PCR primers may anneal to the CpG sequence of interest independently of the methylation status, and further processing steps may be used to determine the status of the CpG. Assays are designed so that the CpG site(s) are located between primer annealing sites. This assay scheme is used in techniques such as bisulphite genomic sequencing, COBRA, Ms-SNuPE. In such assay, DNA can be bisulphite converted before or after amplification.
[0119] Small-scale PCR approaches may be used. Such approaches commonly involve mass partitioning of samples (e.g. digital PCR). These techniques offer robust accuracy and sensitivity in the context of a highly miniaturised system (pico-liter sized droplets), ideal for the subsequent handling of small quantities of DNA obtainable from the potentially small volume of cellular material present in biological samples, particularly urine samples. A variety of such small-scale PCR techniques are widely available. For example, microdroplet-based PCR instruments are available from a variety of suppliers, including RainDance Technologies, Inc. (Billerica, MA; http: / / raindancetech.com / ) and Bio-Rad, Inc. (htp: / / www.bio-rad.com / ). Digital PCR from Qiagen and Thermofisher may also be suitable. Microarray platforms may also be used to carry out small-scale PCR. Such platforms may include microfluidic network-based arrays e.g. available from Fluidigm Corp, (www.fluidigm.com).
[0120] Following amplification of CpG loci, amplified PCR products may be coupled to subsequent analytical platforms in order to determine the methylation status of the CpGs of interest. For example, the PCR products may be directly sequenced to determine the presence or absence of a methylcytosine at the target CpG or analysed by array-based techniques.
[0121] Any suitable sequencing techniques may be employed to determine the sequence of target DNA. In the assays of the present invention the use of high-throughput, so-called “second generation”, “third generation” and “next generation” techniques to sequence bisulphite-treated DNA can be used.
[0122] In second generation techniques, large numbers of DNA molecules are sequenced in parallel. Typically, tens of thousands of molecules are anchored to a given location at high density and sequences are determined in a process dependent upon DNA synthesis. Reactions generally consist of successive reagent delivery and washing steps, e.g. to allow the incorporation of reversible labelled terminator bases, and scanning steps to determine the order of base incorporation. Array-based systems of this type are available commercially e.g. from Illumina, Inc. (San Diego, CA; http: / / www.illumina.com / ).
[0123] Third generation techniques are typically defined by the absence of a requirement to halt the sequencing process between detection steps and can therefore be viewed as realtime systems. For example, the base-specific release of hydrogen ions, which occurs during the incorporation process, can be detected in the context of microwell systems (e.g. see the Ion Torrent system available from Life Technologies; http: / / www.lifetechnologies.com / ). Similarly, in pyrosequencing the base-specific release of pyrophosphate (PPi) is detected and analysed. In nanopore technologies, DNA molecules are passed through or positioned next to nanopores, and the identities of individual bases are determined following movement of the DNA molecule relative to the nanopore. Systems of this type are available commercially e.g. from Oxford Nanopore (https: / / www.nanoporetech.com / ). In an alternative assay, a DNA polymerase enzyme is confined in a “zero-mode waveguide” and the identity of incorporated bases are determined with florescence detection of gamma-labeled phosphonucleotides (see e.g. Pacific Biosciences; http: / / www.pacificbiosciences.com / ). In other assays sequencing steps may be omitted. For example, amplified PCR products may be applied directly to hybridization arrays based on the principle of the annealing of two complementary nucleic acid strands to form a double-stranded molecule. Hybridization arrays may be designed to include probes which are able to hybridize to amplification products of a CpG and allow discrimination between methylated and nonmethylated loci. For example, probes may be designed which are able to selectively hybridize to an CpG locus containing thymine, indicating the generation of uracil following bisulphite conversion of an unmethylated cytosine in the starting template DNA. Conversely, probes may be designed which are able to selectively hybridize to a CpG locus containing cytosine, indicating the absence of uracil conversion following bisulphite treatment. This corresponds with a methylated CpG locus in the starting template DNA.
[0124] Following the application of a suitable detection system to the array, computer- based analytical techniques can be used to determine the methylation status of a CpG. Detection systems may include, e.g. the addition of fluorescent molecules following a methylation status-specific probe extension reaction. Such techniques allow CpG status determination without the specific need for the sequencing of CpG amplification products. Such array-based discriminatory probes may be termed methylation-specific probes.
[0125] Any suitable methylation-discriminatory microarrays may be employed to assess the methylation status of the CpGs described herein. One particular methylation- discriminatory microarray system is provided by Illumina, Inc. (San Diego, CA; http: / / www.illumina.com / ). In particular, the Infinium MethylationEPIC BeadChip array and the Infinium HumanMethylation450 BeadChip array systems may be used to assess the methylation status of CpGs for diagnosing breast cancer as described herein. Such a system exploits the chemical modifications made to DNA following bisulphite treatment of the starting DNA molecule. Briefly, the array comprises beads to which are coupled oligonucleotide probes specific for DNA sequences corresponding to the unmethylated form of a CpG, as well as separate beads to which are coupled oligonucleotide probes specific for DNA sequences corresponding to the methylated form of an CpG. Candidate DNA molecules are applied to the array and selectively hybridize, under appropriate conditions, to the oligonucleotide probe corresponding to the relevant epigenetic form. Thus, a DNA molecule derived from a CpG which was methylated in the corresponding genomic DNA will selectively attach to the bead comprising the methylation-specific oligonucleotide probe, but will fail to attach to the bead comprising the non-methylation- specific oligonucleotide probe. Single-base extension of only the hybridized probes incorporates a labeled ddNTP, which is subsequently stained with a fluorescence reagent and imaged. The methylation status of the CpG is determined by calculating the ratio of the fluorescent signal derived from the methylated and unmethylated sites.
[0126] Infinium HumanMethylation450 and MethylationEPIC BeadChip array systems and custom-built arrays can be used to interrogate CpGs in the assays described herein. Alternative or customised arrays could, however, be employed to interrogate the cancerspecific CpG biomarkers defined herein, provided that they comprise means for interrogating all CpG for a given assay, as defined herein.
[0127] Techniques involving combinations of the above-described assays may also be used. For example, DNA containing CpG sequences of interest may be hybridized to microarrays and then subjected to DNA sequencing to determine the status of the CpG as described above.
[0128] In the assays described above, sequences corresponding to CpG loci may also be subjected to an enrichment process if desired. DNA containing CpG sequences of interest may be captured by binding molecules such as oligonucleotide probes complementary to the CpG target sequence of interest. Sequences corresponding to CpG loci may be captured before or after bisulphite conversion or before or after amplification. Probes may be designed to be complementary to bisulphite converted DNA. Captured DNA may then be subjected to further processing steps to determine the status of the CpG, such as DNA sequencing steps.
[0129] Capture / separation steps may be custom designed. Alternatively, a variety of such techniques are available commercially, e.g. the SureSelect target enrichment system available from Agilent Technologies (http: / / www.agilent.com / home). In this system biotinylated “bait” or “probe” sequences (e.g. RNA) complementary to the DNA containing CpG sequences of interest are hybridized to sample nucleic acids. Streptavidin- coated magnetic beads are then used to capture sequences of interest hybridized to bait sequences. Unbound fractions are discarded. Bait sequences are then removed (e.g. by digestion of RNA) thus providing an enriched pool of CpG target sequences separated from non-CpG sequences. Template DNA may be subjected to bisulphite conversion and target loci amplified by small-scale PCR such as microdroplet PCR using primers which are independent of the methylation status of the CpG. Following amplification, samples may be subjected to a capture step to enrich for PCR products containing the target CpG, e.g. captured and purified using magnetic beads, as described above. Following capture, a standard PCR reaction is carried out to incorporate DNA sequencing barcodes into CpG- containing amplicons. PCR products are again purified and then subjected to DNA sequencing and analysis to determine the presence or absence of a methyl cytosine at the target genomic CpG.
[0130] CpG methylation status may be measured indirectly using a detection system such as fluorescence. A methylation-discriminatory microarray may be used. When calculating the degree of methylation of a given CpG, the Illumina® definition of beta- values may be used. The Illumina® methylation beta- value of a specific CpG site is calculated from the intensity of the methylated (M) and unmethylated (U) alleles, as the ratio of fluorescent signals 0=Max(M,O) / [Max(M,O)+Max(U,O)+lOO]. On this scale, 0<J3< 1 , 0 values of 1 or close to 1 indicate 100% methylation whereas values of 0 or close to 0 indicate 0% methylation.
[0131] As explained in more detail in the Examples below, one particular exemplary technique which the inventors have used is a methylation discriminatory array, such as an Illumina InfiniumM ethylation EPIC BeadChip. These assays utilise probes directed to methylated and unmethylated CpGs at a given locus.
[0132] Another exemplary technique which the inventors have used to determine the methylation status of any one or more CpGs is a fluorescence-based PCR technique referred to as MethyLight. These assays utilise forward and reverse PCR primers specific for sequences encompassing a DMR such as those described herein. The methylation status of one or more CpGs contained within said DMRs may therefore be determined by MethyLight analysis. The assays also utilise detectable probes for specific regions within the one or more CpGs that are to be assayed. The detectable probes are typically designed such that they hybridise only to methylated forms of the one or more CpGs to be assayed.
[0133] Other techniques which may be used to determine the methylation status of any one or more CpGs include the use of restriction enzymes. Such methods may employ methylation-sensitive enzymes which only digest unmethylated DNA and / or methylationdependent enzymes which only cut methylated DNA. Such enzymes may be used to enrich for methylated or unmethylated sequences and provide a read-out of DNA methylation. Accordingly, in any of the methods described and defined herein, steps of determining in the population of DNA molecules in the sample the methylation status of CpGs in a test panel may be performed by methods which comprise subjecting the population of DNA molecules in the sample to endonuclease treatment, such as restriction endonuclease treatment. Any such method may further comprise an amplification step, preferably a PCR amplification step.
[0134] Bioinformatic tools and statistical metrics for CpG-based assays
[0135] Software programs which aid in the in silico analysis of bisulphite converted DNA sequences and in primer design for the purposes of methylation-specific analyses are generally available and have been described previously.
[0136] In risk models for predicting breast cancer, a receiver-operating-characteristic (ROC) curve analysis is often used, in which the area under the curve (AUC) is assessed. Each point on the ROC curve shows the effect of a rule for turning a risk / likelihood estimate into a prediction of the presence or absence of cancer in an individual. The AUC measures how well the model discriminates between case subjects and control subjects. An ROC curve that corresponds to a random classification of case subjects and control subjects is a straight line with an AUC of 50%. An ROC curve that corresponds to perfect classification has an AUC of 100%.
[0137] In any of the methods described herein, the 95% confidence interval for the ROC AUC may be between 0.55 and 1.
[0138] In any of the methods described herein, the interval may be defined as a range having as an upper limit any number between 0.55 and 1. The upper limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77, 0.78, 0.79, 0.80, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89, 0.90, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.97, 0.98, 0.99 or 1.00.
[0139] In any of the methods described herein, the interval may be defined as a range having as a lower limit any number between 0.55 and 1. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77, 0.78, 0.79, 0.80, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89, 0.90, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.97, 0.98, 0.99 or 1.00.
[0140] In any of the methods described herein, the interval range may comprise any of the above lower limit numbers combined with any of the above upper limit numbers as appropriate.
[0141] Preferably, the upper limit number is 1. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 1 and as a lower limit any number between 0.55 and 1. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77, 0.78, 0.79, 0.80, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89, 0.90, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.97, 0.98, 0.99 or 1.00.
[0142] The upper limit number may be 0.99. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.99 and as a lower limit any number between 0.55 and 0.99. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77, 0.78, 0.79, 0.80, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89, 0.90, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.97, 0.98 or 0.99.
[0143] The upper limit number may be 0.98. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.98 and as a lower limit any number between 0.60 and 0.98. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77, 0.78, 0.79, 0.80, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89, 0.90, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.97 or 0.98.
[0144] The upper limit number may be 0.97. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.97 and as a lower limit any number between 0.60 and 0.97. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77, 0.78, 0.79, 0.80, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89, 0.90, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96 or 0.97.
[0145] The upper limit number may be 0.96. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.96 and as a lower limit any number between 0.60 and 0.96. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77, 0.78, 0.79, 0.80, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89, 0.90, 0.91, 0.92, 0.93, 0.94, 0.95 or 0.96.
[0146] The upper limit number may be 0.95. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.95 and as a lower limit any number between 0.60 and 0.95. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77, 0.78, 0.79, 0.80, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89, 0.90, 0.91, 0.92, 0.93, 0.94 or 0.95.
[0147] The upper limit number may be 0.94. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.94 and as a lower limit any number between 0.60 and 0.94. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77, 0.78, 0.79, 0.80, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89, 0.90, 0.91, 0.92, 0.93 or 0.94.
[0148] The upper limit number may be 0.93. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.93 and as a lower limit any number between 0.60 and 0.93. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77, 0.78, 0.79, 0.80, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89, 0.90, 0.91, 0.92 or 0.93.
[0149] The upper limit number may be 0.92. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.92 and as a lower limit any number between 0.60 and 0.92. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77, 0.78, 0.79, 0.80, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89, 0.90, 0.91 or 0.92.
[0150] The upper limit number may be 0.91. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.91 and as a lower limit any number between 0.60 and 0.91. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77, 0.78, 0.79, 0.80, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89, 0.90 or 0.91.
[0151] The upper limit number may be 0.90. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.90 and as a lower limit any number between 0.60 and 0.90. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77, 0.78, 0.79, 0.80, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88, 0.89 or 0.90.
[0152] The upper limit number may be 0.89. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.89 and as a lower limit any number between 0.60 and 0.89. The lower limit number may be 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77, 0.78, 0.79, 0.80, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87, 0.88 or 0.89.
[0153] The upper limit number may be 0.88. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.88 and as a lower limit any number between 0.60 and 0.88. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77, 0.78, 0.79, 0.80, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86, 0.87 or 0.88.
[0154] The upper limit number may be 0.87. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.87 and as a lower limit any number between 0.60 and 0.87. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77, 0.78, 0.79, 0.80, 0.81, 0.82, 0.83, 0.84, 0.85, 0.86 or 0.87.
[0155] The upper limit number may be 0.86. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.86 and as a lower limit any number between 0.60 and 0.86. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77, 0.78, 0.79, 0.80, 0.81, 0.82, 0.83, 0.84, 0.85 or 0.86.
[0156] The upper limit number may be 0.85. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.85 and as a lower limit any number between 0.60 and 0.85. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77, 0.78, 0.79, 0.80, 0.81, 0.82, 0.83, 0.84 or 0.85.
[0157] The upper limit number may be 0.84. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.84 and as a lower limit any number between 0.60 and 0.84. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77, 0.78, 0.79, 0.80, 0.81, 0.82, 0.83 or 0.84.
[0158] The upper limit number may be 0.83. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.83 and as a lower limit any number between 0.60 and 0.83. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77, 0.78, 0.79, 0.80, 0.81, 0.82 or 0.83.
[0159] The upper limit number may be 0.82. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.82 and as a lower limit any number between 0.60 and 0.82. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77, 0.78, 0.79, 0.80, 0.81 or 0.82. The upper limit number may be 0.81. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.81 and as a lower limit any number between 0.60 and 0.81. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77, 0.78, 0.79, 0.80 or 0.81.
[0160] The upper limit number may be 0.80. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.80 and as a lower limit any number between 0.60 and 0.80. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77, 0.78, 0.79 or 0.80.
[0161] The upper limit number may be 0.79. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.79 and as a lower limit any number between 0.60 and 0.79. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77, 0.78 or 0.79.
[0162] The upper limit number may be 0.78. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.78 and as a lower limit any number between 0.60 and 0.78. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76, 0.77 or 0.78.
[0163] The upper limit number may be 0.77. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.77 and as a lower limit any number between 0.60 and 0.77. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75, 0.76 or 0.77.
[0164] The upper limit number may be 0.76. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.76 and as a lower limit any number between 0.60 and 0.76. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, 0.74, 0.75 or 0.76.
[0165] The upper limit number may be 0.75. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.75 and as a lower limit any number between 0.60 and 0.75. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, 0.74 or 0.75.
[0166] The upper limit number may be 0.74. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.74 and as a lower limit any number between 0.60 and 0.74. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73 or 0.74.
[0167] The upper limit number may be 0.73. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.73 and as a lower limit any number between 0.60 and 0.73. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72 or 0.73.
[0168] The upper limit number may be 0.72. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.72 and as a lower limit any number between 0.60 and 0.72. The lower limit number may be 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71 or 0.72.
[0169] The upper limit number may be 0.71. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.71 and as a lower limit any number between 0.60 and 0.71. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69, 0.70 or 0.71.
[0170] The upper limit number may be 0.70. Thus, the 95% confidence ROC AUC interval may be defined as a range having an upper limit of 0.70 and as a lower limit any number between 0.60 and 0.70. The lower limit number may be 0.55, 0.56, 0.57, 0.58, 0.59, 0.60, 0.61, 0.62, 0.63, 0.64, 0.65, 0.66, 0.67, 0.68, 0.69 or 0.70.
[0171] In vitro DNA methylation methods
[0172] The invention provides an in vitro method of assaying DNA and detecting methylation of the DNA therein, the method comprising measuring a methylation status of one or more CpGs, wherein selected from the 30,037 CpGs identified at nucleotide positions 15 to 16 in SEQ ID NOs 1 - 30,037, preferably wherein the DNA is comprised in a sample as described herein, preferably which has been taken from the buccal area of an individual.
[0173] The measuring of methylation status may be performed according to the determining step according to the assay of the invention.
[0174] The test panel of the in vitro method may be defined according to the test panel of the assay of the invention. The in vitro method may comprise PCR that is defined in accordance with the PCR of the assay of the invention.
[0175] Primer and probes may be used in the in vitro method of the invention which are described herein as aspects of the assay of the invention.
[0176] PCR in the in vitro method may comprise the use of a negative control defined according to the assay of the invention.
[0177] Arrays and kits of parts
[0178] The invention also encompasses arrays capable of discriminating between methylated and non-methylated forms of CpGs as defined herein; the arrays may comprise oligonucleotide probes specific for methylated forms of CpGs as defined herein and oligonucleotide probes specific for non-methylated forms of CpGs as defined herein.
[0179] The invention also provides an array capable of discriminating between methylated and non-methylated forms of CpGs; the array comprising oligonucleotide probes specific for a methylated form of each CpG in a CpG panel and oligonucleotide probes specific for a non-methylated form of each CpG in the panel; wherein the panel consists of at least 500 CpGs selected from the CpGs identified at nucleotide position 15 to 16 in SEQ ID NOs 1 to 30,037.
[0180] Preferably the panel comprises the CpGs identified in SEQ ID NOs 1-53, more preferably the CpGs identified in SEQ ID NOs 1-222, and even more preferably comprises or consists of the CpGs identified in SEQ ID NOs 1-585.
[0181] The panel may consist of all CpGs identified in SEQ ID NOs 1 to 30,037.
[0182] In some embodiments the array is not an Infinium MethylationEPIC BeadChip array or an Illumina Infinium HumanMethylation450 BeadChip array.
[0183] Separately or additionally, in some embodiments the number of CpG-specific oligonucleotide probes of the array is 482,000 or less, 480,000 or less, 450,000 or less, 440,000 or less, 430,000 or less, 420,000 or less, 410,000 or less, or 400,000 or less, 375,000 or less, 350,000 or less, 325,000 or less, 300,000 or less, 275,000 or less, 250,000 or less, 225,000 or less, 200,000 or less, 175,000 or less, 150,000 or less, 125,000 or less, 100,000 or less, 75,000 or less, 50,000 or less, 45,000 or less, 40,000 or less, 35,000 or less, 30,000 or less, 25,000 or less, 20,000 or less, 15,000 or less, 10,000 or less, 5,000 or less, 4,000 or less, 3,000 or less or 2,000 or less.
[0184] The CpG panel may comprise any set of CpGs defined in the assays of the invention described herein. The arrays of the invention may comprise one or more oligonucleotides comprising any set of CpGs defined in the assays of the invention, wherein the one or more oligonucleotides are hybridized to corresponding oligonucleotide probes of the array.
[0185] The invention also encompasses a process for making a hybridized array described herein, comprising contacting an array according to the present invention with a group of oligonucleotides comprising any set of CpGs defined in the assays of the invention.
[0186] Any of the arrays as defined herein may be comprised in a kit. The kit may comprise any array as defined herein together with instructions for use.
[0187] The invention further encompasses the use of any of the arrays as defined herein in any of the assays for determining the methylation status of CpGs for the purposes of predicting the presence or absence of breast cancer in an individual.
[0188] Method of treatment and diagnosis
[0189] The term “treatment” as used herein is intended to refer to any intervention or procedure performed on an individual, including a surgical intervention, or a pharmacological intervention such as the administration of a compound or drug, or screening intervention.
[0190] The invention also encompasses the performance of one or more treatment steps following a positive diagnosis of breast cancer using the assays described herein. Said treatments may be considered “therapeutic” treatments.
[0191] The invention thus encompasses a method of treating a breast cancer patient comprising administering chemotherapy, radiation, immunotherapy or any cancer therapy described herein to the patient assessed to be positive for breast cancer based on any of the assays described herein.
[0192] The invention therefore provides a method of treating breast cancer in an individual, the method comprising: i. assessing the presence or absence of breast cancer in an individual according to the invention; and ii. administering one or more therapeutic treatments or measures to the individual based on the assessment.
[0193] The treatments administered to the individual may comprise any treatments considered suitable by a person skilled in the art. For example, the treatments administered to the individual may comprise any one of: intensified screening, preferably wherein the intensified screening comprises one or more of: o a test for a BRCA1 and / or BRCA2 germline mutation; o a breast MRI scan, preferably wherein the scan is repeated about a year after the previous scan; o a mammogram, preferably wherein the mammogram is repeated about a year after the previous mammogram; o a test for CAI 25, preferably wherein the test is repeated three- monthly; o a test for cell-free tumour DNA methylation in plasma / serum, preferably wherein the test is repeated annually; o a test for cell-free tumour DNA methylation in vaginal fluid, preferably wherein the test is repeated annually; o a repeat assay according to the invention described herein, preferably wherein the repeat assay is performed about one year after the previous assay; administration of one or more of Denosumab, “selective estrogen receptor modulators” (SERMs), and “selective progesterone receptor modulators” (SPRMs); a bilateral mastectomy; and / or a bilateral salpingo-oophorectomy.
[0194] In any of the assays described herein, wherein tests are performed as part of intensified screening the tests may be repeated at any suitable interval. For example, a test for CA125 may be performed three-monthly, six-monthly, annually or about once every two, three or four years. A test for cell-free tumour DNA methylation in plasma / serum may be performed three-monthly, six -monthly, annually or about once every two, three or four years. A test for cell-free tumour DNS methylation in vaginal fluid may be performed three-monthly, six-monthly, annually or about once every two, three or four years. A mammogram may be performed three-monthly, six-monthly, annually or about once every two, three or four years.
[0195] Other exemplary treatments comprise one or more surgical procedures, one or more chemotherapeutic agents, one or more cytotoxic chemotherapeutic agents one or more radiotherapeutic agents, one or more immunotherapeutic agents, one or more biological therapeutics, one or more anti-hormonal treatments or any combination of the above following a positive diagnosis of cancer.
[0196] Cancer treatments may be administered to an individual harbouring cancer, in an amount sufficient to treat, cure, alleviate or partially arrest cancer or one or more of its symptoms. Such treatments may result in a decrease in severity, and / or decreased cancer index value, of cancer symptoms, or an increase in frequency or duration of symptom-free periods. A treatment amount adequate to accomplish this is defined as "therapeutically effective amount". Effective amounts for a given purpose will depend on the severity of cancer and / or the individual’s cancer index value as well as the weight and general state of the individual. As used herein, the term "individual" includes any human, preferably the human is a woman. As used herein, “treatment’ is to be considered synonymous with “therapeutic agent”.
[0197] The following therapeutic agents may be administered to an individual based on their cancer risk alone or in combination with any other treatment described herein. The therapeutic agent may be directly attached, for example by chemical conjugation, to an antibody. Methods of conjugating agents or labels to an antibody are known in the art. For example, carbodiimide conjugation (Bauminger & Wilchek (1980) Methods Enzymol. 70, 151-159) may be used to conjugate a variety of agents, including doxorubicin, to antibodies or peptides. The water-soluble carbodiimide, l-ethyl-3-(3- dimethylaminopropyl) carbodiimide (EDC) is particularly useful for conjugating a functional moiety to a binding moiety. Other methods for conjugating a moiety to antibodies can also be used. For example, sodium periodate oxidation followed by reductive alkylation of appropriate reactants can be used, as can glutaraldehyde crosslinking. However, it is recognised that, regardless of which method of producing a conjugate of the invention is selected, a determination must be made that the antibody maintains its targeting ability and that the functional moiety maintains its relevant function.
[0198] A cytotoxic moiety may be directly and / or indirectly cytotoxic. By “directly cytotoxic” it is meant that the moiety is one which on its own is cytotoxic. By “indirectly cytotoxic” it is meant that the moiety is one which, although is not itself cytotoxic, can induce cytotoxicity, for example by its action on a further molecule or by further action on it. The cytotoxic moiety may be cytotoxic only when intracellular and is preferably not cytotoxic when extracellular.
[0199] Cytotoxic chemotherapeutic agents are well known in the art. Cytotoxic chemotherapeutic agents, such as anticancer agents, include: alkylating agents including oz nitrogen mustards such as mechlorethamine (HN2), cyclophosphamide, ifosfamide, melphalan (L-sarcolysin) and chlorambucil; ethylenimines and methylmelamines such as hexamethylmelamine, thiotepa; alkyl sulphonates such as busulfan; nitrosoureas such as carmustine (BCNU), lomustine (CCNU), semustine (methyl-CCNU) and streptozocin (streptozotocin); and triazenes such as decarbazine (DTIC; dimethyltriazenoimidazole- carboxamide); Antimetabolites including folic acid analogues such as methotrexate (amethopterin); pyrimidine analogues such as fluorouracil (5 -fluorouracil; 5-FU), floxuridine (fluorodeoxyuridine; FUdR) and cytarabine (cytosine arabinoside); and purine analogues and related inhibitors such as mercaptopurine (6-mercaptopurine; 6-MP), thioguanine (6-thioguanine; TG) and pentostatin (2’-deoxycoformycin). Natural Products including vinca alkaloids such as vinblastine (VLB) and vincristine; epipodophyllotoxins such as etoposide and teniposide; antibiotics such as dactinomycin (actinomycin D), daunorubicin (daunomycin; rubidomycin), doxorubicin, bleomycin, plicamycin (mithramycin) and mitomycin (mitomycin C); enzymes such as L-asparaginase; and biological response modifiers such as interferon alphenomes. Miscellaneous agents including platinum coordination complexes such as cisplatin (cis-DDP) and carboplatin; anthracenedione such as mitoxantrone and anthracycline; substituted urea such as hydroxyurea; methyl hydrazine derivative such as procarbazine (N -methylhydrazine, MIH); and adrenocortical suppressant such as mitotane (o,p’-DDD) and aminoglutethimide; taxol and analogues / derivatives; and hormone agonists / antagonists such as flutamide and tamoxifen.
[0200] A cytotoxic chemotherapeutic agent may be a cytotoxic peptide or polypeptide moiety which leads to cell death. Cytotoxic peptide and polypeptide moieties are well known in the art and include, for example, ricin, abrin, Pseudomonas exotoxin, tissue factor and the like. Methods for linking them to targeting moieties such as antibodies are also known in the art. Other ribosome inactivating proteins are described as cytotoxic agents in WO 96 / 06641. Pseudomonas exotoxin may also be used as the cytotoxic polypeptide. Certain cytokines, such as TNFa and IL-2, may also be useful as cytotoxic agents.
[0201] Certain radioactive atoms may also be cytotoxic if delivered in sufficient doses. Radiotherapeutic agents may comprise a radioactive atom which, in use, delivers a sufficient quantity of radioactivity to the target site so as to be cytotoxic. Suitable radioactive atoms include phosphorus-32, iodine- 125, iodine-131, indium-i l l, rhenium- 186, rhenium-188 or yttrium-90, or any other isotope which emits enough energy to C destroy neighbouring cells, organelles or nucleic acid. Preferably, the isotopes and density of radioactive atoms in the agents of the invention are such that a dose of more than 4000 cGy (preferably at least 6000, 8000 or 10000 cGy) is delivered to the target site and, preferably, to the cells at the target site and their organelles, particularly the nucleus.
[0202] The radioactive atom may be attached to an antibody, antigen-binding fragment, variant, fusion or derivative thereof in known ways. For example, EDTA or another chelating agent may be attached to the binding moiety and used to attach 11 Un or 90Y. Tyrosine residues may be directly labelled with 1251 or 1311.
[0203] A cytotoxic chemotherapeutic agent may be a suitable indirectly-cytotoxic polypeptide. In a particularly preferred embodiment, the indirectly cytotoxic polypeptide is a polypeptide which has enzymatic activity and can convert a non-toxic and / or relatively non-toxic prodrug into a cytotoxic drug. With antibodies, this type of system is often referred to as ADEPT (Antibody-Directed Enzyme Prodrug Therapy). The system requires that the antibody locates the enzymatic portion to the desired site in the body of the patient and after allowing time for the enzyme to localise at the site, administering a prodrug which is a substrate for the enzyme, the end product of the catalysis being a cytotoxic compound. The object of the approach is to maximise the concentration of drug at the desired site and to minimise the concentration of drug in normal tissues. In a preferred embodiment, the cytotoxic moiety is capable of converting a non-cytotoxic prodrug into a cytotoxic drug.
[0204] In any of the methods of treatment described herein, the one or more treatments that the individual is subjected to may be repeated on one or more occasions. The one or more treatments may be repeated at regular intervals. The repetitive nature of the treatment administration may depend on the particular treatment being administered. Some treatments may require repetitive administration at greater frequency than others. The skilled person would be aware of the frequency of administration required for therapies known in the art. The one or more treatments may be repeated weekly, two weekly, three weekly, four weekly, monthly, three monthly, six monthly, yearly, two yearly, three yearly, four yearly, or five yearly.
[0205] Methods of monitoring
[0206] The invention also provides methods of monitoring the breast cancer status of an individual. The invention is therefore for monitoring the presence of breast cancer in the individual. “Monitoring’ in the context of the present invention may refer to longitudinal assessment of an individual’s breast cancer status. This longitudinal assessment may be carried out according to any of the assays of the invention described herein. This longitudinal assessment may involve performance of any of the assays of the invention described herein to predict the presence breast cancer in an individual at more than one time point over the course of an undetermined time window. The time window may be any period of time whilst the individual is still living. The time window may persist for the lifetime of the individual. The time window may persist until the individual’s breast cancer status, risk of having breast cancer falls below a certain level. The level may be determined by the methylation status of the one or more CpGs in the test panel. The level may be determined by any of the methods for determining methylation described herein. The level may preferably be determined by a mean percent of fully methylated reference (PMR) values for each DMR based on the methylation status of the CpGs from each DMR comprised in the test panel, preferably wherein the sum of the PMR values for each DMR (SPMR) may be used to assess the presence or absence of breast cancer in the individual and therefore the level of risk.
[0207] The invention therefore provides a method of monitoring the breast cancer status in an individual, the method comprising: (a) assessing the presence or absence of breast cancer in an individual by performing the assay according to the invention at a first time point; (b) assessing the presence or absence of breast cancer in the individual by performing the assay according to the invention at one or more further time points; and (c) monitoring any change in the breast cancer status of the individual.
[0208] In any of the methods of monitoring described herein, the steps of assessing the presence or absence of breast cancer in an individual based on a threshold breast cancer risk, for example as described in the method of treatment stratification provided herein. Threshold values can provide an indication of an individual’s breast cancer status or risk of having breast cancer. For example, methylation status of the test panel of CpGs may indicate the presence or absence of breast cancer, or a high or low risk of harbouring breast cancer.
[0209] The invention further encompasses a method of measuring methylation in a patient at multiple time points comprising (a) assessing the presence or absence of breast cancer in an individual by performing any one of the assays of the invention described herein at a first time point; (b) assessing the presence or absence of breast cancer in the individual by performing any one of the assays of the invention described herein at one or more further time points, and (c) detecting differential methylation status between (a) and (b), monitoring any change in the breast cancer status of the individual.
[0210] In any of the methods of monitoring described herein, depending on the risk of the presence of cancer in the individual, one or more treatments are administered to the individual according to any one of the methods of treatment encompassed by the invention and described herein. Different treatments may be administered depending on the stratification of an individual on the basis of their breast cancer status or risk of having breast cancer. The method may further comprise administration of one or more treatments according to the methods of treatment provided herein.
[0211] Methylation status of the CpGs of the test panel, and therefore breast cancer risk, may change between any two or more time points. For this reason, longitudinal monitoring of an individual’s CpG methylation status and cancer status could be of particular benefit to the assessment of, for example, cancer treatment efficacy.
[0212] In any of the methods of monitoring described herein, the one or more further time points may be any suitable time point. Preferably the one or more further time points may of suitable distance apart for sufficiently frequent screening in order to predict any particularly early onset cases of presence of breast cancer in an individual. Preferably the one or more further time points may be of suitable distance apart for assessing the efficacy of one or more treatments. Preferably the one or more further time points may be of suitable distance apart for predicting whether an individual remains free of cancer after a successful course of treatment. The one or more further time points may be about monthly, about two monthly, about three monthly, about four monthly, about five monthly, about six monthly, about seven monthly, about eight monthly, about nine monthly, about ten monthly, about eleven monthly, about yearly, about two yearly, or more than two yearly.
[0213] In any of the methods of monitoring described herein, changes may be made to the one or more treatments wherein a positive or negative responses to the one or more treatments are observed. Treatments may be changed in accordance with the methods of treatments described herein. Treatments may particularly be changed if the individual’s breast cancer status or risk stratification as described herein with respect to the methods of treatment provided herein.
[0214] In any of the methods of monitoring encompassed by the invention, the step of predicting the presence or absence of breast cancer in an individual may involve the use of any one of the arrays described herein. Types of cancers
[0215] The methods described herein may be applied to any breast cancer.
[0216] The breast cancer may be a primary cancer lesion. The breast cancer may be a secondary cancer lesion. The breast cancer may be a metastatic lesion.
[0217] In assays described herein, wherein the assays is for assessing the presence or absence of breast cancer, the breast cancer may be a ductal carcinoma in situ or an invasive ductal carcinoma such as tubular type invasive ductal carcinoma (IDC), medullary type IDC, mucinous type IDC, papillary type IDC or cribriform type IDC.
[0218] The breast cancer may be an invasive carcinoma such as a pleomorphic carcinoma, carcinoma with osteoclast giant cells, carcinoma with choriocarcinoma features, carcinoma with melanotic features. The invasive breast carcinoma may be an invasive lobular carcinoma, tubular carcinoma, invasive cribriform carcinoma, medullary carcinoma, mucinous carcinoma and other tumours with abundant mucin such as mucinous carcinoma, cystadenocarcinoma and columna cell mucinous carcinoma, signet ring cell carcinoma. The invasive breast carcinoma may be a neuroendocrine tumour such as solid neuroendocrine carcinoma (carcinoid of the breast), atypical acarcinoid tumour, small cell / oat cell carcinoma, large cell neuroendocrine carcinoma. The invasive breast carcinoma may be an invasive papillary carcinoma, invasive micropapillary carcinoma, apocrine carcinoma, metaplastic carcinomas such as pure epithelial metaplastic carcinomas including squamous cell carcinoma, adenocarcinoma with spindle cell metaplasia, adenosquamous carcinoma, mucoepidermoid carcinoma, mixed epithelial / mesenchymal metaplastic carcinomas, matrix -producing carcinoma, spindle cell carcinoma, carcinosarcoma, squamous cell carcinoma of mammary origin, metaplastic carcinoma with osteoclastic giant cells. The invasive breast carcinoma may be a lipid-rich carcinoma, secretory carcinoma, oncocytic carcinoma, adenoid cystic carcinoma, acinic cell carcinoma, glycogen-rich clear cell carcinoma, sebaceous carcinoma, inflammatory carcinoma, bilateral breast carcinoma.
[0219] The breast cancer may be a mesenchymal breast tumour. The mesenchymal tumour may include sarcoma. The mesenchymal breast tumour may be a hemangioma, angiomatosis, Hemangiopericytoma, Pseudoangiomatous stromal hyperplasia, Myofibroblastoma, Fibromatosis (aggressive), Inflammatory myofibroblastic tumor, Lipoma Angiolipoma, Granular cell tumour, Neurofibroma, Schwannoma, Angiosarcoma, Liposarcoma, Rhabdomyosarcoma, Osteosarcoma, Leiomyoma, Leiomyosarcoma.
[0220] The breast cancer may be a malignant lymphoma such as Non-Hodgkin lymphoma.
[0221] The breast cancer may be a metastatic tumour in which the primary lesion originated in a tissue other than the breast.
[0222] The breast cancer may be a precursor breast cancer lesion. The precursor breast cancer lesion may be a Lobular neoplasia, lobular carcinoma in situ, Intraductal proliferative lesions, Usual ductal hyperplasia, Flat epithelial hyperplasia, Atypical ductal hyperplasia, Ductal carcinoma in situ, Microinvasive carcinoma, Intraductal papillary neoplasms, Central papilloma, Peripheral papilloma, Atypical papilloma, Intraductal papillary carcinoma, Intracystic papillary carcinoma.
[0223] The breast cancer may be a myoepithelial breast cancer lesion. The myoepithelial breast cancer lesion be myoepitheliosis, adenomyoepithelial adenosis, adenomyoepithelioma, malignant myoepithelioma.
[0224] The breast cancer may be a fibroepithelial breast tumour. The fibroepithelial breast tumour may be a fibroadenoma, phyllodes tumour, periductal stromal sarcoma, mammary hamartoma.
[0225] The breast cancer may be Paget’s disease of the nipple.
[0226] Biological samples
[0227] In any of the assays described herein, in the sample which has been taken from the buccal area of the individual, the sample comprises a population of DNA molecules. The individual may be asymptomatic or symptomatic of a breast cancer. Symptoms of breast cancer will be well known to a person of skill in the art
[0228] The sample may have been self-collected or collected by a medical professional, such as a clinician. The sample may have been obtained using any suitable means, such as a buccal swab. Any suitable device for sample collection from the buccal area may be used. Preferably the sample for example has been obtained by contacting the buccal area with a brush device which collects epithelial cells from the buccal area, e.g. using a Copan Medical Diagnostics buccal swab.
[0229] Alternative to the providing step of the assays of the invention comprising the use of sample which has been taken from the buccal area of the individual, the providing step of the assays of the invention may comprise obtaining a sample from the buccal area of the individual, the sample comprising a population of DNA molecules. Preferably the individual is a woman.
[0230] In any of the assays described herein, the assay may or may not encompass the step of obtaining the sample from the individual. In assays which do not encompass the step of obtaining the sample from the individual, a sample which has previously been obtained from the individual is provided.
[0231] The sample may be provided directly from the individual for analysis or may be derived from stored material, e.g. frozen, preserved, fixed or cryopreserved material.
[0232] In any of the assays described herein, the sample may be self-collected or collected by any suitable medical professional.
[0233] Any of the assays described herein, the sample may comprise cells. The sample may comprise genetic material such as DNA and / or RNA.
[0234] Any of the assays described herein may involve providing a biological sample from the patient as the source of patient DNA for methylation analysis.
[0235] Any of the assays described herein may involve obtaining patient DNA from a biological sample which has previously been obtained from the patient.
[0236] Any of the assays described herein may involve obtaining a biological sample from the patient as the source of patient DNA for methylation analysis. The sample may be selfcollected or collected by any suitable medical professional. Procedures for obtaining a biological sample include biopsy.
[0237] Methods for sample isolation and for the subsequent extraction and isolation of DNA from such cell or tissue samples in preparation for assessing DNA methylation, are well known to those skilled in the art. In the context of the assays or methods described herein, the entirety of a sample may be used, or alternatively cells may be concentrated or cell types may be fractionated in order to only apply subsets of one or more cell types to the present assays or methods. Any suitable methods of concentration or fractionation may be used.
[0238] In any one of the assays described and defined herein the sample from the individual, or the sample which has been taken from the individual, may derive from a tissue which is different from the tissue which harbours the tumour, if a tumour is present in the individual. Accordingly, in any one of the assays described and defined herein the sample from the individual, or the sample which has been taken from the individual, may not comprise nucleic acid, including DNA, which derives from the tumour, i.e. tumourspecific nucleic acid, including tumour-specific DNA. Thus, consistent with the data set out in the examples and the disclosures herein, methylation profiles derived from DNA molecules in the sample are used as surrogate markers for tumour-specific nucleic acid, including tumour-specific DNA, which exists at an anatomical site in the body of the individual which is remote from the anatomical site from which the sample is derived. A skilled person would able to identify the absence of tumour-specific DNA in any given population of sample-specific DNA molecules by routine means, such as determining the absence of known genetic mutations which are characteristic of the particular cancer, by e.g. sequence based screening. However, the performance of any one of the assays described and defined herein does not require any assessment to verify the absence of tumour-specific DNA in any given population of DNA molecules.
[0239] EXAMPLES
[0240] The present invention will now be described with reference to specific Examples, which should not be construed as in any way limiting.
[0241] EXAMPLE 1
[0242] ABSTRACT
[0243] DNA methylation (DNAme) profiles in cervical samples have previously been found to detect breast cancer (BC) but no study has so far systematically compared the suitability of various non-invasive surrogate samples for this purpose. Here, we compare non-tumour DNAme in cervical, buccal, and blood samples for BC detection, performing an epigenome-wide association studies using a total of 1,100 samples from BC cases and controls, comparing similarities and differences in epigenetic changes, and deriving DNAm-based classifiers for BC detection in each sample type (WID-buccal-, cervical-, or blood-BC). Our findings reveal cervical samples exhibit the largest number of differentially methylated sites, followed by buccal samples, whereas no sites remained significant for blood after false discovery rate adjustment. The WID-buccal-, -cervical-, and -blood-BC indices distinguished BC cases from controls with AUCs of 0.75 [95% CI: 0.68-0.82], 0.66 [0.58-0.74], and 0.51 [0.38-0.54] in external validation sets, respectively. Our findings moreover indicate that although epithelial sample DNAme-classifiers (i.e. WID-buccal- and cervical -BC) are able to distinguish between BC cases and control both in the surrogate samples and breast tissue itself (ROC AUC > 0.88), individual sites and directionality of methylation changes are not identical, and buccal sample methylation generally aligned with breast methylation changes more closely Once temporal and mechanistic relationships between epigenetic changes in various surrogate and at-risk tissues have been further explored, these insights may have the potential to change clinical practice and enable non-invasive personalized primary and secondary cancer prevention.
[0244] INTRODUCTION
[0245] Breast Cancer (BC) is the most common and second most fatal cancer affecting women, emphasizing the demand for effective detection methodsx. Current screening or detection approaches predominantly revolve around imaging techniques, primarily mammography. While mammography screening has been proven to reduce breast cancer mortality2it can lead to overdiagnosis and not insignificant false positive rates. Therefore, the combination of existing screening and early detection approaches with molecular biomarkers has been named a key priority in a recent consensus statement for breast cancer prevention .
[0246] In addition to polygenic risk scores4that aim to stratify the population based on heritable risk, various other tools are being proposed for screening and early detection. The majority of these utilize blood-based measurements of circulating cell-free tumor nucleic acids (DNA or RNA) or their modification (e.g., DNA methylation5’6), proteins, and carcinoma antigens (CAs)7. The emphasis on blood for cancer biomarkers is likely based on multiple factors, including but not limited to practicality (blood samples are routinely collected and are also often stored in biobanks, facilitating biomarker development) and the fact that tumors shed material into the bloodstream that can be detected using ‘liquid biopsies’. However, the use of other biospecimens may also offer specific advantages and provide further insights into systemic effects causing, or caused by, cancer. For instance, in previous work, we identified that DNA methylation (DNAme) in cervical samples that are routinely collected for cervical cancer screening could be leveraged to assess the risk of being diagnosed with BC. The cervical methylation classifier, called Women’s cancer risk identifier - Breast Cancer (WID-BC) achieved an area under the curve (AUC) of 0.81 in an external validation set derived from cervical samples of women with breast cancer or healthy age-matched controls8. DNAme is a relatively stable epigenetic modification that plays a crucial role in regulating gene and protein expression without altering the DNA sequence, and can be modified by external exposures. Thus, the epigenome is hypothesized to capture environmental changes, serving as an important link between genes and environment. In the absence of cancer tissue present in anatomically distant cervical samples, our prior work indicated that the observed DNAme changes may be indicative of systemic lifetime exposure that may drive cancer in one tissue (breast), but can be read out in a non-invasively collected ‘surrogate’ sample (cervix), rendering it suitable for detection and screening.
[0247] No study has systematically compared different surrogate tissues in their potential to detect epigenetic field defects associated with BC, although this investigation could yield further insights into whether in addition to blood and cervical samples, buccal samples could for instance be used to detect BC. Buccal samples may represent an underappreciated but promising surrogate sample for DNAme-based cancer detection due to their relative ease of collection and high suitability for self-sampling that has already been employed for other purposes, such as diagnosing celiac disease in children9and detecting COVID-19 viral proteins10. Aligned with the fact that both buccal and breast epithelial cells originate from the ectoderm, previous work indicated that buccal samples shared higher variability at breast-variable sites than peripheral blood cells, suggesting that buccal sample DNAme might be a better proxy indicative of breast field cancerization11. The level of physical activity in women has previously been associated with BC risk12, and previous research has provided preliminary evidence of a relationship between exercise behavior, cardiovascular fitness, and methylation on CpG sites linked to breast cancer in buccal cells. These effects seemed specific to genes associated with breast cancer, rather than affecting the general level of methylation of CpG sites across the methylome13. This early evidence indicates that the epigenome of buccal cells might at least in part mirror the risk of being diagnosed with breast cancer.
[0248] To further investigate potential epigenetic alterations associated with breast cancer in three easy to access surrogate tissues, here we conduct a comparative analysis of the suitability of cervical, buccal, and blood samples (Fig. la). We initially perform an epigenome-wide association study (EWAS) for breast cancer in cervical, buccal, and blood samples from women recently diagnosed with BC and cancer-free age-matched controls. We next investigate whether differentially methylated positions (or regions) we identified in cervical, buccal, and blood samples are shared or whether the different surrogate samples have distinct features associated with breast cancer. We moreover derive and validate an epigenetic classifier for BC detection in each tissue, comparing their diagnostic accuracy in external validation sets and breast tissue itself to assess their potential to reflect epigenetic changes in a different tissue, and evaluate their association with genetic risk
[0249] METHODS
[0250] Study design and data acquisition
[0251] A study outline is shown in Fig. la, and sample numbers and population characteristics of our discovery and validation sets reported in Fig. 5 and 6. Our study population consisted of a case-control study including women with primary BC with at least one poor prognostic feature (defined as >2 cm cancers and / or lymph-node positive and / or hormone-receptor negative and / or grade 3), and age-matched healthy controls (Fig. 5). We utilized cervical, buccal, and blood sample DNAme data from BC cases and controls collected as part of the multicentre ‘FORECEE’ study in a case-control setting, involving 15 recruitment sites across Europe as previously described8, and publicly available datasets (GSE237036). In the FORECEE study, women with symptoms indicative of BC, subsequently confirmed through diagnosis (referred to as “cases”), were approached during outpatient hospital clinics. Healthy volunteers from the general population (referred to as “controls”) were engaged through outreach campaigns, public engagement initiatives, and participation in cervical screening programs. During the FORECEE sample collection, two sets of samples (Discovery set and External validation set) were collected. Samples within the External Validation Set were exclusively reserved for the validation of the index, i.e. there was no overlap between individuals in the discovery / training and those in the validation phase. However, within the discovery / training and validation sets, matched buccal and cervical samples were available from the same individuals. This allowed us to assess whether, when the two sample types are directly compared from the same individuals, classifiers derived from buccal or cervical samples exhibit better performance. Blood samples were not from matched individuals.
[0252] All cervical, buccal, and blood samples were collected at the point of diagnosis and prior to the initiation of any treatment.
[0253] A separate sample collection as part of the FORECEE study involved recruitment of BRCA1 and BRCA2 mutation carriers, as well as from healthy age -matched controls (BRCA Mutation Carrier Set: Fig. 7), and matched samples for cervical, buccal, and blood samples were available and have been previously describedn. For blood external validation, we utilized a publicly available dataset of peripheral blood-derived mononuclear cell (PBMC) methylation from BC cases and controls (GSE237036), described previously14. A cross-tissue validation was conducted using breast methylation data.
[0254] Buccal, cervical, and blood sample collection
[0255] Buccal cells were collected using two Copan 4N6FLOQ Buccal Swabs (Copan Medical Diagnostics, cat #4504C) by firmly brushing the swab head 5-6 times against the buccal mucosa of each cheek. The swabs were re-capped and left to dry out at room temperature within the sampling tube which contains a drying desiccant. Cervical and blood sample collections have been described previously8,n. Briefly, 2.5 ml of venous whole blood was collected in PAX gene blood DNA tubes (BD Biosciences #761165) and stored locally at 4°C. Cervical sample collection was conducted by trained staff using the ThinPrep system (Hologic Inc, cat #70098-002). A description of blood collected as part of GSE237036 (Blood external validation dataset) has been published previously14.
[0256] Breast DNA methylation datasets for cross-tissue validation
[0257] The current study uses three breast tissue DNA methylation datasets for crosstissue validation: the ‘Tissue at risk set’ (EGAS00001005055, n=56, previously described in Barrett et al.3, IlluminaMethylationEPIC), consisting of cancer tissue, matched normal adjacent tissue, normal breast tissue from cancer-free women without a BRCA1 / 2 mutation and normal breast tissue from women with a BRCA1 / 2 mutation; data from The Cancer Genome Atlas (TCGA) (n=889, IlluminaMethylation450k), consisting of matched cancer tissue and normal adjacent tissue; and dataset GSE23243215(n=595, IlluminaMethylationEPIC), consisting of cancer tissue and matched normal adjacent samples and data from health controls (reduction mammoplasty) from the NCI-Maryland Breast Cancer Cohort. . Details of breast methylation datasets are described in Fig. 8.
[0258] DNA methylation array profiling
[0259] Cervical and blood sample preprocessing have been described in previous publications8,1 x. Buccal sample DNAme was specifically generated for the purpose of this study and followed previous procedures. Briefly, buccal sample DNA was normalized to 25 ng / pL and 500 ng total DNA was bisulfite modified using the EZ-96 DNA Methylation- Lightning kit (Zymo Research Corp, cat #D5047) on the Hamilton Star Liquid handling platform. Eight pL of modified DNA was subjected to methylation analysis on the Illumina InfiniumMethylation EPIC BeadChip (Illumina, CA, USA) at UCL Genomics according to the manufacturer’s standard protocol. Buccal discovery and external validation sets were processed on different occasions (Discovery set: January 2019; External validation set: March 2021). For external data, all data from GEO (GSE237036, GSE225845) were obtained via the GEOquery package whereas TCGA data were accessed using TCGAbiolinks
[0260] DNA methylation data preprocessing and analysis
[0261] With the exception of data from TCGA, all methylation microarray data was processed through the same standardized pipeline that has previously been described8,1 x, packaged as eutopsQC (https: / / github.com / chiaraherzog / eutopsQC). In brief, raw .idat fdes are loaded using the R package minfi (current version 1.43.1), samples with median methylated or unmethylated intensity below 9.5 or more than 10% failed probes (detection p-value > 0.01) are removed, and single sample background intensity- and dye bias correction is carried out using the ssNOOB function (minfi). Beta Mixture Quantile Normalisation (BMIQ) is applied to correct for probe type bias (ChAMP, version 2.30.0) and beta values from failed probes are imputed using the impute.knn function function (impute, version 1.72.3). Furthermore, non-CpG probes (n=2932), SNP-related probes as identified by Zhou et al.16(n=82,108), and any probes that map to the Y-chromosome are removed from the dataset. As a result, approximately 10% of the probes on the EPIC array are removed as part of the Quality Control (QC) process by default. Beta distributions for each sample are visualised and inspected to identify any samples that exhibit obvious abnormalities. Additionally, Principal Component Analysis (PCA) plots are examined to investigate potential associations with technical or biological factors which emphasized that the major drivers of variability were sample type and composition, and indicated no obvious batch effects (Fig. 9). Cell type proportions from processed beta matrices were inferred using the epithelial, fibroblast and immune cell reference dataset in EpiDISH version 2.14.1 (ref.m = centEpiFibIC.m, method = ‘RPC’). Immune cell subtypes were inferred using the hierarchical EpiDISH function (ref2.m = centBloodSub.m, maxit = 500, h.CT.idx = 3, method = ‘RPC’).
[0262] Epigenome-wide association study (EWAS)
[0263] Previous studies have highlighted the importance of accounting for cellular heterogeneity in epigenome-wide studies, and the merit of inferring methylation differences in ‘pure’ immune or epithelial cells of a given sample8’17. We conducted our epigenomewide association study with breast cancer adjusting for age and the major variable cell type in each tissue (immune cell proportion (ic) in cervical and buccal samples; neutrophil proportion in blood samples) by including them as covariates in linear models. Additionally, we inferred differential methylation, i.e. delta-beta, in pure epithelial cell samples (ic=0) or pure immune cell samples (ic= 1 ) by fitting linear regression models separately for cases and controls, and computing the differences of intercepts at ic=0 and ic=l, as previously described8’17. To ensure compatibility with future DNAme studies, we restricted our further analysis to probes shared between IlluminaMethylationEPIC version 1 (vl) and the more recent version 2 (v2). Differentially methylated positions (DMPs) were defined as significant sites after false discovery rate (FDR) correction using the Benjamini-Hochberg correction method, with adjusted p value < 0.05.
[0264] Differentially methylated regions (DMR) were identified using the DMRcate method18(R package, version 2.14.1), accounting for age and cell heterogeneity [model. matrix(~type + age + ic), where ic is the immune cell proportion in buccal or cervical samples and neutrophil proportion in blood samples]. Briefly, we set the parameter types to use the EPIC array annotation, perform “differential” methylation analysis, and kept default parameters for lambda. No significant DMRs were found for blood samples, even after increasing the false FDR threshold to 0.2.
[0265] Comparison of directionality in buccal and cervical samples and overrepresentation analysis
[0266] Differentially methylated CpGs associated with breast cancer at p<0.05 in buccal and cervical samples were visualized according to delta-beta in the respective sample, either overall or inferred epithelial- or immune-specific differential methylation. Hypo- or hypermethylated CpGs were defined as a delta-beta of below or above 0, respectively. Subsequently, four quadrants were defined: hypomethylated in both, hypermethylated in both, hypomethylated in buccal and hypermethylated in cervical, or hypermethylated in buccal and hypomethylated in cervical. Overrepresentation analysis of shared or opposing hyper- and hypomethylated CpGs was conducted via cross-tabulation of expected and observed overlaps, with the assumption that at random distribution each quadrant would contain 25% of the total shared significant sites. The odds ratio was computed as follows: [p / ( 1 -p)] / [po / ( 1 -po)] , where p is number of CpGs in quadrant divided by the total number of shared significant CpGs, and po is 0.25. eFORGE and gene ontology analysis.
[0267] For interpretation of groups of CpGs via regulatory elements (DNase I hypersensitive sites) we used the online implementation of eFORGE 2.019(https: / / eforge.altiusinstitute.org / ). Where more than 1,000 sites were present in one quadrant, the top 1,000 sites were selected and analyzed using eFORGE. Gene ontology analysis was conducted only on sites significant after multiple testing correction using the Benjamini-Hochberg method (except for blood where no sites remained after multiple testing correction, and all sites with p<0.05 were used). Genes within the DMPs were extracted and analyzed using the clusterProfiler R package20, with the background list representing all genes represented on the Illumina Methylation EPIC array. Gene sets with a q value of < 0.05 were considered significant.
[0268] Classifier training
[0269] To train classifiers for breast cancer, we used the R package glmnet (version 4.1.8) with parameter values alpha = 0 (ridge penalty), alpha = 0.5 (elastic net), and alpha = 1 (lasso penalty) with binomial response type, as previously described8. Two-thirds of the discovery set (Fig. 5) were used as training data (buccal: n=267; ncaSes= 134, nControis = 133. Cervical: n=2639Ucases 133 •> Ilcontrols 130. Blood: n=208; ncases= 72, nCOntrois = 136). For each tissue type, four ranked lists of CpGs were generated and selected as inputs for the classifier. The ranking of CpGs was determined by delta-beta estimates in epithelial cells (1) and immune cells (2), as well as their combination, each time taking the CpG with the largest epithelial delta-beta, followed by the CpG with the largest immune delta-beta, followed by the next largest epithelial delta-beta and so forth (any duplicates were removed) (3). Finally, a ranked list based on p-values, regardless of cell type, was used (4). Ten-fold cross-validation was used inside the training set by the cv. glmnet function in order to determine the optimal value of the regularization parameter lambda. Classifier performance was evaluated on an internal validation dataset using the AUC metric versus n, the number of ranked candidate CpGs used as inputs during training. The optimal classifier was selected based on the highest AUCs obtained in the internal validation set (buccal: n=1359Ucases 66, nControis=69; cervical: n=1359Ucases 66, nCOntrois=69; blood: n=102; Ucases 30, nControis=72); part of the discovery set, Fig. 5), out-of-bag estimates of the AUC as well as calibration slope and intercept obtained from the rms R package (https: / / hbiostat.Org / r / rms / ). Training AUC values are shown in Fig. 10. Subsequently, the training and internal validation datasets were combined and the classifier was retrained using the entire discovery dataset with alpha and lambda fixed to their optimal values, using the optimal number of ranked input CpGs. In the final index, the top n CpGs are represented as fil, and the regression coefficients from the trained classifier as wl, ...,wn. The WID-BC-index in each tissue is calculated as £ni=l(wi0i-|f) / o, where p and o refer to the mean and standard deviation of ^rii=l Wiff in the respective training dataset, ensuring that the index is normalized to have a zero mean and unit standard deviation within the training dataset. To ensure comparability across different age ranges and cellular compositions, indices are adjusted after fitting linear models to age (and immune cell proportion in buccal and cervical samples, and the neutrophil proportion in blood samples) in control samples (i.e., adjusted indices are residuals).
[0270] Statistical analyses
[0271] Statistical analysis was carried out in R version 4.2.3. Areas under the curve of the receiver operating characteristic (AUROCs) and corresponding 95% confidence intervals were calculated using the pROC R package 1.18.4. Code to reproduce the analysis and figures will be provided under https : / / www. github . com / eutops / systems ,BC .
[0272] Ethics statement
[0273] The multicentre FORECEE study, during which buccal, blood and cervical samples were collected, received ethical approval from UK Health Research Authority (REC 14 / LO / l 633) and all contributing centres, including the NRES Committee London (UK), Ethics Committee of the General University Hospital, Prague (Czech Republic), Comitato Etico degli IRCSS Institute Europeo di Oncologia e Centro Cardiologico Monzino (Itality), Regionale Komiteer for Medisinsk og Helsefaglig Forskningsetikk (Norway), and Ethikkommission bei der LMU Miinchen (Germany). All participants were aged >18 years and provided written informed consent. Each prospective study volunteer was given a Participant Information Sheet, as well as a Consent Form and the rationale for the study was explained.
[0274] RESULTS DNAme changes associated with breast cancer in buccal, cervical, and blood samples
[0275] We initially performed an epigenome -wide analysis of DNAme changes associated with breast cancer in the three surrogate tissues (buccal, cervical, and blood; Fig. la), looking at both individually differentially methylated positions (DMPs) and differentially methylated regions (DMRs), accounting for age and cellular heterogeneity (immune or neutrophil proportion in buccal / cervical and blood samples, respectively). A full overview of sample composition in each dataset is shown in Fig. 11, and principal component plots of the top 20% variable CpGs and their association with immune cell proportion is shown in Fig. 9d, j, g, revealing a strong dependence of the most variable methylation sites on cellular composition. Overall, sample type and composition were the strongest drivers of variability in the first principal components, as expected (Fig. 9a-c). The epigenome -wide study within each surrogate sample type revealed 79,030, 143,986 and 52,147 DMPs significantly associated with breast cancer in buccal, cervical, and blood samples (p<0.05) of which 585, 21,614, and 0 remained significant after Benjamini -Hochberg correction, respectively (Fig. 12). An overview of the genetic location and region of these sites is shown in Fig. 13 and revealed an approximately similar distribution of genomic contexts for differentially methylated sites in buccal and cervical samples, with the majority of CpGs located in Open Sea regions.
[0276] For buccal and cervical samples, for whom matched samples were available, we assessed overlaps of differentially methylated positions that were shared by both tissues at p < 0.05 and their directionality (Fig. le). We found that the directionality was not always shared. Specifically, we observed an underenrichment of shared hypermethylated sites (OR = 0.48 [95% CI: 0.43-0.54]). We also observed that a larger than expected proportion of shared significant sites was hypomethylated in cervical cells but hypermethylated in buccal cells (OR = 2.55 [2.50-2.59]). To improve our understanding of the potential functionalities of the shared sites in each quadrant of Fig. le (common or opposing directionalities), we applied the eFORGE tool which aids in interpretation of groups of CpGs via regulatory elements in specific tissues, specifically DNase I hypersensitive sites (DHS) (Breeze et al. 2016) (Fig. lb, f, h). We also queried the mean methylation of the CpGs in each quadrant in normal breast tissue, normal breast tissue adjacent to a cancer, or breast tumor tissue (Fig. 1c, d, h, i). Specifically, we found that CpGs hypomethylated in buccal cells of breast cancer cases but hypermethylated in cervical sites were enriched for DHS in mammary epithelial cells (Fig. lb) and exhibited hypomethylation in tumor tissue compared to both normal or normal-adjacent tissue (Fig. 1c). No DHS enrichment was found for sites that shared hypermethylation in both cervical and buccal samples, although these sites exhibited significant hypermethylation in breast tumor tissue compared to normal breast (Fig. Id). Sites that were shared hypomethylated in both buccal and cervical samples were enriched for stem cell DHS (Fig. If) whereas sites with buccal hypermethylation and cervical hypomethylation were enriched for fetal tissue DHS (Fig. Ih). Interestingly, breast tissue methylation generally followed methylation changes in buccal cells, both when comparing normal tissue versus tumor but also normal tissue versus healthy tissue adjacent to a cancer (‘normal-adjacent’; Fig. 1c, d, g, i). As methylation differences could be driven by cellular heterogeneity, we also investigated inferred delta-beta in epithelial or immune cell fractions of buccal and cervical cells, respectively (Fig. 14a and b), which revealed that epithelial cell changes (Fig. 14a) followed the overall trend of Fig. le - revealing a lack of correlation - whereas immune cell changes correlated moderately across the two tissues (Fig. 14b; R=0.4, p < 2.2e'16for delta-beta immune versus R = -0.047, p = 7.1e'10for delta-beta epithelial).
[0277] We identified differentially methylated regions (DMRs), which consist of several CpGs in genomic proximity, using the DMRcate method21. This method fits DNAme measurements spatially across the genome, ranking the most differentially methylated regions based on tunable kernel smoothing of the differential methylation (DM) signal. Again, cervical samples exhibited more pronounced changes than buccal samples, with 1,992 and 145 DMRs significantly associated with breast cancer in cervical and buccal samples, respectively. The majority of alterations in cervical samples were associated with a loss of methylation (hypomethylation). We investigated the directionality of differential methylation across regions, and again assessed potential overlaps between buccal and cervical samples (Fig. 14c). Twenty-four DMRs were shared by cervical and buccal samples, e.g. spanning genes LTBP4, RP11-551L14.1 , NCK1, and CCDC88C. These included both sites with the same directionality (e.g., RP11-551L14.1, LTBP4) and opposing directionality of differential methylation (e.g., NCK1, CCDC88C) between BC cases and controls, with some examples shown in Fig. 14d.
[0278] Subsequently, to aid in interpretation of significantly enriched sites (CpGs) in each surrogate sample, we conducted gene ontology enrichment analysis on the loci that remained significant after Benjamini-Hochberg correction (hence only in buccal and cervical samples). The top 10 significant gene ontology enrichments of biological processes, molecular functions, and cellular functions in buccal and cervical samples are illustrated in Fig. 15 and 16, respectively. Notably, several enriched biological and molecular processes in buccal and cervical samples were associated with GTPase activity. GTPases are known to regulate cytoskeletal dynamics, which play a crucial role in oncogenic processes such as cellular migration and cell cycle progression22. Aberrant GTPase activity has been established across all subtypes of breast cancer23. In addition, buccal samples exhibited enrichment in Wnt signaling (Fig. 15). Dysregulated Wnt signaling has been linked to BC, and expression of Wnt genes has been linked to BC cancer aggressiveness24. Additionally, cervical samples also showed an enrichment for CpGs located in genes associated with developmental pathways, such as gland development (Fig. 16).
[0279] Development and validation of DNA methylation-based classifiers for breast cancer based on buccal, cervical, and blood samples
[0280] We next aimed to compare the performance of buccal, cervical, and blood samples for detecting breast cancer based on DNAme. We trained classifiers using ridge or lasso regression, applying several approaches for feature (pre)selection, including ranking the top 30,000 sites by p value in the original EWAS, by top epithelial or immune delta-beta values as previously described8, or inputting the entire beta methylation value matrix (the latter using lasso only). In the internal validation set, we evaluated model fit parameters, including internal validation and out-of-bag estimates of the area under the curve and the calibration intercept and slope. The final model was selected based on the number of input CpGs with the optimum number of slope (closest to 1), intercept (closest to 0), and highest out-of-bag AUC estimates; for all surrogate samples, the optimum fit was found to be for ridge regression, inputting the top 30,000 CpGs ranked by p value (lowest to highest) (Fig. 10). The final model for each surrogate sample was subsequently trained on the entire discovery set (training and internal validation), and index coefficients and values in the discovery set were saved for computation of the index and scaling in external datasets.
[0281] Following hyperparameter optimization, the final WID-indices for each tissue were validated on external validation sets (Fig. 2; population characteristics shown in Fig. 6). Notably, as for the discovery set, within the external validation set matched buccal and cervical samples were available from the same individuals. The WID-buccal-BC index exhibited a slight to moderate dependence on immune cell proportion overall (Fig. 2a) and performance was slightly higher in samples with higher than median immune cell proportion (Fig. 2d, Fig. 17a). There was no significant correlation with age (Fig. 2b) or menopausal status, although the performance appeared slightly higher in pre- compared to postmenopausal women (Fig. 17b, Fig. 18). The WID-buccal-BC index performed significantly better in women with an early age at menarche (defined by median age at menarche (13y) in the External Validation Set; AUC=0.81 for early age versus AUC=0.59 for high age at menarche, respectively; Fig. 17c; DeLong's test for two ROC curves: p = 0.01446). Overall, the WID-buccal-BC index was significantly higher in controls than cases (Fig. 2c, p = 2.179e-09) and after accounting for age and immune cell differences achieved an area under the curve of 0.75 (95% CI: 0.68-0.82) (Fig. 2d). Interestingly the WID-cervical-BC index also exhibited a slightly better performance in samples with higher than median immune cell proportion and post-menopausal women compared to pre -menopausal women (Fig. 2e, d, Fig. 17d, e), but in contrast to the WID-buccal-BC, it performed better in women with later age at menarche (>13) than women with earlier age at menarche (Fig. 17f; AUCs of 0.72 and 0.63, respectively). The WID-cervical-BC was significantly elevated in cases compared to controls (p = 1.92e-04) and achieved an area under the curve of 0.66 (95% CI: 0.58-0.74) (Fig. 2g, h), lower than the buccal index in the same group of women (Fig. 2d; DeLong's test for two ROC curves: p = 0.06). The blood validation set comprised DNAme data derived from peripheral blood mononuclear cells (PBMCs). This data was obtained from a previous report of BC cases and controls with limited phenotypic information (e.g., no information on age or menopausal status)14. The WID-blood-BC exhibited a strong overall dependence on the proportion of neutrophils, although discriminative potential was not influenced by neutrophil proportion (Fig. 17g), and overall exhibited limited discriminative performance AUC: 0.51, 95% CI: 0.38-0.64) (Fig. 2i-j) relative to WID-buccal-BC and WID-cervical- BC.
[0282] Association of the classifier with epidemiological and genetic characteristics
[0283] To evaluate factors associated with the WID-indices in each tissue, we assessed its relationship with various epidemiological, genetic, and sample characteristics (Figure 3, Fig. 18). As we previously identified a dependence of the index on immune cell proportion and to balance for differences in age, we evaluated residuals after fitting models for age and immune cell proportion in controls (‘adjusted indices’, see methods). The WID-buccal-BC index was not significantly different in T1 versus T2 or T3 stages and distinguished across cancer cases from controls regardless of grade, ER, PR, or HER2 status (Fig. 3a-e). HER2 positive (HER2+) cases appeared to exhibit slightly higher values than HER2 negative cancer cases, in line with HER2+ typically corresponding to more aggressive cancer subtypes, but this did not reach significance. Surprisingly, the WID-buccal-BC index did not show any association with a polygenic risk score (PRS313) for breast cancer4obtained from the same women (Fig. 3f). Likewise, women with increased risk for breast cancer due to BRCA1 / 2 mutations did not exhibit elevated WID-buccal-BC values (Fig. 3g, h). Similar findings were obtained for cervical samples (Fig. 3i-p), where the index was not strongly dependent on differences in cancer characteristics or genetic risk factors. The WID-buccal- BC, but not WID-cervical-BC, was associated with nodal status but this association was not linear with increasing nodal stage (Fig. 18f, 1), requiring further investigation. The WID- blood-BC index was slightly elevated in BRCA1 but not BRCA2 mutation carriers compared to controls, but this did not reach significance (Fig. 3q, r).
[0284] Application of buccal and cervical indices in blood samples
[0285] We were intrigued by the finding that buccal and cervical indices consistently tended to exhibit higher AUCs in samples with higher than median immune cell proportion, which indicated that information about BC status was contained in immune cells (Fig. 2d, h, respectively), yet the blood-specific index performed poorly (Fig. 2k). We applied the buccal and cervical indices to blood samples and found that AUCs were low but significantly higher than 0.5 in the FORECEE study, but not in the GSE237036 validation set (Fig. 19).
[0286] Validation in breast cancer tissue
[0287] Lastly, we were curious as to whether indices trained in the three surrogate tissues (buccal, cervical, blood) could distinguish between normal breast tissue from cancer-free women and normal breast tissue from women with breast cancer (i.e. normal adjacent breast tissue) and actual breast cancer tissue, and might perhaps be indicative of factors driving cancer progression in the ‘tissue at risk’ . We identified three relevant datasets using the EPIC or 450k array and evaluated values of the WID-buccaL, cervical-, and blood-BC indices. Across all three datasets, the WID-buccal-BC was consistently higher in breast tissue compared to normal control tissue (Fig. 4a-f) and exhibited AUCs of 0.94, 0.9, and 0.88 respectively. Conversely, the WID-cervical-BC index exhibited significantly lower values in cancer tissue compared to control tissue (i.e., an inverted behavior compared to performance in cervical samples, where cancer cases would exhibit higher values), resulting in AUCs of 0.13, 0.04, and 0.12. The behavior of the WID-blood-BC was not consistent across datasets, although breast cancer tissue tended to exhibit significantly lower values than control tissue (Fig. 4a-d), also exhibiting an inverted directionality compared to the surrogate tissue. Normal tissue adjacent to a breast cancer occasionally exhibited significantly higher (or lower) values compared to controls, but this was not consistent across all datasets (Fig. 4a, e). Lastly, normal breast tissue from BRCA1 / 2 mutation carriers exhibited a significantly lower WID-buccal-BC than control tissue, and significantly higher WID-cervical-BC, indicating it exhibited the opposing directionality compared to breast cancer (Fig. 4a), although sample numbers were small.
[0288] DISCUSSION
[0289] Studies on the identification and optimization of biomarkers for earlier cancer detection and stratification of risk of being diagnosed with cancer are a key area of research, and currently often focus on blood samples due to their routine implementation in clinical workflows and availability in many biobanks. Approaches utilizing cell-free DNA methylation for BC detection have so far shown limited success: for instance, in the recent PATHFINDER study, no primary BC was detected25. Our previous work indicated that samples other than blood could provide valuable information for detection of women’s cancers, importantly cervical samples that contain (hormone-sensitive) epithelial cells with the potential to mirror epigenetic changes in breast tissue8. Additionally, our previous work comparing blood, cervical, and buccal samples as three different types of ‘surrogate’ samples suggested that these tissues share varying levels of information with breast tissuen. Here, we systematically investigated DNAme changes in three surrogate sample types derived from breast cancer cases and controls, and developed and compared DNAm-based classifiers for breast cancer detection. Importantly, in contrast to other approaches, the current study did not investigate cell-free DNA derived from tumor material but intended to identify DNAme changes in surrogate samples with limited or no tumor material that mirror changes in the “at-risk” tissue, such as due to cumulative clinical / lifetime effects.
[0290] To allow for comparability of findings across different sample types, we utilized similar sample numbers across discovery / training sets and, where available, leveraged datasets that consisted of matched surrogate samples from the same individual (i.e., buccal and cervical samples from the same individual). Regrettably, the blood discovery dataset was slightly smaller and no matched data for blood samples from the same individuals for buccal and cervical samples were available. However, we were able to leverage a publicly available blood methylation dataset to perform validation of our index. Although fewer CpGs reached formal significance in an epigenome -wide discovery following Benjamini- Hochberg correction in buccal than cervical samples, our results indicated that given similar training sizes, buccal samples exhibited a higher area under the curve in external validation (0.75 [95% CI: 0.68-0.82]) than cervical samples (0.66 [95% CI: 0.58-0.74]) (Fig. 2d, h). No CpGs were significantly associated with breast cancer in DNAme from blood samples, and although during index training out-of-bag estimates of the AUC of up to 0.9 were achieved, the WID-blood-BC was not able to distinguish breast cancer cases from controls in external validation dataset (AUC: 0.46 [95% CI: 0.33-0.6]) (Fig. 2k). It is worth noting that the discovery set was slightly smaller for blood samples compared with buccal and cervical samples (n=310 for blood, n=402 for buccal / cervical samples), and the external validation set was derived from peripheral mononuclear blood cells (PBMCs) rather than whole blood, which may also have influenced discovery and performance, and it is possible that larger training sizes could lead to slightly improved performance. For instance, the discriminative power of the cervical BC classifier reported in the current work, trained on 402 samples, exhibits a much lower AUC than a previously reported index trained on 1,198 samples8(AUC: 0.66 versus 0.81, respectively), not surprisingly indicating that performance is linked to training size. It is worth noting that cervical samples typically show high variability in the proportion of immune cells (Fig. 11). The pronounced heterogeneity in cell type composition and other factors influencing methylation in cervical samples, in contrast to the less variable immune cell proportion and possibly more homogeneous buccal samples, could make training a classifier in cervical samples more challenging and may therefore necessitate an even larger training set (as previously described in8, n=l,198 samples for training) compared to that required for buccal samples. However, the relatively lower performance of blood compared to buccal and cervical samples is in line with our previous report that cervical and buccal samples share more variability with breast cancer tissue11and, for instance hormone-sensitive epithelial cells in cervical samples may be able to reflect cumulative exposure such as to progesterone8.
[0291] Interestingly, while we previously found that buccal and cervical samples with higher inferred epithelial cell proportions share more information (variability) with breast tissue compared to samples with lower epithelial cell proportions, our current and previous8analysis indicate that detection of BC may be slightly better in samples with higher than median immune cell content (i.e., lower epithelial contents; Fig. 2d, h). This stands in contrast with the finding that blood samples, consisting entirely of immune cells, have generally lower potential for detection (noting certain limitations of the current study design above; Fig. 2k). This could be for several reasons, including a discrepancy of peripheral (blood) and resident (buccal / cervical) immune cells in their relative DNAme information content with relation to breast cancer, or the association of immune content with other unobserved features that can help to detect BC. Future work should more thoroughly assess the potential of purified epithelial, immune, or mixed sample types to detect cancers to better understand the biology surrogate sample-based cancer detection and potential risk prediction.
[0292] Intriguingly, the buccal classifier was consistently elevated in breast tissue, whereas the WID-cervical-BC exhibited a significant reduction, i.e., an inverted directionality. This indicated that although both surrogate tissues reflect changes associated with cancer, they do so in different ways and by focusing on different features. The potential for opposing directionality in changes between the surrogate tissues was already hinted at during the initial assessment of differentially methylated regions: Fig. lb highlighted prominent hypomethylation in cervical, but more balanced methylation changes in buccal samples. Notably, among shared DMRs between buccal and cervical samples, these unexpectedly did not always share the same directionality in buccal and cervical samples (Fig. 14c). This was particularly intriguing as buccal and cervical samples for discovery originated from the same individuals (matched samples), and provided further evidence for tissue-dependent epigenomic alterations in cancer.
[0293] The surrogate-specific classifiers developed in the current manuscript were not elevated in samples from cancer-free women at increased genetic risk, such as due to a BRCA1 / 2 mutation, or associated with a previously described polygenic risk score for breast cancer (Fig. 3). This suggested that the current classifiers are not strongly associated with genetic features but instead potentially reflect either cumulative risk due to lifetime (including potential in-utero) risk factors, mirrored in the epigenome, or systemic changes elicited by the presence of a cancer. In absence of large-scale DNAme validation datasets derived from buccal or cervical samples collected at or preceding diagnosis, it is challenging to investigate the cause or temporal relationship of the DNAme alterations with BC diagnosis. Our current findings indicate that the WID-BC-buccal index is not elevated in individuals with genetic risk (BRCA1 / 2 mutation carriers), but at what time in relation to a BC diagnosis the indices would become elevated remains unclear and will need to be investigated in the future.
[0294] While the purpose of the current study was to compare surrogate tissues and not develop optimized classifiers in any given tissue, our findings moreover suggest that there is considerable potential for buccal DNAm-based detection of BC that should be investigated in future studies with larger and more diverse populations. Buccal sample collection is non- invasive, making it a promising tool for simpler detection or risk stratification for breast cancer with the potential for self-collection of samples. Although the sample size in the external validation set was limited and our dataset focused on BC cases with at least one poor prognostic feature, the lack of dependence on stage or grade observed in our analysis (Fig. 3a, b) indicates that such an approach may be suitable for early detection and / or screening using buccal samples. Future optimisation could also include addition of epidemiological variables or other molecular or imaging features into algorithms. A positive test result could have the consequence of earlier, intensified, and / or more sensitive (i.e., magnetic resonance imaging) breast cancer screening measures, although the clinical utility and impact on stage shifts and mortality will need to be further confirmed in clinical studies.
[0295] Our study has several limitations, including those inherent to case-control design, and a modest sample size for epigenome-wide discovery given practical constraints, outlined above. Nonetheless, we perform external validation of our buccal and cervical BC classifiers with AUCs of 0.75 and 0.66, respectively, and demonstrate that classifiers also validate in breast cancer tissue. Our study also has several notable strengths, including a comparative analysis of three non-invasive samples for breast cancer detection and crosstissue validation. To our knowledge, we describe for the first time that cancer-associated differential methylation derived from the same individuals is tissue-specific, identifying potentially opposing differential methylation in buccal and cervical samples of breast cancer cases compared to controls. Future studies of large-scale collections of buccal or cervical samples from diverse and representative populations will be required to further investigate and optimize non-invasive cancer detectors / risk predictors to serve as reliable biomarkers for breast cancer early detection. Pending further validation, WID-buccal- or - cervical-BC indices could be applied for regular screening in at-risk populations and risk monitoring may be combined with (pharmacological) preventive measures in those deemed high risk to either prevent cancers or capture cancers earlier at the point when they become detectable. Thresholds for ‘high risk’ populations will need to be set in large populationbased studies with several years follow-up to identify the proportion at highest risk of cancer with appropriate sensitivity and specificity.
[0296] REFERENCES
[0297] 1. Sun, Y.-S. et al. Risk Factors and Preventions of Breast Cancer. Int. J. Biol. Sci. 13, 1387-1397 (2017).
[0298] 2. Clift, A. K. et al. The current status of risk-stratified breast screening. Br. J. Cancer 126, 533-550 (2022).
[0299] 3. Pashayan, N. et al. Personalized early detection and prevention of breast cancer: ENVISION consensus statement. Nat. Rev. Clin. Oncol. 17, 687-705 (2020).
[0300] 4. Mavaddat, N. et al. Polygenic Risk Scores for Prediction of Breast Cancer and Breast Cancer Subtypes. Am. J. Hum. Genet. 104, 21-34 (2019). Moss, J. et al. Circulating breast-derived DNA allows universal detection and monitoring of localized breast cancer. Ann. Oncol. 31, 395-403 (2020). Cheng, N. et al. Pre-diagnosis plasma cell-free DNA methylome profiling up to seven years prior to clinical detection reveals early signatures of breast cancer.
[0301] 2023.01.30.23285027 Preprint at https: / / doi.org / 10.1101 / 2023.01.30.23285027 (2023). Li, J. et al. Non-Invasive Biomarkers for Early Detection of Breast Cancer. Cancers 12, 2767 (2020). Barrett, J. E. et al. The WID-BC-index identifies women with primary poor prognostic breast cancer based on DNA methylation in cervical samples. Nat. Commun. 13, 449 (2022). Adriaanse, M. P. M. et al. Human leukocyte antigen typing using buccal swabs as accurate and non-invasive substitute for venipuncture in children at risk for celiac disease. J. Gastroenterol. Hepatol. 31, 1711-1716 (2016). Kalil, M. N. A. et al. Performance Validation of COVID-19 Self-Conduct Buccal and Nasal Swabs RTK- Antigen Diagnostic Kit. Diagnostics 11, 2245 (2021). Herzog, C. et al. DNA methylation at quantitative trait loci (mQTLs) varies with cell type and nonheritable factors and may improve breast cancer risk assessment. Npj Precis. Oncol. 7, 1-10 (2023). Guo, W., Fensom, G. K., Reeves, G. K. & Key, T. J. Physical activity and breast cancer risk: results from the UK Biobank prospective cohort. Br. J. Cancer 122, 726- 732 (2020). Bryan, A. D., Magnan, R. E., Hooper, A. E. C., Harlaar, N. & Hutchison, K. E. Physical activity and differential methylation of breast cancer genes assayed from saliva: a preliminary investigation. Ann. Behav. Med. Publ. Soc. Behav. Med. 45, SO- 98 (2013). Wang, T. et al. A multiplex blood-based assay targeting DNA methylation in PBMCs enables early detection of breast cancer. Nat. Commun. 14, 4724 (2023). GEO Accession viewer. https: / / www.ncbi.nlrn.nih.gov / geo / query / acc.cgi7accA SE225845. Zhou, W., Laird, P. W. & Shen, H. Comprehensive characterization, annotation and innovative use of Infinium DNA methylation BeadChip probes. Nucleic Acids Res. 45, e22 (2017). Barrett, J. E. et al. The DNA methylome of cervical cells can predict the presence of ovarian cancer. Nat Commun 13, 448 (2022). Peters, T. J. et al. De novo identification of differentially methylated regions in the human genome. Epigenetics Chromatin 8, 6 (2015). Breeze, C. E. et al. eFORGE v2.0: updated analysis of cell type-specific signal in epigenomic data. Bioinformatics 35, 4767-4769 (2019). Wu, T. et al. clusterProfiler 4.0: A universal enrichment tool for interpreting omics data. The Innovation 2, 100141 (2021). Peters, T. J. et al. Calling differentially methylated regions from whole genome bisulphite sequencing with DMRcate. Nucleic Acids Res. 49, el09 (2021). Haga, R. B. & Ridley, A. J. Rho GTPases: Regulation and roles in cancer cell biology. Small GTPases 7, 207-221 (2016). Kazmi, N. et al. Rho GTPase gene expression and breast cancer risk: a Mendelian randomization analysis. Sci. Rep. 12, 1463 (2022). Castagnoli, L., Tagliabue, E. & Pupa, S. M. Inhibition of the Wnt Signalling Pathway: An Avenue to Control Breast Cancer Aggressiveness. Int. J. Mol. Sci. 21, 9069 (2020). Schrag, D. et al. Blood-based tests for multicancer early detection (PATHFINDER): a prospective cohort study. The Lancet 402, 1251-1260 (2023). EXAMPLE 2
[0302] Introduction and objective
[0303] To evaluate the performance of DNA methylation data at CpG sites denoted by SEQ ID NOs 1-30,037, we constructed Receiver Operating Characteristic (ROC) curves and calculated the Area under the Curve (AUC) for various combinations. We moreover identified example thresholds by setting the specificity to 75%, although alternative threshold options may be feasible, such as Youden’s Index.
[0304] Data
[0305] Methylation data were based on the Buccal Validation Set described in Example 1 above. Briefly, this dataset consisted of buccal samples from 91 cancer- free controls and 94 breast cancer cases. Data were processed using a standardised preprocessing pipeline described above, and available under https: / / github.io / chiaraherzog / eutopsQC / .
[0306] Classifier construction
[0307] We exemplify thresholds based on three potential scenarios:
[0308] • the WID-buccal-BC, whose derivation is described in Example 1. Briefly, leveraging a training dataset, a linear classifier was derived using penalised regression and the combination with the optimum AUCs and calibration characteristics was chosen (ridge regression at n=30,000, excluding 37 CpGs (namely defined by SEQ ID NOs 30,001-30,037) from SEQ IDs 1-30,037 as they were not present on EPIC version 2)
[0309] • A subset of 25 random CpGs. These CpGs were randomly selected from the first 585
[0310] CpGs using the R sampleQ function. Instead of deriving a classifier by training penalised regression, this classifier was simplify defined as the average methylation levels across these 25 CpG sites.
[0311] • A single CpG. This classifier constitutes the methylation values of a single CpG site
[0312] (SEQ ID number 1).
[0313] ROC curve construction, AUC calculation and threshold selection
[0314] ROC curves were generated based on each of the three classifiers indicated above leveraging the roc() function from the pROC R package, version .1.18.5 and visualised using ggplot2 (3.5.1). Thresholds were selected by finding the threshold at minimum difference from 75%, i.e.. min(l-0.75[sensitivity]). Alternative approaches could leverage optimum thresholds for given scenarios, maximising sensitivity or specificity.
[0315] Unlike in Example 1, values for the classifiers in Figure 20 were not corrected for cellular heterogeneity, such as immune cell proportion. Alternative approaches that may yield higher areas under the curve could include assessment of residuals after accounting for differences in cell type composition.
[0316] Conclusion The AUC analysis above demonstrates the invention’s combination of molecular data leads to a significant discernment of cancer-free controls and cases based on buccal sample methylation, supporting the novelty and utility of the approach described in the patent claims.
Claims
CLAIMS1. An assay for assessing the presence or absence of breast cancer in an individual, the assay comprising: a. providing a sample which has been taken from the buccal area of the individual, the sample comprising a population of DNA molecules; b. determining in the population of DNA molecules in the sample the methylation status of a test panel of one or more CpGs selected from the 30,037 CpGs identified at nucleotide positions 15 to 16 in SEQ ID NOs 1 - 30,037; and c. assessing the presence or absence of breast cancer in the individual based on the methylation status of the CpGs in the test panel.
2. An assay according to claim 1, wherein the test panel of one or more CpGs is selected from the 585 CpGs identified at nucleotide positions 15 to 16 in SEQ ID NOs 1- 585, preferably selected from the 376 CpGs identified at nucleotide positions 15 to 16 in SEQ ID NOs 1 - 376, preferably selected from the 222 CpGs identified at nucleotide positions 15 to 16 in SEQ ID NOs 1 - 222, preferably selected from the 117 CpGs identified at nucleotide positions 15 to 16 in SEQ ID NOs 1 - 117, preferably selected from the 53 CpGs identified at nucleotide positions 15 to 16 in SEQ ID NOs 1 53, and more preferably selected from the 17 CpGs identified at nucleotide positions 15 to 16 in SEQ ID NOs 1 - 17.
3. An assay according to claim 1 or claim 2, wherein the test panel of one or more CpGs comprises two or more CpGs selected from the CpGs identified in the differentially methylated regions (DMRs) defined by SEQ ID NOs 30,038 - 30,182.
4. An assay according to any one of claims 1 to 3, wherein the test panel of one or more CpGs comprises two or more CpGs, or preferably all of the CpGs, identified in a single DMR defined by SEQ ID NOs 30,038 - 30,182, preferably identified in a single DMR defined by SEQ ID NOs 30,038 - 30,099, and more preferably identified in a single DMR defined by SEQ ID NOs 30,038 - 30,051.
5. An assay according to any one of claims 1 to 4, wherein the test panel of one or more CpGs comprises 100, 500, 1,000, 2,000, 5,000, 10,000, 20,000, or all 30,037 CpGs selected from the 30,037 CpGs identified at nucleotide positions 15 to 16 in SEQ ID NOs 1 - 30,037.
6. An assay according to any one of claims 1 to 5, wherein when the one or more CpGs of the test panel are determined to be methylated, the individual is assessed as having breast cancer.
7. An assay according to any one of claims 1 to 6, wherein the step of determining the methylation status of the CpGs in the test panel comprises: i. bisulfite treatment of the test panel of one or more CpGs; and / or ii. an amplification-based technique, preferably wherein the amplificationbased technique is PCR, more preferably further comprising quantification and / or the use of a probe comprising a detection moiety; iii. sequencing the test panel of one or more CpGs.
8. An assay according to any one of claims 1 to 7, wherein at least 25% of the DNA molecules in the sample are derived from immune cells.
9. An assay according to any one of claims 1 to 8, wherein the individual is a woman that did not have breast cancer when the sample was taken from her buccal area.
10. An assay according to any one of claims 1 to 9, wherein the assay is characterised as having an AUC of 0.55 or more, 0.575 or more, 0.60 or more, 0.625 or more, 0.65 or more, 0.675 or more, or 0.70 or more.
11. An in vitro method of assaying DNA and detecting methylation of the DNA therein, the method comprising measuring a methylation status of one or more CpGs, wherein selected from the 30,037 CpGs identified at nucleotide positions 15 to 16 in SEQ ID NOs 1 - 30,037.
12. An in vitro method according to claim 11 , wherein the test panel of one or more CpGs is selected from the 585 CpGs identified at nucleotide positions 15 to 16 in SEQ IDNOs 1-585, preferably selected from the 376 CpGs identified at nucleotide positions 15 to 16 in SEQ ID NOs 1 - 376, preferably selected from the 222 CpGs identified at nucleotide positions 15 to 16 in SEQ ID NOs 1 - 222, preferably selected from the 117 CpGs identified at nucleotide positions 15 to 16 in SEQ ID NOs 1 - 117, preferably selected from the 53 CpGs identified at nucleotide positions 15 to 16 in SEQ ID NOs 1 53, and more preferably selected from the 17 CpGs identified at nucleotide positions 15 to 16 in SEQ ID NOs 1 - 17.
13. An in vitro method according to claim 11 or claim 12, wherein the test panel of one or more CpGs comprises two or more CpGs selected from the CpGs identified in the differentially methylated regions (DMRs) defined by SEQ ID NOs 30,038 - 30,182.
14. An in vitro method according to any one of claims 11 to 13, wherein the test panel of one or more CpGs comprises two or more CpGs, or preferably all of the CpGs, identified in a single DMR defined by SEQ ID NOs 30,038 - 30,182, preferably identified in a single DMR defined by SEQ ID NOs 30,038 - 30,099, and more preferably identified in a single DMR defined by SEQ ID NOs 30,038 - 30,051.
15. An in vitro method according to any one of claims 11 to 14, wherein the test panel of one or more CpGs comprises 100, 500, 1,000, 2,000, 5,000, 10,000, 20,000, or all 30,037 CpGs selected from the 30,037 CpGs identified at nucleotide positions 15 to 16 in SEQ ID NOs 1 - 30,037.
16. A method of treating breast cancer in an individual, the method comprising: i. assessing the presence or absence of breast cancer in an individual according to any one of claims 1 to 10; and ii. administering one or more therapeutic treatments or measures to the individual based on the assessment.
17. A method of monitoring the breast cancer status in an individual, the method comprising: (a) assessing the presence or absence of breast cancer in an individual by performing the assay according to any one of claims 1 to 10 at a first time point; (b) assessing the presence or absence of breast cancer in the individual by performing theassay according to any one of claims 1 to 10 at one or more further time points; and (c) monitoring any change in the breast cancer status of the individual.
Citation Information
Patent Citations
Conjugates of vascular endothelial growth factor with targeted agents
WO1996006641A1
Method of diagnosing bladder cancer
WO2010123354A2
Methods for detecting and predicting breast cancer
WO2021255463A1