Blood-based protein biomarker panels for early and accurate detection of cancer

JP2024500575A5Inactive Publication Date: 2025-08-21THE BRIGHAM & WOMEN S HOSPITAL INC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023562640
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-12-22
Filing Date
2021-12-22
Publication Date
2025-08-21
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Current breast cancer detection methods, such as mammography, suffer from high false-positive and false-negative rates and lack molecular information, making early and accurate detection difficult, especially for heterogeneous breast cancer.

Method used

A blood-based protein biomarker panel using ultrasensitive immunoassays like digital ELISA to detect a combination of biomarkers (e.g., MICA, CA125, CD25, HER3, HSP70, CYR61, LCN2) in blood samples, analyzed through methods like SIMOA, to accurately identify breast cancer and its subtypes.

Benefits of technology

The biomarker panel achieves high accuracy in detecting breast cancer with an AUC of 0.95 and reduces unnecessary invasive procedures by 50% compared to mammography, while also distinguishing between different subtypes with an AUC of 0.96, providing molecular information for timely intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000038_0000
    Figure 00000038_0000
  • Figure 00000038_0001
    Figure 00000038_0001
  • Figure 00000038_0002
    Figure 00000038_0002
Patent Text Reader

Abstract

Methods and compositions for accurate blood biomarker panel-based detection of cancer, such as breast cancer, and for subtyping, for example, using ultrasensitive immunoassays, such as digital ELISA.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] Claiming priority This application claims priority to U.S. Provisional Patent Application No. 63 / 129,432, filed December 22, 2020, the entire contents of which are incorporated herein by reference.

[0002] Federally Sponsored Research or Development This invention was made with Government support under Grant No. W81XWH-11-1-0814 awarded by the Department of Defense. The Government has certain rights in this invention.

[0003] Technical Field Described herein are methods and compositions for accurate blood biomarker panel-based detection of cancer, e.g., breast cancer, and for subtyping using, e.g., ultrasensitive immunoassays, e.g., digital ELISA. [Background technology]

[0004] Breast cancer is the second leading cause of cancer death among women in the United States (1). Summary of the Invention

[0005] Described herein are methods and compositions for accurate blood biomarker panel-based detection of cancer, e.g., breast cancer, for blood samples, and subtyping using, e.g., ultrasensitive immunoassays, e.g., digital ELISA. Thus, provided herein are methods that include obtaining a sample comprising blood (e.g., whole blood, serum, or plasma) from a subject, and determining the levels of at least 2, 3, 4, 5, 10, 15, 20, or all 24 biomarkers listed in Table A in the sample. In some embodiments, the biomarkers include at least MICA, CA125, and CD25. In some embodiments, the biomarkers include at least HER3, HSP70, CYR61, and LCN2. In some embodiments, the biomarkers include at least ER, HER3, HER4, CXCL10, CYR61, P21, MICA, CD25, IL-6, and CD125.

[0006] In some embodiments, the methods include calculating a score for the subject based on the level of the biomarkers, where a score higher than a threshold score indicates that the subject has or is at risk for developing cancer.

[0007] In some embodiments, the methods include calculating a score for a subject based on the levels of the biomarkers and comparing the score to a subtype reference score for a known subtype of breast cancer, and identifying subjects having a score comparable to the subtype reference as having that subtype of breast cancer.

[0008] In some embodiments, the methods include recommending or directing the subject for further evaluation, for example by imaging and / or biopsy.

[0009] In some embodiments, the methods include administering a treatment for breast cancer to a subject identified as having or at risk for developing breast cancer, hi some embodiments, the treatment includes chemotherapy, hormonal therapy, immunotherapy, radiation, or surgical resection.

[0010] In some embodiments, determining the level of the biomarker includes using digital ELISA, e.g., single molecule array (SIMOA); mesoscale discovery (MSD); single molecule counting (SMC); LUMINEX; SOMA scan assay; mass spectrometry (e.g., MALDI-MS), and / or mass cytometry (e.g., CyTOF).

[0011] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention belongs.Methods and materials are described herein for use in the present invention; other suitable methods and materials known in the art can also be used.Materials, methods and examples are for illustrative purposes only and are not intended to be limiting.All publications, patent applications, patents, sequences, database entries and other references mentioned herein are incorporated herein by reference in their entirety.In case of conflict, the present specification, including definitions, will take precedence.

[0012] Other features and advantages of the invention will become apparent from the following detailed description and drawings, and from the claims. [Brief description of the drawings]

[0013] [Figure 1] Selection and initial validation of a biomarker panel in tumor tissue and blood. A. List of biomarkers. B. PCA of mRNA expression for 24 biomarkers measured in various human tumors (9,860 cancer subjects, of which 1,084 are breast cancer subjects) from the TCGA database. Samples were evaluated by RNA-seq. C. Histogram of principal component 1 (data from B). D. PCA of protein levels for 24 biomarkers measured in serum from healthy subjects (n=24) and breast cancer subjects (n=25). Serum samples were measured using the Simoa assay. [Diagram 2]FIG. 1 illustrates the use of blood biomarkers to discriminate between healthy and breast cancer subjects. A. ROC curves for models using a panel of 24 biomarkers + age and age only. B. ROC curves for models using a panel of 4 biomarkers + age. The four biomarkers are HER3, HSP70, CYR61, and LCN2. C. ROC curve for HSP70 + age. 95% confidence intervals are shown in parentheses for panels A-C. D. Decision curves based on 1) panel of 24 biomarkers + age, 2) panel of 4 biomarkers + age, 3) HSP70 + age, 4) age only, 5) all patients classified with cancer (all treated), and 6) no patients classified with cancer (no treatment). [Diagram 3] Figure 1. Subtype analysis using candidate biomarkers. A. Model performance for accurately classifying different breast cancer subtypes as cancer. B. ROC curves for healthy and ER+ breast cancer subjects and healthy and TNBC subjects using a panel of 24 biomarkers + age and a panel of 4 biomarkers + age. C. PCA of mRNA expression levels in breast cancer tumors using our biomarker candidates. Luminal A (n=412), Luminal B (n=174), Normal (n=25), TNBC (n=136), HER2 (n=65). Percent contribution of each marker to the PCA shown in DC. E. ROC curves for a blood protein panel using the 10 most informative markers (ER, HER3, HER4, CXCL10, CYR61, P21, MICA, CD25, IL-6, and CA125) and the top 3 most informative markers (MICA, CA125, and CD25). Hormone positive (n=81) and TNBC (n=10). For panels B and E, 95% confidence intervals are shown in parentheses. [Figure 4]Figure 25. Digital ELISA based on an array of femtoliter-sized wells. (A, B) Standard ELISA reagents are used to capture and label single protein molecules on beads (A) and the beads are loaded into a femtoliter-volume well array (B). (C) SEM of a section of the femtoliter-volume well array after loading of the beads. (D) Fluorescence image of a section of the femtoliter-volume well array after generation of a single enzyme-derived signal. Only a portion of the beads have enzyme activity, indicating a single bound protein molecule. [Figure 5-1] FIG. 1 shows the Simoa assay calibration curve and detection limit. [Figure 5-2] FIG. 1 shows the Simoa assay calibration curve and detection limit. [Figure 5-3] FIG. 1 shows the Simoa assay calibration curve and detection limit. [Figure 5-4] FIG. 1 shows the Simoa assay calibration curve and detection limit. [Figure 6-1] FIG. 1 shows Simoa assay dilution linearity. [Figure 6-2] FIG. 1 shows Simoa assay dilution linearity. [Figure 7-1] FIG. 1 shows Simoa assay spiking and recovery. [Figure 7-2] FIG. 1 shows Simoa assay spiking and recovery. [Figure 7-3] FIG. 1 shows Simoa assay spiking and recovery. [Figure 7-4] FIG. 1 shows Simoa assay spiking and recovery. [Figure 8-1] FIG. 1 shows biomarker levels in cancer and healthy subjects. [Figure 8-2] FIG. 1 shows biomarker levels in cancer and healthy subjects. [Figure 9] FIG. 1 shows a calibration plot for a predictive model. [Figure 10] FIG. 1 shows an XY scatter plot of informative markers. [Figure 11] FIG. 1 shows the correlation between biomarker levels and age in healthy subjects. [Figure 12] FIG. 1 shows variable importance of the model used to discriminate between different subtypes in blood. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0014] Large-scale breast cancer screening programs have been widely implemented because early detection and treatment can improve patient outcomes (2). However, early and accurate detection of breast cancer is challenging due to limitations of traditional detection methods such as mammography, which suffer from high false-positive and false-negative rates (3-11). Furthermore, current screening methods do not provide any disease-related molecular information, which limits their ability to distinguish between benign and negative breast tumors. Because breast cancer is a highly heterogeneous disease, detection methods that provide molecular information are promising for early and accurate detection. Thus, advances in breast cancer detection may reduce patient morbidity by preventing unnecessary invasive biopsies due to false-positives detected by screening. Advances in detection methods also allow timely intervention in cancers requiring treatment, thereby improving patient outcomes.

[0015] Liquid biopsies for cancer detection are particularly promising because they provide molecular information and are minimally invasive (12, 13). Currently, efforts to develop liquid biopsies for breast cancer rely primarily on detecting circulating tumor DNA (ctDNA) and circulating tumor cells (CTCs) (14-16). However, applying these two classes of biomarkers to early cancer detection is challenging because tumors must be relatively large to generate sufficient amounts of ctDNA or CTCs that may be detectable in blood (17-19). Proteins are particularly promising biomarkers because they are directly involved in biological processes that are dysregulated in disease and are abundant in cells. In addition, plasma proteins have been shown to be indicators of health status (20, 21). Previous studies have developed blood tests for breast cancer detection; however, these attempts have limited accuracy, especially for early breast cancer detection (22, 23). Thus, developing tests using circulating proteins may improve the ability to accurately detect breast cancer (24).

[0016] A blood protein biomarker panel for breast cancer detection is described herein. In some embodiments, the method uses analytically robust single molecule array (Simoa) immunoassays (25, 26). Using gene expression data from The Cancer Genome Atlas (TCGA) (27), the inventors showed that the biomarkers can distinguish between breast cancer and other types of cancer in tumor tissue. The inventors then developed and analytically validated assays for these biomarkers in blood, and showed in a small preliminary cohort (n=49) that the panel can distinguish between healthy subjects and breast cancer patients. The inventors then applied the biomarker panel to a second, larger cohort (n=197) of healthy subjects and newly diagnosed, untreated breast cancer patients.

[0017] The results reported here provide evidence that circulating proteins can accurately detect breast cancer. This encouraged critical consideration of detection and screening methods, especially given that most breast cancer subjects had tumors consistent with early-stage disease. For a model using 24 biomarkers plus age, the overall AUC was 0.95 (95% CI 0.92-0.98), and 88% of subjects were correctly classified, with a sensitivity of 87% and a specificity of 90%. This compares favorably to mammography, which has a false-negative rate of approximately 20% (7-11). Furthermore, more than 50% of patients screened annually for 10 years in the United States have false-positive mammograms, requiring further evaluation with biopsy (3-6). Reducing false-positives detected by screening mammograms would reduce unnecessary invasive diagnostic surgical procedures, and overall patient morbidity. Furthermore, the model using 24 biomarkers plus age showed a greater net benefit over a wide range of threshold probabilities compared to the other models, suggesting that diagnostic decisions made with information from a panel of markers may be superior to those made without it.

[0018] We also narrowed the selection of the most informative markers and showed that HER3, HSP70, CYR61, and LCN2 were particularly important biomarkers, with an AUC of 0.87 (95% CI 0.81-0.92) for this 4-biomarker panel.

[0019] Panels of protein biomarkers substantially outperformed any single protein. We observed an AUC of 0.95 for a panel using 24 biomarkers + age, 0.87 for a panel using the four most informative markers + age, and 0.77 for HSP70 and age, which were the best performing single markers. The complete panel had better discrimination, calibration, and improved diagnostic decision than the four biomarker panel using the most informative markers, by net gain (51). Furthermore, the panel substantially outperformed any individual marker. For a given biomarker, there was a large overlap in concentrations in the breast cancer and healthy groups, indicating that the ability to discriminate between breast cancer and healthy subjects depends on the cumulative effect of multiple markers. These results indicate that a complete panel is important for accurate detection of breast cancer.

[0020] Finally, as shown herein, some of the biomarkers may be used to distinguish molecular subtypes of breast cancer: MICA, CA125, and CD25 were the top three most informative protein biomarkers in blood for subtype classification (Figure 12), with an AUC of 0.96 (95% CI 0.91-1.00) using this panel of three markers (Figure 3E).

[0021] The blood tests described herein may be used, for example, individually or in combination with another clinical treatment, such as mammography, to improve the accuracy of breast cancer screening.

[0022] How is it diagnosed? Included herein are methods for diagnosing breast cancer and / or determining the subtype of breast cancer present in a subject. The methods rely on the detection of a biological marker, or a plurality of protein biological markers as described herein, for example, as described in Table A. In some embodiments, the methods provide blood tests for breast cancer detection and diagnosis using circulating protein biomarkers.

[0023] Proteins are factors in cell growth, proliferation, signal transduction, motility, metabolic processes, and regulate tumor development through cell adhesion, invasion, and migration. In addition, proteins regulate the immune system's response to cancer. Therefore, the characterization of proteins involved in breast cancer pathophysiology holds great promise for breast cancer detection and diagnosis. Provided herein is a panel of protein biomarkers associated with breast cancer. These biomarkers are involved in various biological processes, including angiogenesis, growth signaling, and metastasis.

[0024] [Table 1-1]

[0025] [Table 1-2]

[0026] [Table 1-3]

[0027] [Table 1-4]

[0028] [Table 1-5]

[0029] [Table 1-6]

[0030] In some embodiments, the method comprises determining the level of at least 3, 4, 5, 10, 15, 20, or all 24 of the biomarkers of Table A. In some embodiments, the biomarkers comprise at least MICA, CA125, and CD25. In some embodiments, the biomarkers comprise at least HER3, HSP70, CYR61, and LCN2. In some embodiments, the biomarkers comprise at least ER, HER3, HER4, CXCL10, CYR61, P21, MICA, CD25, IL-6, and CA125. In some embodiments, when multiple isoforms of a biomarker are present, a method is used to detect all of the isoforms.

[0031] The method includes obtaining a sample from a subject and assessing the presence and / or level of a breast cancer biomarker in the sample.

[0032] The method may also include comparing the presence and / or level with one or more references, such as a control reference representing a normal level of the breast cancer biomarker, such as the level in an unaffected subject, and / or a disease reference representing the level of a protein associated with breast cancer, such as the level in a subject with breast cancer. In some embodiments, the level provides a differential diagnosis, such as the level in a subject with a known type of breast cancer (e.g., ER+ or TNBC). Suitable reference values ​​may include those shown in Table 1.

[0033] As used herein, the term "sample" refers to the material to be tested for the presence of biological markers using the method of the present invention, and includes whole blood, plasma, or serum. Various methods are well known within the art for identifying and / or isolating and / or purifying biological markers from samples, as needed. An "isolated" or "purified" biological marker is substantially free of cellular material or other contaminants from the cell or tissue source from which the biological marker is derived, i.e., partially or completely altered or removed from its natural state by human intervention. For example, proteins contained in a sample can be isolated according to standard methods, for example using lytic enzymes, chemical solutions, or isolated by protein-binding resins according to the manufacturer's instructions.

[0034] The presence and / or level of protein can be evaluated using methods known in the art.In a preferred embodiment, the method comprises the use of sensitive or ultrasensitive and preferably multiplexed detection methods, including mesoscale discovery (MSD); single molecule array (SIMOA); single molecule counting (SMC); LUMINEX; SOMA scan assay; mass spectrometry (e.g. MALDI-MS) and mass cytometry (e.g. CyTOF) (see, for example, Cohen and Walt, Chem. Rev. 2019, 119, 293-321).

[0035] In some embodiments, protein biomarkers in blood for breast cancer detection are measured using the SIMOA assay (25, 26). The SIMOA assay has several advantages over traditional ELISA, the current gold standard for protein detection in blood. First, SIMOA is 1000 times more sensitive than ELISA, allowing for the quantification of analytes present at low concentrations (25). SIMOA is 10 -12 10 compared to the ability of conventional ELISA to detect only M -19It can detect protein concentrations as low as 1000 μg / mL. Second, due to the high sensitivity of SIMOA, serum samples can be more diluted, thereby reducing nonspecific binding due to matrix effects (53, 54). Third, SIMOA has a wide dynamic range spanning four orders of magnitude in concentration, so a single assay can be used to detect both low and high abundance markers (55). In some embodiments, SIMOA technology achieves this high sensitivity by digitally counting the number of molecules in a sample by labeling and physically isolating each immune complex into a femtoliter-sized well (Figure 4A-D). These advantages provide for the detection and quantification of blood biomarkers for developing robust biomarker panels.

[0036] In some embodiments, mass spectrometry, and in particular matrix-assisted laser desorption / ionization mass spectrometry (MALDI-MS) and surface-enhanced laser desorption / ionization mass spectrometry (SELDI-MS), is used for detection of biomarkers (see U.S. Pat. Nos. 5,118,937; 5,045,694; 5,719,060; 6,225,047).

[0037] In some embodiments, other methods may be used, such as standard electrophoretic and quantitative immunoassay methods for proteins, including, but not limited to, Western blot; enzyme-linked immunosorbent assay (ELISA); enzyme-linked immunospot (ELISPOT); biotin / avidin type assays; protein array detection, such as protein microarrays; radioimmunoassays; immunohistochemistry (IHC); immunoprecipitation assays; flow cytometry / FACS (fluorescence activated cell sorting); proximity ligation assay (PLA); lateral flow assays; surface plasmon resonance (SPR); optical imaging; and mass spectrometry (Kim (2010) Am J Clin Pathol 134: 157-162; Yasun (2012) Anal Chem 84(14): 6008-6015; Brody (2010) Expert Rev Mol Diagn 10(8): 1013-1022; Philips (2014) PLOS One 9(3): e90226;Pfaffe (2011) Clin Chem 57(5): 675-687;Cohen and Walt, Chem. Rev. 2019, 119, 293-321). The methods typically involve revealing labels, such as fluorescent, chemiluminescent, radioactive, and enzyme or dye molecules, that provide a signal either directly or indirectly. As used herein, the term "labeling" refers to the coupling (i.e., physical binding) of a detectable substance, such as a radioactive agent or a fluorophore (e.g., phycoerythrin (PE) or indocyanine (Cy5)), to an antibody or probe, and indirect labeling of a probe or antibody by reactivity with a detectable substance (e.g., horseradish peroxidase, HRP).

[0038] In some embodiments, if the presence and / or level of biomarkers is comparable to the presence and / or level of proteins in disease reference and the subject has one or more symptoms associated with breast cancer, the subject has breast cancer.In some embodiments, if the subject does not have obvious signs or symptoms of breast cancer, but the presence and / or level of one or more proteins evaluated is comparable to the presence and / or level of proteins in disease reference, the subject has breast cancer or is at increased risk of developing breast cancer.In some embodiments, once it is determined that a person has breast cancer or is at increased risk of developing breast cancer, treatment can be administered, for example, as known in the art or as described herein.

[0039] Suitable reference value can be determined by using methods known in the art, for example by using standard clinical test methodology and statistical analysis.Reference value can have any relevant form.In some cases, reference comprises a predetermined value for a meaningful level of biomarker, for example a normal level of biomarker, for example a control reference level representing the level in a subject who is not affected or does not have the risk of developing the disease described herein, and / or a disease reference representing the level of a protein associated with breast cancer, for example the level in a subject who has breast cancer.

[0040] The predetermined level can be a single cut-off (threshold) value, such as a median or mean, or a level that defines the boundaries of upper or lower quartiles, tertiles, or other segments of the clinical trial population that are determined to be statistically different from other segments. It can be a range of cut-off (or threshold) values, such as a confidence interval. It can be established based on a comparison group, such as the association between the risk of developing disease or the presence of disease in one defined group being fold higher or lower (e.g., approximately 2-fold, 4-fold, 8-fold, 16-fold or more) than the risk or presence of disease in another defined group. It can be, for example, a range in which a population of subjects (e.g., control subjects) is divided evenly (or unevenly) into groups, such as low-risk, medium-risk, and high-risk groups, or into quartiles, with the lowest quartile being the subjects with the lowest risk and the highest quartile being the subjects with the highest risk, or into n quartiles (i.e., n regular intervals), with the lowest n quartile being the subjects with the lowest risk and the highest n quartile being the subjects with the highest risk.

[0041] In some embodiments, the predetermined level is the level or occurrence in the same subject, eg, at a different time point, eg, an earlier time point.

[0042] The subject associated with the predetermined value is typically referred to as a reference subject. For example, in some embodiments, the control reference subject does not have breast cancer, does not have a risk of developing breast cancer, or does not subsequently develop breast cancer.

[0043] A disease reference subject is one who has breast cancer (or is at increased risk of developing breast cancer), where increased risk is defined as a risk higher than the subject's risk in the population.

[0044] Thus, in some cases, when a biomarker is decreased in cancer (see Table 1), a level of the biomarker in a subject that is less than or equal to the reference level of the biomarker indicates the presence of breast cancer or risk of developing breast cancer, and a level of the biomarker in a subject that is greater than or equal to the reference level of the biomarker indicates the absence of disease or a normal risk of disease.

[0045] In other cases, when a biomarker is elevated in cancer (see Table 1), a level of the biomarker in a subject that is higher than or equal to the reference level of the biomarker indicates the presence of breast cancer or risk of developing breast cancer, and a level of the biomarker in a subject that is less than or equal to the reference level of the biomarker indicates the absence of disease or a normal risk of disease.

[0046] To build diagnostic models, as described below, the outcome was a binary breast cancer case status (breast cancer vs. healthy). Age and protein markers were modeled as continuous predictors. Values ​​were log-transformed and a logistic regression model was used to classify breast cancer and healthy subjects. To assess the classification accuracy of each particular model, subjects with a predicted probability of at least 50% were designated as predicted to have cancer, and subjects with a predicted probability of less than 50% were predicted to be healthy. The predicted case status of a subject for a given model was then compared to the observed case status.

[0047] Thus, in some embodiments, to assess whether a subject has breast cancer in a clinic, the method may include first log-transforming the biomarker values, and then assigning a predicted probability to generate a probability score, for example, using a logistic regression model. If the subject has a predicted probability score that is higher than a selected threshold, for example at least 50%, the subject will be predicted to have cancer (e.g., assigned to a cancer category). If the predicted probability score is lower than a selected threshold, for example 50%, the subject will be predicted to be healthy (e.g., assigned to a healthy category).

[0048] In some embodiments, the level of biomarkers is used, for example, together with one or more additional variables, for example, age, to calculate a score.The score can be calculated, for example, using an algorithm, such as summation or weighted summation of (normalized) levels of biomarkers.The specific algorithm can be identified using known statistical methods, including PCA, linear regression, SVM (support vector machine), decision tree, KNN (K nearest neighbors), K-means, gradient boosting, or random forest methods.

[0049] For example, in some embodiments, the exemplary model uses logistic regression analysis where each variable (biomarker, X) is assigned a weight (B). In the exemplary equation below, a weight (B) is calculated for each marker, and for each of the biomarkers, for example, for the 24 biomarkers and age (25 in total), there may be a unique B value.

[0050]

number

[0051] In the clinic, the measured biomarker values ​​(X values) can be used to derive a probability score that a patient will have cancer by plugging the measured biomarker values ​​(X) into an equation and then calculating a probability value (P). In some embodiments, the clinical procedure for deriving an individual's probability of having breast cancer is as follows:

[0052] First, blood is collected from the screenee. Second, Simoa is used to measure the blood concentration of each biomarker protein in the panel of the screenee. Third, the predicted probability of the screenee having breast cancer is calculated based on a logistic regression equation with the dependent variable of the natural logarithm of [(probability of having breast cancer) / (probability of not having breast cancer)], and with the dependent variables of age and each biomarker in the panel. The predicted probability can then inform the discussion between the screenee and the doctor regarding the best course of action, such as the decision that no further follow-up is necessary or the decision to perform a confirmatory radiographic imaging.

[0053] For the panel of 24 markers, model parameter estimates based on the Tufts sample in 197 participants, with age measured in years, CA15-3 and CA19-9 measured in units / mL, and all other markers measured in pg / mL, were as follows:

[0054] [Table 2]

[0055] For the panel of four markers identified via cross-validation, model parameter estimates based on the Tufts sample in 197 participants, age measured in years and all markers measured in pg / mL, were as follows:

[0056] [Table 3]

[0057] In some embodiments, the amount by which the level (or score) in a subject is less than the reference level (or score) is sufficient to distinguish the subject from a control subject, and optionally is statistically significantly less than the level (or score) in the control subject. When the level (or score) of a biomarker in a subject is equal to the reference level (or score) of the biomarker, "equal" refers to being approximately equal (e.g., not statistically different).

[0058] The predetermined value may depend on the particular subject population (e.g., human subject) selected. For example, an apparently healthy population will have a different "normal" range of biomarker levels than a population of subjects who have, are likely to have, or are at higher risk of having a disorder described herein. Thus, the selected predetermined value may take into account the category (e.g., sex, age, health status, risk, presence of other diseases) to which the subject (e.g., human subject) belongs. Appropriate ranges and categories can be selected by those skilled in the art through routine experimentation.

[0059] In characterizing the likelihood, or risk, a number of predetermined values ​​may be established.

[0060] Breast cancer is typically classified into one of three main subtypes based on the presence or absence of estrogen or progesterone receptor and human epidermal growth factor 2 (ERBB2; formerly HER3) molecular markers: hormone receptor positive / ERBB2 negative, ERBB2 positive, and triple negative (tumors lacking all three standard molecular markers); see, e.g., Waks and Winer, JAMA. 2019 Jan 22; 321(3): 288-300. The method may also be used to make a differential diagnosis between estrogen receptor positive (ER+) and triple negative breast cancer (TNBC). In these methods, at least MICA, CA125, and CD25, or at least ER, HER3, HER4, CXCL10, CYR61, P21, MICA, CD25, IL-6, and CA125 are determined and used to identify whether a subject has ER+ breast cancer or TNBC. Exemplary coefficients for the 10 and 3 marker panels are as follows:

[0061] [Table 4]

[0062] [Table 5]

[0063] Thus, in some embodiments, the model is used to identify the presence of an ER+ subtype. The model provides the log odds of having an ER+ breast tumor versus having no breast cancer at all, as well as the predicted probability of an individual having ER+ breast cancer compared to having no breast cancer at all. For the triple-negative subtype, the model provides the log odds of having a triple-negative breast tumor versus having no breast cancer at all, as well as the predicted probability of an individual having triple-negative breast cancer compared to having no breast cancer at all.

[0064] The methods can also be used to identify subjects for further evaluation, such as imaging (e.g., mammogram or ultrasound) and / or biopsy, to confirm a cancer diagnosis.

[0065] Treatment Method The methods described herein include methods for treating breast cancer.Generally, the methods include selecting and optionally administering a therapeutically effective amount of breast cancer treatment to a subject determined by the methods described herein to be in need of such treatment.Breast cancer treatments include radiation, surgical resection, chemotherapy, hormone / endocrine therapy, and immunotherapy.

[0066] In some embodiments, if a subject is identified as likely to have TNBC, a treatment is selected and optionally administered that includes administration of a regimen that includes chemotherapy, e.g., platinum compounds, anthracycline-based or anthracycline and taxane-based chemotherapy, and / or antimetabolites (e.g., cyclophosphamide, methotrexate and 5-fluorouracil (CMF), or cyclophosphamide, epirubicin and 5-fluorouracil (CEF)) (see, e.g., Bianchini et al., Nat Rev Clin Oncol. 2016 Nov; 13(11): 674-690; Bergin and Loi, F1000Res. 2019 Aug 2; 8: F1000 Faculty Rev-1342; Kumar and Aggarwal, Arch Gynecol Obstet. 2016 Feb; 293(2): 247-69; Nedeljkovic and Damjanovic, Cells. 2019 Aug 22; 8(9): 957;Al-Mahmood et al., Drug Deliv Transl Res. 2018 Oct; 8(5): 1483-1507;Caparica et al., ESMO Open. 2019 May 13; 4(Suppl 2): ​​e000504).

[0067] In some embodiments, if a subject is identified as likely to have ER+ breast cancer, a treatment is selected and optionally administered that includes endocrine therapy (e.g., tamoxifen, toremifene, fulvestrant, an aromatase inhibitor (AI) (e.g., letrozole (Femara), anastrozole (Arimidex), or exemestane (Aromasin)) or ovarian function suppression, e.g., with oophorectomy or LHRH analogs), and optionally chemotherapy (e.g., those mentioned above or administration of a phosphoinositide 3 kinase (PI3K), mechanistic target of rapamycin (mTOR), or cyclin-dependent kinase (CDK) 4 / 6 inhibitor or a poly(ADP-ribose) polymerase (PAPP) inhibitor) (see Waks and Winer, JAMA. 2019 Jan 22; 321(3): 288-300). EXAMPLES

[0068] The invention is further described in the following examples, which do not limit the scope of the invention described in the claims.

[0069] material and method The following materials and methods were used in the examples presented herein.

[0070] Study design In this study, we sought to develop a blood-based protein biomarker panel for breast cancer detection using an analytically robust Simoa assay. We identified 24 biomarker candidates and developed and analytically validated the Simoa assay. We further confirmed that our selected biomarkers are indicative of breast cancer compared to other cancers using mRNA expression levels in tumor tissue for these biomarkers from TCGA. We then used a first preliminary sample cohort of healthy subjects and breast cancer patients (n=49) to measure the 24 protein biomarker candidates in serum using our developed Simoa assay. We then sought to confirm our results in a second, larger cohort. We began sampling at Tufts Medical Center. All subjects in this cohort were female and over 40 years old. Breast cancer subjects had not previously been treated for breast cancer and generally had tumors consistent with early stage disease. We measured the concentration of 24 biomarkers in serum using the Simoa assay. We developed a model using logistic regression analysis of these 24 biomarkers + age to discriminate between healthy and breast cancer subjects. To narrow the selection of the most important markers, we used a backwards selection process and then developed a model using the four most informative markers + age. As a second analysis, we evaluated the subtypes that were correctly classified as cancer by the two models. We also used TCGA data to identify important biomarkers to discriminate between subtypes using protein biomarkers in blood. We then built a model using logistic regression analysis to determine whether a subject has ER+ or TNBC in serum using protein biomarkers. Informed consent was obtained for all blood samples used in this study.

[0071] Biomarker panel analysis in tumor tissue We obtained mRNA expression data deposited in The Cancer Genome Atlas (TCGA) database (cancergenome.nih.gov / ) and performed principal component analysis (PCA) using the Caret package in R version 3.6.2. A total of 9,860 cancer subjects were analyzed, of which 1,084 were breast cancer subjects. For this analysis, we used 23 of the 24 biomarkers shown in Figure 1A. We did not include CA19-9 in the analysis due to the lack of corresponding mRNA data.

[0072] Simoa Assay Description The Simoa assay is a bead-based immunoassay with a major advance in signal detection by single molecule counting, resulting in ultra-high sensitivity. Antibody-coated capture beads are added in large excess to a sample containing a low concentration of target analyte molecules. Poisson statistics require that either one or zero target protein molecules bind to each bead. The beads are then incubated with biotinylated detection antibodies and streptavidin-β-galactosidase to form an enzyme-labeled immune complex. The beads are then loaded into an array of 50 fL sized wells, where each well can hold only one bead. A fluorogenic substrate is added and the wells are sealed with oil, resulting in a locally high concentration of fluorescent product, thus allowing quantification of single molecules by counting active wells. At high target molecule concentrations, the fluorescence intensity of the array is used to measure the target concentration, thereby expanding the dynamic range of the assay. The signal output is measured with a Simoa instrument using the base units of average enzyme per bead (AEB). All Simoa consumables and reagents were purchased from Quanterix Corp.

[0073] Development of ultrasensitive Simoa assay Capture antibodies were reconstituted and stored according to the instructions provided by the manufacturer. Antibody catalog numbers are provided in Table 1. Antibodies were buffer exchanged by first adding 0.13 mg of antibody solution to an Amicon filter (50K, EMD Millipore) to remove the storage buffer. Bead conjugation buffer (Quanterix) was then added to the filter to a total volume of 500 μL. The filter device was centrifuged at 14,000 x g for 5 minutes. The waste liquid was discarded and the process was repeated twice. The filter was inverted into a new tube and centrifuged at 1000 x g for 2 minutes. The filter was rinsed with 50 μL of bead conjugation buffer and centrifuged at 1000 x g for 2 minutes. The concentration of the antibody was measured using a NanoDrop 2000 spectrophotometer. The antibody was diluted to 0.5 mg / mL in bead conjugation buffer and stored on ice until ready for use. 2.8 × 10 8100 carboxylated 2.7 μm paramagnetic beads (Quanterix) were transferred to a microtube and washed three times with 200 μL of bead wash buffer (Quanterix). The beads were then washed twice with 200 μL of bead conjugation buffer and resuspended in 190 μL of bead conjugation buffer. Fresh 10 mg of 1-ethyl-3-(3-dimethylaminopropyl) carbodiimide hydrochloride (EDC) (ThermoFisher) was reconstituted in 1 mL of bead conjugation buffer immediately before use. To activate the beads, 10 μL of EDC was added to the bead suspension to obtain a final concentration of 0.5 mg / ml and a final volume of 200 μL. The beads were then placed on a rotator for 30 minutes. The activated beads were then washed with 200 μL of bead conjugation buffer. Then, for conjugation, 200 μL of capture antibody solution was added to the beads, vortexed, and placed on a rotator for 120 minutes. The antibody-conjugated beads were then washed twice with 200 μL of bead washing buffer. The beads were then blocked with 200 μL of bead blocking buffer (Quanterix) and placed on a rotator for 30 minutes. The beads were washed with 200 μL of bead washing buffer, washed with 200 μL of bead diluent (Quanterix), and resuspended in 200 μL of bead diluent. The beads were counted using a Beckman Coulter multi-sizer and stored at 4°C.

[0074] Detection antibodies not already biotinylated by the supplier were biotinylated for use in the Simoa assay as previously described. (56) Briefly, antibodies were purified using Amicon filters in biotinylation reaction buffer (Quanterix) in triplicate. Antibody concentrations were measured using a NanoDrop One spectrophotometer. Antibodies were conjugated to biotin using EZ-Link NHS-PEG4 Biotin (Thermo Fisher Scientific) using a 40-fold molar excess and incubated for 30 min. Biotinylated antibodies were then purified using Amicon filters.

[0075] Serum samples were measured using a Simoa HD-1 analyzer with a standard curve. 2 A calibration curve was fitted using 4PL fit with weighting factors. The calibration curve was used to determine the concentrations of unknown samples. The analysis was performed automatically on a Simoa HD-1 analyzer using software provided by Quanterix. The limit of detection (LOD) was calculated as the mean of background plus 3 times the standard deviation.

[0076] Biomarker measurements in exploratory sample cohorts Breast cancer serum samples (n=25) and self-reported healthy serum samples (n=24) were obtained from BioIVT. 24 protein markers were measured in duplicate in the samples using the Simoa assay. The average of the measurements was calculated and the values ​​were log-transformed. Principal component analysis was then performed using the Caret package in R version 3.6.2.

[0077] Subject selection for developing a biomarker panel Breast cancer patients at Tufts Medical Center were screened and diagnosed with breast cancer via standard approach, i.e., mammography followed by biopsy. Patients who had not undergone surgical and / or therapeutic interventions were eligible. Eligible patients consented to donate blood for research upon positive breast cancer diagnosis. Healthy subjects were obtained from Partners biobank, which provides a cohort of curated healthy subjects collected at several different hospitals. All subjects were female and over 40 years of age. Cases are referred to as breast cancer subjects and non-cases are referred to as healthy subjects.

[0078] statistical analysis Blood biomarker levels were analyzed in 197 subjects (100 healthy, 97 cancer). The outcome was binary breast cancer case status (breast cancer vs. healthy). Age and protein markers were modeled as continuous predictors. Each marker had up to three replicates per subject. An individual's final marker measurement was the average of the non-missing replicate measurements. In a given analysis model, if a subject did not have an observed replicate of a particular marker, the individual was first assigned an imputed value for the marker using multiple imputation. If a subject had a biomarker level lower than the LOD for a given assay, the value was assigned as the LOD for that assay. Values ​​were log-transformed and a logistic regression model was used to classify breast cancer and healthy subjects.

[0079] A 5-fold cross-validation was used to identify a subset of "high-performing" markers. To perform cross-validation and marker selection, each of the 197 subjects was randomly assigned to one of five groups. For each of the five folds, one group was eliminated (test set) and the analysis was performed on the combination of the other four groups (training set). Using PROC ADAPTIVEREG in SAS, each fold started with a fold-specific training set from the model of age and all 24 markers, and back-calculated to an intercept-only model with age in the model. The set of predictors that produced the smallest cross-validation error was selected as the fold-specific model. Generalized cross-validation (GCV) is a measure of the predictive accuracy of a fold-specific model. The contribution of each variable to the fold-specific model was measured by its importance, defined as the square root of the GCV value of the fold-specific model with all basis functions related to the variable removed minus the square root of the GCV value of the selected model, and then adjusted to set the maximum importance to 100. Markers with a significance of at least 70 in at least three folds were selected as cross-validated markers.

[0080] We then compared four models that differed by the set of predictors included: first, age only; second, age + HSP70, selected by being the single marker with the largest AUC; third, age + markers selected in the 4-way cross-validation (HSP70, HER3, CYR61, LCN2); and fourth, age + all 24 markers. For each model, discrimination was assessed by AUC. Calibration was assessed using a LOESS-smoothed calibration plot of the observed probability (0 or 1) versus the predicted probability of the outcome. We explored the possibility of improving clinical judgment for each model using decision curves plotting net benefit versus threshold probability. The threshold probability is the probability designated as the cutoff that defines a high probability outcome, i.e., a positive test result.

[0081] To assess the classification accuracy of each particular model, subjects with a predicted probability of at least 50% were designated as predicted to have cancer, and those below 50% were predicted to be healthy. The predicted case status of a subject for a given model was then compared to the observed case status.

[0082] To evaluate our ability to discriminate between different molecular breast cancer subtypes, we first performed a principal component analysis using the mRNA expression levels of 23 biomarkers from the TCGA database. In the TCGA database, tumors are classified as luminal A (n=412), luminal B (n=174), normal (n=25), basal (n=136), and HER2 (n=65). We then used the factoextra package in R to evaluate the contribution of biomarkers to the first two principal components and identified the 10 most informative markers (top markers) and the 10 least informative markers (bottom markers). Using biomarker measurements in breast cancer serum samples, we performed a logistic regression analysis using two panels (top and bottom markers) in ER+ (n=81) and TNBC (n=10) breast cancer subjects. ER+ tumors were defined as having at least 1% positive staining using immunohistochemistry of tissue biopsies. Triple-negative tumors did not have ER, PR, or HER2 expression. We then selected the three most informative markers from the model using the top markers and performed another logistic regression analysis using the three-marker panel. The three markers were identified in R using the varImp function in the caret package.

[0083] Statistical analyses were performed using SAS 9.4 (SAS Institute, Cary, NC) and R version 3.6.2. Figures were generated using GraphPad Prism 7 (San Diego, CA). Decision curves and standard errors for estimating AUC confidence limits were generated using R and SAS macros available online (57, 58).

[0084] result Initial validation of candidate biomarkers for breast cancer detection We selected 24 biomarker candidates for breast cancer detection based on previous studies (28-49) (Figure 1A). We first assessed whether the biomarkers were associated with breast cancer based on gene expression levels in primary tumor tissues. Principal component analysis (PCA) of mRNA expression data deposited in The Cancer Genome Atlas (TCGA) database showed that the biomarkers were able to discriminate breast cancer from all other cancers (Figure 1B-C). We then developed digital ELISA using single molecule array (Simoa) assays for these biomarkers and ensured that the assays were analytically robust by performing rigorous confirmation testing (Figures 5-7, Tables S1-S2). Using these Simoa assays, we tested serum samples from a preliminary cohort of female self-reported healthy subjects (n=24) and breast cancer subjects (n=25) (Figure 1D). The inventors have shown that this panel can easily distinguish between healthy subjects and breast cancer subjects. These results suggest that the panel of 24 biomarkers is promising for breast cancer detection. To confirm this result, the inventors have sought to use these biomarkers to determine whether they can detect breast cancer in blood in a larger cohort of newly diagnosed patients who have not undergone any treatment.

[0085] A Blood Biomarker Panel for Breast Cancer Detection in Newly Diagnosed, Treatment-Naive Cohorts To evaluate our ability to detect breast cancer using a blood biomarker panel, we analyzed serum samples from newly diagnosed patients beginning with sampling at Tufts Medical Center. Patients were screened by mammography and diagnosed with breast cancer by biopsy. Patients with a positive breast cancer diagnosis and no surgical or therapeutic intervention were eligible. Tumor characteristics of breast cancer subjects are shown in Table 3. Tumors were generally consistent with early stage disease, most were small (T0-T2) and node negative (N0), and all were nonmetastatic (M0). The majority of tumors were estrogen receptor (ER) positive, with a median ER measurement of 95% (interquartile range 85%, 98%) using immunohistochemistry on biopsies. Healthy subjects were obtained from Partner's biobank, which provides a cohort of blood samples from selected healthy subjects collected at several different hospitals. These 197 subjects (100 healthy subjects and 97 breast cancer subjects) were all female and at least 40 years old.

[0086] We measured serum biomarker levels in this sample cohort using 24 Simoa assays. Table 1 and Figure 8 provide the age and biomarker distributions for healthy and breast cancer subjects. Age distributions were similar for healthy and breast cancer subjects. We then used logistic regression analysis to examine whether the biomarker panel could discriminate between healthy and breast cancer subjects. As shown in Figure 2A, the model using all 24 biomarkers + age had an area under the curve (AUC) of 0.95 (95% CI 0.92-0.98), while the model using only age was uninformative with an AUC of 0.51 (95% CI 0.43-0.59). The model using all 24 biomarkers + age correctly identified 174 of 197 (88%) subjects, with a sensitivity of 87% and a specificity of 90%.

[0087] We then narrowed the selection of the most informative markers in the model using a cross-validated reduced variable selection process with 24 protein biomarkers + age, resulting in HER3, HSP70, CYR61, and LCN2 as the most informative markers (Table 4). The model using these four biomarkers + age (Figure 2B) had an AUC of 0.87 (95% CI 0.81-0.92). The model correctly identified 165 of 197 subjects (84%) with a sensitivity of 85% and a specificity of 83%. The AUC of the combined cross-validated test set was 0.94 (95% CI 0.92-0.97), indicating that the four biomarker panel was well validated (Table 4). Model calibration is shown in Figure 9. We also evaluated the performance of each of these biomarkers + age against themselves (Table 5). The HSP70+age model (Figure 2C) had an AUC of 0.77 (95% CI 0.71-0.84) and outperformed all other individual marker+age models. Compared to each individual marker+age model, the model using the four biomarkers+age performed substantially better. These results suggest that the panel is important to obtain optical discrimination and that individual markers alone are not sufficient to detect breast cancer. Figure 10 shows an XY scatter plot of the relationship between these four markers. Furthermore, because the risk of breast cancer increases with age, we included age in our model. We show that the concentrations of some biomarkers are correlated with age in healthy subjects (Figure 11).

[0088] We also evaluated the net benefit rate (Figure 2D) (50-52). For all threshold probabilities of approximately 10% or higher, the model using 24 biomarkers + age had a higher net benefit than any alternative model, including the decision to classify all subjects as healthy or, conversely, the decision to classify all subjects as cancer. For threshold probabilities below 10%, the difference in net benefit across the various models was small.

[0089] [Table 6]

[0090] Subtype analysis using candidate biomarkers Breast cancer is a heterogeneous disease consisting of different molecular subtypes, therefore, we sought to evaluate whether our model could accurately classify different breast cancer subtypes as cancer. We identified three subtypes in our breast cancer cohort: ER+ tumors, ER- / HER2+ tumors, and triple-negative breast cancer (TNBC). We investigated the performance of the 24-biomarker panel and the 4-biomarker panel we described in the previous section, and found that both models were able to accurately classify different breast cancer subtypes as cancer (Figure 3A). These results suggest that the panels can be used to broadly detect different subtypes of breast cancer. To further confirm these results, we developed a new model using two biomarker panels and two different groups. The first group consisted of healthy subjects and ER+ breast cancer subjects, and the second group consisted of healthy subjects and TNBC subjects. We found that all four models had high AUC (Figure 3B), suggesting that these biomarker panels can accurately discriminate between healthy and breast cancer subtypes.

[0091] We next wanted to determine whether the 24 protein biomarkers could discriminate between ER+ and TNBC in blood. Due to our small sample size (n=81 for ER+ and n=10 for TNBC), we sought to narrow the selection and identify the most important biomarkers that could discriminate between the different subtypes. To narrow the selection of markers, we examined the mRNA expression levels of the 24 biomarkers in primary tumors by TCGA and observed that ER+ and TNBC subtypes clustered away from each other (Figure 3C). We selected the top 10 markers that contributed most to the principal components (Figure 3D) and developed a model using these 10 protein biomarkers in blood, which provided an AUC of 0.96 (95% CI 0.92-1.00) (Figure 3E). We identified MICA, CA125, and CD25 as the top three most informative protein biomarkers in blood for subtype classification (Figure 12), and using this panel of three markers, we observed an AUC of 0.96 (95% CI 0.91-1.00) (Figure 3E). Overall, our results suggest that protein biomarkers can accurately classify each of several different breast cancer subtypes, and further suggest that a subset of 24 biomarkers can distinguish ER+ from TNBC in blood.

[0092] [Table 7]

[0093] [Table 8]

[0094] Tumor characteristics of breast cancer cases (n=97)

[0095] [Table 9]

[0096] [Table 10]

[0097] [Table 11] reference

[0098] [Table 12-1]

[0099] [Table 12-2]

[0100] [Table 12-3]

[0101] [Table 12-4]

[0102] [Table 12-5]

[0103] [Table 12-6]

[0104] [Table 12-7]

[0105] [Table 12-8]

[0106] Other embodiments Although the present invention has been described in conjunction with the detailed description thereof, it should be understood that the above description is intended to illustrate, and not to limit, the scope of the invention, which is defined by the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.

Claims

1. A biomarker for detecting cancer or determining the risk of developing cancer, comprising at least 2, 3, 4, 5, 10, 15, 20, or 24 proteins listed in Table A in a sample containing blood obtained from a subject.

2. The biomarker described in claim 1, wherein the biomarker includes at least MICA, CA125, and CD25.

3. The biomarker described in claim 1, wherein the biomarker includes at least HER3, HSP70, CYR61, and LCN2.

4. The biomarker of claim 1, wherein the biomarkers include at least ER, HER3, HER4, CXCL10, CYR61, P21, MICA, CD25, IL-6, and CD125.

5. A method for detecting cancer or assessing the risk of developing cancer, comprising determining a level of a biomarker described in any one of claims 1 to 4, calculating a score for a subject based on said level, and confirming that said score is higher than a threshold score.

6. A method for detecting breast cancer of a known subtype, comprising determining a level of a biomarker described in claim 2 or 4, calculating a score for a subject based on said level, comparing said score to a subtype reference score for breast cancer of said known subtype, and confirming that said subject has a score comparable to said subtype reference.

7. The method of claim 5, comprising conducting further evaluation of the subject.

8. The method of claim 7, wherein the further evaluation includes imaging.

9. The method of claim 6, comprising conducting further evaluation of the subject.

10. The method of claim 9, wherein the further evaluation includes imaging.

11. The method of claim 5, wherein determining the level of the biomarker comprises using digital ELISA; mesoscale discovery (MSD); single molecule counting (SMC); LUMINEX; SOMA scan assay; mass spectrometry (optionally MALDI-MS), and / or mass cytometry (optionally CyTOF).

12. The method of claim 11, wherein the digital ELISA uses a single molecule array (SIMOA).

13. The method of claim 6, wherein determining the level of the biomarker comprises using digital ELISA; mesoscale discovery (MSD); single molecule counting (SMC); LUMINEX; SOMA scan assay; mass spectrometry (optionally MALDI-MS), and / or mass cytometry (optionally CyTOF).

14. The method of claim 13, wherein the digital ELISA uses a single molecule array (SIMOA).