Biomarkers for breast cancer detection
Patent Information
- Application Number
- PCT/EP2026/055495
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-28
- Filing Date
- 2026-02-27
- Publication Date
- 2026-09-03
Smart Images

Figure IMGF000012_0001 
Figure IMGF000021_0001 
Figure IMGF000040_0001
Abstract
Description
[0001] BIOMARKERS FOR BREAST CANCER DETECTION
[0002] FIELD OF THE INVENTION
[0003] The present invention relates to the field of molecular diagnostics, specifically to methods and kits for the early and minimally invasive detection of breast cancer. More particularly, the invention concerns the use of a synergistic combination of transcriptomic (mRNA) biomarkers measured in blood samples, and the calculation of a probability score for the diagnosis of breast cancer, including early-stage disease, with high sensitivity and specificity.
[0004] BACKGROUND OF THE INVENTION
[0005] Currently, the state of the art in the early detection of breast cancer is based on imaging methods, in particular mammography. No alternative of validated routine biochemical or molecular tests exists. Histological, genetic and transcriptomic analyses are performed on tumor material obtained at surgery to determine tumor subtypes and stratify patients for optimal therapy. For the follow-up, patients are routinely monitored non-specifically by clinical, hematological, biochemical and molecular tests, mostly to detect unwanted effects of chemotherapy. Relapses are largely detected indirectly, upon clinical symptoms or nonspecific laboratory results.
[0006] Breast cancer is a well-known and well-described pathology affecting approximately 12% of women over an 80-year lifespan. In total, roughly 2 million new cases per year are diagnosed world-wide (nearly 500,000 in Europe and over 6,000 in Switzerland). Breast cancer prognosis and therapy are largely determined by the biological and molecular characteristics of the tumor and the stage at diagnosis. Three main clinically relevant subtypes of breast cancer have been defined: Estrogen receptor positive (ER+), often also progesterone receptor positive (PR+), Human Epidermal Growth Factor Receptor 2 amplified (HER2+) and triple negative breast cancer (TNBC, i.e. ER-, PR- and HER2- breast cancers). Molecular subtypes (e.g. Luminal A, Luminal B, HER2+, basal-like) that overlap largely, but not fully with the clinical subtypes, have been defined based on gene expression profiling [1], Adjuvant treatments, in combination with mastectomy or breast-saving surgery, have improved survival by about 30%, in particular with the introduction of anti-estrogen (e.g. tamoxifen) and anti-HER2 (e.g. Herceptin) based-treatments, for ER+ and HER2+ cancers, respectively. For TNBC there are still no specific therapies due to the lack of defined targets, and radio- and chemotherapies are routinely used instead. Further improvements of adjuvant treatments are becoming rare and are mostly dependent on the identification of patients responding to available treatments [2],Gene expression profiling from tumor biopsies classified breast cancer into different molecular subtypes with distinct features and clinical outcomes: luminal A, luminal B, HER2-enriched, basal-like, claudin-low and normal-like subtypes [1], This molecular classification contributed to refining the diagnosis and selection of the most appropriate therapy. More recently, gene expression signatures (for example commercialized as Mammaprint, Oncotype DX, Prosigna) have been used, in conjunction with clinical-pathological parameters, to predict risk of recurrency and response to chemotherapy with the intent to avoid overtreatment of patients with good prognosis. These signatures, however, cannot be applied for early detection or follow up of patients as they are based on the analysis of tumor tissue [3],
[0007] The current screening test for breast cancer is mammography, which is usually made once every 2 years starting mostly at the age of 50. The main limitation of mammography is that it is based on morphology, hence on the detection of a sizable tumor. Additional imaging-based techniques (i.e. computer tomography (CT), Magnetic Resonance Imaging (MRI), Sonography) may be used in complement to mammography [6], A biopsy is then necessary for formal diagnosis. In some countries, the necessary screening infrastructure is limited, so that mammography is seldomly applied, rendering the early detection difficult. Mammography particularly detects ductal carcinoma in situ (DCIS) or lobular carcinoma in situ (LCIS), non-invasive cancers, leading to over-treatments: for 1000 women screened over 20 years only 5 deaths are prevented while 17 patients are overtreated [5], Because of overtreatment the benefit of mammography is currently under scrutiny. On average, mammography has a sensitivity of 60% (60-90%) and a specificity of 80% (80-95%), with large variations. Thus, mammography misses up to 40% of cancers resulting in interval cancers, and up to 80% false positives [7],
[0008] Alternative methods not based on morphology are under evaluation, including the detection of circulating tumor cells (CTC) or genetic material such as free circulating tumoral DNA (ctDNA), mRNA, micro RNA (miRNA), other RNAs in the systemic circulation (“liquid biopsy”). These approaches require the presence of a significant tumor mass releasing material into the systemic circulation, and at sufficient amounts thus allowing detection with current techniques. ctDNA-based approaches are based on the detection of cancer-associated mutations, copy number variations or epigenetic modifications (e.g. methylation) specific for a given cancer, thereby limiting sensitivity and possibly specificity for early detection purposes [8], The interest in CTC to develop clinical tests has been hampered due to technical difficulties in their detection (very rare events; limited cell surface markers; multistep technologies) and their rather late presence during disease progression. None of these approaches have reached clinical routine yet. Basically, all the current blood-based tests are based on markers generatedby the tumor itself, and thus dependent on already developed and vascularized tumors. Thus, the detection of microscopic primary tumor or metastatic lesions remains a challenge and an unmet need in oncology.
[0009] Several prior art documents have attempted to address the detection or prognosis of breast cancer using molecular biomarkers. For example, WO 2008 / 077165 A1 (ARC AUSTRIA RESEARCH CENTERS GMBH [AT]; LAUSS MARTIN [AT] et al.) [see also 11] describes a method for detecting breast cancer based on the measurement of a large panel of gene expression markers, including CYBRD1, mainly in tumor tissue. The approach is prognostic rather than diagnostic, and does not teach or suggest the specific combination of PGPEP1 , CYBRD1, and OSBPL8, nor its application to early, non-invasive detection in blood.
[0010] XP21116189A - Deepthi Telikicherla et al. : "Overexpression of ribosome binding protein 1 (RRBP1) in breast cancer", Clinical Proteomics, vol. 9, n° 1, 18 juin 2012 (2012-06-18), page 7, ISSN : 1559-0275, DOI : 10.1186 / 1559-0275-9-7 identifies PGPEP1 as a marker overexpressed in breast cancer tissue, but does not disclose its use in blood or in combination with CYBRD1 and OSBPL8 for early detection. XP93301887A -"Gene: CYBRD1 (ENSG00000071967) -", Ensembl database, 1 May 2025 (2025-05-01), and WO 2022 / 152911 A1 (Universite de Fribourg) [see also 4, 11] disclose other gene signatures or detection methods, including blood-based approaches, but none achieve the high diagnostic accuracy or enable early-stage, non-invasive detection as provided by the present invention. WO 2012 / 116248 A1 (MASSACHUSETTS INSTITUTE OF TECHNOLOGY [US]; EINSTEIN COLLEGE OF MEDICINE [US] et al.) describes further molecular approaches but does not disclose the present invention’s combination or its performance.
[0011] While blood-based gene expression signatures have been proposed (see, for example, WO 2008 / 077165 A1 and WO 2022 / 152911 A1), none have demonstrated the combination of high sensitivity, specificity, and early-stage detection achieved by the present invention. In particular, the prior art does not disclose or suggest the synergistic combination of PGPEP1, CYBRD1 , and OSBPL8 measured in blood, nor does it achieve an area under the ROC curve (AUC) of at least 0.90 for the detection of breast cancer, including early-stage disease.
[0012] There remains a need for a minimally invasive, highly sensitive and specific method for the early detection of breast cancer, in particular at early stages namely 0-I-II based on the TNM classification (T 1 / T2), which overcomes the limitations of imaging-based screening and tissuebased molecular signatures, and which is suitable for routine screening and follow-up.BRIEF DESCRIPTION OF THE INVENTION
[0013] The present invention is partially based upon the discovery that a small panel of biomarkers in the blood can specifically identify and distinguish subjects with breast cancer from subjects without such lesions. Accordingly, the invention provides the unique advantages associated with early detection of breast cancer or tumor in the patient, including less aggressive and more effective therapies, increased life span, decreased morbidity and mortality, decreased exposure to radiation during screening and repeated screenings and a minimally invasive diagnostic model. Importantly, the methods of the invention allow for a patient to avoid invasive procedures, thus increasing patient's compliance.
[0014] The invention mainly uses blood samples to detect biomarkers generated by the immune system when cancer cells are encountered. Surprisingly, using the reaction of the organism toward cancer, instead of signals created by the cancer itself, enables early and sensitive detection of the disease.
[0015] The invention relates to the detection of breast cancer in a patient by measuring the variations of expression levels of genetic elements (transcriptome) within circulating immune cells (leukocytes) from the blood of a patient with breast cancer at an early stage of development.
[0016] In summary, the invention provides a method for early and minimally-invasive detection of breast cancer in a patient, comprising: (a) measuring in a biological sample obtained from the patient the expression level of transcriptomic (mRNA) biomarkers of a panel comprising the combination of at least three genes consisting of PGPEP1, CYBRD1, and OSBPL8; (b) calculating a probability score based on the measurement of step (a), wherein the probability score P is calculated by the formula:
[0017]
[0018] where x_ni is a measured value for the biomarker n and subject i and (p0, Pi, ..., pn) is a vector of coefficients with p0a panel-specific constant, and pn is the corresponding logistic regression coefficient of the biomarker n, (c) ruling out breast cancer for the patient if the score in step (b) is lower than a pre-determined threshold, or ruling in the likelihood of breast cancer for the patient if the score in step (b) is higher than a pre-determined threshold, wherein a probability score value Pyi closer to 1.0 represents the likelihood of breast cancer for said patient and a probability score value Pyi closer to 0.0 represents the likelihood that said patient has not developed breast cancer, and wherein the combination of PGPEP1, CYBRD1, and OSBPL8provides a diagnostic performance with an area under the ROC curve (AUC) of at least 0.90, said performance being significantly superior to any combination of only two of these genes, and wherein the method enables detection of early-stage (T1 / T2) breast cancer.
[0019] The present invention also provides a kit for the detection of breast cancer in a patient from a biological sample, preferably a blood sample, comprising at least one probe or primer for measuring the expression level of each of PGPEP1, CYBRD1, and OSBPL8, and instructions for calculating a probability score as defined above. The biomarker panel in the kit can be advantageously completed with one or more genes selected from the group consisting of PSD4, GALK1, IPMK, GAPDH, HNRNPK, CGGBP1, NRDC, THRA, VMP1, ADD1, ARHGAP9, SLC43A3, IGF1R, B3GNTL1 , WWC3, and TOM1L2.
[0020] Another object of the invention is a method for monitoring disease progression or response to therapy in a breast cancer patient, comprising measuring the expression levels of the same panel of biomarkers in biological samples obtained at different time points, calculating the probability score as described above, and comparing the scores to assess disease progression or response to therapy.
[0021] Yet a further object of the invention, is to provide a method for early and minimally-invasive detection of primary breast cancer in a patient, comprising measuring the expression levels of the same panel of biomarkers, calculating the probability score as described above, and ruling out or ruling in primary breast cancer for the patient based on a pre-determined threshold, wherein a probability score value Pyi closer to 1.0 represents the likelihood of primary breast cancer for said patient and a probability score value Pyi closer to 0.0 represents the likelihood that said patient has not developed primary breast cancer.
[0022] Other objects and advantages of the invention will become apparent to those skilled in the art from a review of the ensuing detailed description, which proceeds with reference to the following illustrative drawings, and the attendant claims.
[0023] BRIEF DESCRIPTION OF THE FIGURES
[0024] Figure 1: is a heatmap representing all genes that are differentially expressed between Healthy donors and Primary Breast cancer patients using sequencing data, showing that the transcriptomic of the two groups is highly different.
[0025] Figure 2 is a heatmap of the genes of interest only, based on the sequencing data, showing the difference of expression level between the two groups.Figure 3 is a heatmap representing the expression level of all genes of interest between Healthy donors and Primary Breast cancer patients using PCR data.
[0026] Figure 4 represents A) AUC / SPEC I SENSI density graph and B) ROC curve of the presented signature in Healthy donors and Primary Breast cancer patients using PCR data, showing the very good performance of the signature combination to separate the two groups.
[0027] Figure 5 shows the significant performance of the Z-Scoring of the combined genes of the signature in both Healthy donors and Primary Breast cancer patients groups using PCR data.
[0028] Figure 6 is a heatmap representing the level of expression of the 4 genes composing the claim from WO 2022 / 152911 A1 between Healthy donors and Primary Breast cancer patients of the GENOA study, using sequencing data.
[0029] Figure 7 illustrates the protein signature obtained with the 4 main protein biomarkers used according to WO 2022 / 152911 A1, showing that these proteins are not able to separate Healthy donors and Primary Breast cancer patients.
[0030] Figure 8: represents the heatmap of the mandatory gene expression level by PCR in the different patient group (P = primary breast cancer; H = healthy donors), showing differential gene expression of the core genes of the signature in healthy donors and Primary cancer patients.
[0031] DETAILED DESCRIPTION OF THE INVENTION
[0032] The present invention mainly uses blood samples to detect biomarkers generated by the immune system when cancer cells are encountered. One object of the invention relates to a method to detect and predict breast cancers at any stage of the disease, including early-stage breast cancer. The invention also discloses a breast cancer detection kit directed to the detection of breast cancer in a human patient, preferably a female patient.
[0033] One advantage of the present invention is, for example, to identify early-stage or primary breast cancer in a patient by means of expression markers.
[0034] Surprisingly, the invention provides for a method and a kit for the identification of early-stage primary breast cancer, by means of expression profiling: (1.) the transcriptomic expression level of a marker panel of at least 3 genes PGPEP1, CYBRD1, and OSBPL8 is measured, where subsequently a probability score is calculated based on the measurement of the transcriptomic expression of the marker panel, and where a pre-determined score decides onwhether breast cancer or breast cancer can be ruled out or in. The probability score P is calculated by the formula:
[0035] LOG(P(y_i=1) / (1-P(y_i=1))) = p0+ Pi Xn + ... + pnxni,
[0036] where x_ni is a measured value for the biomarker n and subject i and (p0, Pi, ..., pn) is a vector of coefficients with p0a panel-specific constant, and pn is the corresponding logistic regression coefficient of the biomarker n. A probability score value Pyi closer to 1.0 represents the likelihood of breast cancer (or primary breast cancer) for said patient and a probability score value Pyi closer to 0.0 represents the likelihood that said patient has not developed breast cancer (or primary breast cancer).
[0037] According to the invention, the marker panel of these at least 3 genes (PGPEP1, CYBRD1, and OSBPL8) can be advantageously completed with one or more genes selected from the group comprising or consisting of PSD4, GALK1 , IPMK, GAPDH, HNRNPK, CGGBP1 , NRDC, THRA, VMP1, ADD1, ARHGAP9, SLC43A3, IGF1R, B3GNTL1, WWC3, and TOM1L2.
[0038] The combination of PGPEP1, CYBRD1, and OSBPL8 provides a diagnostic performance with an area under the ROC curve (AUC) of at least 0.90, said performance being significantly superior to any combination of only two of these genes, and enables detection of early-stage (T 1 / T2) breast cancer.
[0039] Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, suitable methods and materials are described below. All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety. The publications and applications discussed herein are provided solely for their disclosure prior to the filing date of the present application. Nothing herein is to be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. In addition, the materials, methods, and examples are illustrative only and are not intended to be limiting.
[0040] In the case of conflict, the present specification, including definitions, will control.
[0041] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as is commonly understood by one of skill in art to which the subject matter herein belongs. As used herein, the following definitions are supplied in order to facilitate the understanding of the present invention.The term “comprise” is generally used in the sense of include, that is to say permitting the presence of one or more features or components.
[0042] As used in the specification and claims, the singular forms “a”, “an” and “the” include plural references unless the context clearly dictates otherwise.
[0043] "Treating" or "treatment" as used herein with regard to a condition may refer to preventing the condition, slowing the onset or rate of development of the condition, reducing the risk of developing the condition, preventing or delaying the development of symptoms associated with the condition, reducing or ending symptoms associated with the condition, generating a complete or partial regression of the condition, or some combination thereof.
[0044] A "biomarker" used herein refers to a molecular indicator of a specific biological property; a biochemical feature or facet that can be used to detect breast cancer. "Biomarker" encompasses, without limitation, proteins, nucleic acids, and metabolites, together with their polymorphisms, mutants, isoform variants, related metabolites, derivatives, precursors including nucleic acids and pro-proteins, cleavage products, protein-ligand complexes, post-translationally modified variants (such as cross-linking or glycosylation), fragments, and degradation products, as well as any multi-unit nucleic acid, protein, and glycoprotein structures comprised of any of the biomarkers as constituent subunits of the fully assembled structure, and other analytes or sample-derived measures.
[0045] "Measuring", "measurement", "detection" and "detecting" refer to assessing the presence, absence, quantity or amount (which can be an effective amount) of either a given substance within a clinical or subject-derived sample, including qualitative or quantitative concentration levels of such substances, or otherwise evaluating the values or categorization of a patient's clinical parameters.
[0046] "Altered", "an increase" or "a decrease" refers to a detectable change or difference between the measured biomarker and the reference value from a reasonably comparable state, profile, measurement, or the like. One skilled in the art should be able to determine a reasonable measurable change. Such changes may be all or none. They may be incremental and need not to be linear. They may be by orders of magnitude. A change may be an increase or decrease by 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 100%, or more, or any value in between 0% and 100%. Alternatively, the change may be 1-fold, 1.5-fold, 2-fold, 3-fold, 4-fold, 5-fold or more, or any values in between 1-fold and five-fold. The change may be statistically significant with a p value of 0.1 , 0.05, 0.001 , or 0.0001.
[0047] A “patient” can be one who has not been previously diagnosed or identified as having breast tumor. A patient can be a healthy subject who is classified as low risk for developing a breast cancer condition. Alternatively, a patient can be one who has a risk of developing breast cancer or tumor. A risk factor is anything that affects the subject's chance of getting a disease such as breast cancer or tumor.
[0048] More exactly, a patient can be one who has been previously diagnosed with or identified as suffering from or having breast cancer, and optionally, but need not have already undergone treatment for the breast cancer. A patient can also be one who is not suffering from breast cancer. A patient can also be one who has been diagnosed with or identified as suffering from breast cancer, but who shows improvements in the disease (such as, for example, a decrease in tumor size) as a result of receiving one or more treatments for breast cancer. Alternatively, a patient can also be one who has not been previously diagnosed or identified as having breast cancer. For example, a patient can be one who exhibits one or more risk factors for breast cancer, or a subject who does not exhibit risk factors for breast cancer, or a subject who is asymptomatic for breast cancer. A patient can also be one who is suffering from or is at risk of developing breast cancer. Preferably the patient is a female patient.
[0049] By “asymptomatic” patient it is intended a patient showing no breast cancer symptoms. This definition may comprise the detection of an early breast cancer in a patient that is mammography negative. One advantage of the invention is the detection or diagnosis of breast cancer at an early stage.
[0050] However, the goal of the present invention is not limited to the detection of breast cancer in a patient having a negative mammography but also the detection of breast cancer independently from mammography and preferably before mammography becomes positive. This is also true for the screening test or kit according to the invention. The kit is used to detect a potential breast cancer in a patient who does not have any symptom of the disease.
[0051] Preferably, the patient has not been previously diagnosed or identified as having breast cancer, however said patient is at risk of developing breast cancer. Thus, the asymptomatic subject in that case is a patient carrying a breast cancer but who experiences no symptoms.
[0052] Preferably, the patient is a women at high risk for BC (frequent screening, starting early in life):
[0053] High risk definition (Lifetime risk :40-70% vs 12%):
[0054] -dense breast tissue.-previous high grade / premalignant lesions.
[0055] -BRCA1 / 2mut, familiarity, high polygenic risk score
[0056] -previous chest irradiation.
[0057] -hormonal replacement therapy (HRT / EPT), early menarche I late menopause.
[0058] -obesity & alcohol consumption.
[0059] -Socioeconomic & access disparities (migrants, lower education / income groups, rural residents).
[0060] A "biological sample" in the context of the present invention is a biological sample isolated from a patient and can include, by way of example and not limitation, whole blood, serum, plasma, blood cells, peripheral blood mononuclear cells, endothelial cells, circulating tumor cells, tissue biopsies, lymphatic fluid, ascites fluid, interstitial fluid, bone marrow, cerebrospinal fluid (CSF), saliva, mucous, sputum, sweat, urine, or any other secretion, excretion, or other bodily fluids. Preferably, the biological sample used is whole blood, from which white blood cells (leucocytes) are used.
[0061] In the context of the present invention, the collection of a peripheral blood sample is referred to as a “minimally invasive” procedure. This term distinguishes the invention from truly “non-invasive” methods (e.g., mammography, external imaging, ultrasound), which do not penetrate the skin or enter the body. A blood draw requires a simple venipuncture causing minimal physical trauma and discomfort, yet it remains a simple and safe procedure compared to more invasive techniques such as tissue biopsy or surgery.
[0062] The term "primer" refers to a strand of nucleic acid having usually between 8 and 25 nucleotides that serves as a starting point for DNA replication.
[0063] The terms "probe" and "hydrolysis probe" refer to a short strand of nucleic acid designed to hybridize to a region within the amplicon, and which is dual labeled with a reporter dye and a quenching dye. The close proximity of the quencher suppresses the fluorescence of the reporter dye. The probe relies on the 5'-3' exonuclease activity of Taq polymerase, which degrades a hybridized non-extendible DNA probe during the extension step of the polymerase chain reaction (PCR). Once the Taq polymerase has degraded the probe, the fluorescence of the reporter increases at a rate that is proportional to the amount of template present.
[0064] The term "gene expression" means the production of a protein or a functional mRNA from its gene.The terms "signature", "classifier", "model" and "predictor" are used interchangeably. They refer to an algorithm that discriminates between disease states with a predetermined level of statistical significance. A two-class classifier is an algorithm that uses data points from measurements from a sample and classifies the data into one of two groups. In certain embodiments, the data used in the classifier is the relative expression of nucleic acids or proteins in a biological sample. Protein or nucleic acid expression levels in a patient can be compared to levels in patients previously diagnosed as disease free or with a specified condition.
[0065] A "reference” or “baseline level / value" as used herein can be used interchangeably and is meant to be relative to a number or value derived from population studies, including without limitation, such subjects having similar age range, disease status (e.g., stage), patients in the same or similar ethnic group, or relative to the starting sample of a patient undergoing treatment for cancer. Such reference values can be derived from statistical analyses and / or risk prediction data of populations obtained from mathematical algorithms and computed indices of breast cancer. Reference indices can also be constructed and used utilizing algorithms and other methods of statistical and structural classification.
[0066] In some embodiments of the present invention, the reference or baseline value is the expression level of a particular biomarker of interest in a control sample derived from one or more healthy patients who have not been diagnosed with any breast cancer.
[0067] In some embodiments of the present invention, the reference or baseline value is the expression level of a particular biomarker of interest in a sample obtained from the same patient prior to any cancer treatment. In other embodiments of the present invention, the reference or baseline value is the expression level of a particular biomarker of interest in a sample obtained from the same patient during a cancer treatment. Alternatively, the reference or baseline value is a prior measurement of the expression level of a particular gene of interest in a previously obtained sample from the same patient or from a patient having similar age range, disease status (e.g., stage) to the tested patient.
[0068] The term "ruling out" as used herein is meant that the patient is selected not to receive a treatment protocol.The term "ruling in" as used herein is meant that the patient is selected to receive a treatment protocol.
[0069] The term "normalization" or "normalizer" as used herein refers to the expression of a differential value in terms of a standard value to adjust for effects which arise from technical variation due to sample handling, sample preparation and mass spectrometry measurement rather than biological variation of protein concentration in a sample. For example, when measuring the expression of a differentially expressed protein (nucleic acid), the absolute value for the expression of the protein (nucleic acid) can be expressed in terms of an absolute value for the expression of a standard protein (nucleic acid) that is substantially constant in expression. This prevents the technical variation of sample preparation and PCR measurement from impeding the measurement of protein (nucleic acid) concentration levels in the sample.
[0070] The term "score" or "scoring" refers to calculating a probability likelihood (or a probability value) by the model (e.g., a logistic regression model) for a sample. For the present invention, values closer to 1.0 are used to represent the likelihood that a sample is derived from a patient with a breast cancer condition, values closer to 0.0 represent the likelihood that a sample is derived from a patient without a breast cancer condition.
[0071] A "pre-determined score" refers to a probability threshold that has been determined during the modeling / training phase by, for instance, logistic regression and receiver operating characteristic ROC analysis, and that defines the likelihood of breast tumor and / or diagnosis of breast tumor. A skilled artisan can readily determine such score according to any methods available in the art.
[0072] “Transcriptomic (mRNA) markers or biomarkers” are specific mRNA molecules and their quantity present in a given cell or tissue or fluid that correlate with a particular cell population or tissue type, state of activation, differentiation, or function. In disease conditions, transcriptomic (mRNA) markers may have diagnostic, prognostic or predictive values.
[0073] As used herein, “primary breast cancer” refers to a breast cancer that has not metastasized to distant organs and is diagnosed at its initial site of origin in the breast tissue. Primary breast cancer includes both non-invasive (e.g., ductal carcinoma in situ, DCIS) and invasive forms (e.g., invasive ductal carcinoma, invasive lobular carcinoma) prior to the development of distant metastases.As used herein, “early-stage breast cancer” refers to breast cancer diagnosed at stage 0-1-11 based on the TNM classification, meaning tumors that are <2 cm (T 1) or >2 cm but <5 cm (T2) in greatest dimension, with or without limited lymph node involvement, and without distant metastasis.
[0074] As used herein, “monitoring” refers to the repeated measurement and evaluation of biomarker expression levels in biological samples from a patient over time, in order to assess disease progression, recurrence, or response to therapy. Monitoring may be performed at regular intervals during or after treatment.
[0075] As used herein, “disease progression” refers to the worsening of a patient’s condition, including an increase in tumor size, the appearance of new lesions, or the recurrence of cancer after initial treatment, as determined by clinical, imaging, or molecular criteria.
[0076] As used herein, “response to therapy” refers to a change in the patient’s disease status following treatment, including reduction in tumor size, decrease in biomarker levels, or other clinical or molecular evidence of therapeutic efficacy.
[0077] As used herein, “panel” refers to a defined set of two or more biomarkers (e.g., genes, transcripts, or proteins) whose combined measurement is used for diagnostic, prognostic, or monitoring purposes.
[0078] As used herein, “supervised machine learning algorithm” refers to any computational method that uses labeled training data to learn a predictive model, including but not limited to logistic regression, support vector machines, random forests, and neural networks.
[0079] One object of the invention is to provide a method for early and minimally invasive detection of breast cancer in a patient, comprising:
[0080] (a) measuring in a biological sample obtained from the patient the expression level of transcriptomic (mRNA) biomarkers of a panel comprising the combination of at least three genes consisting of PGPEP1 , CYBRD1 , and OSBPL8;
[0081] (b) calculating a probability score based on the measurement of step (a), wherein the probability score P is calculated by the formula:
[0082]
[0083] where x_ni is a measured value for the biomarker n and subject i and (p0, Pi, pn) is a vector of coefficients with p0a panel-specific constant, and pn is the corresponding logistic regression coefficient of the biomarker n,
[0084] (c) ruling out breast cancer for the patient if the score in step (b) is lower than a pre-determined threshold, or ruling in the likelihood of breast cancer for the patient if the score in step (b) is higher than a pre-determined threshold,
[0085] wherein a probability score value Pyi closer to 1.0 represents the likelihood of breast cancer for said patient and a probability score value Pyi closer to 0.0 represents the likelihood that said patient has not developed breast cancer,
[0086] and wherein the combination of PGPEP1, CYBRD1, and OSBPL8 provides a diagnostic performance with an area under the ROC curve (AUC) of at least 0.90, said performance being significantly superior to any combination of only two of these genes, and wherein the method enables detection of early-stage (T 1 / T2) breast cancer.
[0087] According to the invention, the biomarker panel of these at least three genes (PGPEP1, CYBRD1, and OSBPL8) can be advantageously completed with one or more genes selected from the group consisting of PSD4, GALK1, IPMK, GAPDH, HNRNPK, CGGBP1, NRDC, THRA, VMP1, ADD1, ARHGAP9, SLC43A3, IGF1R, B3GNTL1, WWC3, and TOM1L2.
[0088] The probability score can be calculated according to any method known in the art. For example, the probability score may be calculated using a logistic regression model or another supervised machine learning algorithm applied to the measurement by an algorithm.
[0089] In some embodiments, the method is validated on an independent cohort using RT-qPCR.
[0090] In some embodiments, the method is used for screening asymptomatic women at risk of developing primary breast cancer.
[0091] In some embodiments, the likelihood of breast cancer is further determined by the sensitivity, specificity, negative predictive value (NPV), or positive predictive value (PPV) associated with the score.
[0092] In some embodiments, the method is suitable for detecting invasive breast cancer at stage I or II.In some embodiments, the method is suitable for distinguishing breast cancer patients from healthy controls with a sensitivity of at least 85% and a specificity of at least 85%.
[0093] Preferably, the biological sample is selected from the group consisting of whole blood, peripheral blood mononuclear cells, leukocytes, serum, plasma, or any other blood-derived sample, but may also include circulating tumor cells, lymphatic fluid, bone marrow, cerebrospinal fluid (CSF), saliva, or any other patient’s body fluid.
[0094] The method according to the invention may be used alone or in combination with other diagnostic modalities, such as imaging or additional biomarker panels, to further improve diagnostic accuracy.
[0095] According to one embodiment, said breast cancer is a carcinoma.
[0096] According to a further embodiment of the invention, the detection is an early detection of breast cancers at any stage of the disease.
[0097] In some embodiment, the patient is at risk of developing primary breast cancer.
[0098] In an embodiment, when breast cancer is ruled out, the patient does not receive a treatment protocol.
[0099] In another embodiment, when breast cancer is ruled in, the patient is to receive a treatment protocol. Preferably, said treatment protocol is a biopsy, a surgery, a chemotherapy, a radiotherapy, a hormonal therapy, a targeted therapy, an immunotherapy, or any combination thereof.
[0100] Advantageously, the patient subject can be treated by one of the following therapeutic modalities, or combinations thereof, referred herein as at least one breast cancer-modulating agent.
[0101] Surgical therapy: it consists in the removal of the tumor and some surrounding healthy tissue during an operation. Surgery is also used to examine the nearby axillary lymph nodes to determine disease spreading. Surgery can be performed as lumpectomy, the removal of the tumor and a small cancer-free margin of healthy tissue around the tumor. Mastectomy, the removal of the entire breast.Radiation therapy: it consists in the use of high-energy x-rays or other particles to destroy cancer cells and lower the risk of recurrence in the breast. Radiotherapy can be applied as external-beam therapy, given from a machine outside the body; as intra-operatively, when radiation treatment is given using a probe in the operating room, as brachytherapy, when given by placing radioactive sources into the tumor.
[0102] Chemotherapy: it consists in the use of drugs to kill cancer cells or keeping the cancer cells from dividing and growing. It may be given before surgery to shrink a large tumor, to make surgery easier, and / or to reduce the risk of recurrence. This is called neoadjuvant chemotherapy. Chemotherapy can be given after surgery to reduce the risk of recurrences. It can be given once a week, once every 2 weeks, once every 3 weeks, or once every 4 weeks. There are many types of chemotherapy used to treat breast cancer. Exemplary breast cancerchemotherapy agents include, but are not limited to: Docetaxel (Taxotere); Paclitaxel (Taxol); Doxorubicin; Epirubicin (Ellence); Pegylated liposomal doxorubicin (Doxil); Capecitabine (Xeloda); Carboplatin (available as a generic drug); Cisplatin (available as a generic drug); Cyclophosphamide (available as a generic drug); Eribulin (Halaven); Fluorouracil (5-Fll); Gemcitabine (Gemzar); Ixabepilone (Ixempra); Methotrexate (Rheumatrex, Trexall); Proteinbound paclitaxel (Abraxane); Vinorelbine (Navelbine).
[0103] Hormonal therapy: it consists in the administration of agents suppressing hormonal signalling. It is an effective treatment for most tumors expressing oestrogen and progesterone receptors (ER+PR+). It can be given before surgery, typically for at least 3 to 6 months, and or after surgery and continued for 5 to 10 years. Blocking the hormones can help prevent a cancer recurrence and death from breast cancer when used either alone or after chemotherapy. Exemplary hormonal therapy agents include, but are not limited to: Tamoxifen (Novaldex) blocks oestrogen from binding to its receptor; Aromatase inhibitors, including anastrozole (Arimidex), exemestane (Aromasin), and letrozole (Femara) inhibit synthesis of estrogen; Ovarian suppression using gonadotropin or luteinizing releasing hormone (GnRH or LHRH) agonist to stop the ovaries from making oestrogen such as Goserelin (Zoladex) and leuprolide (Eligard, Lupron); Ovarian ablation, the surgical removal of the ovaries.
[0104] HER2-targeted therapy: it is used to treat HER2 breast cancers. Many HER2 inhibitors are available including, but no limited to: Trastuzumab (Herceptin); Pertuzumab (Perjeta), Neratinib (Nerlynx), Ado-trastuzumab emtansine or T-DM1 (Kadcyla) are antibodies given intravenously. Lapatinib (Tykerb), Tucatinib (Tukysa) are orally available HER2 tyrosine kinase inhibitors.Other targeted drugs used in breast cancer therapy include: Olaparib (Lynparza), PARP inhibitor; abemaciclib (Verzenio), palbociclib (Ibrance), and ribociclib (Kisqali), CDK4 / 6 inhibitors; PI3K inhibitor Alpelisib (Piqray); Sacituzumab govitecan-hziy (Trodelvy); Entrectinib (Rozyltrek) and larotrectinib (Vitrakvi); Talazoparib (Talzenna).
[0105] Immunotherapy: it is used to stimulate the immune system to attack cancer cells, using drugs called immune checkpoint inhibitors. The following drugs are used for recurrent, advanced or metastatic breast cancer: Pembrolizumab (Keytruda), PD-L1 inhibitor; Ipilimumab (Yervoy), CTLA-4 inhibitor; Dostarlimab (Jemperli), PD-1 inhibitor.
[0106] Also provided is a kit for the detection of breast cancer in a patient from a biological sample, preferably a blood sample. The kit comprises at least one probe or primer for measuring the expression level of each of PGPEP1, CYBRD1, and OSBPL8, and instructions for calculating a probability score as defined in the present invention.
[0107] The kit is designed to be used with a method for early and minimally invasive detection of breast cancer, wherein the probability score is calculated based on the measurement of the transcriptomic (mRNA) biomarkers of the panel comprising the combination of at least three genes consisting of PGPEP1, CYBRD1, and OSBPL8. The probability score P is calculated by the formula:
[0108] LOG(P(y_i=1) / (1-P(y_i=1))) = p0+ Pi Xn + ... + pnxni,
[0109] where x_ni is a measured value for the biomarker n and subject i and (p0, Pi, ..., pn) is a vector of coefficients with p0a panel-specific constant, and pn is the corresponding logistic regression coefficient of the biomarker n. A probability score value Pyi closer to 1.0 represents the likelihood of breast cancer for said patient and a probability score value Pyi closer to 0.0 represents the likelihood that said patient has not developed breast cancer.
[0110] In some embodiments, the kit further comprises at least one probe or primer for measuring the expression level of one or more additional genes selected from the group consisting of PSD4, GALK1, IPMK, GAPDH, HNRNPK, CGGBP1, NRDC, THRA, VMP1, ADD1, ARHGAP9, SLC43A3, IGF1R, B3GNTL1 , WWC3, and TOM1L2.
[0111] The kit may further comprise reagents for RNA extraction, reverse transcription, and quantitative PCR (RT-qPCR), as well as positive and negative control samples, and instructions for use. The kit may also include software or access to an online platform forcalculating the probability score and interpreting the results according to the thresholds defined in the present invention.
[0112] The kit is suitable for use in the early and non-invasive detection of breast cancer, for distinguishing early-stage (T1 / T2) breast cancer from healthy controls, and for monitoring disease progression or response to therapy in breast cancer patients.
[0113] Preferably, the kit further comprises one or more probes, reference samples for performing measurement quality controls.
[0114] In practice, a score, indicative of primary breast cancer, can be calculated based on the results obtained from the kit of the invention.
[0115] Ideally, the kit of the invention further comprises one or more plastic containers and reagents for performing test reactions and optionally instructions for use.
[0116] The actual measurement of levels of the biomarkers of the invention, namely the transcriptomic (mRNA) markers, can be determined at the nucleic acid or protein level using any method known in the art. For example, at the nucleic acid level, the biomarkers can be measured by extracting ribonucleic acids from the sample and performing any type of quantitative PCR on the reverse-transcribed nucleic acids. Another way to detect the biomarkers can also be by a whole transcriptome analysis based on high-throughput sequencing methodologies, e.g., RNA-seq, or hybridization-based technologies including microarrays.
[0117] Another way to detect the biomarkers can also be by a whole transcriptome analysis based on high-throughput sequencing methodologies, e.g., RNA-seq.
[0118] Ideally, the kit of the invention further comprises one or more plastic container and reagents for performing test reactions and optionally instructions for use.
[0119] In case breast cancer is ruled in, the subject is to receive a treatment protocol. Preferably, said treatment protocol is a biopsy, a surgery, a chemotherapy, a radiotherapy, or any combination thereof.
[0120] In another aspect, the invention provides a method for monitoring disease progression or response to therapy in a breast cancer patient. This method comprises:(a) measuring in a biological sample obtained from the patient the expression level of transcriptomic (mRNA) biomarkers of a panel comprising the combination of at least three genes consisting of PGPEP1 , CYBRD1 , and OSBPL8;
[0121] (b) calculating a probability score based on the measurement of step (a), wherein the probability score P is calculated by the formula:
[0122] LOG(P(y_i=1) / (1-P(y_i=1))) = p0+ Pi Xn + ... + pnxni,
[0123] where x_ni is a measured value for the biomarker n and subject i and (p0, Pi, ..., pn) is a vector of coefficients with p0a panel-specific constant, and pn is the corresponding logistic regression coefficient of the biomarker n;
[0124] (c) comparing the probability score obtained at different time points to assess disease progression or response to therapy,
[0125] wherein a probability score value Pyi closer to 1.0 represents the likelihood of breast cancer for said patient and a probability score value Pyi closer to 0.0 represents the likelihood that said patient has not developed breast cancer.
[0126] This method may be used to monitor patients during or after treatment, to detect recurrence, or to evaluate the effectiveness of a therapeutic intervention. The method can be applied to any of the biological samples described herein, and may be validated using RT-qPCR or other suitable techniques.
[0127] In a further aspect, the invention provides a method for early and minimally invasive detection of primary breast cancer in a patient, comprising:
[0128] (a) measuring in a biological sample obtained from the patient the expression level of transcriptomic (mRNA) biomarkers of a panel comprising the combination of at least three genes consisting of PGPEP1 , CYBRD1 , and OSBPL8;
[0129] (b) calculating a probability score based on the measurement of step (a), wherein the probability score P is calculated as described above;
[0130] (c) ruling out primary breast cancer for the patient if the score in step (b) is lower than a predetermined threshold, or ruling in the likelihood of primary breast cancer for the patient if the score in step (b) is higher than a pre-determined threshold,wherein a probability score value Pyi closer to 1.0 represents the likelihood of primary breast cancer for said patient and a probability score value Pyi closer to 0.0 represents the likelihood that said patient has not developed primary breast cancer.
[0131] In some embodiments, the method is suitable for screening asymptomatic women at risk of developing primary breast cancer, for distinguishing primary breast cancer patients from healthy controls, or for detecting early-stage (T1 / T2) primary breast cancer.
[0132] In all aspects, the method may further comprise measuring the expression of one or more additional genes selected from the group consisting of PSD4, GALK1, IPMK, GAPDH, HNRNPK, CGGBP1, NRDC, THRA, VMP1, ADD1, ARHGAP9, SLC43A3, IGF1R, B3GNTL1, WWC3, and TOM1L2. The probability score may be calculated using a logistic regression model or another supervised machine learning algorithm. The likelihood of breast cancer or primary breast cancer may also be determined by the sensitivity, specificity, negative predictive value (NPV), or positive predictive value (PPV) associated with the score. The method may be validated on an independent cohort using RT-qPCR.
[0133] The biological sample is preferably selected from the group consisting of whole blood, peripheral blood mononuclear cells, leukocytes, serum, plasma, or any other blood-derived sample, but may also include circulating tumor cells, lymphatic fluid, bone marrow, cerebrospinal fluid (CSF), saliva, or any other patient’s body fluid.
[0134] The method and kit of the invention are suitable for use in population screening, differential diagnosis of breast cancer versus other conditions, and longitudinal monitoring of patients during and after therapy.
[0135] The actual measurement of levels of the biomarkers of the invention, namely the transcriptomic (mRNA) markers, can be determined at the nucleic acid or protein level using any method known in the art. For example, at the nucleic acid level, the biomarkers can be measured by extracting ribonucleic acids from the sample and performing any type of quantitative PCR on the reverse-transcribed nucleic acids. Another way to detect the biomarkers can also be by a whole transcriptome analysis based on high-throughput sequencing methodologies, e.g., RNA-seq, or hybridization-based technologies including microarrays.
[0136] By way of example, other methods that can be used for measuring the biomarkers may involve any other method of quantification known in the art of nucleic acids, such as, but not limited to, amplification of specific sequences, oligonucleotide probes, hybridization of target genes withcomplementary probes, fragmentation by restriction endonucleases and study of the resulting fragments (polymorphisms), pulsed field gels techniques, isothermic multiple-displacement amplification, rolling circle amplification or replication, immuno-PCR, among others known to those skilled in the art.
[0137] By using information provided by database entries for the biomarker sequences, biomarker expression levels can be detected and measured using techniques well known to one of ordinary skill in the art. For example, biomarker sequences within the sequence database entries, or within the sequences disclosed herein, can be used to construct probes and primers for detecting biomarker mRNA sequences in methods which specifically, and, preferably, quantitatively amplify specific nucleic acid sequences such as reverse-transcription based realtime polymerase chain reaction (RT-qPCR).
[0138] Those skilled in the art will be familiar with numerous specific immunoassay and nucleic acid amplification assay formats and variations thereof which may be useful for carrying out the embodiments of the invention disclosed herein.
[0139] Preferably, expression levels of the biomarkers of the present invention are detected by RT-qPCR, and in particular by real-time PCR, as described further herein.
[0140] In general, total RNA can be isolated from the target sample, such as peripheral blood, PBMC, or any leukocyte population, using any isolation procedure. This RNA can then be used to generate first strand copy DNA (cDNA) using any procedure, for example, using random primers, oligo-dT primers or random-oligo-dT primers which are oligo-dT primers coupled on the 3'-end to short stretches of specific sequence covering all possible combinations. The cDNA can then be used as a template in quantitative PCR.
[0141] In real-time PCR, the quantification of PCR products relies, for example, on increases in fluorescence, released at each amplification cycle of the reaction, for example, by a probe that hybridizes to a portion of the amplification product. Fluorescence approaches used in real-time quantitative PCR are typically based on a fluorescent reporter dye such as FAM, fluorescein, HEX, TET, etc. and a quencher such as TAMRA, DABSYL, Black Hole, etc. When the quencher is separated from the probe during the extension phase of PCR, the fluorescence of the reporter can be measured. Systems like Universal ProbeLibrary, Molecular Beacons, Taqman Probes, Scorpion Primers or Quantinova kits and probes and others use this approachto perform real-time quantitative PCR. Alternatively, fluorescence can be measured from DNA-intercalating fluorochromes such as Sybr Green.
[0142] The abundance of target RNA molecules can be performed by real-time PCR in a relative or absolute manner. Relative methods can be based on the threshold cycle determination (Ct) or, in the case of the Roche's PCR instruments, the crossing point (Cp). Relative RNA molecule abundance is then calculated by the delta Ct (delta Cp) method by subtracting Ct (Cp) value of one or more housekeeping genes. Alternatively, absolute measurements can be performed by determining the copy number of the target RNA molecule by the mean of standard curves.
[0143] The application of logistic regression to biological problems is routine in the art. Various statistical analysis software can be used for building logistic regression models. Fitted logistic regression models are tested by asking whether the model can correctly predict the clinical outcome using patient data other than that with which the logistic regression model was fitted but having a known clinical outcome. After training, the model output from 0 (control) to 1 (cancer) can be calculated in blind fashion by the average error of all N predictions (a validation group). Based on the output values, the receiver operating characteristic (ROC) curve can be built to calculate the outcome of clinical prediction: specificity and sensitivity of breast cancer detection. They are statistical measures of the performance of a binary classification test. Sensitivity measures the proportion of actual positives which are correctly identified as such (e.g., the percentage of patients who are correctly identified as having the condition). Specificity measures the proportion of negatives which are correctly identified (e.g., the percentage of healthy women who are correctly identified as not having the condition). A perfect predictor would be described as 100% sensitive (i.e. , predicting all patients from the sick group as sick) and 100% specific (i.e., not predicting anyone from the healthy group as sick). However, any predictor will possess a minimum error bound.
[0144] One embodiment of the present disclosure is a predictive model comprising a combination / profile of peripheral blood mononuclear cell biomarkers detecting breast tumours preferably with sensitivity equal to or above 60%, preferably equal to or above 70%, more preferably equal to or above 80% and even more preferably equal to or above 85% and specificity equal to or above 84%, preferably equal to or above 85%.
[0145] The term "sensitivity of a test" refers to the probability that a test result will be positive when the disease is present in the patient (true positive rate). This is derived from the number of patients with the disease who have a positive test result (true positive) divided by the totalnumber of patients with the disease, including those with true positive results and those patients with the disease who have a negative result, i.e. , false negative.
[0146] The term "specificity of a test" refers to the probability that a test result will be negative when the disease is not present in the patient (true negative rate). This is derived from the number of patients without the disease who have a negative test result (true negative) divided by all patients without the disease, including those with a true negative result and those patients without the disease who have a positive test result, e.g. false positive. While the sensitivity, specificity, true or false positive rate, and true or false negative rate of a test provide an indication of a test's performance, e.g. relative to other tests, to make a clinical decision for an individual patient based on the test's result, the clinician requires performance parameters of the test with respect to a given population.
[0147] The term "positive predictive value" (PPV) refers to the probability that a positive result correctly identifies a patient who has the disease, which is the number of true positives divided by the sum of true positives and false positives.
[0148] The term "negative predictive value" (NPV) refers to the probability that a negative test correctly identifies a patient without the disease, which is the number of true negatives divided by the sum of true negatives and false negatives. Like the PPV, it also is inherently impacted by the prevalence of the disease and pre-test probability of the population intended to be tested. A positive result from a test with a sufficient PPV can be used to rule in the disease for a patient, while a negative result from a test with a sufficient NPV can be used to rule out the disease, if the disease prevalence for the given population, of which the patient can be considered a part, is known.
[0149] A "Receiver Operating Characteristics (ROC) curve" as used herein refers to a plot of the true positive rate (sensitivity) against the false positive rate (specificity) for a binary classifier system as its discrimination threshold is varied. A ROC curve can be represented equivalently by plotting the fraction of true positives out of the positives (TPR=true positive rate) versus the fraction of false positives out of the negatives (FPR=false positive rate). Each point on the ROC curve represents a sensitivity / specificity pair corresponding to a particular decision threshold. AUC (Area Under the Curve) represents the area under the ROC curve. The AUC is an overall indication of the diagnostic accuracy of 1) a biomarker or a panel of biomarkers and 2) a ROC curve. AUC is determined by the "trapezoidal rule." For a given curve, the data points are connected by straight line segments, perpendiculars are erected from the abscissa to eachdata point, and the sum of the areas of the triangles and trapezoids so constructed is computed. In certain embodiments of the methods provided herein, a biomarker protein has an AUC in the range of about 0.75 to 1.0. In certain of these embodiments, the AUC is in the range of about 0.8 to 0.85, 0.9 to 0.95, or 0.95 to 1.0.
[0150] The methods provided herein are minimally invasive and pose little or no risk of adverse effects. As such, they may be used to diagnose, monitor and provide clinical management of patients who do not exhibit any symptoms of a breast cancer condition, and subjects classified as low risk for developing a breast cancer condition. Similarly, the methods disclosed herein may be used as a strictly precautionary measure to diagnose healthy patients who are classified as low risk for developing a breast cancer condition.
[0151] The present invention is based on the quantitative detection of multiple transcripts (mRNA) in peripheral blood leukocytes, their association and analysis.
[0152] Biomarkers are modulated by the presence of a breast cancer at any stage.
[0153] Molecular techniques (including but not limited to, Polymerase Chain Reaction - PCR, RNASeq, hybridization) are used to detect mRNA.
[0154] Results are analyzed by an algorithm calculating a score to identify patients, preferably with breast cancer at any stage, from patients without breast cancer (diagnostic purpose) or to stratify patients with relapsing breast cancer from patients with non-relapsing breast cancer after initial therapy (prognostic purpose).
[0155] Patients with a positive result indicating the presence of a breast cancer or a breast cancer relapse will be subjected to further diagnostic validation steps and treated accordingly.
[0156] The invention requires minimally invasive sample collection (peripheral blood) processes than tumor biopsy, which is considered nowadays as the gold-standard and required for definitive diagnosis.
[0157] The sensitivity of Applicant’s approach is better compared to tumor-derived markers, as is it is based on a natural amplification reaction of the immune system elicited by the presence of a minute number of cells, resulting in the generation of a measurable signal revealed by the detection kit according to the invention. The present invention leads to a more sensitive detection compared to detection via signals directly released by cancer cells. The cancer canbe detected earlier thereby increasing likelihood of a successful treatment, improved survival and quality of life.
[0158] Besides, the invention can be combined with existing methods of detection in particular with those based on tumor-derived material.
[0159] The list of the transcriptomic biomarkers (mRNA) of the invention is given in Table 1:
[0160] Group 1
[0161] Gene
[0162]
[0163] Group 3
[0164] Gene
[0165]
[0166]
[0167] Table 1
[0168] It is of interest that the method and kit of the invention detect all primary breast cancer subtypes at the earliest possible stages (non-invasive, minimally invasive, size < 2 cm).
[0169] Early detection of primary breast cancer is associated with decreased breast cancer-related mortality as it allows to remove the lesion as a smaller size and start an adjuvant therapy, if necessary, thereby reducing the risk of local invasion and metastatic spreading.
[0170] Thus, in one aspect, the invention provides a method for early and minimally invasive detection of primary breast cancer in a patient, comprising:
[0171] (a) measuring in a biological sample obtained from the patient the expression level of transcriptomic (mRNA) biomarkers of a panel comprising the combination of at least three genes consisting of PGPEP1 , CYBRD1 , and OSBPL8;
[0172] (b) calculating a probability score based on the measurement of step (a), wherein the probability score P is calculated by the formula:
[0173] LOG(P(y_i=1) / (1-P(y_i=1))) = p0+ Pi Xn + ... + pnxni,
[0174] where x_ni is a measured value for the biomarker n and subject i and (p0, Pi, ..., pn) is a vector of coefficients with p0a panel-specific constant, and pn is the corresponding logistic regression coefficient of the biomarker n;(c) ruling out primary breast cancer for the patient if the score in step (b) is lower than a predetermined threshold, or ruling in the likelihood of primary breast cancer for the patient if the score in step (b) is higher than a pre-determined threshold,
[0175] wherein a probability score value Pyi closer to 1.0 represents the likelihood of primary breast cancer for said patient and a probability score value Pyi closer to 0.0 represents the likelihood that said patient has not developed primary breast cancer.
[0176] Preferably, the primary breast cancer detection method further comprises measuring the expression of one or more additional transcriptomic (mRNA) biomarkers selected from the group consisting of PSD4, GALK1 , IPMK, GAPDH, HNRNPK, CGGBP1 , NRDC, THRA, VMP1 , ADD1, ARHGAP9, SLC43A3, IGF1R, B3GNTL1, WWC3, and TOM1L2.
[0177] Also envisioned is the use of a primary breast cancer detection kit for the detection of primary breast cancer in a patient from a biological sample, preferably a peripheral blood sample, said kit comprising at least one probe or primer for measuring the expression level of transcriptomic (mRNA) biomarkers of a panel of at least three genes comprising the combination of PGPEP1 , CYBRD1, and OSBPL8. According to the invention, the marker panel in the kit can be advantageously completed with one or more genes selected from the group consisting of PSD4, GALK1, IPMK, GAPDH, HNRNPK, CGGBP1, NRDC, THRA, VMP1, ADD1, ARHGAP9, SLC43A3, IGF1R, B3GNTL1, WWC3, and TOM1L2. Preferably, this kit for use in the detection of primary breast cancer further comprises additional transcriptomic (mRNA) markers and cell surface biomarkers specific for the immune response elicited by breast cancer at its different stages as identified above.
[0178] In another aspect, the invention provides a method of treating breast cancer in a patient, the method comprising:
[0179] (a) measuring in a biological sample obtained from the patient the expression level of transcriptomic (mRNA) biomarkers of a panel comprising the combination of at least three genes consisting of PGPEP1 , CYBRD1 , and OSBPL8;
[0180] (b) calculating a probability score based on the measurement of step (a), wherein the probability score P is calculated as described above;
[0181] (c) ruling out breast cancer for the patient if the score in step (b) is lower than a pre-determined threshold, or ruling in the likelihood of breast cancer for the patient if the score in step (b) is higher than a pre-determined threshold,wherein a probability score value Pyi closer to 1.0 represents the likelihood of breast cancer for said patient and a probability score value Pyi closer to 0.0 represents the likelihood that said patient has not developed breast cancer;
[0182] (e) administering to said patient at least one breast cancer-modulating agent when the patient is identified as suffering from breast cancer by step (b).
[0183] Preferably, the marker panel further comprises one or more genes selected from the group consisting of PSD4, GALK1, IPMK, GAPDH, HNRNPK, CGGBP1, NRDC, THRA, VMP1, ADD1, ARHGAP9, SLC43A3, IGF1R, B3GNTL1, WWC3, and TOM1L2.
[0184] According to one embodiment, the reference value comprises an index value, a value derived from one or more breast cancer risk prediction algorithms or computed indices, a value derived from a patient not suffering from breast cancer, or a value derived from a patient diagnosed with or identified as suffering from breast cancer.
[0185] According to another embodiment, the patient comprises one who has been previously diagnosed as having breast cancer, one who has not been previously diagnosed as having breast cancer, or one who is asymptomatic for the breast cancer.
[0186] The invention can be applied to various patient populations, including women at increased genetic risk, younger or older age groups, and patients with a family history of breast cancer. High or increased genetic risk categories include:
[0187] -dense breast tissue.
[0188] -previous high grade / premalignant lesions.
[0189] -BRCA1 / 2mut, familiarity, high polygenic risk score
[0190] -previous chest irradiation.
[0191] -hormonal replacement therapy (HRT / EPT), early menarche I late menopause.
[0192] -obesity & alcohol consumption.
[0193] As a minimally invasive blood-based test, the invention poses minimal risk to patients and is compatible with current regulatory requirements for in vitro diagnostic devices.
[0194] An “effective amount” can be the total amount or levels of biomarkers that are detected in a sample, or it can be a “normalized” amount, e.g., the difference between biomarkers detected in a sample and background noise. Normalization methods and normalized values will differ depending on the method by which the biomarkers are detected.Advantages of the invention and technical effect
[0195] The present invention provides several significant advantages and technical effects over the prior art in the field of breast cancer detection and monitoring:
[0196] Minimal and synergistic gene signature:
[0197] The invention identifies a minimal panel of three transcriptomic biomarkers (PGPEP1, CYBRD1, and OSBPL8) whose combined measurement provides a synergistic effect, resulting in diagnostic performance that is significantly superior to any individual gene or pairwise combination. This synergy is demonstrated by the high sensitivity and specificity achieved only with the three-gene combination, as shown in the experimental examples. The use of a minimal three-gene panel facilitates standardization and automation in clinical laboratories, enabling rapid and cost-effective implementation. The calculation of a single probability score simplifies result interpretation for clinicians.
[0198] Early and minimally invasive detection:
[0199] The method enables the early detection of breast cancer, including primary and early-stage (O-ll-ll) disease, using a simple blood test. This approach is minimally invasive, and can be applied to asymptomatic patients or those at risk, thereby improving patient compliance and facilitating population screening.
[0200] High diagnostic accuracy:
[0201] The three-gene signature achieves an area under the ROC curve (AUC) of at least 0.90, with sensitivity and specificity both exceeding 85% in clinical cohorts. This level of accuracy is not attainable with any single gene or with previously described gene panels, as demonstrated by comparative studies with the prior art.
[0202] Robustness and reproducibility:
[0203] The signature has been validated across independent patient cohorts and using different analytical platforms (RNA sequencing and RT-PCR), ensuring its robustness and reproducibility in clinical practice.
[0204] The measurement of gene expression levels can be performed using various analytical platforms, including RT-qPCR, next-generation sequencing, or microarray technologies, providing flexibility for clinical laboratories.Technical superiority over prior art:
[0205] Unlike prior art documents that disclose large lists of candidate genes without teaching which combinations are optimal, the present invention provides a specific, minimal, and validated combination that is both necessary and sufficient for reliable diagnosis. Comparative data show that prior art signatures fail to achieve the same level of performance, especially for early-stage disease.
[0206] The superior diagnostic performance of the three-gene signature is illustrated in Table 2 and Figure 8, while Table 3 and Figure 6 demonstrate the comparative advantage over prior art signatures.
[0207] Versatility for monitoring and therapy guidance:
[0208] The method may also be used for monitoring disease progression or response to therapy by tracking changes in the probability score over time, offering a valuable tool for patient management beyond initial diagnosis.
[0209] Potential for Kit development and clinical implementation:
[0210] The invention is readily adaptable to kit formats, enabling standardized and widespread clinical use. The kit can include all necessary reagents and instructions for measuring the gene signature and calculating the probability score.
[0211] In summary, the invention addresses the unmet need for a reliable, minimally invasive, and highly accurate method for early breast cancer detection and monitoring. The technical effects and advantages are supported by robust experimental data and are not rendered obvious by the prior art.
[0212] Technical effects:
[0213] The experimental results presented in Examples 1 to 3, including the new data provided by the inventors, clearly demonstrate the unexpected and synergistic effect of the specific three-gene combination (PGPEP1, CYBRD1, and OSBPL8) for the early and non-invasive detection of breast cancer.
[0214] In Example 1, the selection and validation of the gene panel using both sequencing and PCR data from the GENOA clinical study shows that this combination achieves high sensitivity and specificity, with an AUG of 0.967, for distinguishing primary breast cancer patients from healthy donors. The statistical robustness of the signature is further supported by the z-score normalization and the results of the Welch Two sample t-test.Example 2 provides a direct comparison with the prior art signature described in WO 2022 / 152911 A1. The results demonstrate that the prior art gene panel fails to achieve adequate discrimination between healthy donors and breast cancer patients in early-stage disease, whereas the present invention’s signature performs significantly better in all key metrics (specificity, sensitivity, AUG). This highlights the inventive step of the claimed combination, which is not suggested or rendered obvious by the prior art.
[0215] Most importantly, Example 3 and the additional data provided by the inventors (A4) establish that neither individual genes nor any pairwise combination of the three mandatory genes is sufficient to reach the required diagnostic thresholds for clinical utility (>85% sensitivity and specificity). Only the full three-gene panel achieves robust and reproducible performance, with all other combinations missing a significant number of positive cases or yielding lower specificity. This demonstrates a true synergistic effect, which is both unexpected and non-obvious in view of the prior art, where large gene lists are disclosed without teaching or suggesting the optimal minimal combination.
[0216] Furthermore, the examples show that the claimed method is effective for early-stage (T1 / T2) and primary breast cancer, and that the probability score calculated according to the invention provides a reliable and clinically meaningful diagnostic tool. The ability to monitor disease progression or response to therapy using the same panel further supports the broad applicability and technical advantage of the invention.
[0217] The robustness of the invention is further supported by the reproducibility of the results across independent patient cohorts and different analytical platforms (e.g., RNA sequencing and RT-PCR). The method provides consistent diagnostic performance regardless of the biological sample type (whole blood, PBMC, serum, plasma), and the probability score threshold can be adjusted to optimize sensitivity and specificity for different clinical settings.
[0218] In summary, the examples provide clear experimental evidence that the invention solves the technical problem of providing a minimal, synergistic, and highly accurate gene signature for early and non-invasive breast cancer detection, which is not rendered obvious by the prior art. The inventive step is supported by the unexpected performance of the three-gene combination, the failure of alternative combinations, and the comparative data versus existing signatures.
[0219] Algorithms:Preferably, mathematical algorithms can be used to combine information from results of multiple individual biomarkers into a single measurement or index. Patients identified as having an increased risk of breast cancer can optionally be selected to receive treatment regimens, such as the ones defined above.
[0220] Applicants preferably applies the use of a classification algorithms, and methods of risk index construction, utilizing pattern recognition features, including established techniques such as the Kth-Nearest Neighbor, Boosting, Decision Trees, Neural Networks, Bayesian Networks, Support Vector Machines, and Hidden Markov Models are encompassed by or within the ambit of the present invention, as is the use of such combination to create single numerical “risk indices” or “risk scores” encompassing information from multiple biomarkers inputs.
[0221] Mathematical model:
[0222] Biomarkers of interest selection:
[0223] Several methods can be combined to identify the biomarkers of interest. The statistical analyzes specific to each biomarker measurement method make it possible to measure the "differential values" of each of the biomarkers between the samples of the different groups (healthy, healthy cancer, relapse, metastatic). Statistical significance for each candidate can be determined by calculating p-value or fold-change.
[0224] Since some of the selected biomarkers might be highly correlated, a solution is to use penalized logistic regression (L1 and / or L2) to select it during the probability score calculation (see Example 4).
[0225] Probability score:
[0226] The probability score can be calculated according to any method known in the art. For example, the probability score is calculated from a logistic regression prediction model applied to the measurements, e.g. biomarkers normalized amounts or biomarkers indices. The probability score may be calculated by:
[0227]
[0228] where xni,i is a measured value for the biomarker n and subject i and ( ?0, plt..., pn) is a vector of coefficients with ?0a panel-specific constant, and n is the corresponding logistic regression coefficient of the biomarker n.Penalized logistic regression models may be validated directly on a training set or by nonoverlapped bootstrap method: X random datasets were drawn with replacement from training set; each dataset had the same size as the training set. The model may be re-fit at each bootstrap and validated with the out-of-bag samples.
[0229] Fitted logistic regression models are tested by asking whether the model can correctly predict the clinical outcome using patient data other than that with which the logistic regression model was fitted but having a known clinical outcome. After training, the model output from 0 (control) to 1 (cancer) can be calculated in blind fashion by the average error of all N predictions (a validation group).
[0230] Various statistical analysis software can be used for building logistic regression models. Penalized regressions are applied to identify the marker combination, followed by GLM model for performance testing.
[0231] In our settings, candidate genes from the sequencing dataset were selected with a p value < 0.01, a Log fold change > 0.5, and AUG between the clinical states > 0.8. Then LASSO regression method was used to obtain the most predictive gene to the diagnostic status. The R packages “glmnet” and “glm” were applied to perform a lasso regression analysis. The most selected genes (>40%) in the models with the best performance (AUG > 0.9) are considered as candidate biomarkers. A non-penalized logistic regression was then performed to keep only the significant variables and evaluate signature performances. In order to build confidence intervals for these performances, the previous leave-n-out operation was repeated on this model and the distributions of the performance indicators are used to construct empirical confidence intervals.
[0232] Performances (sensitivity, specificity) estimation
[0233] The application of logistic regression to biological problems is routine in the art. Based on the output values, the receiver operating characteristic (ROC) curve can be built to calculate the outcome of clinical prediction: specificity and sensitivity of CRC cancer detection.
[0234] They are statistical measures of the performance of a binary classification test. Sensitivity measures the proportion of actual positives which are correctly identified as such (e.g., the percentage of sick people who are correctly identified as having the condition). Specificity measures the proportion of negatives which are correctly identified (e.g., the percentage of healthy people who are correctly identified as not having the condition). A perfect predictor would be described as 100% sensitive (i.e. , predicting all people from the sick group as sick)and 100% specific (i.e., not predicting anyone from the healthy group as sick). However, any predictor will possess a minimum error bound.
[0235] Definitions:
[0236] “Penalized logistic regression”:
[0237] Penalized logistic regression is based on mathematical equation derived from logistic regression. More specifically, penalized logistic regression is a ridge regression for logistic model with L2-norm or L1-norm penalty. To estimate the parameters in this method a quadratic (L2) or / and L1-norm penalty is added on the log-likelihood that should be maximized. To choose the best value of A1 and A2, the cross-validation is used with the AIC criteria. To fit the penalized logistic model, the following algorithms (packages in R Cran, statistical software) can be used: glmpath (Park M. Y and Hastie T. (2006) An L1 Regularization-path Algorithm for Generalized Linear Models. A generalization of the LARS algorithm for GLMs and the Cox proportional hazard model), penalized (Goeman, J. (2010) L1 (lasso) and L2 (ridge) penalized estimation in GLMs and in the Cox model) and glmnet (Hasti, T., Tibshirani and R., Friedman, J. (2010). Lasso and elastic-net regularized generalized linear models) with different tuning parameters.
[0238] “Sensitivity, specificity, PPV, NPV”:
[0239] The term "sensitivity of a test" refers to the probability that a test result will be positive when the disease is present in the patient (true positive rate). This is derived from the number of patients with the disease who have a positive test result (true positive) divided by the total number of patients with the disease, including those with true positive results and those patients with the disease who have a negative result, i.e., false negative.
[0240] The term "specificity of a test" refers to the probability that a test result will be negative when the disease is not present in the patient (true negative rate). This is derived from the number of patients without the disease who have a negative test result (true negative) divided by all patients without the disease, including those with a true negative result and those patients without the disease who have a positive test result, e.g. false positive.
[0241] The term "positive predictive value" (PPV) refers to the probability that a positive result correctly identifies a patient who has the disease, which is the number of true positives divided by the sum of true positives and false positives.The term "negative predictive value" or "NPV" refers to the probability that a negative test correctly identifies a patient without the disease, which is the number of true negatives divided by the sum of true negatives and false negatives. Like the PPV, it also is inherently impacted by the prevalence of the disease and pre-test probability of the population intended to be tested. A positive result from a test with a sufficient PPV can be used to rule in the disease for a patient, while a negative result from a test with a sufficient NPV can be used to rule out the disease, if the disease prevalence for the given population, of which the patient can be considered a part, is known.
[0242] “Significant”:
[0243] “Statistically significant” means that the alteration is greater than what might be expected to happen by chance alone. Statistical significance can be determined by any method known in the art. For example, statistical significance can be determined by p-value. The p-value is a measure of probability that a difference between groups during an experiment happened by chance. (P (z zObserved)). For example, a p-value of 0.01 means that there is a 1 in 100 chance the result occurred by chance. The lower the p-value, the more likely it is that the difference between groups was caused by treatment. An alteration is considered to be statistically significant if the p-value is at least 0.05. Preferably, the p-value is 0.04, 0.03, 0.02, 0.01, 0.005, 0.001 or less. As noted below, and without any limitation of the invention, achieving statistical significance generally, but not always, requires that combinations of several biomarkers be used together in panels and combined with mathematical algorithms in order to achieve a statistically significant biomarkers index.
[0244] “ROC, AUG”:
[0245] A Receiver Operating Characteristics (“ROC”) curve can be used as an indicator that allows representation of the sensitivity and specificity of a test, assay, or method over the entire range of test (or assay) cut points with just a single value (see, e.g., Shultz, “Clinical Interpretation Of Laboratory Procedures,” chapter 14 in Teitz, Fundamentals of Clinical Chemistry, Burtis and Ashwood (eds.), 4thedition 1996, W.B. Saunders Company, pages 192-199; and Zweig et al.). “ROC Curve Analysis: An ROC curve is an x-y plot of sensitivity on the y-axis, on a scale of zero to one (e.g., 100%), against a value equal to one minus specificity on the x-axis, on a scale of zero to one (e.g., 100%).
[0246] Thus, a ROC curve is a plot of the true positive rate against the false positive rate for that test, assay, or method. To construct the ROC curve for the test, assay, or method in question, subjects can be assessed using a perfectly accurate or “gold standard” method that is independent of the test, assay, or method in question to determine whether the subjects aretruly positive or negative for the disease, condition, or syndrome (for example, coronary angiography is a gold standard test for the presence of coronary atherosclerosis). The subjects can also be tested using the test, assay, or method in question, and for varying cut points, the subjects are reported as being positive or negative according to the test, assay, or method. The sensitivity (true positive rate) and the value equal to one minus the specificity (which value equals the false positive rate) are determined for each cut point, and each pair of x-y values is plotted as a single point on the x-y diagram. The “curve” connecting those points is the ROC curve. The ROC curve is often used in order to determine the optimal single clinical cut-off or treatment threshold value where sensitivity and specificity are maximized. Such a situation represents the point on the ROC curve that describes the upper left corner of the single largest rectangle which can be drawn under the curve.
[0247] The total area under the curve (“AUC”) is the indicator that allows representation of the sensitivity and specificity of a test, assay, or method over the entire range of cut points with just a single value. The maximum AUC is one (a perfect test) and the minimum area is one half (e.g. the area where there is no discrimination of normal versus disease). The closer the AUC is to one, the better is the accuracy of the test. It should be noted that implicit in all ROC and AUC is the definition of the disease and the post-test time horizon of interest.
[0248] As defined herein, a “high degree of diagnostic accuracy” means a test or assay wherein the AUC (area under the ROC curve for the test or assay) is at least 0.70, preferably at least 0.75, more preferably at least 0.80, preferably at least 0.85, more preferably at least 0.90, and most preferably at least 0.95.
[0249] Normalization, Benjamini correction for multiple testing:
[0250] Use of Benjamini-Hochberg method for calculating the false discovery rate in a clinical or diagnostic assay (Am. J. Public Health 86(5): 628-629).
[0251] Normalization refers to a collection of processes that are used to adjust data means or variances for effects resulting from systematic non-biological differences between arrays, subarrays (or print-tip groups), and dye-label channels. An array is defined as the entire set of target probes on the chip or solid support. A subarray or print-tip group refers to a subset of those target probes deposited by the same print-tip, which can be identified as distinct, smaller arrays of proves within the full array. The dye-label channel refers to the fluorescence frequency of the target sample hybridized to the chip. Experiments where two differently dye-labeled samples are mixed and hybridized to the same chip are referred to in the art as “dualdye experiments”, which result in a relative, rather than absolute, expression value for eachtarget on the array, often represented as the log of the ratio between “red” channel and “green channel.” Normalization can be performed according to ratiometric or absolute value methods. Ratiometric analyses are mainly employed in dual-dye experiments where one channel or array is considered in relation to a common reference. A ratio of expression for each target probe is calculated between test and reference sample, followed by a transformation of the ratio into log 2(ratio) to symmetrically represent relative changes. Absolute value methods are used frequently in single-dye experiments or dual-dye experiments where there is no suitable reference for a channel or array. Relevant “hits” are defined as expression levels or amounts that characterize a specific experimental condition. Usually, these are nucleic acids or proteins in which the expression levels differ significantly between different experimental conditions, usually by comparison of the expression levels of a nucleic acid or protein in the different conditions and analyzing the relative expression (“fold change”) of the nucleic acid or protein and the ratio of its expression level in one set of samples to its expression in another set.
[0252] Those skilled in the art will appreciate that the invention described herein is susceptible to variations and modifications other than those specifically described. It is to be understood that the invention includes all such variations and modifications without departing from the spirit or essential characteristics thereof. The invention also includes all the steps, features, compositions and compounds referred to or indicated in this specification, individually or collectively, and any and all combinations or any two or more of said steps or features. The present disclosure is therefore to be considered as in all aspects illustrated and not restrictive, the scope of the invention being indicated by the appended Claims, and all changes which come within the meaning and range of equivalency are intended to be embraced therein.
[0253] Various references are cited throughout this specification, each of which is incorporated herein by reference in its entirety.
[0254] The foregoing description will be more fully understood with reference to the following Examples. Such Examples, are, however, exemplary of methods of practicing the present invention and are not intended to limit the scope of the invention.
[0255] Examples
[0256] Example 1: Evaluation of the score with the combination of the 3 mandatory genes This example illustrates the process of identifying and validating the optimal gene combinationfor early and non-invasive detection of breast cancer, using patient data from the GENOA clinical study.
[0257] Consequently, to develop the particular gene combination of the invention, different groups of patients from the GENOA clinical study were compared, by sequencing the total mRNA of the leukocytes present in their blood circulation. There were healthy patients, patients with primary breast cancer, patients with metastatic breast cancer and patients in remission after primary breast cancer. After analysis, a long list of genes that were expressed very differently with high stability (OR reproducibility) in patients with primary breast cancer compared to all other groups, especially healthy donors (HD), with highly significant P value, was found (Figure 1). The best genes able to separate between the two groups were investigated using a mathematical penalized linear regression model (LASSO), with further performance validation using an algorithmic GLM model on the more promising genes (Figure 2). A total of 120 patients were tested. Several of these genes were then validated in a larger cohort of patients using the PCR technology (Figure 3), proving the quality of the sequencing results. The best gene combination was finally determined on these PCR data using a mathematical penalized linear regression model (LASSO), with further performance validation using an algorithmic GLM model (Figure 4).
[0258] A smaller list of target genes was obtained, with number of highly correlated genes. This PCR signature was validated using z-scoring calculation, meaning that the data were transformed to be normalized. The z-score indicates how many standard deviations a measurement is from the mean (see formula z) and helps to understand how exceptional or normal a particular measurement is compared to the overall average. Here under is the z-formula, where x is the observed value, . the mean, and a the standard deviation:
[0259]
[0260] The z-score was calculated for each of the gene and a global mean is the z-score per patient. The results did show a high performance of separation of primary breast cancer patient from healthy donors (Figure 5). From this list, AUC of each gene was calculated, to understand the contribution of each gene to the signature performance to discriminate primary breast cancer patient from the rest. A minimal signature, core of the performance, was selected for the invention.
[0261] Performance of the core signature, calculated on the PCR relative expression data of the full cohort, can be observed in Table 2, underlining the excellent performance of this signature.
[0262]
[0263]
[0264] Table 2
[0265] The Welch Two sample t-test performed on the PCR dataset between primary breast cancer patient and healthy donors using this signature of 3 genes give a significant p value of 0.02591 , representing the high capabilities of this mandatory gene combination to identify breast cancer patients from healthy subjects. Out of the 28 breast cancer patients, the 28 were identified as breast cancer with a minimum of 95.8% of confidence, and out of the 26 healthy donors, only 2 were identified as wrong positives subjects, showing a very high specificity of the combination. These results are represented in the heatmap of figure 8.
[0266] Conclusion:
[0267] These results demonstrate that the combination of PGPEP1, CYBRD1, and OSBPL8 provides high sensitivity and specificity for distinguishing primary breast cancer patients from healthy donors. The statistical analyses confirm the robustness and reproducibility of this gene signature, supporting its use as a minimal, synergistic panel for early and non-invasive breast cancer detection as claimed in the invention.
[0268] Example 2: Comparison of the score of this combination towards WO 2022 / 152911 A1 This example compares the diagnostic performance of the present invention’s gene signature with that of a prior art signature (WO 2022 / 152911 A1), highlighting the superiority of the claimed combination for early breast cancer detection.
[0269] Applicant attempted to test in a comparative study the performance of the main genes as described in WO 2022 / 152911 A1 (Universite de Fribourg) towards the 3 main genes described in the present invention. The described in WO 2022 / 152911 A1 (Universite de Fribourg) was claiming the strategical approach of the test, as well as a list of target genes, as well as proteins, that should be combined for breast cancer screening or follow-up.
[0270] However, the genes described in WO 2022 / 152911 were not even expressed differently between HD and P in the GENOA study, showing that the combination of these genes is not good in this specific early-stage situation. Looking for the performance of the WO 2022 / 152911 A1 transcriptomic signature (SOX4, TNFS10, CD3G, NR3C2), even if the genes are not differentially expressed, comparing to the invention’s signature, it is clear that the present invention’s signature is performing much better, as also represented in the Figure 6, where the heatmap of the WO 2022 / 152911 A1 signature sows no clear group separation, compared to the figure 2 from the present invention’s signature.
[0271] SPECIFICITY
[0272] SENSITIVITY
[0273] AUC
[0274] AIC
[0275]
[0276] Table 3
[0277] Applicant can just conclude that they are ultimately not relevant in the early stage of the disease (preliminary data were mainly in metastatic progression). Regarding proteins’ signature, Figure 7 shows the 4 main proteins that were used in WO 2022 / 152911 are also not useful to separate healthy donors (HD) from breast cancer patient (P) at an early stage. Applicant compared healthy patients and primary breast cancer patients in the GENOA study, to illustrate the fact that this protein signature does not work in early detection of breast cancer. It can be concluded that the signature according to the present invention gives much better results in breast cancer early screening conditions.
[0278] Conclusion:
[0279] The comparative analysis clearly shows that the gene signature of the present invention outperforms the prior art signature (WO 2022 / 152911 A1) in terms of specificity, sensitivity, and AUC, particularly for early-stage breast cancer detection. This highlights the inventive step and clinical advantage of the selected three-gene panel.
[0280] Example 3: Evaluation of diagnostic performance using various mandatory gene combinations
[0281] This example evaluates the diagnostic accuracy of individual genes, pairs, and the full three-gene combination, demonstrating the necessity of the claimed panel for achieving robust clinical performance.
[0282] To assess whether the selected combination of mandatory genes delivers optimal diagnostic accuracy for early breast cancer detection, performance metrics were calculated for both individual genes and their combinations. Polymerase chain reaction (PCR) data from healthy donors and primary breast cancer patients were utilized in this evaluation.
[0283] Specificity is defined as the test’s capacity to correctly identify true negatives (healthy individuals), whereas sensitivity measures the ability to accurately detect true positives(patients with the disease). Clinical relevance requires a threshold of 85% for both specificity and sensitivity, indicating robust diagnostic effectiveness.
[0284] Analysis presented in Table 4 demonstrates that the three-gene combination consistently surpassed this threshold, achieving specificity of 0.92 and sensitivity of 0.89. In contrast, individual gene analyses did not meet these criteria: for instance, PGPEP1 alone yielded a specificity of 0.5 and sensitivity of 0.84, while neither OSBPL8 nor CYBRD1 individually exceeded 85% in both metrics. Accordingly, it is evident that combining multiple genes is essential to attain the requisite diagnostic performance.
[0285]
[0286] Table 5 further illustrates the outcomes for combinations comprising only two of the three genes. None of these double gene combinations performed as well as the full three-gene set; notably, each missed four of the twenty-eight primary breast cancer cases, rendering them suboptimal for first-line detection where maximal sensitivity is critical and no positive cases should be missed.
[0287] SPECIFICITY SENSITIVITY AUC
[0288] T-test (Hd vs P)
[0289]
[0290] Table 5Conclusion:
[0291] The data confirm that only the combination of all three mandatory genes achieves the required diagnostic thresholds for both sensitivity and specificity. Neither individual genes nor any pairwise combination is sufficient. This demonstrates the necessity and synergy of the three-gene panel for reliable early detection of breast cancer, as reflected in the claims.References
[0292] [1] Sotiriou, C. et al. Breast cancer classification and prognosis based on gene expression profiles from a population-based study. Proc Natl Acad Sci 100, 10393-10398 (2003).
[0293] [2] Sleeman, J. P. et al. Concepts of metastasis in flux: The stromal progression model. Semin Cancer Biol 22, 174-186 (2012).
[0294] [3] Cardoso, F. et al. 70-Gene Signature as an Aid to Treatment Decisions in Early-Stage Breast Cancer. N Engl J Med 375, 717-729 (2016).
[0295] [4] Lorusso, G. & Ruegg, C. The tumor microenvironment and its contribution to tumor evolution toward metastasis. Histochem Cell Biol 130, 1091-1103 (2008).
[0296] [5] Salgado, R. et al. The evaluation of tumor-infiltrating lymphocytes (TILs) in breast cancer: recommendations by an International TILs Working Group 2014. Annals of Oncology 26, 259-271 (2015).
[0297] [6] van den Ende, C., Oordt-Speets, A. M., Vroling, H. & van Agt, H. M. E. Benefits and harms of breast cancer screening with mammography in women aged 40-49 years: A systematic review. Int. J. Cancer 141, 1295-1306 (2017).
[0298] [7] Fitzjohn, J., Zhou, C. & Chase, J. G. Critical Assessment of Mammography Accuracy. IFAC-PapersOnLine 56, 5620-5625 (2023).
[0299] [8] Connal, S. et al. Liquid biopsies: the future of cancer early detection. J Transl Med 21, 118 (2023).
[0300] [9] https: / / www.cancer.org / content / dam / cancer-org / research / cancer-facts-and- statistics / breast-cancer-facts-and-figures / breast-cancer-facts-and-figures-2019-2020.pdf.
[0301]
[0010] Duffy, M. J., McDermott, E. W. & Crown, J. Blood-based biomarkers in breast cancer: From proteins to circulating tumor cells to circulating tumor DNA. Tumour Biol. 40, 101042831877616 (2018).
[0302]
[0011] Cattin, S. et al. Bevacizumab specifically decreases elevated levels of circulating KIT+CD11b+ cells and IL-10 in metastatic breast cancer patients. Oncotarget 7, 11137— 11150 (2016). 11. Cattin, S. et al. Bevacizumab specifically decreases elevated levels ofcirculating KIT+CD11b+ cells and IL-10 in metastatic breast cancer patients. Oncotarget7, 11137-11150 (2016).
Claims
45CLAIMS1. A method for early and minimally invasive detection of breast cancer in a patient, comprising:(a) measuring in a biologic sample obtained from the patient the expression level of transcriptomic (mRNA) biomarkers of a panel comprising the combination of at least 3 genes consisting of PGPEP1 , CYBRD1 and OSBPL8; (b) calculating a probability score based on the measurement of step (a) wherein the probability score P is calculated by the formula:where xni, is a measured value for the biomarker n and subject i and (?0,pn) is a vector of coefficients with ?0a panel-specific constant, and pn is the corresponding logistic regression coefficient of the biomarker n,(c) ruling out breast cancer for the patient if the score in step (b) is lower than a pre-determined threshold, or ruling in the likelihood of breast cancer for the patient if the score in step (b) is higher than a pre-determined threshold, wherein a probability score value Pyi closer to 1.0 represents the likelihood of breast cancer for said patient and a probability score value Pyi closer to 0.0 represents the likelihood that said patient has not developed breast cancer and wherein the combination of PGPEP1, CYBRD1 and OSBPL8 provides a diagnostic performance with an area under the ROC curve (AUC) of at least 0.90, said performance being significantly superior to any combination of only two of these genes, and wherein the method enables detection of early-stage (T1 / T2) breast cancer.
2. The method of claim 1 , wherein the biological sample is selected from the group consisting of whole blood, peripheral blood mononuclear cells, leukocytes, serum, plasma, or any other blood-derived sample.
3. The method of claim 1 or 2, wherein the probability score is calculated using a logistic regression model or another supervised machine learning algorithm.
464. The method of any one of claims 1 to 3, wherein the method is validated on an independent cohort using RT-qPCR.
5. The method of any one of claims 1 to 4, wherein the method is used for screening asymptomatic women at risk of developing primary breast cancer.
6. The method of any one of claims 1 to 5, wherein the likelihood of breast cancer is further determined by the sensitivity, specificity, negative predictive value (NPV), or positive predictive value (PPV) associated with the score.
7. The method of any one of claims 1 to 6, wherein the method is suitable for detecting breast cancer at stage (04 1).
8. The method of any one of claims 1 to 7, wherein the method is suitable for distinguishing breast cancer patients from healthy controls with a sensitivity of at least 85% and a specificity of at least 85%.
9. The method of any one of claims 1 to 8, wherein the method further comprises measuring the expression of one or more additional genes selected from the group consisting of PSD4, GALK1, IPMK, GAPDH, HNRNPK, CGGBP1, NRDC, THRA, VMP1, ADD1, ARHGAP9, SLC43A3, IGF1R, B3GNTL1, WWC3, and TOM1L2.
10. A kit for the detection of breast cancer in a patient from a biological sample, preferably a blood sample, comprising at least one probe or primer for measuring the expression level of each of PGPEP1, CYBRD1, and OSBPL8, and instructions for calculating a probability score as defined in claim 1.
11. The kit of claim 10, further comprising at least one probe or primer for measuring the expression level of one or more additional genes selected from the group consisting of PSD4, GALK1, IPMK, GAPDH, HNRNPK, CGGBP1, NRDC, THRA, VMP1, ADD1, ARHGAP9, SLC43A3, IGF1R, B3GNTL1, WWC3, and TOM1L2.
12. Use of the kit of claim 10 or 11 , for the early and non4nvasive detection of breast cancer in a patient.4713. Use of the kit of claim 10 or 12 for distinguishing early-stage (O-l-ll) breast cancer from healthy controls.
14. A method for monitoring disease progression or response to therapy in a breast cancer patient, comprising:(a) measuring in a biological sample obtained from the patient the expression level of transcriptomic (mRNA) biomarkers of a panel comprising the combination of at least 3 genes consisting of PGPEP1 , CYBRD1 and OSBPL8; (b) calculating a probability score based on the measurement of step (a) wherein the probability score P is calculated by the formula:LOG(P(y_i=1) / (1-P(y_i=1))) = p0+ Pi Xn + ... + pnxni,where x_ni is a measured value for the biomarker n and subject i and (p0, Pi, ..., pn) is a vector of coefficients with p0a panel-specific constant, and pn is the corresponding logistic regression coefficient of the biomarker n,(c) comparing the probability score obtained at different time points to assess disease progression or response to therapy,wherein a probability score value Pyi closer to 1.0 represents the likelihood of breast cancer for said patient and a probability score value Pyi closer to 0.0 represents the likelihood that said patient has not developed breast cancer.
15. A method for early and minimally invasive detection of primary breast cancer in a patient, comprising:(a) measuring in a biological sample obtained from the patient the expression level of transcriptomic (mRNA) biomarkers of a panel comprising the combination of at least 3 genes consisting of PGPEP1 , CYBRD1 and OSBPL8; (b) calculating a probability score based on the measurement of step (a) wherein the probability score P is calculated by the formula:LOG(P(y_i=1) / (1-P(y_i=1))) = p0+ Pi Xn + ... + pnxni,where x_ni is a measured value for the biomarker n and subject i and (p0, Pi, ..., pn) is a vector of coefficients with p0a panel-specific constant, and pn is the corresponding logistic regression coefficient of the biomarker n,(c) ruling out primary breast cancer for the patient if the score in step (b) is lower than a pre-determined threshold, or ruling in the likelihood of primary breast cancer for the patient if the score in step (b) is higher than a pre-determined threshold,wherein a probability score value Pyi closer to 1.0 represents the likelihood of primary breast cancer for said patient and a probability score value Pyi closer to 0.0 represents the likelihood that said patient has not developed primary breast cancer.