Methods for predicting clinical outcomes in cancer
A gene expression-based method using paraffin-embedded biopsies predicts breast cancer outcomes, enabling personalized treatment by calculating a risk score from normalized gene expression levels, addressing inefficiencies in current diagnostic methods.
Patent Information
- Application Number
- JP2023213416
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2009-11-23
- Filing Date
- 2023-12-18
- Publication Date
- 2026-03-03
- Estimated Expiration
- 2030-11-19
AI Technical Summary
Current diagnostic methods for breast cancer are inefficient and often rely on non-standardized reagents and subjective interpretations, leading to variable results and the inability to accurately predict disease recurrence and long-term survival, necessitating a more precise method for individualized treatment.
A set of genes whose expression levels are associated with clinical outcomes in breast cancer, using archival paraffin-embedded biopsies to predict prognosis and tailor treatment, compatible with various biopsy methods, and employing a method that includes normalizing gene expression levels to calculate a risk score.
Enables accurate prediction of breast cancer recurrence and long-term survival, allowing for personalized treatment strategies and improved treatment success rates by stratifying patients based on risk.
Smart Images

Figure 0007823012000001 
Figure 0007823012000002 
Figure 0007823012000003
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 61 / 263,763, filed November 23, 2009, which is incorporated herein by reference in its entirety. [Background technology]
[0002] Oncologists have numerous treatment options available to them, including various combinations of treatment regimens characterized as "standard of care." The absolute benefit of adjuvant therapy is greater in patients with poor prognosis, resulting in the selection of only these so-called "high-risk" patients for adjuvant chemotherapy. See, e.g., S. Paik et al., J Clin Oncol. 24(23):3726-34 (2006). Therefore, to maximize the likelihood of a favorable treatment outcome, it is necessary to assign patients the best available cancer treatment, and for this assignment to occur as quickly as possible after diagnosis.
[0003] Our healthcare system today is rife with inefficiencies and wasteful spending—one example is the effectiveness rate of many oncology drugs, which only work about 25% of the time. Many of these cancer patients experience toxic side effects from expensive treatments that may not work. This disparity between high treatment costs and low treatment effectiveness often results from treating a particular diagnosis with one method across diverse patient populations. However, with the advent of genetic profiling tools, genomic testing, and advanced diagnostic methods, this is beginning to change.
[0004] In particular, once a patient is diagnosed with breast cancer, there is a strong need for methods that allow physicians to predict the expected course of the disease, such as the likelihood of cancer recurrence and the patient's long-term survival, and to select the most appropriate treatment options accordingly. Recognized prognostic and predictive factors for breast cancer include age, tumor size, axillary lymph node status, tumor histology, pathological grade, and hormone receptor status. However, molecular diagnostics have been demonstrated to identify many more low-risk breast cancer patients than is possible with standard prognostic indicators. S. Paik, The Oncologist 12(6):631-635 (2007).
[0005] Despite recent advances, targeting pathogenetically distinct tumor types with specific treatment regimens and ultimately individualizing tumor treatment to achieve optimal outcomes remains a challenge in breast cancer treatment. Accurate prediction of prognosis and clinical outcome may enable oncologists to tailor the administration of adjuvant chemotherapy so that women at higher risk of recurrence or poor prognosis receive more aggressive treatment. Furthermore, accurate stratification of patients based on risk may greatly advance our understanding of the expected absolute benefit from treatment, thus increasing the success rate of clinical trials of novel breast cancer therapies.
[0006] Currently, most diagnostic tests used in clinical practice are not quantitative and often rely on immunohistochemistry (IHC). This method often produces variable results between laboratories, in part because reagents are not standardized and in part because interpretation is subjective and not easily quantified. Other RNA-based molecular diagnostics require fresh-frozen tissue, which poses numerous challenges, including incompatibility with current clinical practice and specimen transport regulations. Fixed, paraffin-embedded tissue is more readily available, and methods for detecting RNA in fixed tissue are well established. However, these methods typically cannot test a large number of genes (DNA or RNA) from small amounts of material. Therefore, fixed tissue has traditionally been rarely used outside of IHC detection of proteins. [Prior art documents] [Non-patent literature]
[0007] [Non-Patent Document 1] S. Paik et al., J Clin Oncol. 24(23):3726-34 (2006) [Non-patent document 2] S. Paik, The Oncologist 12(6):631-635 (2007) Summary of the Invention [Means for solving the problem]
[0008] The present invention provides a set of genes whose expression levels are associated with a particular clinical outcome in cancer. For example, assuming the patient receives standard treatment, the clinical outcome can be a good prognosis or a bad prognosis. Clinical outcome may be defined by a clinical endpoint, such as disease- or recurrence-free survival, metastasis-free survival, overall survival, etc.
[0009] The present invention accommodates the use of archival paraffin-embedded biopsies for assaying any marker in the set, and is therefore compatible with the most widely available types of biopsy material. The present invention is also compatible with several different tumor tissue collection methods, such as core biopsy or fine needle aspiration. The tissue sample may contain cancer cells.
[0010] In one aspect, the present invention relates to a method for predicting clinical outcome in a cancer patient, comprising the steps of: (a) obtaining, from a tissue sample obtained from the patient's tumor, an expression level of an expression product (e.g., an RNA transcript) of at least one prognostic gene listed in Tables 1-12; (b) normalizing the expression level of the expression product of the at least one prognostic gene to obtain a normalized expression level; and (c) calculating a risk score based on the normalized expression value, wherein increased expression of prognostic genes in Tables 1, 3, 5, and 7 is positively correlated with a favorable prognosis, and increased expression of prognostic genes in Tables 2, 4, 6, and 8 is negatively associated with a favorable prognosis. In some embodiments, the tumor is estrogen receptor-positive. In other embodiments, the tumor is estrogen receptor-negative.
[0011] In one aspect, the disclosure provides a method for predicting clinical outcome of a cancer patient, comprising the steps of: (a) obtaining an expression level of an expression product (e.g., RNA transcript) of at least one prognostic gene from a tissue sample obtained from the patient's tumor, wherein the at least one prognostic gene is selected from GSTM2, IL6ST, GSTM3, C8orf4, TNFRSF11B, NAT1, RUNX1, CSF1, ACTR2, LMNB1, TFRC, LAPTM4B, ENO1, CDC20, and IDH2; and (b) determining an expression level of at least one prognostic gene. (c) normalizing the expression levels of the expression products of the genes to obtain normalized expression levels; and (c) calculating a risk score based on the normalized expression values, wherein increased expression of prognostic genes selected from GSTM2, IL6ST, GSTM3, C8orf4, TNFRSF11B, NAT1, RUNX1, and CSF1 is positively correlated with a favorable prognosis, and increased expression of prognostic genes selected from ACTR2, LMNB1, TFRC, LAPTM4B, ENO1, CDC20, and IDH2 is negatively associated with a favorable prognosis. In some embodiments, the tumor is estrogen receptor positive. In other embodiments, the tumor is estrogen receptor negative.
[0012] In various embodiments, normalized expression levels (as determined by assaying the levels of the expression products of the genes) of at least 2, or at least 5, or at least 10, or at least 15, or at least 20, or at least 25 prognostic genes are determined. In alternative embodiments, the normalized expression level of at least one of the genes in Tables 16-18 that is co-expressed with the prognostic genes is obtained.
[0013] In another embodiment, the risk score is determined using normalized expression levels of at least one stromal group or transferrin receptor group gene, or a gene co-expressed with a stromal group or transferrin receptor group gene.
[0014] In another embodiment, the cancer is breast cancer.In another embodiment, the patient is a human patient.
[0015] In yet another embodiment, the cancer is ER-positive breast cancer.
[0016] In yet another embodiment, the cancer is ER-negative breast cancer.
[0017] In further embodiments, the expression product comprises RNA. For example, the RNA may be exonic RNA, intronic RNA, or short RNA (e.g., microRNA, siRNA, promoter-associated small RNA, shRNA, etc.). In various embodiments, the RNA is fragmented RNA.
[0018] In a different embodiment, the invention relates to an array comprising polynucleotides that hybridize to an RNA transcription of at least one of the prognostic genes listed in Tables 1-12.
[0019] In yet another aspect, the present invention provides a method for preparing a patient-specific genomic profile, comprising the steps of: (a) obtaining, from a tissue sample obtained from the patient's tumor, an expression level of an expression product (e.g., an RNA transcript) of at least one prognostic gene listed in Tables 1-12; (b) normalizing the expression level of the expression product of the at least one prognostic gene to obtain a normalized expression level; and (c) calculating a risk score based on the normalized expression value, wherein increased expression of prognostic genes in Tables 1, 3, 5, and 7 is positively correlated with a favorable prognosis, and increased expression of prognostic genes in Tables 2, 4, 6, and 8 is negatively associated with a favorable prognosis. In some embodiments, the tumor is estrogen receptor-positive; in other embodiments, the tumor is estrogen receptor-negative.
[0020] In various embodiments, the method can further include providing a report, which can include a prediction of the patient's likelihood of risk of having a particular clinical outcome.
[0021] The present invention further provides a computer-implemented method for classifying a cancer patient based on risk of cancer recurrence, comprising the steps of: (a) classifying, on a computer, the patient as having a good prognosis or a poor prognosis based on an expression profile comprising measurements of expression levels of expression products of a plurality of prognostic genes in a tumor tissue sample obtained from the patient, wherein the plurality of genes comprises at least three different prognostic genes listed in any of Tables 1 to 12, wherein a good prognosis predicts freedom from recurrence or metastasis within a predetermined time period after initial diagnosis, and a poor prognosis predicts freedom from recurrence or metastasis within the predetermined time period after initial diagnosis; and (b) calculating a risk score based on the expression levels. DETAILED DESCRIPTION OF THE INVENTION
[0022] definition Unless otherwise defined, scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Singleton et al., "Dictionary of Microbiology and Molecular Biology," 2nd ed., J. Wiley & Sons (New York, NY 1994), and March, "Advanced Organic Chemistry Reactions, Mechanisms and Structure," 4th ed., John Wiley & Sons (New York, NY 1992), provide those skilled in the art with a general guide to many of the terms used herein.
[0023] One skilled in the art will recognize many methods and materials similar or equivalent to those described herein, which could be used in the practice of the present invention. Indeed, the present invention is not intended to be limited to the methods and materials described. For purposes of the present invention, the following terms are defined below.
[0024] "Prognostic factors" are variables related to the natural history of cancer that influence a patient's recurrence rate and outcome after cancer onset. Clinical parameters associated with a worse prognosis include, for example, lymph node involvement and high-grade tumors. Prognostic factors are often used to stratify patients into subgroups with different underlying risks of relapse.
[0025] The term "prognosis" is used herein to refer to a prediction of the likelihood of cancer-related death or progression, including recurrence, metastatic spread, and drug resistance, of a neoplastic disease such as breast cancer. The term "good prognosis" refers to a desirable or "positive" clinical outcome. For example, in the context of breast cancer, a good prognosis can be a prediction of no recurrence or no metastasis within 2, 3, 4, 5, or more years from the initial diagnosis of breast cancer. The terms "poor prognosis" and "unfavorable prognosis" are used interchangeably herein to refer to an unfavorable clinical outcome. For example, in the context of breast cancer, a poor prognosis can be a prediction of no recurrence or no metastasis within 2, 3, 4, 5, or more years from the initial diagnosis of breast cancer.
[0026] The term "prognostic gene" is used herein to refer to a gene whose expression is positively or negatively correlated with a favorable prognosis for cancer patients treated with standard treatment. A gene can be both a prognostic gene and a predictive gene, depending on the correlation of the gene expression level with the corresponding endpoint. For example, using the Cox proportional hazards model, if a gene is only a prognostic gene, its hazard ratio (HR) will not change when measured in patients treated with standard treatment or in patients treated with novel interventions.
[0027] The term "predictive gene" is used herein to refer to a gene whose expression correlates positively or negatively with a beneficial response to a treatment. For example, the treatment may include chemotherapy.
[0028] The terms "risk score" or "risk classification" are used interchangeably herein and refer to the level of risk (or likelihood) that a patient will experience a particular clinical outcome. Patients may be classified into risk groups or may be classified with a level of risk based on the methods of the present disclosure, e.g., high, moderate, or low risk. A "risk group" is a group of subjects or individuals who have a similar risk level for a particular clinical outcome.
[0029] Clinical outcomes can be defined using various endpoints. The term "long-term" survival is used herein to refer to survival over a specific period of time, for example, at least 3 years, more preferably at least 5 years. The term "recurrence-free survival" (RFS) is used herein to refer to survival over the period (usually in years) from randomization to first cancer recurrence or death due to cancer recurrence. The term "overall survival" (OS) is used herein to refer to the period (in years) from randomization to death from any cause. The term "disease-free survival" (DFS) is used herein to refer to survival over the period (usually in years) from randomization to first cancer recurrence or death from any cause.
[0030] The calculation of the measures listed above may actually vary from study to study depending on the definition of events that are censored or not considered.
[0031] The term "biomarker" as used herein refers to a gene, gene product, or its expression level as measured using the gene product.
[0032] The term "microarray" refers to an ordered arrangement of hybridizable array elements, preferably polynucleotide probes, on a substrate.
[0033] As used herein, the term "normalized expression level," when used with respect to a gene, refers to a normalized value determined for a normalized gene product level, e.g., the RNA expression level of the gene or the polypeptide expression level of the gene.
[0034] The term “C t " as used herein refers to the threshold cycle, the cycle number in a quantitative polymerase chain reaction (qPCR) when the fluorescence generated in a reaction well exceeds a defined threshold, i.e., the point at which a sufficient number of amplicons have accumulated in the reaction to reach the defined threshold.
[0035] The terms "gene product" or "expression product" are used herein to refer to the RNA transcription products (transcripts) of a gene, including mRNA, and the polypeptide translation products of such RNA transcripts. A gene product can be, for example, unspliced RNA, mRNA, splice variant mRNA, microRNA, fragmented RNA, a polypeptide, a post-translationally modified polypeptide, a splice variant polypeptide, etc.
[0036] The term "RNA transcript" as used herein refers to the RNA transcription product of a gene, including, for example, mRNA, unspliced RNA, splice variant mRNA, microRNA, and fragmented RNA. "Fragmented RNA" as used herein refers to a mixture of intact RNA and RNA that has been degraded as a result of sample processing (e.g., fixation, sectioning of tissue blocks, etc.).
[0037] Unless otherwise indicated, each gene name used herein corresponds to the official symbol assigned to the gene and provided by Entrez Gene (URL: www.ncbi.nlm.nih.gov / sites / entrez) as of the filing date of this application.
[0038] The terms "correlated" and "associated" are used interchangeably herein and refer to the strength of the association between two measurements (or measured entities). The present disclosure provides genes and gene subsets whose expression levels are associated with specific outcome measures. For example, an increased expression level of a gene may be positively correlated (positively associated) with an increased likelihood of a good clinical outcome for a patient, e.g., an increased likelihood of long-term survival without cancer recurrence and / or metastasis-free survival. Such a positive correlation can be statistically demonstrated in various ways, for example, by a low hazard ratio (e.g., HR<1.0). In another example, an increased expression level of a gene may be negatively correlated (negatively associated) with an increased likelihood of a good clinical outcome for a patient. In this case, for example, a patient may have a decreased likelihood of long-term survival without cancer recurrence and / or cancer metastasis. Such a negative correlation indicates that the patient is more likely to have a poor prognosis, e.g., a high hazard ratio (e.g., HR>1.0). "Correlated" is also used herein to refer to the strength of association between the expression levels of two different genes, such that the expression level of a first gene can be replaced in a given algorithm by the expression level of a second gene, taking into account the correlation of their expressions. Such "correlated expression" of two genes that can be replaced in an algorithm typically refers to gene expression levels that are positively correlated with each other; for example, if increased expression of a first gene is positively correlated with an outcome (e.g., an increased likelihood of a good clinical outcome), a second gene that is co-expressed with the first gene and exhibits correlated expression will also be positively correlated with the same outcome.
[0039] The term "recurrence," as used herein, refers to local or distant (metastatic) recurrence of cancer. For example, breast cancer may come back as a local recurrence (near the treated breast or tumor surgery site) or as a distant recurrence within the body. The most common sites of recurrence for breast cancer include lymph nodes, bone, liver, or lung.
[0040] The term "polynucleotide," when used in the singular or plural, generally refers to any polyribonucleotide or polydeoxyribonucleotide, which may be unmodified RNA or DNA, or modified RNA or DNA. Thus, for example, polynucleotides as defined herein include, without limitation, single- and double-stranded DNA, DNA containing single- and double-stranded regions, single- and double-stranded RNA, and RNA containing single- and double-stranded regions, and hybrid molecules containing DNA and RNA, which may be single-stranded or, more typically, double-stranded, or may contain single- and double-stranded regions. Additionally, the term "polynucleotide," as used herein, refers to triple-stranded regions containing RNA or DNA or both RNA and DNA. The strands in such regions may be from the same molecule or from different molecules. Such regions may include all of one or more of the molecules, but more typically, only a portion of a molecule. One of the molecules in a triple-helical region is often an oligonucleotide. The term "polynucleotide" specifically includes cDNA. The term includes DNAs (including cDNAs) and RNAs that contain one or more modified bases. Thus, DNAs or RNAs with backbones modified for stability or for other reasons are "polynucleotides" as that term is intended herein. Additionally, DNAs or RNAs that contain unusual bases, such as inosine, or modified bases, e.g., tritiated bases, are included within the scope of the term "polynucleotide" as defined herein. In general, the term "polynucleotide" encompasses all chemically, enzymatically, and / or metabolically modified forms of unmodified polynucleotides, as well as the chemical forms of DNA and RNA characteristic of viruses and cells, including simple and complex cells.
[0041] The term "oligonucleotide" refers to a relatively short polynucleotide, including, without limitation, single-stranded deoxyribonucleotides, single- or double-stranded ribonucleotides, RNA:DNA hybrids, and double-stranded DNA. Oligonucleotides, such as single-stranded DNA probe oligonucleotides, are often synthesized by chemical methods, for example, using commercially available automated oligonucleotide synthesizers. However, oligonucleotides may also be produced by a variety of other methods, such as in vitro recombinant DNA-mediated methods and DNA expression in cells and organisms.
[0042] The term "amplification" refers to the process of generating multiple copies of a gene or RNA transcript in a particular sample or cell line. The replicated region (the stretch of amplified polynucleotide) is often referred to as an "amplicon." Typically, the amount of messenger RNA (mRNA) produced, i.e., the level of gene expression, also increases proportionally to the number of copies made of a particular expressed gene.
[0043] The term "estrogen receptor (ER)" refers to the estrogen receptor status of a cancer patient. If a significant number of estrogen receptors are present in the cancer cells, the tumor is ER-positive, while ER-negative indicates that the cells do not have a significant number of receptors. The definition of "significant" varies depending on the examination site and method (e.g., immunohistochemistry, PCR). The ER status of a cancer patient can be assessed by various known means. For example, the ER level in breast cancer is determined by measuring the expression level of the gene encoding the estrogen receptor in a breast tumor sample obtained from the patient.
[0044] The term "tumor," as used herein, refers to all neoplastic cell growth and proliferation, whether malignant or benign, and all pre-cancerous and cancerous cells and tissues.
[0045] The terms "cancer" and "cancerous" refer to or describe the physiological condition in mammals that is typically characterized by unregulated cell growth. Examples of cancer include, but are not limited to, breast cancer, ovarian cancer, colon cancer, lung cancer, prostate cancer, hepatocellular carcinoma, gastric cancer, pancreatic cancer, cervical cancer, liver cancer, bladder cancer, urinary tract cancer, thyroid cancer, renal cancer, carcinoma, melanoma, and brain cancer.
[0046] The gene subset identified herein as the "stromal group" includes genes synthesized primarily by stromal cells and involved in the stromal reaction, as well as genes co-expressed with stromal group genes. "Stromal cells" are defined herein as connective tissue cells that constitute the supportive structure of biological tissues. Stromal cells include fibroblasts, immune cells, pericytes, endothelial cells, and inflammatory cells. "Stromal reaction" refers to the connective tissue-forming reaction of host tissues at the site of a primary tumor or invasion. See, e.g., E. Rubin and J. Farber, "Pathology," pp. 985-986 (2nd ed. 1994). The stromal group includes, for example, CDH11, TAGLN, ITGA4, INHBA, COLIA1, COLIA2, FN1, CXCL14, TNFRSF1, CXCL12, C10ORF116, RUNX1, GSTM2, TGFB3, CAV1, DLC1, TNFRSF10, F3, and DICER1, as well as co-expressed genes identified in Tables 16-18.
[0047] The gene subsets identified herein as the "metabolic group" include genes associated with cellular metabolism, such as transport proteins for iron transport, cellular iron homeostasis pathways, and homeostatic biochemical metabolic pathways, as well as genes co-expressed with metabolic group genes. The metabolic group includes, for example, TFRC, ENO1, IDH2, ARF1, CLDN4, PRDX1, and GBP1, as well as the co-expressed genes identified in Tables 16-18.
[0048] The gene subset identified herein as the "immune group" includes genes involved in cellular immunoregulatory functions such as T cell and B cell intracellular trafficking, lymphocyte-associated or lymphocyte markers, and interferon-regulated genes, as well as genes co-expressed with immune group genes. The immune group includes, for example, CCL19 and IRF1, and the co-expressed genes identified in Tables 16-18.
[0049] The subset of genes identified herein as the "proliferation group" includes genes associated with cell development and division, cell cycle and mitotic regulation, angiogenesis, cell replication, nuclear transport / stability, wnt signaling, apoptosis, and genes co-expressed with proliferation group genes, including, for example, PGF, SPC25, AURKA, BIRC5, BUB1, CCNB1, CENPA, KPNA, LMNB1, MCM2, MELK, NDC80, TPX2M, and WISP1, as well as the co-expressed genes identified in Tables 16-18.
[0050] The term "co-expressed," as used herein, refers to a statistical correlation between the expression level of one gene and the expression level of another gene. Pairwise co-expression can be calculated by various methods known in the art, for example, by calculating the Pearson correlation coefficient or the Spearman correlation coefficient. Co-expressed gene cliques can also be identified using graph theory.
[0051] As used herein, the terms "gene clique" and "clique" refer to a subgraph of a graph in which every vertex is connected by an edge to every other vertex in the subgraph.
[0052] As used herein, a "maximal clique" is a clique to which no other vertices can be added and still be a clique.
[0053] The "pathology" of cancer includes any phenomenon that compromises the health of the patient, including, without limitation, abnormal or uncontrolled cell growth, metastasis, interference with the normal function of neighboring cells, release of abnormal levels of cytokines or other secretory products, suppression or exacerbation of inflammatory or immune responses, neoplasia, precancerous conditions, malignant tumors, invasion of surrounding or distant tissues or organs, e.g., lymph nodes, etc.
[0054] A "computer-based system" refers to a system of hardware, software, and data storage media used to analyze information. The minimum hardware of a patient computer-based system includes a central processing unit (CPU) and hardware for data input, data output (e.g., a display), and data storage. One of ordinary skill in the art will readily recognize that any currently available computer-based system and / or its components are suitable for use in connection with the methods of the present disclosure. A data storage medium may include any product containing a record of this information as described above, or a memory access device capable of accessing such a product.
[0055] "Recording" data, programming, or other information on a computer-readable medium refers to a method for storing information, using any such method as known in the art. Any convenient data storage structure can be selected based on the means used to access the stored information. A variety of data processor programs and formats can be used for storage, such as word processing text files, database formats, etc.
[0056] A "processor" or "computing means" refers to any combination of hardware and / or software capable of performing the required functions. For example, a suitable processor may be a programmable digital microprocessor, such as those available in the form of an electronic controller, mainframe, server, or personal computer (desktop or portable). If the processor is programmable, suitable programming may be transmitted to the processor from a remote location or may be pre-stored on a computer program product (such as a portable or fixed computer-readable storage medium, whether magnetic, optical, or solid-state based). For example, a magnetic medium or optical disk may carry the programming, which may be read by a suitable reader in communication with each processor at its corresponding station.
[0057] As used herein, "graph theory" refers to the field of study in computer science and mathematics that represents situations using diagrams that include a set of points and lines connecting some of those points. This diagram is called a "graph," and the points and lines are called the "vertices" and "edges" of the graph. In gene co-expression analysis, genes (or their equivalent identifiers, e.g., array probes) can be represented as nodes or vertices of the graph. If a measure of similarity between two genes (e.g., correlation coefficient, mutual information, and alternating conditional expectation) is higher than a significance threshold, the two genes are said to be co-expressed, and an edge can be drawn in the graph. Once co-expression edges have been drawn for all possible gene pairs for a given test, all maximal cliques are calculated. The resulting maximal clique is defined as a gene clique. A gene clique is a computationally co-expressed group of genes that meets a predefined criterion.
[0058] The "stringency" of a hybridization reaction can be readily determined by one of ordinary skill in the art and is generally an empirical calculation dependent on probe length, washing temperature, and salt concentration. Generally, longer probes require higher temperatures for proper annealing, while shorter probes require lower temperatures. Hybridization generally depends on the ability of denatured DNA to reanneal when complementary strands are present in an environment below their melting temperature. The higher the desired homology between the probe and hybridizable sequence, the higher the relative temperature that can be used. As a result, higher relative temperatures tend to increase the stringency of the reaction conditions, while lower temperatures tend not to. For further details and explanation of the stringency of hybridization reactions, see Ausubel et al., "Current Protocols in Molecular Biology," Wiley Interscience Publishers, (1995).
[0059] "Stringent conditions" or "high stringency conditions," as defined herein, typically refer to: (1) low ionic strength and high temperature wash conditions, e.g., 0.015 M sodium chloride / 0.0015 M sodium citrate / 0.1% sodium dodecyl sulfate at 50°C; (2) denaturing agents such as formamide during hybridization, e.g., 50% (v / v) formamide and 0.1% bovine serum albumin / 0.1% Ficoll / 0.1% polyvinylpyrrolidone / 50 mM sodium phosphate buffer at pH 6.5, plus 750 mM sodium chloride, 75 mM sodium citrate, at 42°C; or (3) 50% formamide, 5x SSC (0.75M NaCl, 0.075 M sodium citrate), 50 mM sodium phosphate (pH 6.8), 0.1% sodium pyrophosphate, 5x Denhardt's solution, sonicated salmon sperm DNA (50 μg / ml), 0.1% SDS, and 10% dextran sulfate, followed by washes with 0.2x SSC (sodium chloride / sodium citrate) at 42°C and 50% formamide at 55°C, followed by a high stringency wash consisting of 0.1x SSC with EDTA at 55°C.
[0060] "Moderately stringent conditions" can be identified as described by Sambrook et al., "Molecular Cloning: A Laboratory Manual," New York: Cold Spring Harbor Press, 1989, and may include the use of less stringent wash solutions and hybridization conditions (e.g., temperature, ionic strength, and % SDS) than those described above. An example of moderately stringent conditions is overnight incubation at 37°C in a solution containing 20% formamide, 5x SSC (150 mM NaCl, 15 mM trisodium citrate), 50 mM sodium phosphate (pH 7.6), 5x Denhardt's solution, 10% dextran sulfate, and 20 mg / ml denatured, sheared salmon sperm DNA, followed by washing the filter with 1x SSC at approximately 37-50°C. One of skill in the art will know how to adjust temperature, ionic strength, etc., as needed to accommodate factors such as probe length.
[0061] In the context of the present invention, reference to "at least one," "at least two," "at least five," etc. of the genes listed in any particular gene set means any one or any combination of the listed genes.
[0062] The term "node-negative" cancer, eg, "node-negative" breast cancer, is used herein to refer to cancer that has not spread to the lymph nodes.
[0063] The terms "splicing" and "RNA splicing" are used interchangeably and refer to RNA processing that removes introns and joins exons to produce a mature mRNA with a contiguous coding sequence that is transported into the cytoplasm of a eukaryotic cell.
[0064] Theoretically, the term "exon" refers to any segment of an interrupted gene that is expressed in the mature RNA product (B. Lewin, "Genes IV," Cell Press, Cambridge Mass., 1990). Theoretically, the term "intron" refers to any segment of DNA that is transcribed but removed from the transcript by splicing together with the exons on either side of it. Operationally, exon sequences appear in the mRNA sequence of a gene as defined by a reference sequence number. Operationally, intron sequences are intervening sequences within the genomic DNA of a gene that are surrounded by exon sequences and have GT and AG splice consensus sequences at their 5' and 3' boundaries.
[0065] Gene expression assays The present disclosure provides methods which employ, unless otherwise indicated, conventional techniques of molecular biology (including recombinant techniques), microbiology, cell biology, and biochemistry, which techniques are within the skill of the art. Such techniques are fully explained in such references as "Molecular Cloning: A Laboratory Manual," 2nd ed. (Sambrook et al., 1989); "Oligonucleotide Synthesis" (M.J. Gait, ed., 1984); "Animal Cell Culture" (R.I. Freshney, ed., 1987); "Methods in Enzymology" (Academic Press, Inc.); "Handbook of Experimental Immunology," 4th ed. (D.M. Weir & C.C. Blackwell, eds., Blackwell Science Inc., 1987); "Gene Transfer Vectors for Mammalian Cells" (J.M. Miller & M.P. Calos, eds., 1987); "Current Protocols in Molecular Biology" (F.M. Ausubel et al., eds., 1987); and "PCR: The Polymerase Chain Reaction," (Mullis et al., eds., 1994).
[0066] 1. Gene Expression Profiling Gene expression profiling methods include methods based on polynucleotide hybridization analysis, polynucleotide sequencing, and proteomics. The most commonly used methods known in the art for quantifying mRNA expression in a sample include Northern blotting and in situ hybridization (Parker & Barnes, Methods in Molecular Biology 106:247-283 (1999)); RNase protection assay (Hod, Biotechniques 13:852-854 (1992)); and PCR-based methods, such as reverse transcription polymerase chain reaction (RT-PCR) (Weis et al., Trends in Genetics 8:263-264 (1992)). Alternatively, antibodies capable of recognizing specific duplexes, including DNA duplexes, RNA duplexes, and DNA-RNA hybrid duplexes or DNA-protein duplexes, may be used.
[0067] 2. PCR-based gene expression profiling methods a. Reverse transcriptase PCR (RT-PCR) Of the techniques listed above, the most sensitive and most flexible quantitative method is RT-PCR, which can be used to compare mRNA levels in different samples of normal and tumor tissues with or without drug treatment, thereby characterizing gene expression patterns, distinguishing between closely related mRNAs, and analyzing RNA structure.
[0068] The first step is the isolation of mRNA from a target sample. The starting material is typically total RNA isolated from human tumors or tumor cell lines and their corresponding normal tissues or cell lines. Thus, RNA can be isolated from various primary tumors, including tumors of the breast, lung, colon, prostate, brain, liver, kidney, pancreas, spleen, thymus, testis, ovary, uterus, etc., or tumor cell lines, as well as pooled DNA from healthy donors. When the source of mRNA is a primary tumor, mRNA can be extracted, for example, from frozen or archived, paraffin-embedded and fixed (e.g., formalin-fixed) tissue samples.
[0069] General methods for mRNA extraction are known in the art and are disclosed in standard molecular biology textbooks, including Ausubel et al., "Current Protocols of Molecular Biology," John Wiley and Sons (1997). Methods for RNA extraction from paraffin-embedded tissues are disclosed in Rupp and Locker, Lab Invest. 56:A67 (1987) and De Andres et al., BioTechniques 18:42044 (1995). In particular, RNA isolation can be performed using purification kits, buffer sets, and proteases from commercial manufacturers, such as Qiagen, according to the manufacturer's instructions. For example, total RNA from cultured cells can be isolated using Qiagen RNeasy mini-columns. Other commercially available RNA isolation kits include the MasterPure™ Complete DNA and RNA Purification Kit (EPICENTRE®, Madison, WI) and the Paraffin Block RNA Isolation Kit (Ambion, Inc.). Total RNA from tissue samples can be isolated using RNA Stat-60 (Tel-Test). RNA prepared from tumors can be isolated, for example, by cesium chloride density gradient centrifugation.
[0070] In some cases, it may be appropriate to amplify RNA before initiating expression profiling. Often, only a very limited amount of clinical specimen is available for molecular analysis. This may be due to the tissue having already been used for other laboratory analyses, or due to the original specimen being very small, as in the case of needle biopsies or very small primary tumors. When tissue is limited in quantity, generally only a small amount of total RNA can be recovered from the specimen, resulting in only a limited number of genomic markers that can be analyzed in the specimen. RNA amplification compensates for this limitation by faithfully replicating the original RNA sample as a much larger amount of RNA with the same relative composition. Unlimited genomic analysis can be performed using amplified copies of this original RNA specimen to discover biomarkers associated with the clinical characteristics of the original biological specimen. This effectively immortalizes clinical research specimens for the purposes of genomic analysis and biomarker discovery.
[0071] Because RNA cannot serve as a template for PCR, the first step in gene expression profiling by real-time RT-PCR (RT-PCR) is reverse transcription of the RNA template into cDNA, followed by its exponential amplification in a PCR reaction. The two most commonly used reverse transcriptases are avian myeloblastosis virus reverse transcriptase (AMV-RT) and Moloney murine leukemia virus reverse transcriptase (MMLV-RT). The reverse transcription step is typically primed using specific primers, random hexamers, or oligo-dT primers, depending on the context and goal of expression profiling. For example, extracted RNA can be reverse transcribed using a GeneAmp RNA PCR kit (Perkin Elmer, CA, USA) according to the manufacturer's instructions. The resulting cDNA can then be used as a template in a subsequent PCR reaction. For further details, see, e.g., Held et al., Genome Research 6:986-994 (1996).
[0072] While various thermostable DNA-dependent DNA polymerases can be used in the PCR step, Taq DNA polymerase is typically used, which possesses 5'-3' nuclease activity but lacks 3'-5' proofreading endonuclease activity. Thus, TaqMan® PCR typically utilizes the 5'-nuclease activity of Taq or Tth polymerase to hydrolyze hybridization probes bound to its target amplicon; however, any enzyme with equivalent 5' nuclease activity can be used. A typical amplicon in a PCR reaction is generated using two oligonucleotide primers. A third oligonucleotide, the probe, is designed to detect the nucleotide sequence located between the two PCR primers. The probe is non-developable by the Taq DNA polymerase enzyme and is labeled with a reporter and a quencher fluorescent dye. When the two dyes are positioned close to each other on the probe, any laser-induced emission from the reporter dye is quenched by the quencher dye. During the amplification reaction, the Taq DNA polymerase enzyme cleaves the probe in a template-dependent manner. The resulting probe fragments dissociate in solution, and the signal from the released reporter dye is free from the quenching effect of the second fluorophore. One molecule of reporter dye is liberated with each new molecule synthesized, and detection of the unquenched reporter dye provides the basis for quantitative interpretation of the data.
[0073] TaqMan® RT-PCR can be performed using commercially available instruments, such as the ABI PRISM 7900® Sequence Detection System™ (Perkin-Elmer-Applied Biosystems, Foster City, CA, USA) or the LightCycler® 480 Real-Time PCR System (Roche Diagnostics, GmbH, Penzberg, Germany). In a preferred embodiment, the 5' nuclease procedure is performed on a real-time quantitative PCR instrument such as the ABI PRISM 7900® Sequence Detection System™. This system consists of a thermocycler, a laser, a charge-coupled device (CCD), a camera, and a computer. The system amplifies samples in a 384-well format on the thermocycler. During amplification, laser-induced fluorescent signals are collected in real time for all 384 wells via fiber optic cables and detected by the CCD. The system includes software for running the instrument and for data analysis.
[0074] 5'-nuclease assay data was initially C t , or threshold cycle. As mentioned above, fluorescence values are recorded during each cycle, and these fluorescence values correspond to the amount of product amplified up to that point in the amplification reaction. The point at which the fluorescence signal is first recorded as statistically significant is called the threshold cycle (C t )
[0075] To minimize errors and the effects of sample-to-sample variation, RT-PCR is usually performed using an internal standard. An ideal internal standard is expressed at a consistent level across different tissues and is unaffected by experimental handling. The RNAs most commonly used to normalize gene expression patterns are mRNAs for the housekeeping genes glyceraldehyde-3-phosphate dehydrogenase (GAPDH) and β-actin.
[0076] The steps of a typical gene expression profiling protocol using fixed, paraffin-embedded tissue as an RNA source, including mRNA isolation, purification, primer extension, and amplification, have been described in various published journal articles. M. Cronin, Am J Pathol 164(1):35-42 (2004). Briefly, a typical method begins with cutting approximately 10 μm thick sections of a paraffin-embedded tumor tissue sample. RNA is then extracted, and protein and DNA are removed. After analysis of RNA concentration, RNA repair and / or amplification steps may be included, if necessary, and the RNA is reverse-transcribed using gene-specific primers, followed by RT-PCR.
[0077] b. Design of intron-based PCR primers and probes PCR primers and probes can be designed based on the exon or intron sequences present in the mRNA transcript of the gene of interest. Before primer / probe design can be performed, the target gene sequence must be mapped to the human genome assembly to identify intron-exon boundaries and the overall gene structure. This can be done using publicly available software such as Primer3 (Whitehead Inst.) and Primer Express® (Applied Biosystems).
[0078] If necessary or desirable, nonspecific signals can be reduced by masking repetitive sequences in the target sequence. An exemplary tool for accomplishing this is the Repeat Masker program available online from Baylor College of Medicine, which screens a DNA sequence against a library of repetitive elements and returns the query sequence with the repetitive elements masked. The masked intron and exon sequences can then be used to design primer and probe sequences for the desired target sites using any commercially available or otherwise publicly available primer / probe design package, such as Primer Express (Applied Biosystems); MGB Assay-by-Design (Applied Biosystems); or Primer3 (Steve Rozen and Helen J. Skaletsky (2000), "Primer3 on the WWW for general users and for biologist programmers." Rrawetz S, Misener S (eds.), Bioinformatics Methods and Protocols: Methods in Molecular Biology, Humana Press, Totowa, NJ, pp. 365-386).
[0079] Other factors that can affect PCR primer design include primer length, melting temperature (Tm), G / C content, specificity, complementary primer sequence, and 3' end sequence. In general, optimal PCR primers are generally 17-30 bases long, contain about 20-80%, e.g., about 50-60%, G+C bases, and exhibit a Tm of 50-80°C, e.g., about 50-70°C.
[0080] For further guidance on PCR primer and probe design, see, e.g., Dieffenbach, CW et al., "General Concepts for PCR Primer Design," in PCR Primer, A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York, 1995, pp. 133-155; Innis and Gelfand, "Optimization of PCRs," in PCR Protocols, A Guide to Methods and Applications, CRC Press, London, 1994, pp. 5-11; and Plasterer, TN, "Primerselect: Primer and probe design," Methods Mol. Biol. 70:520-527 (1997), the entire disclosures of which are hereby expressly incorporated by reference.
[0081] Table A provides further information regarding primer, probe, and amplicon sequences associated with the examples disclosed herein.
[0082] c. MassARRAY system In the MassARRAY-based gene expression profiling method developed by Sequenom, Inc. (San Diego, CA), after RNA isolation and reverse transcription, the resulting cDNA is spiked with a synthetic DNA molecule (competitor) that matches the target cDNA region at every position except a single base and serves as an internal standard. The cDNA / competitor mixture is PCR-amplified and subjected to post-PCR shrimp alkaline phosphatase (SAP) enzymatic treatment, resulting in the dephosphorylation of remaining nucleotides. After inactivation of the alkaline phosphatase, the PCR products from the competitor and cDNA are subjected to primer extension, which generates distinct mass signals for the competitor- and cDNA-derived PCR products. After purification, these products are sorted onto a chip array, which is preloaded with the components required for analysis by matrix-assisted laser desorption / ionization time-of-flight mass spectrometry (MALDI-TOF MS). The cDNA present in the reaction mixture is then quantified by analyzing the peak area ratios of the resulting mass spectra. For further details, see, e.g., Ding and Cantor, Proc. Natl. Acad. Sci. USA 100:3059-3064 (2003).
[0083] d. Other PCR-based methods Other PCR-based techniques include, for example, differential display (Liang and Pardee, Science 257:967-971 (1992)); amplified fragment length polymorphism (iAFLP) (Kawamoto et al., Genome Res. 12:1305-1312 (1999)); BeadArray™ technology (Illumina, San Diego, CA; Oliphant et al., Discovery of Markers for Disease (Biotechniques Supplement), June 2002; Ferguson et al., Analytical Chemistry 72:5618 (2000)); and BeadsArray for Detection of Gene Expression (BADGE) (Yang et al., Genome Res. 2002), which uses the commercially available Luminex 100 LabMAP system and multicolor-coded microspheres (Luminex Corp., Austin, TX) in a rapid assay of gene expression. Res. 11:1888-1898 (2001); and high-coverage expression profiling (HiCEP) analysis (Fukumura et al., Nucl. Acids. Res. 31(16)e94 (2003)).
[0084] 3. Microarray Differential gene expression can also be identified or confirmed using microarray technology. Thus, the expression profile of breast cancer-related genes can be measured using microarray technology in either fresh or paraffin-embedded tumor tissues. In this method, polynucleotide sequences of interest (including cDNAs and oligonucleotides) are arranged on a plate or array on a microchip substrate. The arrayed sequences are then hybridized with specific DNA probes derived from cells or tissues of interest. Just like the RT-PCR method, the source of mRNA is typically total RNA isolated from human tumors or tumor cell lines and corresponding normal tissues or cell lines. Thus, RNA can be isolated from various primary tumors or tumor cell lines. When the source of mRNA is a primary tumor, mRNA can be extracted, for example, from frozen or archived paraffin-embedded and fixed (e.g., formalin-fixed) tissue samples that are routinely prepared and stored in daily clinical practice.
[0085] In a specific embodiment of the microarray technique, PCR-amplified inserts of cDNA clones are applied to a high-density array substrate. Preferably, at least 10,000 nucleotide sequences are applied to the substrate. The genes arranged in the microarray, immobilized on a microchip with 10,000 elements each, are suitable for hybridization under stringent conditions. Fluorescently labeled cDNA probes can be generated by incorporating fluorescent nucleotides through reverse transcription of RNA extracted from the tissue of interest. The labeled cDNA probes applied to the chip hybridize specifically with each DNA spot on the array. After stringent washing to remove nonspecifically bound probes, the chip is scanned by confocal laser microscopy or another detection method, such as a CCD camera. Quantifying hybridization of each arrayed element allows assessment of the corresponding mRNA abundance. Separately labeled cDNA probes with dual-color fluorescence generated from two RNA sources are hybridized to the array in a pairwise manner. Thus, the relative abundance of transcripts from the two sources, each corresponding to a specific gene, is determined simultaneously. Miniaturized hybridization allows for the convenient and rapid evaluation of expression patterns of large numbers of genes. Such methods have the sensitivity necessary to detect low-abundance transcripts expressed at a few copies per cell and have been shown to reproducibly detect at least approximately two-fold differences in expression levels (Schena et al., Proc. Natl. Acad. Sci. USA 93(2):106-149 (1996)). Microarray analysis can be performed using commercially available equipment, such as Affymetrix GenChip technology or Agilent microarray technology, following the manufacturer's protocols.
[0086] The deployment of microarray methods for large-scale analysis of gene expression allows for the systematic search for molecular markers of cancer classification and outcome prediction in various tumor types.
[0087] 4. Gene expression analysis by nucleic acid sequencing Nucleic acid sequencing technologies are a preferred method for analyzing gene expression. The underlying principle of these methods is that the number of times a cDNA sequence is detected in a sample is directly related to the relative expression of the mRNA corresponding to that sequence. These methods are sometimes referred to by the term Digital Gene Expression (DGE), reflecting the discrete, numerical nature of the data obtained. Early methods applying this principle were Serial Analysis of Gene Expression (SAGE) and Massively Parallel Signature Sequencing (MPSS). See, for example, S. Brenner et al., Nature Biotechnology 18(6):630-634 (2000). More recently, the advent of "next-generation" sequencing technologies has made DGE simpler, more high-throughput, and less expensive. As a result, DGE is now available to many more laboratories, allowing the expression of many more genes to be screened in many more individual patient samples than previously possible. See, for example, J. Marioni, Genome Research 18(9):1509-1517 (2008); R. Morin, Genome Research 18(4):610-621 (2008); A. Mortazavi, Nature Methods 5(7):621-628 (2008); N. Cloonan, Nature Methods 5(7):613-619 (2008).
[0088] 5. RNA Isolation from Body Fluids Methods for isolating RNA for expression analysis from blood, plasma, and serum (see, e.g., Tsui NB et al. (2002) 48, 1647-53 and references cited therein) and from urine (see, e.g., Boom R et al. (1990) J Clin Microbiol. 28, 495-503 and references cited therein) have been described.
[0089] 6. Immunohistochemistry Immunohistochemical methods are also suitable for detecting the expression levels of the prognostic markers of the present invention. Therefore, antibodies or antisera, preferably polyclonal antisera, and most preferably monoclonal antibodies specific to each marker, are used to detect expression. Antibodies can be detected by directly labeling the antibody itself, for example, with a radioactive label, a fluorescent label, a hapten label such as biotin, or an enzyme such as horseradish peroxidase or alkaline phosphatase. Alternatively, an unlabeled primary antibody is used in conjunction with a labeled secondary antibody, for example, an antiserum, a polyclonal antiserum, or a monoclonal antibody specific to the primary antibody. Immunohistochemical protocols and kits are known in the art and are commercially available.
[0090] 7. Proteomics The term "proteome" is defined as the totality of proteins present in a sample (e.g., a tissue, an organism, or a cell culture) at a given time point. Proteomics includes, inter alia, the study of global changes in protein expression in a sample (also referred to as "expression proteomics"). Proteomics typically involves the following steps: (1) separating individual proteins in a sample by 2-D gel electrophoresis (2-D PAGE); (2) identifying the individual proteins recovered from the gel, for example, by mass spectrometry or N-terminal sequencing, and (3) analyzing the data using bioinformatics. Proteomics methods are useful complements to other gene expression profiling methods and can be used alone or in combination with other methods to detect the products of the prognostic markers of the present invention.
[0091] 8. Overview of mRNA Isolation, Purification, and Amplification The steps of a typical gene expression profiling protocol using fixed, paraffin-embedded tissue as an RNA source, including mRNA isolation, purification, primer extension, and amplification, are provided in various published journal articles (e.g., T.E. Godfrey et al., J. Molec. Diagnostics 2:84-91
[2000] ; K. Specht et al., Am. J. Pathol. 158:419-29
[2001] ). Briefly, a typical method begins with cutting approximately 10 μm-thick sections of a paraffin-embedded tumor tissue sample. RNA is then extracted, and protein and DNA are removed. After analysis of the RNA concentration, RNA repair and / or amplification steps may be included, if necessary, and the RNA is reverse-transcribed using gene-specific primers, followed by RT-PCR. Finally, the data is analyzed to identify the best treatment option or options available to the patient based on the characteristic gene expression patterns identified in the examined tumor sample and the predicted likelihood of cancer recurrence.
[0092] 9. Normalization Expression data used in the methods disclosed herein can be normalized. Normalization can be used to correct for (normalize out) differences in the amount of RNA assayed and variations in the quality of the RNA used, for example, to reduce C t This refers to a method for removing unwanted sources of systematic variation in measurements. In the context of RT-PCR experiments involving archival fixed, paraffin-embedded tissue samples, known sources of systematic variation include the degree of RNA degradation relative to the age of the patient sample and the type of fixative used to preserve the sample. Other sources of systematic variation arise from laboratory processing conditions.
[0093] Assays can provide normalization by incorporating the expression of specific normalizing genes, which do not have significantly different expression levels under relevant conditions. Exemplary normalizing genes include housekeeping genes such as PGK1 and UBB. (See, e.g., E. Eisenberg et al., Trends in Genetics 19(7):362-365 (2003)). Normalization can be achieved by using the mean or median signal (C ) of all or a large subset of the assayed genes. T ) (global normalization approach). Generally, normalization genes, also referred to as reference genes, should be genes that are known not to be significantly differently expressed in colorectal cancer compared to non-cancerous colorectal tissue and that are not significantly affected by various sample and process conditions, thus providing for normalization removal of external effects.
[0094] Unless otherwise specified, the normalized expression level of each mRNA / test tumor / patient can be expressed as a percentage of the expression level measured in the reference set.A sufficiently large number (for example, 40) of tumors in the reference set will provide a distribution of the normalized levels of each mRNA species.The levels measured in the specific tumor sample being analyzed will fall within this range at some percentile, which can be determined by methods known in the art.
[0095] In exemplary embodiments, one or more of the following genes are used as a reference for normalizing expression data: AAMP, ARF1, EEF1A1, ESD, GPS1, H3F3A, HNRPC, RPL13A, RPL41, RPS23, RPS27, SDHA, TCEA1, UBB, YWHAZ, B-actin, GUS, GAPDH, RPLPO, and TFRC. For example, the calibrated weighted average C for each of the prognostic genes is t The measurements may be normalized to the average value of at least three reference genes, at least four reference genes, or at least five reference genes.
[0096] Those skilled in the art will recognize that normalization can be achieved in a variety of ways, and that the techniques described above are intended to be exemplary only and not exhaustive.
[0097] Report Results The disclosed methods are suitable for generating reports summarizing the prediction or forecast of clinical outcomes obtained from the disclosed methods. A "report," as used herein, is an electronic or tangible document containing report elements that provide information of interest regarding a probability or risk assessment and its results. A subject report includes at least a probability or risk assessment, e.g., an indicator regarding the risk of breast cancer recurrence, including local recurrence and metastasis of breast cancer. A subject report may include one or more assessments or estimates of disease-free survival, recurrence-free survival, metastasis-free survival, and overall survival. A subject report may be generated entirely or partially electronically and may be presented, for example, on an electronic display (e.g., a computer monitor). A report may further include one or more of the following: 1) information regarding the testing facility; 2) service provider information; 3) patient data; 4) sample data; 5) an interpretive report, which may include various information, including a) indicators; b) test data (where the test data may include normalized levels of one or more genes of interest); and 6) other characteristics.
[0098] Accordingly, the present disclosure provides methods for generating reports and the resulting reports. The report may include a summary of expression levels of RNA transcripts, or expression products of such RNA transcripts, for specific genes in cells obtained from a patient's tumor. The report may include information regarding the patient's prognostic covariates. The report may include an estimation that the patient is at high risk of recurrence. This estimation may be in the form of a score or patient stratification scheme (e.g., low, intermediate, or high risk of recurrence). The report may include information relevant to aiding decisions regarding appropriate surgery (e.g., partial or total mastectomy) or treatment for the patient.
[0099] Thus, in some embodiments, the methods of the present disclosure further comprise generating a report containing information regarding a predicted clinical outcome for the patient, e.g., risk of recurrence. For example, the methods disclosed herein can further comprise generating or outputting a report providing the subject's risk assessment, which can be provided in electronic form (e.g., an electronic display on a computer monitor) or in tangible form (e.g., a report printed on paper or other tangible medium).
[0100] A report is provided to the user containing information regarding the patient's estimated prognosis (e.g., the likelihood that a patient with breast cancer will respond to surgery and / or treatment with a favorable prognosis or positive clinical outcome). The likelihood assessment is hereinafter referred to as a "risk report," or simply a "risk score." The person or entity generating the report (the "report generator") may also perform the likelihood assessment. The report generator may also perform one or more of sample collection, sample processing, and data generation; for example, the report generator may also perform one or more of the following: a) sample collection; b) sample processing; c) measuring levels of risk genes; d) measuring levels of reference genes; and e) determining normalized levels of risk genes. Alternatively, an entity other than the report generator may perform one or more of sample collection, sample processing, and data generation.
[0101] For clarity, it should be noted that the term "user," used synonymously with "client," is intended to refer to the person or entity to whom the report is sent, and may be the same person or entity that does one or more of the following: a) collects the sample; b) processes the sample; c) provides the sample or the processed sample; and d) generates data (e.g., levels of risk genes; levels of one or more reference gene products; normalized levels of risk genes ("prognostic genes") used in probability assessment. In some cases, the person or persons or entities that provide the sample collection and / or sample processing and / or data generation may be different persons from the person who receives the results and / or report; however, to avoid confusion, both are referred to herein as "user" or "client." In certain embodiments, for example, when the methods are all performed on a single computer, the user or client provides the data input and views the data output. A "user" may be a medical professional (e.g., a clinician, a laboratory technician, a physician (e.g., an oncologist, a surgeon, a pathologist), etc.).
[0102] In embodiments in which a user performs only a portion of the method, the individual who views the data output after computerized data processing according to the disclosed method (e.g., pre-publication results providing a completed report), or the individual who views the "unfinished" report and provides manual intervention and completion of the interpretive report, is referred to herein as a "reviewer." The reviewer may be located at a location remote from the user (e.g., a service provided remotely from a medical facility where the user may be located).
[0103] Where government regulations or other restrictions apply (e.g., requirements by health insurance, malpractice insurance, or liability insurance), all results, whether fully or partially electronically generated, are subject to quality control routines before release to users.
[0104] Clinical utility The gene expression assays and information provided by practicing the methods disclosed herein will help physicians make more informed treatment decisions and tailor cancer treatment to the needs of individual patients, thereby maximizing therapeutic benefit and minimizing patient exposure to unnecessary treatments that may provide little or no significant benefit and often involve significant risks due to toxic side effects.
[0105] Single or multi-analyte gene expression tests can be used to measure the expression levels of one or more genes involved in each of several related physiological processes or constituent cellular characteristics. One or more expression levels may be used to calculate such a quantitative score, and such scores may be organized into subgroups (e.g., tertiles) where all patients within a given range are classified as belonging to a certain risk category (e.g., low, moderate, or high). Gene grouping may be performed, at least in part, based on gene contribution information by physiological function or constituent cellular characteristic, such as in the groups discussed above.
[0106] The usefulness of a genetic marker in predicting cancer may not be specific to that marker. Alternative markers with expression patterns comparable to the selected marker gene may be used in place of or in addition to the test marker. Due to such gene co-expression, substitution of expression level values should have little effect on the overall prognostic usefulness of the test. Closely similar expression patterns of two genes may result from both genes being involved in the same process and / or being under common regulatory control in colon tumor cells. Thus, the present disclosure contemplates the use of such co-expressed genes or gene sets as a substitute for or in addition to the prognostic methods of the present disclosure.
[0107] The molecular assays and related information provided by the methods disclosed herein for predicting clinical outcomes in cancer, e.g., breast cancer, have utility in many areas, including treating cancer through drug development and appropriate use, stratifying cancer patients for inclusion in (or exclusion from) clinical trials, assisting patients and physicians in treatment decisions, and providing economic benefits by targeting treatments based on individualized genomic profiles. For example, a recurrence score may be used on samples collected from patients in a clinical trial, and the test results may be used in conjunction with patient outcomes to determine whether patient subgroups are more or less likely to show absolute benefit from a new drug compared to the group as a whole or other subgroups. Furthermore, such methods can be used to identify clinical data subsets of patients who are expected to benefit from adjuvant therapy. Additionally, if the test results indicate an increased likelihood of a patient having a poor clinical outcome if treated with surgery alone, the patient will be more likely to be enrolled in the clinical trial, and if the test results indicate a decreased likelihood of a patient having a poor clinical outcome if treated with surgery alone, the patient will be less likely to be enrolled in the clinical trial.
[0108] Statistical analysis of gene expression levels Those skilled in the art will recognize that there are many statistical methods that can be used to determine whether there is a significant relationship between the outcome of interest (for example, the probability of survival, the probability of responding to chemotherapy) and the expression level of the marker gene as described herein.This relationship can be presented as a continuous recurrence score (RS), or patients can be stratified into risk groups (for example, low, intermediate, high).For example, a Cox proportional hazards regression model can be fitted to specific clinical endpoints (for example, RFS, DFS, OS).One assumption of the Cox proportional hazards regression model is the proportional hazards assumption, that is, the assumption that the effect parameter is multiplied by the base hazard.
[0109] Coexpression analysis The present disclosure provides genes that are co-expressed with specific prognostic and / or predictive genes identified as having a significant correlation with recurrence and / or treatment benefit. Genes often work together in a coordinated manner, i.e., they are co-expressed, to carry out specific biological processes. Groups of co-expressed genes identified for a disease process such as cancer can serve as biomarkers for disease progression and response to treatment. Such co-expressed genes can be assayed instead of, or in addition to, assaying the prognostic and / or predictive genes with which they are co-expressed.
[0110] Those skilled in the art will recognize that numerous currently known or later developed co-expression analysis methods are within the scope and spirit of the present invention. These methods may incorporate, for example, correlation coefficients, co-expression network analysis, clique analysis, etc., and may be based on expression data from RT-PCR, microarrays, sequencing, and other similar techniques. For example, pairwise correlation analysis based on the Pearson or Spearman correlation coefficient can be used to identify gene expression clusters. (See, e.g., Pearson K. and Lee A., Biometrika 2, 357 (1902); C. Spearman, Amer. J. Psychol 15:72-101 (1904); J. Myers, A. Well, "Research Design and Statistical Analysis," 508 (2nd ed., 2003)). Generally, a correlation coefficient of 0.3 or greater with a sample size of at least 20 is considered statistically significant. (See, e.g., G. Norman, D. Streiner, "Biostatistics: The Bare Essentials," pp. 137-138 (3rd ed. 2007).) In one embodiment disclosed herein, co-expressed genes were identified using a Spearman correlation value of at least 0.7.
[0111] computer program The values from the assays described above, such as expression data, recurrence scores, treatment scores, and / or benefit scores, can be calculated and stored manually. Alternatively, the above steps may be performed, in whole or in part, by a computer program product. Accordingly, the present invention provides a computer program product including a computer-readable storage medium having a computer program stored thereon. When read by a computer, the program can perform relevant calculations based on values obtained from the analysis of one or more biological samples from an individual (e.g., gene expression levels, normalization of values from assays, thresholding, and conversion to scores and / or graphical representations of likelihood of recurrence / response to chemotherapy, gene co-expression analysis, or clique analysis, etc.). The computer program product has stored therein a computer program for performing the calculations.
[0112] The present disclosure provides a system for executing the programs described above, which generally includes: a) a central computing environment; b) an input device operably connected to the computing environment for receiving patient data, where the patient data may include, for example, expression levels or other values obtained from assays using patient-derived biological samples, or microarray data as described in detail above; c) an output device connected to the computing environment for providing information to a user (e.g., medical personnel); and d) an algorithm executed by the central computing environment (e.g., a processor), where the algorithm is performed based on the data received by the input device, and the algorithm calculates risk, risk score, or treatment group classification, gene co-expression analysis, thresholding, or other functions described herein. The methods provided by the present invention may also be automated in whole or in part.
[0113] Manual and Computer-Aided Methods and Products The methods and systems described herein can be implemented in numerous ways. In one particularly useful embodiment, the method includes the use of a communications infrastructure, such as the Internet. Several embodiments are discussed below. It should also be understood that the present disclosure may be implemented in various forms of hardware, software, firmware, processors, or combinations thereof. The methods and systems described herein may be implemented as a combination of hardware and software. The software may be implemented as an application program tangibly embodied on a program storage device, or various portions of the software may be implemented in a user's computing environment (e.g., as an applet) and a reviewer's computing environment, where the reviewer may be located at an associated remote site (e.g., at a service provider's facility).
[0114] For example, during or after data entry by a user, some of the data processing may be performed in the user's computing environment. For example, the user's computing environment may be programmed to provide a test code defined to represent a potential "risk score," where the score is sent as a processed or partially processed response to the reviewer's computing environment in the form of a test code for subsequent execution of one or more algorithms in the reviewer's computing environment to provide a result and / or generate a report. The risk score may be a numeric score (a representative number, e.g., likelihood of recurrence based on a validation study population) or a non-numeric score (e.g., low, medium, or high) representing a number or range of numbers.
[0115] An application program for executing the algorithms described herein may be uploaded to and executed by a machine comprising any suitable architecture. Generally, the machine includes a computer platform having hardware such as one or more central processing units (CPUs), a random access memory (RAM), and one or more input / output (I / O) interfaces. The computer platform also includes an operating system and microinstruction code. The various methods and functions described herein may be either part of the microinstruction code or part of the application program (or a combination thereof) that is executed via the operating system. In addition, various other peripheral devices, such as an additional data storage device and a printing device, may be connected to the computer platform.
[0116] As a computer system, the system generally includes a processor unit that operates to receive information, which may include test data (e.g., levels of risk genes, levels of one or more reference gene products; normalized levels of genes; and may also include other data, such as patient data). This received information may be at least temporarily stored in a database, and the data may be analyzed to generate a report, as described above.
[0117] Some or all of the input and output data may also be transmitted electronically; certain output data (e.g., reports) may be transmitted electronically or by telephone (e.g., by facsimile, e.g., using a faxback or similar device). Exemplary output receiving devices may include display elements, printers, facsimile machines, etc. Electronic forms of transmission and / or display may include email, interactive television, etc. In particularly advantageous embodiments, all or a portion of the input data and / or all or a portion of the output data (e.g., typically at least the final report) are stored on a web server, accessed by a typical browser, preferably with controlled confidential access. The data may be accessed by or transmitted to medical personnel as needed. The input and output data, including all or a portion of the final report, may be used to populate the patient's medical record, which may reside in a confidential database at the medical facility.
[0118] Systems used in the methods described herein generally include at least one computer processor (e.g., where the method is performed entirely at a single site) or at least two networked computer processors (e.g., where data is input by a user (also referred to herein as a "client") and transmitted to a second computer processor at a remote site for analysis, where the first and second computer processors are connected by a network, e.g., via an intranet or the Internet). The system may also include one or more user components for input; and one or more reviewer components for viewing data, generated reports, and manual intervention. Additional components of the system may include one or more server components; and one or more databases for data storage (e.g., a database of report elements, e.g., interpretive report elements, or a relational database (RDB) that may contain user data input and data output). The computer processor may be a processor typically found in a personal desktop computer (e.g., IBM, Dell, Macintosh), portable computer, mainframe, minicomputer, or other computing device.
[0119] A networked client / server architecture can be chosen as needed, for example, a traditional two-tier or three-tier client-server model. A relational database management system (RDMS) provides the interface to the database, either as part of the application server component or as a separate component (RDB machine).
[0120] In one example, the architecture is provided as a database-centric client / server architecture, where client applications generally request services from an application server, which results in requests to a database (or database server) to populate the report with various report elements as required, particularly interpretive report elements, especially interpretive text and alerts. One or more servers (e.g., part of the application server machine or as separate RDB / relational database machines) respond to the client requests.
[0121] The input client component may be a complete, standalone personal computer that provides a complete set of power and functionality to run applications. The client component typically operates under any desired operating system and includes a communications element (e.g., a modem or other hardware for connecting to a network), one or more input devices (e.g., a keyboard, mouse, keypad, or other device used to send information or commands), a storage element (e.g., a hard drive or other computer-readable and computer-writable storage medium), and a display element (e.g., a monitor, television, LCD, LED, or other display device that conveys information to the user). A user enters input commands into a computer processor via the input device. Typically, the user interface is a graphical user interface (GUI) written for a web browser application.
[0122] The server component(s) may be a personal computer, minicomputer, or mainframe and provide data management, information sharing among clients, network management, and security. The applications and any databases used may be on the same server or on different servers.
[0123] Other computing configurations for the client and one or more servers are contemplated, including processing on a single machine such as a mainframe, a collection of machines, or other suitable configuration. Generally, the client and server machines work together to accomplish the processing of the present disclosure.
[0124] If used, the database(s) are typically connected to the database server component and may be any device capable of holding data. For example, the database may be any computer magnetic or optical storage device (e.g., CD-ROM, internal hard drive, tape drive). The database may be located remotely from the server component (accessed via a network, modem, etc.) or may be located locally to the server component.
[0125] When used in the present systems and methods, a database may be a relational database, organized and accessed according to the relationships between data items. A relational database generally consists of multiple tables (entities). Rows in a table represent records (collections of information about individual items), and columns represent fields (particular attributes of the records). In its simplest conception, a relational database is a collection of data entries related to each other by at least one common field.
[0126] Additional workstations equipped with computers and printers may be used at the service point to input data and, in some embodiments, generate appropriate reports if required. One or more computers may have shortcuts (e.g., on the desktop) to launch applications to facilitate initiating data entry, submission, analysis, receiving reports, etc., as required.
[0127] computer-readable storage medium The present disclosure also contemplates a computer-readable storage medium (e.g., CD-ROM, memory key, flash memory card, diskette, etc.) having stored thereon a program that, when executed in a computing environment, provides implementation of an algorithm and performs all or part of the results of the response likelihood assessment as described herein. Where the computer-readable medium includes a complete program for performing the methods described herein, the program includes program instructions for collecting, analyzing, and generating output, and generally includes computer-readable code devices for interacting with a user as described herein, processing the data together with analytical information, and generating a unique printed or electronic medium for the user.
[0128] Where the storage medium provides a program that provides implementation of a portion of the methods described herein (e.g., user-side aspects of the methods (e.g., data entry capabilities, report reception capabilities, etc.)), the program provides for transmission of data entered by a user to a computing environment at a remote site (e.g., via the Internet, via an intranet, etc.). Processing of the data or completion of processing is performed at the remote site, and a report is generated. After receipt of the report and completion of any manual intervention required to provide the completed report, the completed report is then transmitted back to the user as an electronic document or a printed document (e.g., a paper report faxed or mailed). A storage medium containing a program according to the present disclosure may be packaged with instructions (e.g., for installing, using, etc. the program) recorded on a suitable carrier or a web address where such instructions can be obtained. The computer-readable storage medium may also be provided in combination with one or more reagents (e.g., primers, probes, arrays, or other such kit components) for performing a responsiveness assessment.
[0129] All aspects of the invention may be practiced such that a limited number of additional genes that co-express with the disclosed genes, as evidenced, for example, by a statistically significant Pearson and / or Spearman correlation coefficient, are included in the prognostic or predictive test in addition to and / or instead of the disclosed genes.
[0130] Having described the invention, the same may be more readily understood by reference to the following examples, which are provided by way of illustration and are not intended to limit the invention in any way. [Example]
[0131] Example 1: This study incorporated breast cancer tumor samples from 136 patients diagnosed with breast cancer (the "Providence Study"). Biostatistical modeling studies of a prototype dataset demonstrated that amplified RNA is a useful carrier for biomarker identification studies. This was validated in this study by incorporating known breast cancer biomarkers along with candidate prognostic genes in tissue samples. Known biomarkers were shown to be associated with clinical outcomes in amplified RNA based on the criteria outlined in this protocol.
[0132] Test Design Please refer to the original Providence Phase II study protocol for biopsy specimen information. This study examined statistical associations between 384 test candidate biomarkers and clinical outcomes in amplified samples derived from 25 ng of extracted mRNA from fixed, paraffin-embedded tissue samples from 136 specimens in the original Providence Phase II study. Expression levels of candidate genes were normalized using reference genes. Several reference genes were analyzed in this study: AAMP, ARF1, EEF1A1, ESD, GPS1, H3F3A, HNRPC, RPL13A, RPL41, RPS23, RPS27, SDHA, TCEA1, UBB, YWHAZ, B-actin, GUS, GAPDH, RPLPO, and TFRC.
[0133] The 136 samples were split into three automated RT plates, each with 2 x 48 and 40 samples, and three RT positive and negative controls. Quantitative PCR assays were performed in 384 wells without replicates using QuantiTect Probe PCR Master Mix® (Qiagen). Plates were analyzed on a Light Cycler® 480, and after data quality control, all samples from RT plate 3 were repeated to generate new RT-PCR data. Median crossing points (C) for the five reference genes were calculated. P ) (the point at which detection rises above background signal) for each individual candidate gene P The data was normalized by subtracting the values from the overall sample C. This normalization was performed for each of the samples, thereby P This resulted in final data adjusted for differences in the mean and mean values. This data set was used for the final data analysis.
[0134] Data analysis A standard z-test was performed for each gene (S. Darby, J. Reissland, Journal of the Royal Statistical Society 144(3):298-331 (1981)). This returns a z-score (a measure of the sample's distance from the mean in standard deviation units), a p-value, and residuals, along with other statistics and parameters from the model. If the z-score is negative, expression is positively correlated with a favorable prognosis; if it is positive, expression is negatively correlated with a favorable prognosis. The p-value was used to obtain a q-value using a library q-value. Genes with weak expression, with little correlation, were excluded from the calculation of the distribution used for the q-value. A Cox proportional hazards model test was performed to confirm that survival times were consistent with the event vector for gene expression. This returned a hazard ratio (HR), which estimates the effect of each gene's expression (individually) on the risk of a cancer-related event. The resulting data are presented in Tables 1-6. An HR<1 indicates that expression of the gene is positively associated with a good prognosis, whereas an HR>1 indicates that expression of the gene is negatively associated with a good prognosis.
[0135] Example 2: Test Design Amplified samples were derived from 25 ng of extracted mRNA from fixed, paraffin-embedded tissue samples obtained from 78 evaluable cases from a Phase II breast cancer trial conducted at Rush University Medical Center. Three of the samples did not provide enough RNA for amplification with 25 ng, so amplification was repeated a second time with 50 ng of RNA. This study also analyzed several reference genes used for normalization: AAMP, ARF1, EEF1A1, ESD, GPS1, H3F3A, HNRPC, RPL13A, RPL41, RPS23, RPS27, SDHA, TCEA1, UBB, YWHAZ, β-actin, RPLPO, TFRC, GUS, and GAPDH.
[0136] Assays were performed in 384 wells with no replicates using QuantiTect Probe PCR Master Mix. Plates were analyzed on a Light Cycler 480 instrument. This data set was used for final data analysis. C for five reference genes P The median C of each individual candidate gene P The data were normalized by subtracting the α-value from the α-value. This normalization was performed for each sample, thereby obtaining an overall sample C value. P The final data were adjusted for differences in
[0137] Data analysis There were 34 samples with a mean CP score above 35. However, none of the samples were excluded from the analysis because they were deemed to be sufficiently informative to be retained in the study. Principal component analysis (PCA) was used to determine whether there were plate effects that contributed to the variability between different RT plates. The first principal component correlated well with the median expression values, indicating that the majority of the variability between samples was driven by expression levels. There was also no unexpected variability between plates.
[0138] Data for other variables Groups - Patients were divided into two groups (cancer / non-cancer). There were few differences in overall gene expression between the two, as evidenced by minimal differences in median CP (0.7) within each group.
[0139] Sample storage age – Samples varied widely in their overall gene expression, but the shorter the storage age, the greater the C P The values tended to be lower.
[0140] Overall sample gene expression from instrument to instrument was consistent. One instrument had a slightly higher C compared to the other three. P Median values are shown, but the variation was well within the acceptable range.
[0141] Overall sample gene expression from RT plate to RT plate was also highly consistent. C for each of the three RT plates (two automated RT plates and one manual plate containing replicate samples) P All medians are 1C of each other P was within the range.
[0142] Univariate analysis of genes significantly different between study groups These genes were analyzed using z-tests and Cox proportional hazards models as described in Example 1. The resulting data can be seen in Tables 7-12.
[0143] Example 3: The statistical correlation between clinical outcomes and the expression levels of the genes identified in Examples 1 and 2 was examined in a breast cancer gene expression dataset maintained by the Swiss Institute of Bioinformatics (SIB). Further information regarding the SIB database, study datasets, and processing methods is provided in P. Wirapati et al., Breast Cancer Research 10(4):R65 (2008). Univariate Cox proportional hazards analysis was performed to confirm the relationship between clinical outcomes (DFS, MFS, OS) of breast cancer patients and the expression levels of genes identified as significant in the amplified RNA studies described above. Both fixed-effect and random-effect models were included in the meta-analysis. These models are further described in L. Hedges and J. Vevea, Psychological Methods 3(4):486-504 (1998) and K. Sidik and J. Jonkman, Statistics in Medicine 26:1964-1981 (2006), the contents of which are incorporated herein by reference. Validation results for all genes identified as having a statistically significant association with breast cancer clinical outcomes are shown in Table 13. In these tables, "Est" indicates the estimated coefficient of the covariate (gene expression); "SE" is the standard error; "t" is the t-score for this estimator (i.e., Est / SE); and "fe" is the fixed-effects estimator from the meta-analysis. Several gene families (including metabolic, proliferation, immune, and stromal group genes) with significant statistical associations with clinical outcomes in breast cancer were confirmed using the SIB dataset. For example, Table 14 includes analyses of genes in the metabolic group, and Table 15 includes analyses of genes in the stromal group.
[0144] Example 4: Coexpression analysis was performed from six breast cancer datasets using microarray data. "Processed" expression values were available from the GEO website; however, further processing was required. If expression values were RMA, they were median-normalized to the sample level. If expression values were MAS 5.0, they were (1) scaled to 10 if <10; (2) log-transformed to base e; and (3) median-normalized to the sample level.
[0145] Generation of Correlated Pairs: A rank matrix was generated by arranging the expression values for each sample in descending order. A correlation matrix was then created by calculating the Spearman correlation value for each pair of probe IDs. Probe pairs with a Spearman value ≥ 0.7 were considered co-expressed. Redundant or overlapping correlated pairs were identified in multiple datasets. For each correlation matrix generated from the sequence dataset, significant probe pairs occurring in > 1 dataset were identified. This served to filter "non-significant" pairs from the analysis as well as provide additional evidence for "significant" pairs present in multiple datasets. Depending on the number of datasets included in each tissue-specific analysis, only pairs appearing in a minimum number or percentage of datasets were included.
[0146] We generated co-expression cliques using the Bron-Kerbosch algorithm for finding maximum cliques in undirected graphs. This algorithm generates three node sets: compsub, candidate, and not. compsub contains the set of nodes that are expanded or contracted one by one depending on the traversal direction of the tree search. candidate consists of all nodes that are candidates for addition to compsub. not contains the set of nodes added to compsub that are excluded from the expansion. The algorithm consists of five steps: selecting a candidate; adding the candidate to compsub; creating the new set by removing all nodes not connected to the candidate from the original candidate and not sets; recursively calling the expansion operator on the new candidate and not sets; and upon return, removing the candidate from compsub and placing it in the original not set.
[0147] There was a depth-first search with pruning, and the choice of candidate nodes affected the runtime of the algorithm. The runtime was optimized by selecting nodes in descending order of frequency in the pair. Also, while recursive algorithms generally cannot be run in a multithreaded manner, the extension operator for the first recursion level was multithreaded. Being at the top of the recursion tree ensured that the data between threads was independent, so the threads ran in parallel.
[0148] Clique mapping and normalization: Because co-expression pairs and clique members are at the probe level, probe IDs must first be mapped to genes (or Refseq) before they can be analyzed. Affymetrix gene map information was used to map all probe IDs to gene names. A probe can map to multiple genes, and a gene can be represented by multiple probes. The data for each clique is verified by manually calculating correlation values for each pair from a single clique.
[0149] The results of this co-expression analysis are shown in Tables 16 to 18.
[0150] Table A-1
[0151] Table A-2
[0152] Table A-3
[0153] Table A-4
[0154] Table A-5
[0155] Table A-6
[0156] Table A-7
[0157] Table A-8
[0158] Table A-9
[0159] Table A-10
[0160] Table A-11
[0161] Table A-12
[0162] Table A-13
[0163] Table A-14
[0164] Table A-15
[0165] Table A-16
[0166] Table A-17
[0167] Table A-18
[0168] Table A-19
[0169] Table A-20
[0170] Table A-21
[0171] Table A-22
[0172] Table A-23
[0173] Table A-24
[0174] Table A-25
[0175] Table A-26
[0176] Table A-27
[0177] Table A-28
[0178] Table 1
[0179] Table 2
[0180] Table 3
[0181] Table 4
[0182] Table 5
[0183] Table 6
[0184] Table 7
[0185] Table 8
[0186] Table 9
[0187] Table 10
[0188] Table 11
[0189] Table 12
[0190] Table 13-1
[0191] Table 13-2
[0192] Table 13-3
[0193] Table 13-4
[0194] Table 13-5
[0195] Table 13-6
[0196] Table 13-7
[0197] Table 13-8
[0198] Table 13-9
[0199] Table 13-10
[0200] Table 14
[0201] Table 15-1
[0202] Table 15-2
[0203] Table 16-1
[0204] Table 16-2
[0205] Table 16-3
[0206] Table 17-1
[0207] Table 17-2
[0208] Table 17-3
[0209] Table 18-1
[0210] Table 18-2
[0211] Table 18-3
[0212] [Table 18-4]
[0213] [Table 18-5]
[0214] [Table 18-6]
[0215] [Table 18-7]
[0216] [Table 18-8]
[0217] [Table 18-9]
[0218] [Table 18-10]
[0219] Illustrative Embodiments 1. A method for predicting clinical outcome in a patient diagnosed with estrogen receptor (ER)-positive breast cancer, comprising: (a) quantitatively measuring the level of IL6ST RNA transcripts in a tissue sample obtained from a breast cancer tumor of a patient; (b) normalizing the level of the IL6ST RNA transcript to obtain a normalized IL6ST expression level; (c) comparing the normalized IL6ST expression level to a normalized IL6ST expression level obtained from a breast cancer reference set; (d) determining that the patient is likely to have a good prognosis if the normalized IL6ST expression level is increased relative to a normalized IL6ST expression level obtained from a breast cancer reference set, where a good prognosis is a reduced likelihood of recurrence or metastasis or an increased overall survival; A method comprising: 2. The method of embodiment 1, further comprising generating a report based on the normalized IL6ST expression level. 3. The method of embodiment 1 or 2, wherein the tissue sample is a fixed, paraffin-embedded tissue sample. 4. The method of any one of embodiments 1 to 3, wherein the level of the IL6ST RNA transcript is measured using a PCR-based method. 5. The method of any one of embodiments 1 to 4, further comprising measuring the level of an RNA transcript of P2RY5. 6. The method of any one of embodiments 1 to 5, wherein the tissue sample is obtained by core biopsy or fine needle aspiration. 7. The method of any one of embodiments 1 to 6, wherein the level of the RNA transcript of IL6ST is the crossing point (Cp) value and the normalized IL6ST expression level is the normalized Cp value. 8. The method of any one of embodiments 1 to 6, wherein the level of the RNA transcript of IL6ST is a threshold cycle (Ct) value and the normalized IL6ST expression level is a normalized Ct value. 9. The method of any one of embodiments 1 to 8, wherein the favorable prognosis is a reduced likelihood of recurrence or metastasis. 10. A method for predicting clinical outcome in a human patient diagnosed with estrogen receptor (ER)-positive breast cancer, comprising: (a) extracting RNA from a breast cancer tissue sample obtained from a patient; (b) reverse transcribing the IL6ST RNA transcript to produce IL6ST cDNA; (c) amplifying IL6ST cDNA; (d) producing an amplicon of an RNA transcript of IL6ST; (e) quantitatively assaying the level of IL6ST RNA transcript amplicons; (f) normalizing the level of the IL6ST RNA transcript amplicon to provide a normalized IL6ST amplicon level; (g) comparing said normalized IL6ST amplicon level to a normalized IL6ST amplicon level obtained from a breast cancer reference set; (h) determining that the patient is likely to have a good prognosis if the normalized IL6ST amplicon level is increased relative to a normalized IL6ST amplicon level obtained from a breast cancer reference set, wherein a good prognosis is a reduced likelihood of recurrence or metastasis or an increased overall survival; A method comprising: 11. The method of embodiment 10, further comprising generating a report based on the normalized IL6ST amplicon level. 12. The method of embodiment 10 or 11, wherein the tissue sample is a fixed, paraffin-embedded tissue. 13. The method of any one of embodiments 10 to 12, wherein the IL6ST cDNA is amplified using a PCR-based method. 14. The method of any one of embodiments 10 to 13, further comprising amplifying RNA transcripts of P2RY5 in a breast cancer tissue sample obtained from the patient. 15. The method of any one of embodiments 10 to 14, wherein the level of the amplicon of the RNA transcript of IL6ST is a crossing point (Cp) value, and the normalized IL6ST amplicon level is a normalized Cp value. 16. The method of any one of embodiments 10 to 14, wherein the level of the amplicon of the RNA transcript of IL6ST is a threshold cycle (Ct) value, and the normalized IL6ST amplicon level is a normalized Ct value. 17. The method of any one of embodiments 10 to 16, wherein the favorable prognosis is a reduced likelihood of recurrence or metastasis. 18. A computer program product for prognostic classification of human patients diagnosed with estrogen receptor (ER)-positive breast cancer, the computer program product comprising a computer-readable storage medium having encoded thereon a computer program for use in cooperation with a computer having a memory and a processor, the computer program product being loaded into the one or more memory units of the computer and causing the one or more processor units of the computer to: (a) receiving data comprising the level of IL6ST RNA transcripts in a tissue sample obtained from a breast cancer tumor of the patient; (b) normalizing the level to obtain a normalized IL6ST expression level; (c) comparing said normalized IL6ST expression level to a normalized IL6ST expression level obtained from a breast cancer reference set; (d) classifying the patient as likely to have a good prognosis if the normalized IL6ST expression level is increased relative to a normalized IL6ST level obtained from a breast cancer reference set, where a good prognosis is a reduced likelihood of recurrence or metastasis or an increased overall survival. A computer program product that causes the
Claims
1. 1. A method for predicting clinical outcome in a human patient diagnosed with breast cancer, comprising: (a) obtaining expression levels of RNA transcripts of at least two prognostic genes from a tissue sample obtained from the patient's tumor, wherein the at least two prognostic genes include NAT1 and one or more of C8orf4, TNFRSF11B, RUNX1, and PSD3; (b) normalizing the expression levels of the RNA transcripts of the at least two prognostic genes to obtain normalized expression levels; (c) determining that the patient is likely to have a good prognosis if the normalized expression levels of the at least two prognostic genes are increased relative to the normalized expression levels of the at least two prognostic genes obtained from a breast cancer reference set; A method comprising:
2. 2. The method of claim 1, wherein the favorable prognosis is a reduced likelihood of recurrence or metastasis or an increased likelihood of overall survival.
3. 2. The method of claim 1, wherein the favorable prognosis is an increased likelihood of long-term survival without cancer recurrence and / or metastasis-free survival.
4. The method of claim 1 , further comprising generating a report based on the normalized expression levels of the at least two prognostic genes for the patient.
5. 5. The method of claim 1, wherein the tissue sample is a fixed, paraffin-embedded tissue.
6. The method of claim 1 , wherein the expression level is obtained using a PCR-based method.
7. 7. The method of claim 1, wherein the at least two prognostic genes comprise at least two genes in any of the stromal, metabolic, immune, or proliferation groups.
8. 8. The method of claim 1, wherein the at least two prognostic genes comprise at least four genes in any two of the stromal, metabolic, immune, or proliferation groups.
9. 9. The method of claim 1, wherein the at least two prognostic genes include NAT1 and PSD3.
10. 10. The method of claim 1, wherein the at least two prognostic genes include NAT1 and one or more of C8orf4, TNFRSF11B, and RUNX1.
11. 11. The method of any one of claims 1 to 10, wherein the at least two prognostic genes further comprise one or more of IL6ST, GSTM3 and CSF1.
12. 12. The method of any one of claims 1 to 11, wherein the breast cancer is estrogen receptor (ER) positive breast cancer.
13. 1. A method for providing an index for classifying human breast cancer patients according to prognosis, comprising: (a) receiving a first data structure comprising respective levels of RNA transcripts of at least three different prognostic genes in a tissue sample obtained from a patient's tumor, wherein the at least three different prognostic genes comprise NAT1 and two or more of C8orf4, TNFRSF11B, RUNX1, and PSD3; (b) normalizing the level of the RNA transcript of each of the at least three prognostic genes to obtain a normalized expression level; (c) determining a similarity of the normalized expression level of each of the at least three prognostic genes to a respective control level of expression of the at least three prognostic genes obtained from a second data structure to obtain a patient similarity value, wherein the second data structure is based on expression levels from a plurality of cancer tumors; (d) comparing said patient similarity value to a selected threshold of similarity of each normalized expression level of each of said at least three prognostic genes to each control level of expression of said at least three prognostic genes; (e) providing an indication for classifying the patient as having a first outcome if the patient similarity value exceeds a threshold similarity value, and for classifying the patient as having a second outcome if the patient similarity value does not exceed a threshold similarity value; A method comprising:
14. 14. The method of claim 13, wherein one of the at least three prognostic genes comprises at least three different prognostic genes selected from IL6ST, GSTM3, C8orf4, TNFRSF11B, RUNX1 and CSF1.
15. 15. A computer program comprising computer code means for carrying out steps (b) and (c) of the method of any one of claims 1 to 14, said computer program being executed on a computer.
16. 16. The computer program of claim 15 embodied on a computer-readable medium.
Citation Information
Patent Citations
Materials and methods for breast cancer classification
JP2008534002A
Diagnosis and prognosis of breast cancer patients
JP2009131262A
Genes involved estrogen metabolism
US20070275398A1
Methods for diagnosing and treating breast cancer based on a HER / ER ratio
US20080153098A1