Method for molecular profiling of tumors

By analyzing cancer patient samples through immunohistochemistry, microarray, and FISH, the method identifies personalized treatments based on molecular profiles, enhancing treatment efficacy for refractory and metastatic cancer patients.

JP2025106337APending Publication Date: 2025-07-15CARIS MPI INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025054683
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2009-07-29
Filing Date
2025-03-28
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Existing treatment plans for cancer patients, particularly those with refractory and metastatic cancer, often fail to account for individual molecular profiles, leading to limited treatment options and low response rates to novel anticancer agents.

Method used

Perform immunohistochemical, microarray, and fluorescence in situ hybridization analyses on patient samples to determine expression and mutation profiles, comparing them to a rule database to identify candidate treatments with known biological activity against cancer cells.

Benefits of technology

This approach identifies candidate treatments that can extend progression-free survival, disease-free survival, and overall survival of cancer patients by providing personalized treatment options beyond standard therapies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025106337000055
    Figure 2025106337000055
  • Figure 2025106337000056
    Figure 2025106337000056
  • Figure 2025106337000057
    Figure 2025106337000057
Patent Text Reader

Abstract

To provide methods of molecular profiling of a sample from an individual to identify treatments for an individual with a condition or disease, such as cancer, using a result of the molecular profiling.SOLUTION: A method of identifying a candidate treatment for a subject in need of the treatment comprises: performing, on a sample from the subject, an immunohistochemistry (IHC) analysis, a microarray analysis, a fluorescent in-situ hybridization (FISH) analysis, and DNA sequencing; and comparing the IHC expression profile, microarray expression profile, FISH mutation profile and sequencing mutation profile against a rules database comprising a mapping of treatments whose biological activity against cancer cells is known.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross-reference This application claims the benefit of U.S. Provisional Patent Application No. 61 / 151,758, filed Feb. 11, 2009; U.S. Provisional Patent Application No. 61 / 170,565, filed Apr. 17, 2009; and U.S. Provisional Patent Application No. 61 / 229,686, filed Jul. 29, 2009, all of which are hereby incorporated by reference in their entirety.

Background Art

[0002] Background A patient's medical condition is typically treated with a treatment plan or therapy selected based on clinically-based criteria; that is, based on the determination that the patient has been diagnosed with a particular disease (the diagnosis being made from classical diagnostic assays), a treatment or treatment plan is selected for the patient. The molecular mechanisms underlying various medical conditions have been the subject of research for many years, but the specific application of the individual's molecular profile in determining treatment plans and therapies for affected individuals has been disease-specific and not widely promoted.

[0003] Some treatment plans are determined using a combination of molecular profiling and clinical characterization of the patient, such as a physician's findings (e.g., codes according to the International Classification of Diseases and the date on which such codes were determined), clinical test results, X-rays, biopsy results, patient descriptions, and any other medical information on which a physician typically relies to make a diagnosis in a particular disease. However, some treatment plans, although associated with treating a particular type of medical condition, may be effective for different medical conditions, so there is a risk that the combination of selection materials based on molecular profiling and clinical characterization may overlook an effective treatment plan for a particular individual (such as the diagnosis of a particular type of cancer).

[0004] Patients with refractory and metastatic cancer are of particular interest to treating physicians. The majority of metastatic cancer patients will ultimately exhaust their treatment options for those tumors. After standard primary and secondary (and in some cases, tertiary and higher) treatments have been pursued for those tumors, these patients have very limited options. These patients can participate in Phase I and Phase II clinical trials of novel anticancer agents, but usually must meet very strict eligibility criteria to do so. Studies have shown that when patients participate in this type of trial, novel anticancer agents can result in response rates ranging from an average of 5% - 10% in the Phase I setting to 12% in the Phase II setting. These patients also have the option of choosing to receive best supportive care to treat their symptoms.

[0005] In recent years, there has been a surge of interest in the development of novel anticancer agents that more specifically target cell surface receptors or upregulated or amplified gene products. This approach has had some success (e.g., trastuzumab for HER2 / neu in breast cancer cells, rituximab for CD20 in lymphoma cells, bevacizamab for VEGF, and cetuximab for EGFR). However, patients' tumors still ultimately progress with these treatments. If more targets or molecular findings, such as molecular mechanisms, genes, gene expression proteins, and / or their combinations, etc., are measured in patients' tumors, there is a possibility of finding additional targets or molecular findings that can be exploited by using specific therapeutic agents. By identifying multiple agents that can treat multiple targets or underlying mechanisms, viable alternative treatments to existing treatment regimens can be provided to cancer patients.

[0006] One or more individual profiles are often identified by molecular profiling analysis that drive more informed and effective individualized treatment options, which can lead to improved patient care and treatment outcomes. The present invention provides methods and systems for identifying treatments for individuals by performing molecular profiling of samples derived from the individuals. SUMMARY OF THE INVENTION

[0007] The present invention provides methods and systems for performing molecular profiling and using the results of the molecular profiling to identify treatments for individuals. In some embodiments, the treatment is one that was not initially identified as a treatment for the disease.

[0008] In one aspect, the present invention provides a method for identifying a candidate treatment for a subject in need thereof, comprising: performing immunohistochemistry (IHC) analysis on a sample from the subject to determine an IHC expression profile for at least five proteins; performing microarray analysis on the sample to determine a microarray expression profile for at least ten genes; performing fluorescence in situ hybridization (FISH) analysis on the sample to determine a FISH mutation profile for at least one gene; performing DNA sequencing on the sample to determine a sequencing mutation profile for at least one gene; and comparing the IHC expression profile, the microarray expression profile, the FISH mutation profile, and the sequencing mutation profile to a rule database. The rule database includes a mapping for treatments whose biological activity is known for cancer cells that: i) overexpress or underexpress one or more proteins included in the IHC expression profile; ii) overexpress or underexpress one or more genes included in the microarray expression profile; iii) have no mutations or have one or more mutations in one or more genes included in the FISH mutation profile; and / or iv) have no mutations or have one or more mutations in one or more genes included in the sequencing mutation profile. A candidate treatment is identified if: i) the comparison indicates that the treatment should have biological activity against the cancer; and ii) the comparison does not indicate a contraindication to the treatment for treating the cancer.

[0009] In some embodiments, IHC expression profiling includes assaying one or more of SPARC, PGP, Her2 / neu, ER, PR, c-kit, AR, CD52, PDGFR, TOP2A, TS, ERCC1, RRM1, BCRP, TOPO1, PTEN, MGMT, and MRP1.

[0010] In some embodiments, microarray expression profiling comprises assaying one or more of ABCC1, ABCG2, ADA, AR, ASNS, BCL2, BIRC5, BRCA1, BRCA2, CD33, CD52, CDA, CES2, DCK, DHFR, DNMT1, DNMT3A, DNMT3B, ECGF1, EGFR, EPHA2, ERBB2, ERCC1, ERCC3, ESR1, FLT1, FOLR2, FYN, GART, GNRH1, GSTP1, HCK, HDAC1, HIF1A, HSP90AA1, IL2RA, HSP90AA1, KDR, KIT, LCK, LYN, MGMT, MLH1, MS4A1, MSH2, NFKB1, NFKB2, OGFR, PDGFC, PDGFRA, PDGFRB, PGR, POLA1, PTEN, PTGS2, RAF1, RARA, RRM1, RRM2, RRM2B, RXRB, RXRG, SPARC, SRC, SSTR1, SSTR2, SSTR3, SSTR4, SSTR5, TK1, TNF, TOP1, TOP2A, TOP2B, TXNRD1, TYMS, VDR, VEGFA, VHL, YES1, and ZAP70.

[0011] In some embodiments, FISH mutation profiling comprises assaying EGFR and / or HER2.

[0012] In some embodiments, sequencing mutation profiling comprises assaying one or more of KRAS, BRAF, c-KIT, and EGFR.

[0013] In another aspect, the present invention provides a method for identifying a candidate treatment for a subject in need thereof, comprising: performing immunohistochemistry (IHC) analysis on a sample derived from the subject to determine an IHC expression profile for at least 5 of SPARC, PGP, Her2 / neu, ER, PR, c-kit, AR, CD52, PDGFR, TOP2A, TS, ERCC1, RRM1, BCRP, TOPO1, PTEN, MGMT, and MRP1; performing microarray analysis on the sample to determine a microarray expression profile for at least 5 of ABCC1, ABCG2, ADA, AR, ASNS, BCL2, BIRC5, BRCA1, BRCA2, CD33, CD52, CDA, CES2, DCK, DHFR, DNMT1, DNMT3A, DNMT3B, ECGF1, EGFR, EPHA2, ERBB2, ERCC1, ERCC3, ESR1, FLT1, FOLR2, FYN, GART, GNRH1, GSTP1, HCK, HDAC1, HIF1A, HSP90AA1, IL2RA, HSP90AA1, KDR, KIT, LCK, LYN, MGMT, MLH1, MS4A1, MSH2, NFKB1, NFKB2, OGFR, PDGFC, PDGFRA, PDGFRB, PGR, POLA1, PTEN, PTGS2, RAF1, RARA, RRM1, RRM2, RRM2B, RXRB, RXRG, SPARC, SRC, SSTR1, SSTR2, SSTR3, SSTR4, SSTR5, TK1, TNF, TOP1, TOP2A, TOP2B, TXNRD1, TYMS, VDR, VEGFA, VHL, YES1, and ZAP70; performing fluorescence in situ hybridization (FISH) analysis on the sample to determine a FISH mutation profile for EGFR and / or HER2; performing DNA sequencing on the sample to determine a sequencing mutation profile for at least 1 of KRAS, BRAF, c-KIT, and EGFR; and comparing the IHC expression profile, the microarray expression profile, the FISH mutation profile, and the sequencing mutation profile against a rule database.The rules database includes mapping for treatments whose biological activity has been determined for cancer cells that i) overexpress or underexpress one or more proteins included in the IHC expression profile; ii) overexpress or underexpress one or more genes included in the microarray expression profile; iii) have no mutations or have one or more mutations in one or more genes included in the FISH mutation profile; and / or iv) have no mutations or have one or more mutations in one or more genes included in the sequencing mutation profile. Candidate treatments are identified if i) the comparison step indicates that the treatment should have biological activity against the cancer; and ii) the comparison step does not indicate a contraindication to the treatment for the treatment of the cancer. In some embodiments, IHC expression profiling is performed for at least 50%, 60%, 70%, 80%, or 90% of the biomarkers described. In some embodiments, microarray expression profiling is performed for at least 50%, 60%, 70%, 80%, or 90% of the biomarkers described.

[0014] In a third aspect, the present invention provides a method for identifying a candidate treatment for cancer in a subject in need thereof, comprising the steps of performing immunohistochemistry (IHC) analysis on a sample derived from the subject to determine an IHC expression profile for at least a protein group consisting of SPARC, PGP, Her2 / neu, ER, PR, c-kit, AR, CD52, PDGFR, TOP2A, TS, ERCC1, RRM1, BCRP, TOPO1, PTEN, MGMT, and MRP1; performing microarray analysis on the sample to determine a microarray expression profile for at least a gene group consisting of ABCC1, ABCG2, ADA, AR, ASNS, BCL2, BIRC5, BRCA1, BRCA2, CD33, CD52, CDA, CES2, DCK, DHFR, DNMT1, DNMT3A, DNMT3B, ECGF1, EGFR, EPHA2, ERBB2, ERCC1, ERCC3, ESR1, FLT1, FOLR2, FYN, GART, GNRH1, GSTP1, HCK, HDAC1, HIF1A, HSP90AA1, IL2RA, HSP90AA1, KDR, KIT, LCK, LYN, MGMT, MLH1, MS4A1, MSH2, NFKB1, NFKB2, OGFR, PDGFC, PDGFRA, PDGFRB, PGR, POLA1, PTEN, PTGS2, RAF1, RARA, RRM1, RRM2, RRM2B, RXRB, RXRG, SPARC, SRC, SSTR1, SSTR2, SSTR3, SSTR4, SSTR5, TK1, TNF, TOP1, TOP2A, TOP2B, TXNRD1, TYMS, VDR, VEGFA, VHL, YES1, and ZAP70; performing fluorescence in situ hybridization (FISH) analysis on the sample to determine a FISH mutation profile for at least a gene group consisting of EGFR and HER2; performing DNA sequencing on the sample to determine a sequencing mutation profile for at least a gene group consisting of KRAS, BRAF, c-KIT, and EGFR; and comparing the IHC expression profile, the microarray expression profile, the FISH mutation profile, and the sequencing mutation profile against a rule database.The rules database includes mapping for treatments that have been found to have biological activity against cancer cells that i) overexpress or underexpress one or more proteins included in the IHC expression profile; ii) overexpress or underexpress one or more genes included in the microarray expression profile; iii) do not have mutations or have one or more mutations in one or more genes included in the FISH mutation profile; and / or iv) do not have mutations or have one or more mutations in one or more genes included in the sequencing mutation profile. Candidate treatments are identified if i) the comparison step indicates that the treatment should have biological activity against the cancer; and ii) the comparison step does not indicate any contraindications to the treatment for treating the cancer.

[0015] In some aspects of the methods of the invention, the sample includes formalin-fixed paraffin-embedded (FFPE) tissue, fresh frozen (FF) tissue, or tissue contained in a solution that preserves nucleic acid or protein molecules. In some aspects, none of microarray analysis, FISH mutation analysis, or sequencing mutation analysis is performed. For example, the method may not be carried out unless the sample passes a quality control test. In some aspects, the quality control test includes the A260 / A280 ratio or the Ct value of RT-PCR of RPL13a mRNA. For example, the quality control test may require an A260 / A280 ratio lower than 1.5 or an RPL13a Ct value higher than 30.

[0016] In some aspects, microarray expression profiling is performed using a low density microarray, expression microarray, comparative genomic hybridization (CGH) microarray, single nucleotide polymorphism (SNP) microarray, proteomics array, or antibody array.

[0017] The method of the present invention may require assaying specific markers including additional markers. In some embodiments, IHC expression profiling is performed with respect to at least SPARC, TOP2A, and / or PTEN. Microarray expression profiling may be performed with respect to at least CD52. IHC expression profiling further consists of assaying one or more of DCK, EGFR, BRCA1, CK 14, CK 17, CK 5 / 6, E-cadherin, p95, PARP-1, SPARC, and TLE3. In some embodiments, IHC expression profiling further consists of assaying Cox-2 and / or Ki-67. In some embodiments, microarray expression profiling further consists of assaying HSPCA. In some embodiments, FISH mutation profiling further consists of assaying c-Myc and / or TOP2A. Sequencing mutation profiling may include assaying PI3K.

[0018] According to the method of the present invention, several genes and gene products can be assayed. For example, genes used in IHC expression profiling, microarray expression profiling, FISH mutation profiling, and sequencing mutation profiling include one or more independently selected from ABCC1, ABCG2, ACE2, ADA, ADH1C, ADH4, AGT, androgen receptor, AR, AREG, ASNS, BCL2, BCRP, BDCA1, BIRC5, B-RAF, BRCA1, BRCA2, CA2, caveolin, CD20, CD25, CD33, CD52, CDA, CDK2, CDW52, CES2, CK 14, CK 17, CK 5 / 6, c-KIT, c-Myc, COX-2, cyclin D1, DCK, DHFR, DNMT1, DNMT3A, DNMT3B, E-cadherin, ECGF1, EGFR, EPHA2, epiregulin, ER, ERBR2, ERCC1, ERCC3, EREG, ESR1, FLT1, folate receptor, FOLR1, FOLR2, FSHB, FSHPRH1, FSHR, FYN, GART, GNRH1, GNRHR1, GSTP1, HCK, HDAC1, Her2 / Neu, HGF, HIF1A, HIG1, HSP90, HSP90AA1, HSPCA, IL13RA1, IL2RA, KDR, KIT, K-RAS, LCK, LTB, lymphotoxin β receptor, LYN, MGMT, MLH1, MRP1, MS4A1, MSH2, Myc, NFKB1, NFKB2, NFKBIA, ODC1, OGFR, p53, p95, PARP-1, PDGFC, PDGFR, PDGFRA, PDGFRB, PGP, PGR, PI3K, POLA, POLA1, PPARG, PPARGC1, PR, PTEN, PTGS2, RAF1, RARA, RRM1, RRM2, RRM2B, RXRB, RXRG, SPARC, SPARC MC, SPARC PC, SRC, SSTR1, SSTR2, SSTR3, SSTR4, SSTR5, survivin, TK1, TLE3, TNF, TOP1, TOP2A, TOP2B, TOPO1, TOPO2B, topoisomerase II, TS, TXN, TXNRD1, TYMS, VDR, VEGF, VEGFA, VEGFC, VHL, YES1, and ZAP70.

[0019] In some embodiments, microarray expression analysis includes identifying whether a gene is upregulated or downregulated with statistical significance compared to a reference. Statistical significance can be determined by a p-value of 0.05, 0.01, 0.005, 0.001, 0.0005, or less than 0.0001. The p-value can also be corrected for multiple comparisons. Correction for multiple comparisons can include the Bonferroni correction or a modified method thereof.

[0020] In some embodiments, IHC analysis includes determining whether 30% or more of the sample has a staining intensity of +2 or more.

[0021] The rules included in the rule database used by the method of the present invention can be based on the effectiveness of various treatments specific to the target gene or gene product. The rule database can include the rules described in Table 1 and / or Table 2 herein.

[0022] In some embodiments of the method of the present invention, a prioritized list of candidate treatments is identified. Prioritization can include ordering the treatments from highest to lowest priority according to a treatment based on either microarray analysis and IHC analysis or FISH analysis; a treatment based on IHC analysis rather than microarray analysis; and a treatment based on microarray analysis rather than IHC analysis.

[0023] In some aspects of the methods of the present invention, the candidate treatment comprises administration of one or more candidate therapeutic agents. The one or more candidate therapeutic agents are 5-fluorouracil, abarelix, alemtuzumab, aminoglutethimide, Anastrazole, aromatase inhibitors (Anastrazole, letrozole), asparaginase, aspirin, ATRA, azacitidine, bevacizumab, bexarotene, bicalutamide, bortezomib, calcitriol, capecitabine, carboplatin, celecoxib, cetuximab, chemoendocrine therapy, colecalciferol, cisplatin, carboplatin, cyclophosphamide, cyclophosphamide / vincristine, cytarabine, dasatinib, decitabine, doxorubicin, epirubicin, epirubicin, erlotinib, etoposide, exemestane, fluoropyrimidine, flutamide, fulvestrant, gefitinib, gefitinib and trastuzumab, gemcitabine, gonadorelin, goserelin, hydroxyurea, imatinib, irinotecan, ixabepilone, lapatinib, letrozole, leuprolide, liposomal doxorubicin, medroxyprogesterone, megestrol, methotrexate, mitomycin, nab-paclitaxel, octreotide, oxaliplatin, paclitaxel, panitumumab, pegaspargase, pemetrexed, pentostatin, sorafenib, sunitinib, tamoxifen, tamoxifen-based therapy, temozolomide, topotecan, toremifene, trastuzumab, VBMCP / cyclophosphamide, vincristine, or any combination thereof.One or more candidate therapeutic agents may also be 5FU, bevacizumab, capecitabine, cetuximab, cetuximab + gemcitabine, cetuximab + irinotecan, cyclophosphamide, diethylstibesterol, doxorubicin, erlotinib, etoposide, exemestane, fluoropyrimidine, gemcitabine, gemcitabine + etoposide, gemcitabine + pemetrexed, irinotecan, irinotecan + sorafenib, lapatinib, lapatinib + tamoxifen, letrozole, letrozole + capecitabine, mitomycin, nab-paclitaxel, nab-paclitaxel + gemcitabine, nab-paclitaxel + trastuzumab, oxaliplatin, oxaliplatin + 5FU + trastuzumab, panitumumab, pemetrexed, sorafenib, sunitinib, sunitinib, sunitinib + mitomycin, tamoxifen, temozolomide, temozolomide + bevacizumab, temozolomide + sorafenib, trastuzumab, vincristine, or any combination thereof.

[0024] In an aspect of the method of the invention, the sample comprises cancer cells. The cancer may be metastatic cancer. The cancer may be refractory to previous treatment. The previous treatment may be standard care for the cancer. Optionally, the subject has been previously treated with one or more therapeutic agents for treating the cancer. Optionally, the subject has not been previously treated with the one or more identified candidate therapeutic agents.

[0025] In some embodiments, the cancer includes prostate cancer, lung cancer, melanoma, small cell carcinoma (esophagus / retroperitoneum), cholangiocarcinoma, mesothelioma, head and neck cancer (SCC), pancreatic cancer, pancreatic neuroendocrine cancer, small cell carcinoma, gastric cancer, peritoneal pseudomyxoma, anal canal cancer (SCC), vaginal cancer (SCC), cervical cancer, renal cancer, eccrine sweat gland adenocarcinoma, salivary gland adenocarcinoma, uterine soft tissue sarcoma (uterus), GIST (stomach), or anaplastic thyroid cancer. In some embodiments, the cancer includes cancer of the accessory sinuses, middle ear, and inner ear, adrenal gland, appendix, hematopoietic system, bone and joint, spinal cord, breast, cerebellum, cervix, connective soft tissue, uterine body, esophagus, eye, nose, globe, fallopian tube, extrahepatic bile duct, oral cavity, intrahepatic bile duct, kidney, appendix-colon, larynx, lip, liver, lung and bronchus, lymph node, cerebrum, spinal cord, nasal cartilage, retina, eye, oropharynx, endocrine gland, female genitalia, ovary, pancreas, penis and scrotum, pituitary gland, pleura, prostate, rectum renal pelvis, ureter, peritoneum, salivary gland, skin, small intestine, stomach, testis, thymus, thyroid, tongue, unknown, bladder, uterus, vagina, labia, and vulva. In some embodiments, the sample includes cells selected from the group consisting of fat, adrenal cortex, adrenal gland, adrenal-medulla, appendix, bladder, blood, blood vessel, bone, osteochondral, brain, breast, cartilage, cervix, colon, sigmoid colon, dendritic cell, skeletal muscle, endometrium, esophagus, fallopian tube, fibroblast, gallbladder, kidney, larynx, liver, lung, lymph node, melanocyte, mesothelial lining, myoepithelial cell, osteoblast, ovary, pancreas, parotid gland, prostate, salivary gland, accessory nasal tissue, skeletal muscle, skin, small intestine, smooth muscle, stomach, synovium, joint lining tissue, tendon, testis, thymus, thyroid, uterus, and uterine body. In some embodiments, the cancer includes breast cancer, colorectal cancer, ovarian cancer, lung cancer, non-small cell lung cancer, cholangiocarcinoma, mesothelioma, sweat gland cancer, or GIST cancer.

[0026] Using the method of the present invention, the progression-free survival (PFS) or disease-free survival (DFS) of a subject can be extended. For example, the PFS or DFS can be extended by at least about 10%, about 15%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or at least about 100% compared to previous treatments. In addition, using the method of the present invention to select a candidate treatment, the lifespan of a patient can be extended. For example, the lifespan of a patient can be extended by at least 1 week, 2 weeks, 3 weeks, 4 weeks, 1 month, 5 weeks, 6 weeks, 7 weeks, 8 weeks, 2 months, 9 weeks, 10 weeks, 11 weeks, 12 weeks, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months, 9 months, 10 months, 11 months, 12 months, 13 months, 14 months, 15 months, 16 months, 17 months, 18 months, 19 months, 20 months, 21 months, 22 months, 23 months, 24 months, 2 years, 2.5 years, 3 years, 4 years, or at least 5 years. [Inventive Concept 1001] A method for identifying a candidate treatment for a subject in need thereof, comprising the following steps: (a) performing immunohistochemical (IHC) analysis on a sample from the subject to determine an IHC expression profile for at least five proteins; (b) performing microarray analysis on the sample to determine a microarray expression profile for at least ten genes; (c) performing fluorescence in situ hybridization (FISH) analysis on the sample to determine a FISH mutation profile for at least one gene; (d) performing DNA sequencing on the sample to determine a sequencing mutation profile for at least one gene; and (e) i. cancer cells that overexpress or underexpress one or more proteins included in the IHC expression profile, ii. cancer cells that overexpress or underexpress one or more genes included in the microarray expression profile, iii. cancer cells that do not have mutations or have one or more mutations in one or more genes included in the FISH mutation profile, and / or iv. cancer cells that do not have mutations or have one or more mutations in one or more genes included in the sequencing mutation profile comparing the IHC expression profile, microarray expression profile, FISH mutation profile, and sequencing mutation profile against a regular database containing mapping for treatments for which the biological activity against such cancer cells is known; and (f) i. as indicated by the comparison in step (e), the treatment should have biological activity against the cancer, and ii. identifying a candidate treatment when the comparison in step (e) does not indicate any contraindications to the treatment for treating the cancer. [Invention 1002] A method for identifying a candidate treatment for a subject in need thereof, comprising the following steps: (a) performing immunohistochemical (IHC) analysis on a sample from the subject to determine an IHC expression profile for at least five of SPARC, PGP, Her2 / neu, ER, PR, c-kit, AR, CD52, PDGFR, TOP2A, TS, ERCC1, RRM1, BCRP, TOPO1, PTEN, MGMT, and MRP1; (b) performing microarray analysis on the sample to determine a microarray expression profile for at least five of ABCC1, ABCG2, ADA, AR, ASNS, BCL2, BIRC5, BRCA1, BRCA2, CD33, CD52, CDA, CES2, DCK, DHFR, DNMT1, DNMT3A, DNMT3B, ECGF1, EGFR, EPHA2, ERBB2, ERCC1, ERCC3, ESR1, FLT1, FOLR2, FYN, GART, GNRH1, GSTP1, HCK, HDAC1, HIF1A, HSP90AA1, IL2RA, HSP90AA1, KDR, KIT, LCK, LYN, MGMT, MLH1, MS4A1, MSH2, NFKB1, NFKB2, OGFR, PDGFC, PDGFRA, PDGFRB, PGR, POLA1, PTEN, PTGS2, RAF1, RARA, RRM1, RRM2, RRM2B, RXRB, RXRG, SPARC, SRC, SSTR1, SSTR2, SSTR3, SSTR4, SSTR5, TK1, TNF, TOP1, TOP2A, TOP2B, TXNRD1, TYMS, VDR, VEGFA, VHL, YES1, and ZAP70; (c) performing fluorescence in situ hybridization (FISH) analysis on the sample to determine a FISH mutation profile for EGFR and / or HER2; (d) performing DNA sequencing on the sample to determine a sequencing mutation profile for at least one of KRAS, BRAF, c-KIT, and EGFR; and (e) i. cancer cells that overexpress or underexpress one or more proteins included in the IHC expression profile, ii. cancer cells that overexpress or underexpress one or more genes included in the microarray expression profile, iii. cancer cells that do not have a mutation or have one or more mutations in one or more genes included in the FISH mutation profile, and / or iv. cancer cells that do not have mutations or have one or more mutations in one or more genes included in the sequencing mutation profile comparing the IHC expression profile, microarray expression profile, FISH mutation profile, and sequencing mutation profile to a regular database containing mapping for treatments for which the biological activity against has been determined; and (f) i. as shown by the comparison in step (e), the treatment should have biological activity against the cancer, and ii. identifying a candidate treatment when the comparison in step (e) does not indicate a contraindication to the treatment for the treatment of the cancer. [Invention 1003] A method for identifying a candidate treatment for cancer in a subject in need thereof, comprising the following steps: (a) performing immunohistochemical (IHC) analysis on a sample from the subject to determine an IHC expression profile for at least a protein group consisting of SPARC, PGP, Her2 / neu, ER, PR, c-kit, AR, CD52, PDGFR, TOP2A, TS, ERCC1, RRM1, BCRP, TOPO1, PTEN, MGMT, and MRP1; (b) performing microarray analysis on the sample to determine a microarray expression profile regarding at least a gene group consisting of ABCC1, ABCG2, ADA, AR, ASNS, BCL2, BIRC5, BRCA1, BRCA2, CD33, CD52, CDA, CES2, DCK, DHFR, DNMT1, DNMT3A, DNMT3B, ECGF1, EGFR, EPHA2, ERBB2, ERCC1, ERCC3, ESR1, FLT1, FOLR2, FYN, GART, GNRH1, GSTP1, HCK, HDAC1, HIF1A, HSP90AA1, IL2RA, HSP90AA1, KDR, KIT, LCK, LYN, MGMT, MLH1, MS4A1, MSH2, NFKB1, NFKB2, OGFR, PDGFC, PDGFRA, PDGFRB, PGR, POLA1, PTEN, PTGS2, RAF1, RARA, RRM1, RRM2, RRM2B, RXRB, RXRG, SPARC, SRC, SSTR1, SSTR2, SSTR3, SSTR4, SSTR5, TK1, TNF, TOP1, TOP2A, TOP2B, TXNRD1, TYMS, VDR, VEGFA, VHL, YES1, and ZAP70; (c) performing fluorescence in situ hybridization (FISH) analysis on the sample to determine a FISH mutation profile regarding at least a gene group consisting of EGFR and HER2; (d) performing DNA sequencing on the sample to determine a sequencing mutation profile regarding at least a gene group consisting of KRAS, BRAF, c-KIT, and EGFR; and (e) i. cancer cells that overexpress or underexpress one or more proteins included in the IHC expression profile, ii. cancer cells that overexpress or underexpress one or more genes included in the microarray expression profile, iii. cancer cells that do not have a mutation or have one or more mutations in one or more genes included in the FISH mutation profile, and / or iv. cancer cells that do not have mutations in one or more genes included in the array-determined mutation profile or have one or more mutations but for which the biological activity against them has been determined comparing the IHC expression profile, microarray expression profile, FISH mutation profile, and sequencing mutation profile against a regular database that includes mapping for treatments for which the biological activity against cancer cells has been determined; and (f) i. showing by the comparison in step (e) that the treatment should have biological activity against the cancer and ii. identifying a candidate treatment when the comparison in step (e) does not indicate a contraindication to the treatment for treating the cancer. [Invention 1004] The method according to Invention 1001, 1002, or 1003, wherein the sample comprises formalin-fixed paraffin-embedded (FFPE) tissue, fresh frozen (FF) tissue, or tissue contained in a solution that preserves nucleic acid or protein molecules. [Invention 1005] The method according to Invention 1001, 1002, or 1003, wherein none of microarray analysis, FISH mutation analysis, or sequencing mutation analysis is performed. [Invention 1006] The method according to Invention 1001, 1002, or 1003, wherein the sample passes a quality control test. [Invention 1007] The method according to Invention 1006, wherein the quality control test includes the A260 / A280 ratio or the Ct value of RT-PCR of RPL13a mRNA. [Invention 1008] The method according to Invention 1007, wherein the quality control test includes an A260 / A280 ratio lower than 1.5 or an RPL13a Ct value higher than 30. [Invention 1009] The method according to Invention 1002, wherein IHC expression profiling is performed for at least 50%, 60%, 70%, 80%, or 90% of the biomarkers described in step (a). [Invention 1010] The method of the present invention 1002, wherein microarray expression profiling is performed for at least 50%, 60%, 70%, 80%, or 90% of the biomarkers described in step (b). [The present invention 1011] The method of the present invention 1001, 1002, or 1003, wherein microarray expression profiling is performed using a low-density microarray, an expression microarray, a comparative genomic hybridization (CGH) microarray, a single nucleotide polymorphism (SNP) microarray, a proteomics array, or an antibody array. [The present invention 1012] The method of the present invention 1001, 1002, or 1003, wherein IHC expression profiling is performed for at least SPARC, TOP2A, and / or PTEN. [The present invention 1013] The method of the present invention 1001, 1002, or 1003, wherein microarray expression profiling is performed for at least CD52. [The present invention 1014] The method of the present invention 1002, wherein IHC expression profiling further comprises assaying one or more of DCK, EGFR, BRCA1, CK 14, CK 17, CK 5 / 6, E-cadherin, p95, PARP-1, SPARC, and TLE3. [The present invention 1015] The method of the present invention 1002, wherein IHC expression profiling further comprises assaying Cox-2 and / or Ki-67. [The present invention 1016] The method of the present invention 1002, wherein microarray expression profiling further comprises assaying HSPCA. [The present invention 1017] The method of the present invention 1002, wherein FISH mutation profiling further comprises assaying c-Myc and / or TOP2A. [The present invention 1018] The method of the present invention 1002, wherein sequencing mutation profiling further comprises assaying PI3K. [The present invention 1019] The method of the present invention 1001, wherein the genes used for IHC expression profiling, microarray expression profiling, FISH mutation profiling, and sequencing mutation profiling independently include one or more of ABCC1, ABCG2, ACE2, ADA, ADH1C, ADH4, AGT, androgen receptor, AR, AREG, ASNS, BCL2, BCRP, BDCA1, BIRC5, B-RAF, BRCA1, BRCA2, CA2, caveolin, CD20, CD25, CD33, CD52, CDA, CDK2, CDW52, CES2, CK 14, CK 17, CK 5 / 6, c-KIT, c-Myc, COX-2, cyclin D1, DCK, DHFR, DNMT1, DNMT3A, DNMT3B, E-cadherin, ECGF1, EGFR, EPHA2, epiregulin, ER, ERBR2, ERCC1, ERCC3, EREG, ESR1, FLT1, folate receptor, FOLR1, FOLR2, FSHB, FSHPRH1, FSHR, FYN, GART, GNRH1, GNRHR1, GSTP1, HCK, HDAC1, Her2 / Neu, HGF, HIF1A, HIG1, HSP90, HSP90AA1, HSPCA, IL13RA1, IL2RA, KDR, KIT, K-RAS, LCK, LTB, lymphotoxin β receptor, LYN, MGMT, MLH1, MRP1, MS4A1, MSH2, Myc, NFKB1, NFKB2, NFKBIA, ODC1, OGFR, p53, p95, PARP-1, PDGFC, PDGFR, PDGFRA, PDGFRB, PGP, PGR, PI3K, POLA, POLA1, PPARG, PPARGC1, PR, PTEN, PTGS2, RAF1, RARA, RRM1, RRM2, RRM2B, RXRB, RXRG, SPARC, SPARC MC, SPARC PC, SRC, SSTR1, SSTR2, SSTR3, SSTR4, SSTR5, survivin, TK1, TLE3, TNF, TOP1, TOP2A, TOP2B, TOPO1, TOPO2B, topoisomerase II, TS, TXN, TXNRD1, TYMS, VDR, VEGF, VEGFA, VEGFC, VHL, YES1, and ZAP70. [The present invention 1020] The method of the present invention 1001, wherein the IHC expression profiling comprises assaying one or more of SPARC, PGP, Her2 / neu, ER, PR, c-kit, AR, CD52, PDGFR, TOP2A, TS, ERCC1, RRM1, BCRP, TOPO1, PTEN, MGMT, and MRP1. [The present invention 1021] The method of the present invention 1001, wherein the microarray expression profiling comprises assaying one or more of ABCC1, ABCG2, ADA, AR, ASNS, BCL2, BIRC5, BRCA1, BRCA2, CD33, CD52, CDA, CES2, DCK, DHFR, DNMT1, DNMT3A, DNMT3B, ECGF1, EGFR, EPHA2, ERBB2, ERCC1, ERCC3, ESR1, FLT1, FOLR2, FYN, GART, GNRH1, GSTP1, HCK, HDAC1, HIF1A, HSP90AA1, IL2RA, HSP90AA1, KDR, KIT, LCK, LYN, MGMT, MLH1, MS4A1, MSH2, NFKB1, NFKB2, OGFR, PDGFC, PDGFRA, PDGFRB, PGR, POLA1, PTEN, PTGS2, RAF1, RARA, RRM1, RRM2, RRM2B, RXRB, RXRG, SPARC, SRC, SSTR1, SSTR2, SSTR3, SSTR4, SSTR5, TK1, TNF, TOP1, TOP2A, TOP2B, TXNRD1, TYMS, VDR, VEGFA, VHL, YES1, and ZAP70. [The present invention 1022] The method of the present invention 1001, wherein the FISH mutation profiling comprises assaying EGFR and / or HER2. [The present invention 1023] The method of the present invention 1001, wherein the sequencing mutation profiling comprises assaying one or more of KRAS, BRAF, c-KIT, and EGFR. [The present invention 1024] The method of the present invention 1001, 1002, or 1003, wherein the microarray expression analysis comprises identifying whether a gene is upregulated or downregulated with statistical significance compared to a reference. [Invention 1025] The method of Invention 1024, wherein statistical significance is determined by a p-value of 0.05, 0.01, 0.005, 0.001, 0.0005, or 0.0001 or less. [Invention 1026] The method of Invention 1025, wherein the p-value is corrected for multiple comparisons. [Invention 1027] The method of Invention 1026, wherein the correction for multiple comparisons includes the Bonferroni correction or a modified method thereof. [Invention 1028] The method of Invention 1001, 1002, or 1003, wherein the IHC analysis includes determining whether 30% or more of the sample has a staining intensity of +2 or more. [Invention 1029] The method of Invention 1001, 1002, or 1003, wherein the rule database includes the rules described in Table 1 and / or Table 2. [Invention 1030] The method of Invention 1001, 1002, or 1003, wherein the rules included in the rule database are based on the effectiveness of various treatments specific to the target gene or gene product. [Invention 1031] The method of Invention 1001, 1002, or 1003, wherein a priority list of candidate treatments is identified. [Invention 1032] The method of Invention 1031, wherein the prioritization follows the treatments based on microarray analysis and either IHC analysis or FISH analysis; the treatment based on IHC analysis rather than microarray analysis; and the treatment based on microarray analysis rather than IHC analysis, and orders the treatments from the highest priority to the lowest priority. [Invention 1033] The method of Invention 1001, 1002, or 1003, wherein the treatment includes the administration of one or more candidate therapeutic agents. [Invention 1034] The method of the present invention 1033, wherein one or more candidate therapeutic agents include 5-fluorouracil, abarelix, alemtuzumab, aminoglutethimide, Anastrazole, aromatase inhibitors (Anastrazole, letrozole), asparaginase, aspirin, ATRA, azacitidine, bevacizumab, bexarotene, bicalutamide, bortezomib, calcitriol, capecitabine, carboplatin, celecoxib, cetuximab, chemoendocrine therapy, colecalciferol, cisplatin, carboplatin, cyclophosphamide, cyclophosphamide / vincristine, cytarabine, dasatinib, decitabine, doxorubicin, epirubicin, epirubicin, erlotinib, etoposide, exemestane, fluoropyrimidine, flutamide, fulvestrant, gefitinib, gefitinib and trastuzumab, gemcitabine, gonadorelin, goserelin, hydroxyurea, imatinib, irinotecan, ixabepilone, lapatinib, letrozole, leuprolide, liposomal doxorubicin, medroxyprogesterone, megestrol, methotrexate, mitomycin, nab-paclitaxel, octreotide, oxaliplatin, paclitaxel, panitumumab, pegaspargase, pemetrexed, pentostatin, sorafenib, sunitinib, tamoxifen, tamoxifen-based therapy, temozolomide, topotecan, toremifene, trastuzumab, VBMCP / cyclophosphamide, vincristine, or any combination thereof. [The present invention 1035] The method of the present invention 1033, wherein one or more candidate therapeutic agents include 5FU, bevacizumab, capecitabine, cetuximab, cetuximab + gemcitabine, cetuximab + irinotecan, cyclophosphamide, diethylstilbestrol, doxorubicin, erlotinib, etoposide, exemestane, fluoropyrimidine, gemcitabine, gemcitabine + etoposide, gemcitabine + pemetrexed, irinotecan, irinotecan + sorafenib, lapatinib, lapatinib + tamoxifen, letrozole, letrozole + capecitabine, mitomycin, nab-paclitaxel, nab-paclitaxel + gemcitabine, nab-paclitaxel + trastuzumab, oxaliplatin, oxaliplatin + 5FU + trastuzumab, panitumumab, pemetrexed, sorafenib, sunitinib, sunitinib, sunitinib + mitomycin, tamoxifen, temozolomide, temozolomide + bevacizumab, temozolomide + sorafenib, trastuzumab, vincristine, or any combination thereof. [The present invention 1036] The method of the present invention 1001 or 1002, wherein the sample contains cancer cells. [The present invention 1037] The method of the present invention 1001, 1002, or 1003, wherein the subject has been previously treated with one or more therapeutic agents for treating cancer. [The present invention 1038] The method of the present invention 1001, 1002, or 1003, wherein the subject has not been previously treated with one or more candidate therapeutic agents identified in step (f). [The present invention 1039] The method of the present invention 1001, 1002, or 1003, wherein the cancer includes metastatic cancer. [The present invention 1040] The method of the present invention 1001, 1002, or 1003, wherein the cancer is refractory to previous treatment. [The present invention 1041] The method of the present invention 1040, wherein the previous treatment includes standard care for cancer. [The present invention 1042] The method of the present invention 1001, 1002, or 1003, wherein the cancer comprises prostate cancer, lung cancer, melanoma, small cell carcinoma (esophagus / retroperitoneum), cholangiocarcinoma, mesothelioma, head and neck cancer (SCC), pancreatic cancer, pancreatic neuroendocrine cancer, small cell carcinoma, gastric cancer, peritoneal pseudomyxoma, anal canal cancer (SCC), vaginal cancer (SCC), cervical cancer, renal cancer, eccrine sweat gland adenocarcinoma, salivary gland adenocarcinoma, uterine soft tissue sarcoma (uterus), GIST (stomach), or anaplastic thyroid cancer. [The present invention 1043] The method of the present invention 1001, 1002, or 1003, wherein the cancer comprises cancer of the accessory sinuses, middle ear, and inner ear, adrenal gland, appendix, hematopoietic system, bone and joint, spinal cord, breast, cerebellum, cervix, connective soft tissue, uterine body, esophagus, eye, nose, eyeball, fallopian tube, extrahepatic bile duct, oral cavity, intrahepatic bile duct, kidney, appendix - colon, larynx, lip, liver, lung and bronchus, lymph node, cerebrum, spinal cord, nasal cartilage, retina, eye, oropharynx, endocrine gland, female genitalia, ovary, pancreas, penis and scrotum, pituitary gland, pleura, prostate, rectum renal pelvis, ureter, peritoneum, salivary gland, skin, small intestine, stomach, testis, thymus, thyroid, tongue, unknown, bladder, uterus, vagina, labia, and vulva. [The present invention 1044] The method of the present invention 1001, 1002, or 1003, wherein the sample comprises cells selected from the group consisting of fat, adrenal cortex, adrenal gland, adrenal - medulla, appendix, bladder, blood, blood vessel, bone, osteochondral, brain, breast, cartilage, cervix, colon, sigmoid colon, dendritic cell, skeletal muscle, endometrium, esophagus, fallopian tube, fibroblast, gallbladder, kidney, larynx, liver, lung, lymph node, melanocyte, mesothelial inner layer, myoepithelial cell, osteoblast, ovary, pancreas, parotid gland, prostate, salivary gland, accessory nasal tissue, skeletal muscle, skin, small intestine, smooth muscle, stomach, synovium, joint lining tissue, tendon, testis, thymus, thyroid, uterus, and uterine body. [The present invention 1045] The method of the present invention 1001, 1002, or 1003, wherein the cancer comprises breast cancer, colorectal cancer, ovarian cancer, lung cancer, non - small cell lung cancer, cholangiocarcinoma, mesothelioma, sweat gland cancer, or GIST cancer. [The present invention 1046] The method of the present invention 1001, 1002, or 1003, wherein the progression - free survival (PFS) or disease - free survival (DFS) of the subject is prolonged. [The present invention 1047] The method of the present invention 1046, wherein PFS or DFS is extended by at least about 10%, about 15%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or at least about 100% compared to previous treatment. [The present invention 1048] The method of the present invention 1001, 1002, or 1003, wherein the selection of a candidate treatment extends the lifespan of the subject. [The present invention 1049] The method of the present invention 1048, wherein the lifespan of the patient is extended by at least 1 week, 2 weeks, 3 weeks, 4 weeks, 1 month, 5 weeks, 6 weeks, 7 weeks, 8 weeks, 2 months, 9 weeks, 10 weeks, 11 weeks, 12 weeks, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months, 9 months, 10 months, 11 months, 12 months, 13 months, 14 months, 15 months, 16 months, 17 months, 18 months, 19 months, 20 months, 21 months, 22 months, 23 months, 24 months, 2 years, 2.5 years, 3 years, 4 years, or at least 5 years.

[0027] Incorporation by reference All publications and patent applications cited herein are hereby incorporated by reference as if each individual publication or patent application were specifically and individually indicated to be incorporated by reference.

Brief description of the drawings

[0028] A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description, which describes exemplary embodiments in which the principles of the present invention are utilized, and to the accompanying drawings.

Figure 1

Figure 2

Figure 3A

Figure 3B

Figure 3C

Figure 3D

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26A

Figure 26B

Figure 26C

Figure 26D

Figure 26E

Figure 26F

Figure 26G

Figure 26H

Figure 27A

Figure 27B

Figure 27C

Figure 27D

Figure 27E

Figure 27F

Figure 27G

Figure 27H

Figure 28A

Figure 28B

Figure 28C

Figure 28D

Figure 28E

Figure 28F

Figure 28G

Figure 28H

Figure 28I

Figure 28J

Figure 28K

Figure 28L

Figure 28M

Figure 28N

Figure 28O

Figure 29

Figure 30A

Figure 30B

Figure 30C

Figure 30D

Figure 30E

Figure 30F

Figure 30G

Figure 30H

Figure 30I

Figure 30J

Figure 30K

Figure 30L

Figure 30M

Figure 30N

Figure 30O

Figure 31

Figure 32

Figure 33

Figure 34

Figure 35

Figure 36

Figure 37

Figure 38

Figure 39

Figure 40A

Figure 40B

Figure 40C

Figure 40D

Figure 40E

Figure 40F

Figure 40G

Figure 40H

Figure 40I

Figure 40J

DETAILED DESCRIPTION OF THE INVENTION

[0029] Detailed Description of the Invention The present invention provides a method and system for identifying therapeutic targets by using molecular profiling. The molecular profiling approach provides a method for selecting for an individual a candidate treatment that can conveniently alter the clinical course of an individual having a condition or disease such as cancer. The molecular profiling approach can provide a clinical benefit to an individual, such as a longer progression-free survival (PFS), a longer disease-free survival (DFS), a longer overall survival (OS), or an extension of lifespan, when treating using the molecular profiling approach rather than using a conventional approach for selecting a treatment plan. Molecular profiling can suggest a candidate treatment when the disease is refractory to current therapies, such as after cancer has developed resistance to standard care.

[0030] Molecular profiling can be performed by any known means for detecting molecules in a biological sample. Profiling can be performed in any applicable biological sample. The sample typically derives from an individual in whom a disease or disorder is suspected or known, and examples include, but are not limited to, a biopsy sample from a cancer patient. Molecular profiling of the sample can also be performed by a number of techniques for assessing the amount or status of biological factors such as DNA sequences, mRNA sequences, or proteins. Such techniques include, without limitation, immunohistochemistry (IHC), in situ hybridization (ISH), fluorescence in situ hybridization (FISH), various types of microarrays (mRNA expression arrays, protein arrays, etc.), various types of sequencing (Sanger, pyrosequencing, etc.), comparative genomic hybridization (CGH), next-generation sequencing, Northern blot, Southern blot, immunoassays, and any other suitable techniques under development for assaying the presence or amount of a biomolecule of interest. Any one or more of these methods can be used simultaneously or sequentially with each other.

[0031] Molecular profiling is used to select candidate therapies for disorders in a subject. For example, a candidate therapy may be a therapy that has been found to be effective against cells that differentially express genes as identified by molecular profiling techniques. Differential expression can include both overexpression and underexpression of biological products such as genes, mRNAs, or proteins compared to a control. The control may be similar to the sample but may include cells that do not have the disease. The control may be from the same patient, such as a normal adjacent part of the same organ as the diseased cells, or the control may be from healthy tissue from other patients. The control may be a control found in the same sample, such as a housekeeping gene or its product (e.g., mRNA or protein). For example, a control nucleic acid may be one that has been found not to vary according to the cancerous or non-cancerous state of the cell. The expression level of the control nucleic acid can be used to normalize the signal levels in the test population and the reference population. Exemplary control genes include, but are not limited to, β-actin, glyceraldehyde 3-phosphate dehydrogenase, and ribosomal protein P1. Multiple controls or multiple types of controls can be used. The causes of differential expression can vary. For example, the gene copy number may increase in a cell, thereby causing an increase in the expression of the gene. Alternatively, the transcription of a gene may be altered, for example, by chromatin remodeling, differential methylation, differential expression or activity of transcription factors, etc. For example, translation may also be altered by differential expression of factors that degrade mRNA, translate mRNA, or inhibit translation, such as microRNAs or siRNAs. In some embodiments, differential expression includes differential activity. For example, a protein may carry a mutation that increases the activity of the protein, such as constitutive activation, which may contribute to the diseased state. Molecular profiling that reveals changes in activity can be used to guide treatment selection.

[0032] If multiple drug targets are revealed to be differentially expressed by molecular profiling, decision rules can be introduced to prioritize the selection of a particular treatment. For example, any such rules that help prioritize treatment can be used, such as the direct results of molecular profiling, predicted efficacy, prior history with the same or other treatments, predicted side effects, availability, cost, drug interactions, and other factors considered by the treating physician. The physician can ultimately decide on the course of treatment. Thus, molecular profiling can select candidate treatments based on the individual characteristics of diseased cells, such as tumor cells, and other individual factors in the subject in need of treatment, as opposed to relying on traditional, uniform approaches taken to target therapies for certain symptoms. In some cases, the recommended treatment may not typically be used to treat the disease or disorder from which the subject suffers. In some cases, the recommended treatment may be used after standard care has ceased to provide adequate efficacy.

[0033] Nucleic acids include deoxyribonucleotides or ribonucleotides, and polymers thereof in single-stranded or double-stranded form, and their complements. Nucleic acids can include known nucleotide analogs or modified backbone residues or linkages that are synthetic, natural, and non-natural, have binding properties similar to reference nucleic acids, and are metabolized in a manner similar to reference nucleotides. Examples of such analogs include, but are not limited to, phosphorothioates, phosphoramidates, methylphosphonic acids, chiral-methylphosphonic acids, 2-O-methyl ribonucleotides, peptide nucleic acids (PNAs). Nucleic acid sequences can include, in addition to the explicitly recited sequences, conservatively modified variants thereof (e.g., degenerate codon substitutions) and complementary sequences. Specifically, degenerate codon substitutions can be obtained by generating sequences in which the third position of one or more selected (or all) codons is substituted with a mixture of bases and / or deoxyinosine residues (Batzer et al., Nucleic Acid Res. 19:5081 (1991); Ohtsuka et al., J. Biol. Chem. 260:2605-2608 (1985); Rossolini et al., Mol. Cell Probes 8:91-98 (1994)). The term nucleic acid can be used interchangeably with gene, cDNA, mRNA, oligonucleotide, and polynucleotide.

[0034] A particular nucleic acid sequence may implicitly encompass the nucleic acid sequences encoding that particular sequence, "splice variants", and cleavage forms. Similarly, a particular protein encoded by a nucleic acid may encompass any protein encoded by a splice variant or cleavage form of that nucleic acid. A "splice variant" is, as the name implies, the product of alternative splicing of a gene. After transcription, the initial nucleic acid transcript may be spliced so that different (alternative) nucleic acid splice products encode different polypeptides. The mechanisms for producing splice variants are diverse and include alternative splicing of exons. Alternative polypeptides derived from the same nucleic acid by read-through transcription are also encompassed by this definition. Any product of a splicing reaction, including recombinant splice products, is included in this definition. A nucleic acid may be cleaved at the 5' or 3' end. A polypeptide may be cleaved at the N-terminal or C-terminal end. Cleavage forms of nucleic acid or polypeptide sequences may occur naturally or may be produced recombinantly.

[0035] The terms "gene variant" and "nucleotide variant" are used interchangeably herein to refer to changes or alterations to a reference human gene or cDNA sequence at a particular locus, including, but not limited to, deletions, insertions, inversions, and substitutions of nucleotide bases in coding and non-coding regions. Deletions may be of a single nucleotide base, a portion or region of the nucleotide sequence of a gene, or the entire gene sequence. Insertions may be of one or more nucleotide bases. Gene variants or nucleotide variants may occur in transcriptional regulatory regions, untranslated regions of mRNA, exons, introns, exon / intron junctions, etc. Gene variants or nucleotide variants may potentially result in premature stop codons, frameshifts, amino acid deletions, changes in gene transcript splice patterns, or changes in amino acid sequences.

[0036] Alleles or gene alleles generally include a natural gene having a reference sequence or a gene containing a particular nucleotide variant.

[0037] A haplotype refers to a combination of gene (nucleotide) variants in a region of mRNA or genomic DNA on one chromosome found in an individual. Thus, a haplotype typically includes several genetically linked polymorphic variants that are inherited together as a unit.

[0038] As used herein, the term "amino acid variant" refers to an amino acid change to a reference protein sequence resulting from a gene variant or nucleotide variant relative to a reference human gene encoding the reference human protein. The term "amino acid variant" is intended to encompass not only single amino acid substitutions, but also deletions, insertions, and other significant changes in the amino acid sequence of the reference protein.

[0039] As used herein, the term "genotype" means the nucleotide characteristics of a particular nucleotide variant marker (or locus) in one or both alleles of a gene (or a particular chromosomal region). With respect to a particular nucleotide position of a gene of interest, the nucleotide or its equivalent at that locus in one or both alleles forms the genotype of the gene at that locus. A genotype may be homozygous or heterozygous. Thus, "genotyping" means determining the genotype, i.e., the nucleotides at a particular locus. Genotyping can also be performed by determining amino acid variants at specific positions of a protein that can be used to infer the corresponding nucleotide variants.

[0040] The term "locus" refers to a specific position or site in a gene sequence or protein. Thus, one or more contiguous nucleotides may be present at a particular locus, or one or more amino acids may be present at a specific locus in a polypeptide. Further, a locus may refer to a specific position of a gene in which one or more nucleotides are deleted, inserted, or inverted.

[0041] As used herein, the terms "polypeptide", "protein", and "peptide" are used interchangeably and refer to a chain of amino acid residues linked by covalent peptide bonds. The amino acid chain may be of any length of at least two amino acids, including full-length proteins. Unless otherwise specified, polypeptides, proteins, and peptides include their various modified forms, including but not limited to glycosylated forms, phosphorylated forms, etc. A polypeptide, protein, or peptide may also be referred to as a gene product.

[0042] A list of genes and gene products that can be assayed by molecular profiling techniques is provided herein. The list of genes may be presented in connection with molecular profiling techniques for detecting gene products (e.g., mRNA or protein). One of ordinary skill in the art will understand that this means detecting the gene products of the genes described. Similarly, the list of gene products may be presented in connection with molecular profiling techniques for detecting gene sequences or copy numbers. One of ordinary skill in the art will understand that this means detecting the genes corresponding to the gene products, including, for example, the DNA encoding the gene products. As will be understood by one of ordinary skill in the art, "biomarker" or "marker" includes genes and / or gene products depending on the context.

[0043] The terms "label" and "detectable label" can refer to any composition detectable by spectroscopic, photochemical, biochemical, immunochemical, electrical, optical, chemical methods, or similar methods. Such labels include biotin for staining with a labeled streptavidin conjugate, magnetic beads (e.g., DYNABEADS (trademark)), fluorescent dyes (e.g., fluorescein, Texas Red, rhodamine, green fluorescent protein, etc.), radiolabels (e.g., 3 H, 125 I, 35 S, 14 C, or 32P), enzymes (e.g., horseradish peroxidase, alkaline phosphatase, and others commonly used in ELISA), and colorimetric labels such as colloidal gold or colored glass or plastic (e.g., polystyrene, polypropylene, latex, etc.) beads. Patents teaching the use of such labels include U.S. Patent Nos. 3,817,837; 3,850,752; 3,939,350; 3,996,345; 4,277,437; 4,275,149; and 4,366,241. Means for detecting such labels are well known to those skilled in the art. Thus, for example, radiolabels can be detected using a photographic film or a scintillation counter, and fluorescent markers can be detected using a photodetector for detecting the emitted light. Enzyme labels are typically detected by providing a substrate to the enzyme and detecting the reaction product produced by the action of the enzyme on the substrate, and colorimetric labels are detected by simply visualizing the colored label. Labels can include, for example, ligands that bind to labeled antibodies, fluorophores, chemiluminescent agents, enzymes, and antibodies that can function as specific binding partner members of labeled ligands. An overview of labels, labeling procedures, and detection of labels can be found in Polak and Van Noorden Introduction to Immunocytochemistry, 2nd ed., Springer Verlag, NY (1997); and Haugland Handbook of Fluorescent Probes and Research Chemicals, a combined handbook and catalogue Published by Molecular Probes, Inc. (1996).

[0044] Detectable labels include nucleotides (labeled or unlabeled), compomers, sugars, peptides, proteins, antibodies, compounds, binding components such as conductive polymers, biotin, mass tags, colorimetric agents, luminescent agents, chemiluminescent agents, light-scattering agents, fluorescent tags, radioactive tags, charge tags (electric or magnetic charge), volatile tags and hydrophobic tags, biomolecules (e.g., antibody / antigen, antibody / antibody, antibody / antibody fragment, antibody / antibody receptor, antibody / protein A or protein G, hapten / anti-hapten, biotin / avidin, biotin / streptavidin, folic acid / folic acid-binding protein, vitamin B12 / intrinsic factor, chemical reaction groups / complementary chemical reaction groups (e.g., members of sulfhydryl / maleimide, sulfhydryl / haloacetyl derivative, amine / isotriocyanate, amine / succinimidyl ester, and amine / halogenated sulfonyl), etc.), but are not limited thereto.

[0045] As used herein, the term "antibody" includes natural and non-natural antibodies, including, for example, single-chain antibodies, chimeric antibodies, bifunctional antibodies, and humanized antibodies, as well as antigen-binding fragments thereof (e.g., Fab', F(ab')2, Fab, Fv, and rIgG). See also Pierce Catalog and Handbook, 1994-1995 (Pierce Chemical Co., Rockford, Ill). See also, for example, Kuby, J., Immunology, 3.sup.rd Ed., W. H. Freeman & Co., New York (1998). Such non-natural antibodies can be constructed using solid-phase peptide synthesis, produced recombinantly, or obtained, for example, by screening combinatorial libraries consisting of variable heavy and variable light chains as described in Huse et al., Science 246:1275-1281 (1989), which is incorporated herein by reference. These and other methods for making, for example, chimeric antibodies, humanized antibodies, CDR-grafted antibodies, single-chain antibodies, and bifunctional antibodies are well known to those of skill in the art. See, for example, Winter and Harris, Immunol. Today 14:243-246 (1993); Ward et al., Nature 341:544-546 (1989); Harlow and Lane, Antibodies, 511-52, Cold Spring Harbor Laboratory publications, New York, 1988; Hilyard et al., Protein Engineering: A practical approach (IRL Press 1992); Borrebaeck, Antibody Engineering, 2d ed. (Oxford University Press 1995), each of which is incorporated herein by reference.

[0046] Unless otherwise specified, the term "antibody" can include both polyclonal and monoclonal antibodies. The term "antibody" also includes genetically engineered forms such as chimeric antibodies (e.g., humanized mouse antibodies) and heteroconjugate antibodies (e.g., bispecific antibodies). The term also refers to recombinant single-chain Fv fragments (scFv). The term "antibody" also includes bivalent or bispecific molecules, diabodies, triabodies, and tetra-bodies. Bivalent and bispecific molecules are described, for example, in Kostelny et al. (1992) J Immunol 148:1547, Pack and Pluckthun (1992) Biochemistry 31:1579, Hollinger et al. (1993) Proc Natl Acad Sci USA. 90:6444, Gruber et al. (1994) J Immunol :5368, Zhu et al. (1997) Protein Sci 6:781, Hu et al. (1997) Cancer Res. 56:3055, Adams et al. (1993) Cancer Res. 53:4026, and McCartney, et al. (1995) Protein Eng. 8:301.

[0047] Typically, an antibody has a heavy chain and a light chain. The heavy and light chains each contain a constant region and a variable region (these regions are also known as "domains"). The light and heavy chain variable regions contain four framework regions interrupted by three hypervariable regions, also called complementarity determining regions (CDRs). The scope of the framework regions and CDRs is defined. The sequences of the framework regions of different light or heavy chains are relatively conserved within a species. The framework region of an antibody, i.e., the combined framework regions of the constituent light and heavy chains, helps to position and align the CDRs in three-dimensional space. The CDRs are mainly responsible for binding to the antigen epitope. The CDRs of each chain are typically numbered in order from the N-terminus and are called CDR1, CDR2, and CDR3, and are typically identified by the chain in which the particular CDR is located. Thus, V H CDR3 is located in the variable domain of the heavy chain of the antibody in which it is found, and V L CDR1 is the CDR1 derived from the variable domain of the light chain of the antibody in which it is found. V H References to V refer to the variable region of the immunoglobulin heavy chain of an antibody, including the heavy chain of an Fv, scFv, or Fab. V L References to V refer to the variable region of the immunoglobulin light chain, including the light chain of an Fv, scFv, dsFv, or Fab.

[0048] The term "single-chain Fv" or "scFv" refers to an antibody in which the variable domains of the heavy and light chains of a conventional two-chain antibody are linked to form one chain. Typically, a linker peptide is inserted between the two chains to allow proper folding and the creation of an active binding site. A "chimeric antibody" is an immunoglobulin molecule in which (a) the constant region or a portion thereof has been altered, replaced, or exchanged such that the antigen-binding site (variable region) is linked to a different or modified class, effector function, and / or species of constant region, or to an entirely different molecule, such as an enzyme, toxin, hormone, growth factor, drug, etc., that confers new properties to the chimeric antibody; or (b) the variable region or a portion thereof has been altered, replaced, or exchanged with a variable region having a different or modified antigen specificity.

[0049] A "humanized antibody" is an immunoglobulin molecule that contains minimal sequences derived from non-human immunoglobulins. A humanized antibody contains residues derived from the complementarity-determining regions (CDRs) of a non-human species such as a mouse, rat, or rabbit (donor antibody) that have the desired specificity, affinity, and capacity, with residues derived from the recipient's CDRs replaced. In some cases, the Fv framework residues of the human immunoglobulin are replaced with the corresponding non-human residues. A humanized antibody may also contain residues not found in the recipient antibody, the transferred CDRs, or the framework sequences. Generally, a humanized antibody contains substantially all of at least one, typically two, variable domains, with all or substantially all of the CDR regions corresponding to those of the non-human immunoglobulin and all or substantially all of the framework (FR) regions being those of the human immunoglobulin consensus sequence. A humanized antibody optimally also contains at least a portion of the immunoglobulin constant region (Fc), typically that of a human immunoglobulin (Jones et al., Nature 321:522-525 (1986); Riechmann et al., Nature 332:323-327 (1988); and Presta, Curr. Op. Struct. Biol. 2:593-596 (1992)). Humanization can essentially be carried out by replacing the corresponding sequences of a human antibody with rodent CDRs or CDR sequences according to the methods of Winter and co-workers (Jones et al., Nature 321:522-525 (1986); Riechmann et al., Nature 332:323-327 (1988); Verhoeyen et al., Science 239:1534-1536 (1988)). Thus, such humanized antibodies are chimeric antibodies (U.S. Patent No. 4,816,567), where substantially complete but non-human variable domains are replaced by the corresponding sequences from a non-human species.

[0050] The terms "epitope" and "antigenic determinant" refer to the site on an antigen to which an antibody binds. An epitope can be formed from both contiguous amino acids or non-contiguous amino acids juxtaposed by the tertiary folding of a protein. Epitopes formed from contiguous amino acids are typically retained even when exposed to a denaturing solvent, whereas epitopes formed by tertiary folding are typically lost upon treatment with a denaturing solvent. An epitope typically contains at least 3, more commonly at least 5 or 8 - 10 amino acids in its own unique spatial higher-order structure. Methods for determining the spatial higher-order structure of an epitope include, for example, X-ray crystallography and two-dimensional nuclear magnetic resonance. See, for example, Epitope Mapping Protocols in Methods in Molecular Biology, Vol. 66, Glenn E. Morris, Ed (1996).

[0051] The terms "primer", "probe", and "oligonucleotide" are used interchangeably herein to refer to relatively short nucleic acid fragments or sequences. They can include DNA, RNA, or hybrids thereof, or chemically modified analogs or derivatives thereof. Typically, they are single-stranded. However, they can also be double-stranded having two complementary strands that can be separated by denaturation. Usually, primers, probes, and oligonucleotides have a length of about 8 nucleotides to about 200 nucleotides, preferably about 12 nucleotides to about 100 nucleotides, more preferably about 18 to about 50 nucleotides. For various molecular biology applications, they can be labeled with a detectable marker or modified using conventional modalities.

[0052] The term "isolated," when used with respect to a nucleic acid (e.g., genomic DNA, cDNA, mRNA, or fragments thereof), is intended to mean that the nucleic acid molecule exists in a form substantially separated from other native nucleic acids with which it is normally associated. Since naturally occurring chromosomes (or their viral equivalents) contain long nucleic acid sequences, an isolated nucleic acid may be a nucleic acid molecule that has only a portion of the nucleic acid sequence in the chromosome and does not have one or more other portions present on the same chromosome. More specifically, an isolated nucleic acid may contain native nucleic acid sequences that flank the nucleic acid in a naturally occurring chromosome (or its viral equivalent). An isolated nucleic acid may be substantially separated from other native nucleic acids on different chromosomes of the same organism. An isolated nucleic acid may also be a composition in which a particular nucleic acid molecule is significantly enriched such that it constitutes at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or at least 99% of the total nucleic acids in the composition.

[0053] An isolated nucleic acid may be a hybrid nucleic acid in which a particular nucleic acid molecule is covalently bound to one or more nucleic acid molecules that are not nucleic acids that are naturally adjacent to the particular nucleic acid. For example, an isolated nucleic acid may be present in a vector. In addition, the particular nucleic acid may have a nucleotide sequence that is the same as a native nucleic acid, or a modified form or mutant (mutein) thereof having one or more mutations such as nucleotide substitutions, deletions / insertions, inversions, etc.

[0054] An isolated nucleic acid can be prepared from a recombinant host cell (in which the nucleic acid is amplified and / or expressed by recombination), or may be a chemically synthesized nucleic acid having a native nucleotide sequence or an artificially modified form thereof.

[0055] As used herein, the term "isolated polypeptide" is defined as a polypeptide molecule that exists in a form other than its naturally occurring form. Thus, an isolated polypeptide may be a non-natural polypeptide. For example, an isolated polypeptide may be a "hybrid polypeptide". An isolated polypeptide may also be a polypeptide derived from a natural polypeptide by addition, deletion, or substitution of amino acids. An isolated polypeptide may also be a "purified polypeptide" as used herein to mean a composition or preparation in which a particular polypeptide molecule is significantly enriched such that it constitutes at least 10% of the total protein content in the composition. As will be apparent to those skilled in the art, a "purified polypeptide" can be obtained from natural cells or recombinant host cells by standard purification techniques or can be obtained by chemical synthesis.

[0056] The terms "hybrid protein", "hybrid polypeptide", "hybrid peptide", "fusion protein", "fusion polypeptide", and "fusion peptide" are used interchangeably herein to mean a non-natural or isolated polypeptide in which a particular polypeptide molecule is covalently bound to one or more other polypeptide molecules that are not naturally linked to the particular polypeptide. Thus, a "hybrid protein" may be two natural proteins or fragments thereof that are linked together by a covalent bond. A "hybrid protein" may also be a protein formed by covalently bonding two artificial polypeptides together. Although not necessarily, typically two or more polypeptide molecules are linked or "fused" together by peptide bonds that form a single non-branched polypeptide chain.

[0057] The term "high stringency hybridization conditions", when used in reference to nucleic acid hybridization, includes hybridization overnight at 42°C in a solution containing 50% formamide, 5×SSC (750 mM NaCl, 75 mM sodium citrate), 50 mM sodium phosphate, pH 7.6, 5×Denhardt's solution, 10% dextran sulfate, and 20 μg / ml denatured and sheared salmon sperm DNA, and washing of the hybridization filter in 0.1×SSC at approximately 65°C. The term "moderately stringent conditions", when used in reference to nucleic acid hybridization, includes hybridization overnight at 37°C in a solution containing 50% formamide, 5×SSC (750 mM NaCl, 75 mM sodium citrate), 50 mM sodium phosphate, pH 7.6, 5×Denhardt's solution, 10% dextran sulfate, and 20 μg / ml denatured and sheared salmon sperm DNA, and washing of the hybridization filter in 1×SSC at approximately 50°C. It should be noted that, as will be apparent to those skilled in the art, many other hybridization methods, solutions, and temperatures can be used to achieve comparable stringent hybridization conditions.

[0058] For the purpose of comparing two different nucleic acid or polypeptide sequences, one sequence (the test sequence) may be described as being identical to the other sequence (the comparison sequence) at a particular percentage. The percentage identity can be determined by the algorithm of Karlin and Altschul, Proc. Natl. Acad. Sci. USA, 90:5873-5877 (1993), which is incorporated into various BLAST programs. The percentage identity can be determined by the "BLAST 2 Sequences" tool available on the National Center for Biotechnology Information (NCBI) website. See Tatusova and Madden, FEMS Microbiol. Lett., 174(2):247-250 (1999). For pairwise comparison between DNAs, the BLASTN program is used with the default parameters (match: 1; mismatch: -2; open gap: 5 penalty; extension gap: 2 penalty; gap x_dropoff: 50; expectation value: 10; and word size: 11, with filter). For pairwise comparison between protein sequences, the BLASTP program can be used with the default parameters (matrix: BLOSUM62; gap open: 11; gap extension: 1; x_dropoff: 15; expectation value: 10.0; and word size: 3, with filter). The percentage identity of two sequences is calculated by aligning the test sequence and the comparison sequence using BLAST, determining the number of amino acids or nucleotides in the aligned test sequence that are identical to the amino acids or nucleotides in the same position of the comparison sequence, and dividing the number of identical amino acids or nucleotides by the number of amino acids or nucleotides in the comparison sequence. When BLAST is used to compare two sequences, the sequences are thereby aligned and a percentage identity over the defined alignment region is obtained. If the two sequences are aligned over their entire lengths, the percentage identity obtained by BLAST is the percentage identity of the two sequences.If BLAST does not align two sequences over their entire lengths, the number of identical amino acids or nucleotides in the unaligned regions of the test and comparison sequences is considered zero, and the percentage identity is calculated by adding the number of identical amino acids or nucleotides in the aligned regions and dividing that number by the length of the comparison sequence. To compare sequences, various versions of the BLAST program can be used, such as BLAST 2.1.2 or BLAST+ 2.2.22.

[0059] The subject can be any animal that can benefit from the methods of the present invention, including, for example, humans, as well as non-human mammals such as primates, rodents, horses, dogs, and cats. The subject can include, without limitation, eukaryotes, most preferably, primates such as chimpanzees or humans, cows; dogs; cats; rodents such as guinea pigs, rats, mice; mammals such as rabbits; or birds; reptiles; or fish. The subjects specifically contemplated for treatment using the methods described herein include humans. The subject can be referred to as an individual or a patient.

[0060] Treatment of a disease or an individual according to the present invention is an approach for obtaining beneficial or desired medical outcomes, including clinical outcomes, which does not necessarily mean a cure. For the purposes of the present invention, beneficial or desired clinical outcomes include, but are not limited to, reduction or improvement of one or more symptoms, whether or not detectable, decrease in the degree of the disease, stable (i.e., non-worsening) state of the disease, prevention of spread of the disease, delay or slowing of disease progression, improvement or alleviation of the pathological condition, and remission (regardless of whether partial or total). Treatment also includes prolonging the survival period as compared to the survival period predicted if not receiving treatment or if receiving a different treatment. Treatment may include administration of a therapeutic agent, which may be an agent that exerts a cytotoxic, cytostatic, or immunomodulatory effect on diseased cells, such as, for example, cancer cells or other cells that may promote a pathological condition. There is no limitation on the therapeutic agents selected by the method of the present invention. Any therapeutic agent that may be seen to have a relationship between molecular profiling and the potential efficacy of the agent may be selected. Therapeutic agents include, without limitation, small molecules, protein therapies, antibody therapies, viral therapies, gene therapies, etc. Cancer treatment or therapy includes apoptotic and non-apoptotic cancer therapies, including, without limitation, chemotherapy, hormone therapy, radiation therapy, immunotherapy, and combinations thereof. Chemotherapeutic agents include, for example, therapeutic agents and combinations of therapeutic agents that treat, for example, kill, cancer cells. Examples of various types of chemotherapeutic agents include, without limitation, alkylating agents (e.g., nitrogen mustard derivatives, ethyleneimines, alkyl sulfonates, hydrazines and triazines, nitrosoureas, and metal salts), plant alkaloids (e.g., vinca alkaloids, taxanes, podophyllotoxins, and camptothecin analogs), antitumor antibiotics (e.g., anthracyclines, chromomycins, etc.), antimetabolites (e.g., folic acid antagonists, pyrimidine antagonists, purine antagonists, and adenosine deaminase inhibitors), topoisomerase I inhibitors, topoisomerase II inhibitors, and various antineoplastic agents (e.g., ribonucleotide reductase inhibitors, corticosteroid inhibitors, enzymes, antimicrotubule agents, and retinoids).

[0061] Samples used in this specification include, for example, any relevant samples that can be used for molecular profiling, such as biopsy samples or tissue sections such as tissues collected during surgery or other procedures, autopsy samples, and frozen sections collected for histological purposes. Such samples include blood and blood fractions or blood products (e.g., serum, buffy coat, plasma, platelets, red blood cells, etc.), sputum, buccal cell tissue, cultured cells (e.g., primary cultures, explants, and transformed cells), feces, urine, other biological fluids or body fluids (e.g., prostatic fluid, gastric juice, intestinal juice, renal fluid, lung fluid, cerebrospinal fluid, etc.). The sample may be processed according to techniques understood by those skilled in the art. The sample may be, without limitation, fresh, frozen or fixed. In some embodiments, the sample includes formalin-fixed paraffin-embedded (FFPE) tissue or fresh frozen (FF) tissue. The sample may include cultured cells including a primary cell line or immortalized cell line derived from the subject sample. The sample may also refer to an extract from a sample derived from the subject. For example, the sample may contain DNA, RNA, or protein extracted from a tissue or body fluid. For such purposes, many techniques and commercially available kits are available. Fresh samples from an individual can be treated with agents for preserving RNA prior to further processing such as cell lysis and extraction. The sample may include frozen samples collected for other purposes. The sample may be accompanied by relevant information such as the age, gender, and clinical symptoms of the subject; the source of the sample; and the method of sample collection and storage. The sample is typically collected from the subject.

[0062] Biopsy includes the process of collecting tissue samples for diagnostic or prognostic evaluation, and the tissue specimens themselves. Any biopsy technique known in the art can be applied to the molecular profiling method of the present invention. The biopsy technique to be applied may depend, among other factors, particularly on the tissue type to be evaluated (e.g., colon, prostate, kidney, bladder, lymph node, liver, bone marrow, blood cells, lung, breast, etc.), the size and type of the tumor (e.g., solid or suspension, blood or ascites). Representative biopsy techniques include, but are not limited to, excisional biopsy, incisional biopsy, needle biopsy, surgical biopsy, bone marrow biopsy. "Excisional biopsy" refers to the collection of the entire tumor mass and a small peripheral margin of the normal tissue surrounding it. "Incisional biopsy" refers to the collection of a wedge-shaped tissue including the cross-sectional diameter of the tumor. Molecular profiling can use "core needle biopsy" of the tumor mass or "fine needle aspiration biopsy" generally used to obtain a cell suspension from within the tumor mass. Biopsy techniques are discussed, for example, in Harrison's Principles of Internal Medicine, Kasper, et al., eds., 16th ed., 2005, Chapter 70 and Part V in its entirety.

[0063] Standard molecular biology techniques that are known in the art and not specifically described are generally performed according to the methodologies described in Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York (1989), and Ausubel et al., Current Protocols in Molecular Biology, John Wiley and Sons, Baltimore, Md. (1989), and Perbal, A Practical Guide to Molecular Cloning, John Wiley & Sons, New York (1988), and Watson et al., Recombinant DNA, Scientific American Books, New York, and Birren et al (eds) Genome Analysis: A Laboratory Manual Series, Vols. 1-4 Cold Spring Harbor Laboratory Press, New York (1998), as well as U.S. Patent Nos. 4,666,828; 4,683,202; 4,801,531; 5,192,659, and 5,272,057, which are incorporated herein by reference. Polymerase chain reaction (PCR) can generally be performed as described in PCR Protocols: A Guide To Methods And Application, Academic Press, San Diego, Calif. (1990).

[0064] Gene expression profiling In some aspects of the present invention, biomarkers are evaluated by gene expression profiling. Methods of gene expression profiling include methods based on hybridization analysis of polynucleotides and methods based on sequencing of polynucleotides. Commonly used methods known in the art for quantifying mRNA expression in a sample include Northern blotting and in situ hybridization (Parker & Barnes (1999) Methods in Molecular Biology 106:247-283); RNase protection assay (Hod (1992) Biotechniques 13:852-854); and reverse transcription polymerase chain reaction (RT-PCR) (Weis et al. (1992) Trends in Genetics 8:263-264). Alternatively, antibodies that can recognize specific double-stranded molecules including DNA duplexes, RNA duplexes, and DNA-RNA hybrid duplexes, or DNA-protein duplexes can also be used. Representative methods of gene expression analysis based on sequencing include serial analysis of gene expression (SAGE) and expression analysis by massively parallel signature sequencing (MPSS).

[0065] Reverse transcription PCR (RT-PCR) RT-PCR can be used to measure the RNA level of the biomarker of the present invention, such as the mRNA or miRNA level. RT-PCR can be used to compare such RNA levels of the biomarker of the present invention in different sample populations, with or without drug treatment, in normal and tumor tissues, to characterize gene expression patterns, to identify closely related RNAs, and to analyze RNA structures.

[0066] The first step is the isolation of RNA, such as mRNA, from the sample. The starting material may be total RNA isolated from human tumors or tumor cell lines, and the corresponding normal tissues or cell lines, respectively. Thus, RNA can be isolated from a sample, such as tumor cells or a tumor cell line, and compared to pooled DNA from healthy donors. If the source of mRNA is a primary tumor, mRNA can be extracted, for example, from a frozen tissue sample, or a paraffin-embedded and fixed (e.g., formalin-fixed) archival tissue sample.

[0067] General methods for mRNA extraction are well known in the art and are disclosed in standard textbooks of molecular biology, including Ausubel et al. (1997) Current Protocols of Molecular Biology, John Wiley and Sons. Methods for extracting RNA from paraffin-embedded tissues are disclosed, for example, in Rupp & Locker (1987) Lab Invest. 56:A67, and De Andres et al., BioTechniques 18:42044 (1995). In particular, RNA isolation can be performed using purification kits, buffer sets, and proteases from commercial manufacturers, such as Qiagen, according to the manufacturer's instructions (QIAGEN Inc., Valencia, CA). For example, total RNA can be isolated from cultured cells using a Qiagen RNeasy mini column. Many RNA isolation kits are commercially available and can be used in the methods of the present invention.

[0068] In an alternative method, the first step is the isolation of miRNAs from the target sample. The starting material is typically total RNA isolated from human tumors or tumor cell lines, and their corresponding normal tissues or cell lines, respectively. Thus, RNA can be isolated from various primary tumors or tumor cell lines, along with pooled DNA from healthy donors. When the miRNA source is a primary tumor, miRNAs can be extracted, for example, from frozen tissue samples, or paraffin-embedded and fixed (e.g., formalin-fixed) archival tissue samples.

[0069] General methods for miRNA extraction are well-known in the art and are disclosed in standard textbooks of molecular biology, including Ausubel et al. (1997) Current Protocols of Molecular Biology, John Wiley and Sons. Methods for extracting RNA from paraffin-embedded tissues are disclosed, for example, in Rupp & Locker (1987) Lab Invest. 56:A67, and De Andres et al., BioTechniques 18:42044 (1995). In particular, RNA isolation can be performed using purification kits, buffer sets, and proteases from commercial manufacturers such as Qiagen, according to the manufacturer's instructions. For example, total RNA can be isolated from cultured cells using Qiagen's RNeasy mini-columns. Many RNA isolation kits are commercially available and can be used in the methods of the present invention.

[0070] Regardless of whether the RNA contains mRNA, miRNA, or other RNA types, gene profiling by RT-PCR can involve the reverse transcription of the RNA template into cDNA, followed by amplification in a PCR reaction. Commonly used reverse transcriptases include, but are not limited to, avian myeloblastosis virus reverse transcriptase (AMV-RT) and Moloney murine leukemia virus reverse transcriptase (MMLV-RT). The reverse transcription step is typically primed using specific primers, random hexamers, or oligo-dT primers, depending on the context and purpose of the expression profiling. For example, the extracted RNA can be reverse transcribed using the Gene Amp RNA PCR kit (Perkin Elmer, Calif., USA) according to the manufacturer's instructions. The induced cDNA can then be used as a template in the subsequent PCR reaction.

[0071] In the PCR step, various thermostable DNA-dependent DNA polymerases can be used, but typically Taq DNA polymerase, which has 5'-3' nuclease activity but lacks 3'-5' proofreading exonuclease activity, is used. TaqMan PCR typically utilizes the 5'-nuclease activity of Taq or Tth polymerase to hydrolyze a hybridization probe that is bound to its target unit replication sequence, but any enzyme with equivalent 5' nuclease activity can be used. Two oligonucleotide primers are used to generate a unit replication sequence typical of a PCR reaction. To detect the nucleotide sequence located between these two PCR primers, a third oligonucleotide, i.e., a probe, is designed. This probe cannot be extended by the Taq DNA polymerase enzyme and is labeled with a reporter fluorescent dye and a quencher fluorescent dye. When the two dyes are in close proximity to each other on the probe, any laser-induced emission from the reporter dye is quenched by the quencher dye. During the amplification reaction, the Taq DNA polymerase enzyme cleaves the probe in a template-dependent manner. The resulting probe fragments dissociate in solution, and the signal from the released reporter dye is liberated from the quenching effect of the second fluorophore. Since one molecule of the reporter dye is released each time a new molecule is synthesized, detecting the unquenched reporter dye provides a basis for quantitatively interpreting the data.

[0072] TaqMan® RT-PCR can be performed using commercially available equipment, such as the ABI PRISM 7700® Sequence Detection System® (Perkin-Elmer-Applied Biosystems, Foster City, Calif., USA), or the Lightcycler (Roche Molecular Biochemicals, Mannheim, Germany). In one particular embodiment, the 5' nuclease procedure is performed on a real-time quantitative PCR apparatus such as the ABI PRISM 7700® Sequence Detection System®. This system consists of a thermocycler, a laser, a charge-coupled device (CCD), a camera, and a computer. This system amplifies samples in 96-well format on the thermocycler. During amplification, for all 96 wells, laser-induced fluorescence signals are collected in real time via an optical fiber cable and detected by the CCD. This system includes software for operating the instrument and analyzing the data.

[0073] TaqMan data are first expressed as Ct, the threshold cycle. As described above, fluorescence values are recorded during each cycle and represent the amount of product amplified up to that point in the amplification reaction. The threshold cycle (Ct) is the point at which the fluorescence signal is first recorded as being statistically significant.

[0074] To minimize the effects of error and variation between samples, RT-PCR is usually performed using an internal standard. The ideal internal standard is expressed at a constant level between different tissues and is not affected by experimental treatment. The RNAs most frequently used to normalize gene expression patterns are the mRNAs of the housekeeping genes glyceraldehyde-3-phosphate dehydrogenase (GAPDH) and β-actin.

[0075] Real-time quantitative PCR (similarly, quantitative real-time polymerase chain reaction, qRT-PCR, or q-PCR) is a more recent variation of the RT-PCR technique. Q-PCR can measure the accumulation of PCR products through a dual-labeled fluorescent generating probe (i.e., TaqMan™ probe). Real-time PCR is compatible with both quantitative competitive PCR in which an internal competitor for each target sequence is used for normalization, and quantitative comparative PCR which uses a normalization gene contained within the sample or a housekeeping gene for RT-PCR. See, for example, Held et al. (1996) Genome Research 6:986-994.

[0076] Immunohistochemistry (IHC) IHC is the process of identifying the location of antigens (e.g., proteins) within cells of a tissue that bind to antibodies specific for the antigens within the tissue. The antigen-binding antibody can be conjugated or fused to a tag that enables its detection, for example, by visualization. In some embodiments, the tag is an enzyme such as alkaline phosphatase or horseradish peroxidase that can catalyze a chromogenic reaction. The enzyme can be fused to the antibody or non-covalently conjugated, for example, using a biotin-avidin system. Alternatively, the antibody can be tagged with a fluorophore such as fluorescein, rhodamine, DyLight Fluor, or Alexa Fluor. The antigen-binding antibody can be directly tagged or the antigen-binding antibody itself can be recognized by a detection antibody carrying a tag. One or more proteins can be detected using IHC. The expression of the gene product can be related to its staining intensity compared to a control level. In some embodiments, the gene product is considered differentially expressed if its staining changes at least 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.2, 2.5, 2.7, 3.0, 4, 5, 6, 7, 8, 9, or 10-fold in the sample compared to the control.

[0077] Microarray The biomarkers of the present invention can also be identified, confirmed, and / or measured using microarray techniques. Thus, using microarray technology, expression profile biomarkers can be measured in either fresh tumor tissue or paraffin-embedded tumor tissue. In this method, the polynucleotide sequences of interest are plated or arrayed on a microchip substrate. The arrayed sequences are then hybridized with specific DNA probes derived from the cells or tissues of interest. The source of mRNA can be total RNA isolated from samples such as human tumors or tumor cell lines, and corresponding normal tissues or cell lines, for example. Thus, RNA can be isolated from various primary tumors or tumor cell lines. When the source of mRNA is a primary tumor, for example, mRNA can be extracted from frozen tissue samples routinely prepared and stored in daily clinical practice, or paraffin-embedded and fixed (e.g., formalin-fixed) archival tissue samples.

[0078] Using microarray technology, the expression profile of biomarkers can be measured in either fresh tumor tissue, paraffin-embedded tumor tissue, or body fluid. In this method, the polynucleotide sequences of interest are plated or arrayed on a microchip substrate. The arrayed sequences are then hybridized with specific DNA probes derived from the cells or tissues of interest. Similar to the RT-PCR method, the source of miRNA is typically total RNA isolated from human tumors or tumor cell lines, and corresponding normal tissues or cell lines, including body fluids such as serum, urine, tears, and exosomes. Thus, RNA can be isolated from various sources. When the source of miRNA is a primary tumor, for example, miRNA can be extracted from frozen tissue samples routinely prepared and stored in daily clinical practice.

[0079] In certain embodiments of the microarray technique, PCR amplified inserts of cDNA clones are applied to a substrate of a high density array. In one aspect, at least 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 1,500, 2,000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, 15,000, 20,000, 25,000, 30,000, 35,000, 40,000, 45,000, or at least 50,000 nucleotide sequences are applied to the substrate. Each sequence may correspond to a different gene, or multiple sequences for one gene may be arrayed. The microarrayed genes immobilized on the microchip are suitable for hybridization under stringent conditions. Fluorescently labeled cDNA probes can be generated by incorporating fluorescent nucleotides by reverse transcription of RNA extracted from a tissue of interest. The labeled cDNA probes applied to the chip specifically hybridize to each spot of DNA on the array. After performing stringent washes to remove non-specifically bound probes, the chip is scanned by a confocal laser electron microscope or by another detection method such as a CCD camera. By quantifying the hybridization of each arrayed element, the corresponding mRNA abundance can be evaluated. Using two-color fluorescence, separately labeled cDNA probes made from two RNA sources are hybridized to the array in pairs. In this way, the relative abundances of transcripts from two sources corresponding to each specific gene are measured simultaneously. Miniaturization of hybridization enables convenient and rapid assessment of the expression patterns of a large number of genes. Such methods have been shown to have the sensitivity required to detect rare transcripts expressed at a few copies per cell and to reproducibly detect differences in expression levels of at least approximately two-fold (Schena et al. (1996) Proc. Natl. Acad. Sci. USA 93(2):106-149).Microarray analysis can be performed by commercially available instruments according to the protocols of manufacturers including, but not limited to, Affymetrix GeneChip technology (Affymetrix, Santa Clara, CA), Agilent (Agilent Technologies, Inc., Santa Clara, CA) microarray technology, or Illumina (Illumina, Inc., San Diego, CA) microarray technology.

[0080] The development of microarray methods for large-scale analysis of gene expression has made it possible to systematically search for molecular markers for cancer classification and outcome prediction in various tumor types.

[0081] In some embodiments, the Agilent Whole Human Genome Microarray Kit (Agilent Technologies, Inc., Santa Clara, CA) is used. This system can analyze over 41,000 unique human genes and transcripts, all annotated publicly. This system is used according to the manufacturer's instructions.

[0082] In some embodiments, the Illumina Whole Genome DASL assay (Illumina Inc., San Diego, CA) is used. This system provides a way to simultaneously profile over 24,000 transcripts in a high-throughput manner from minimal RNA input from fresh frozen (FF) tissue sources and formalin-fixed paraffin-embedded (FFPE) tissue sources.

[0083] Microarray expression analysis involves identifying whether a gene or gene product is upregulated or downregulated compared to a reference. The identification can be performed using statistical tests to determine the statistical significance of any differential expression observed. In some embodiments, the statistical significance is determined using parametric statistical tests. Parametric statistical tests can include, for example, fractional factorial designs, analysis of variance (ANOVA), t-tests, least squares, Pearson correlation, simple linear regression, non-linear regression, multiple linear regression, or multiple non-linear regression. Alternatively, parametric statistical tests can include one-way ANOVA, two-way ANOVA, or repeated measures ANOVA. In other embodiments, the statistical significance is determined using non-parametric statistical tests. Examples include, but are not limited to, the Wilcoxon signed-rank test, Mann-Whitney test, Kruskal-Wallis test, Friedman test, Spearman's rank correlation coefficient, Kendall's tau analysis, and non-parametric regression tests. In some embodiments, the statistical significance is determined with a p-value of less than about 0.05, 0.01, 0.005, 0.001, 0.0005, or 0.0001. The microarray systems used in the methods of the present invention can assay thousands of transcripts, but data analysis only needs to be performed on the transcripts of interest, thus reducing the problem of multiple comparisons associated with performing multiple statistical tests. The p-value can also be corrected for multiple comparisons using, for example, the Bonferroni correction, its modifications, or other techniques known to those skilled in the art, such as the Hochberg correction, Holm-Bonferroni correction, Sidak correction, or Dunnett correction. The degree of differential expression can also be considered. For example, a gene can be considered to be differentially expressed if the fold change in expression compared to a control level is at least 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.2, 2.5, 2.7, 3.0, 4, 5, 6, 7, 8, 9, or 10-fold different in the sample compared to the control. Differential expression takes into account both overexpression and underexpression. A gene or gene product can be considered to be upregulated or downregulated if the differential expression meets a statistical threshold, a fold change threshold, or both.For example, the criteria for identifying differential expression can include both a p-value of 0.001 and a fold change of at least 1.5-fold (up or down). One of ordinary skill in the art will understand that such statistical measures and threshold measures can be adapted to determine differential expression by any of the molecular profiling techniques disclosed herein.

[0084] The various methods of the present invention use many types of microarrays to detect the presence and potentially the quantity of biological entities in a sample. The array typically includes addressable components that can detect the presence of entities in the sample, for example, via binding events. Microarrays include, without limitation, DNA microarrays such as cDNA microarrays, oligonucleotide microarrays, and SNP microarrays, microRNA arrays, protein microarrays, antibody microarrays, tissue microarrays, cell microarrays (also referred to as transfection microarrays), compound microarrays, and glycan arrays (glycoarrays). DNA arrays typically include addressable nucleotide sequences that can bind to sequences present in the sample. MicroRNA can be detected using microRNA arrays, for example, the MMChips by the University of Louisville, or commercially available systems by Agilent. Protein microarrays can be used to identify protein - protein interactions, or identify targets of bioactive small molecules, including, without limitation, identifying substrates of protein kinases, or identifying transcriptional factor protein activation. A protein array can include an array of nucleotide sequences that bind to various protein molecules, usually antibodies, or proteins of interest. Antibody microarrays include antibodies spotted on a protein chip that are used, for example, as capture molecules for detecting proteins or other biological substances from a sample derived from a cell or tissue lysate solution. For example, for diagnostic purposes, antibody arrays can be used to detect biomarkers from body fluids, such as serum or urine. Tissue macroarrays include separate tissue cores constructed in an array format to enable multiple histological analyses. Cell microarrays, also referred to as transfection microarrays, include various capture agents such as antibodies, proteins, or lipids that can interact with cells and facilitate their capture at addressable positions. Compound microarrays include an array of compounds and can be used to detect proteins or other biological substances that bind to the compounds. Glycan arrays (glycoarrays) include an array of glycans and can, for example, detect proteins that bind to glycan components.Those skilled in the art will understand that similar techniques or improved methods can be used in accordance with the method of the present invention.

[0085] Gene expression analysis by massively parallel signature sequencing (MPSS) This method, described by Brenner et al. (2000) Nature Biotechnology 18:630-634, is a sequencing approach that combines non-gel-based signature sequencing with in vitro cloning of millions of templates on separate microbeads. First, a microbead library of DNA templates is constructed by in vitro cloning. Next, a planar array of template-containing microbeads is assembled at high density within a flow cell. The free ends of the cloned templates on each microbead are simultaneously analyzed using a fluorescence-based signature sequencing method that does not require separation of DNA fragments. This method has been shown to provide hundreds of thousands of gene signature sequences simultaneously and accurately from a cDNA library in a single operation.

[0086] MPSS data have many applications. The expression levels of almost all transcripts can be quantitatively measured; the abundance of the signatures represents the expression levels of the genes in the analyzed tissue. Quantitative methods for analyzing tag frequencies and detecting differences between libraries have been published and incorporated into the public database of SAGE (trademark) data and are applicable to MPSS data. The availability of the complete genome sequence allows for direct comparison of signatures with the genome sequence, further expanding the utility of MPSS data. Since the targets of MPSS analysis are not preselected (similar to microarrays), MPSS data can characterize the full complexity of the transcriptome. This is similar to sequencing millions of ESTs at once, and genomic sequence data can be used so that the source of the MPSS signatures can be readily identified by computer means.

[0087] Serial analysis of gene expression (SAGE) Serial analysis of gene expression (SAGE) is a method that enables the simultaneous and quantitative analysis of a large number of gene transcripts without the need to provide individual hybridization probes for each transcript. First, short sequence tags (e.g., about 10-14 bp) that contain sufficient information to uniquely identify a transcript are generated, provided that the tags are obtained from unique positions within each transcript. Next, many transcripts are ligated together to form long continuous molecules, which can be sequenced to simultaneously reveal the identities of multiple tags. By measuring the abundance of individual tags and identifying the genes corresponding to each tag, the expression pattern of any population of transcripts can be quantitatively evaluated. See, for example, Velculescu et al. (1995) Science 270:484-487; and Velculescu et al. (1997) Cell 88:243-51.

[0088] DNA copy number profiling Any method that has the ability to determine the DNA copy number profile of a particular sample can be used for molecular profiling according to the present invention, provided that the resolution is sufficient to identify the biomarkers of the present invention. Those skilled in the art will recognize and be able to use several different platforms to evaluate genome-wide copy number changes at a resolution sufficient to identify the copy number of one or more biomarkers of the present invention. Some of the platforms and techniques are described in the following embodiments.

[0089] In some embodiments, copy number profile analysis involves the amplification of genomic DNA by whole genome amplification methods. Whole genome amplification methods can use strand displacement polymerases and random primers.

[0090] In some aspects of these embodiments, copy number profile analysis involves hybridization of whole genome amplified DNA with a high density array. In more specific aspects, the high density array has 5,000 or more different probes. In another specific aspect, the high density array has 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, or 1,000,000 or more different probes. In another specific aspect, the different probes on the array are each oligonucleotides having a length of about 15 to 200 bases. In another specific aspect, the different probes on the array are each oligonucleotides having a length of about 15 to 200, 15 to 150, 15 to 100, 15 to 75, 15 to 60, or 20 to 55 bases.

[0091] In some embodiments, a microarray is used to assist in determining the copy number profile of a sample, such as cells derived from a tumor. A microarray typically includes a plurality of oligomers (e.g., polynucleotides or oligonucleotides of DNA or RNA, or other polymers) synthesized or deposited on a substrate (e.g., a glass support) in an array pattern. The support-bound oligomers are "probes" that function to hybridize or bind to a sample substance (e.g., nucleic acids prepared or obtained from a tumor sample) in a hybridization experiment. The reverse situation is also applicable: the sample can be bound to the microarray substrate and the oligomer probes are present in the hybridization solution. In use, the array surface is contacted with one or more targets under conditions that promote specific high-affinity binding of the target to one or more of the probes. In some settings, the sample nucleic acids are labeled with a detectable label, such as a fluorescent tag, so that the hybridized sample and probes are detectable by a scanning device. DNA array technology provides the ability to analyze DNA copy number profiles using a large number (hundreds of thousands) of different oligonucleotides. In some embodiments, the substrate used for the array is a surface-derivatized glass or silica, or a polymer film surface (see, e.g., Z. Guo, et al., Nucleic Acids Res, 22, 5456-65 (1994); U. Maskos, E. M. Southern, Nucleic Acids Res, 20, 1679-84 (1992), and E. M. Southern, et al., Nucleic Acids Res, 22, 1368-73 (1994), each incorporated herein by reference). Modification of the surface of the array substrate can be accomplished by a number of techniques.For example, the surface of silica or a metal oxide can be derivatized with a bifunctional silane, i.e., a silane having a first functional group that enables covalent bonding to the surface (e.g., a Si-halogen or Si-alkoxy group as seen in --SiCl3 or --Si(OCH3)3, respectively) and a second functional group that can impart a desired chemical and / or physical modification to the surface, to covalently or non-covalently attach ligands and / or polymers or monomers for a biological probe array. Silylation derivatization and other surface derivatizations known in the art (see, e.g., U.S. Patent No. 5,624,711 to Sundberg, U.S. Patent No. 5,266,222 to Willis, and U.S. Patent No. 5,137,765 to Farnsworth, each of which is incorporated herein by reference). Other steps for preparing the array are described in U.S. Patent No. 6,649,348 to Bass et. al., assigned to Agilent Corp., which discloses a DNA array made by an in situ synthesis method.

[0092] Polymer array synthesis is also well described in the literature including: WO 00 / 58516, U.S. Patent Nos. 5,143,854, 5,242,974, 5,252,743, 5,324,633, 5,384,261, 5,405,783, 5,424,186, 5,451,683, 5,482,867, 5,491,074, 5,527,681, 5,550,215, 5,571,639, 5,578,832, 5,593,839, 5,599,695, 5,624,711, 5,631,734, 5,795,716, 5,831,070, 5,837,832, 5,856,101, 5,858,659, 5,936,324, 5,968,740, 5,974,164, 5,981,185, 5,981,956, 6,025,601, 6,033,860, 6,040,193, 6,090,555, 6,136,269, 6,269,846, and 6,428,752, 5,412,087, 6,147,205, 6,262,216, 6,310,189, 5,889,165, and 5,959,098, PCT Application Nos. PCT / US99 / 00730 (International Publication No. WO99 / 36760) and PCT / US01 / 04285 (International Publication No. WO 01 / 58593), all of which are hereby incorporated by reference in their entirety for all purposes.

[0093] Nucleic acid arrays useful in the present invention include, but are not limited to, those commercially available from Affymetrix (Santa Clara, Calif.) under the trade name GeneChip(™). Exemplary arrays are shown on the website at affymetrix.com. Another microarray provider is Illumina, Inc. of San Diego, Calif., and exemplary arrays are shown on the website at illumina.com.

[0094] In some embodiments, the method of the invention provides for the preparation of a sample. Depending on the microarray and experiment to be performed, the sample nucleic acid can be prepared in several ways by methods known to those of skill in the art. In some aspects of the invention, the sample can be amplified by a number of mechanisms before, or simultaneously with, genotyping (analysis of copy number profiles). The most common amplification procedure employed involves PCR. See, for example, PCR Technology: Principles and Applications for DNA Amplification (Ed. H. A. Erlich, Freeman Press, NY, N.Y., 1992); PCR Protocols: A Guide to Methods and Applications (Eds. Innis, et al., Academic Press, San Diego, Calif., 1990); Mattila et al., Nucleic Acids Res. 19, 4967 (1991); Eckert et al., PCR Methods and Applications 1, 17 (1991); PCR (Eds. McPherson et al., IRL Press, Oxford); and U.S. Patent Nos. 4,683,202, 4,683,195, 4,800,159, 4,965,188, and 5,333,675, each of which is hereby incorporated by reference in its entirety for all purposes. In some embodiments, the sample can be amplified on the array (see, for example, U.S. Patent No. 6,300,070, which is hereby incorporated by reference herein).

[0095] Other suitable amplification methods include ligase chain reaction (LCR) (e.g., Wu and Wallace, Genomics 4, 560 (1989), Landegren et al., Science 241, 1077 (1988), and Barringer et al. Gene 89:117 (1990)), transcription amplification (Kwoh et al., Proc. Natl. Acad. Sci. USA 86, 1173 (1989), and WO88 / 10315), self-sustained sequence replication (Guatelli et al., Proc. Nat. Acad. Sci. USA, 87, 1874 (1990), and WO90 / 06995), selective amplification of target polynucleotide sequences (U.S. Patent No. 6,410,276), consensus sequence primed polymerase chain reaction (CP-PCR) (U.S. Patent No. 4,437,975), arbitrarily primed polymerase chain reaction (AP-PCR) (U.S. Patents Nos. 5,413,909, 5,861,245), and nucleic acid-based sequence amplification (NABSA). (See, e.g., U.S. Patents Nos. 5,409,818, 5,554,517, and 6,063,603, each of which is incorporated herein by reference). Other amplification methods that can be used are described in U.S. Patents Nos. 5,242,794, 5,494,810, 4,998,617, and U.S. Patent Application No. 09 / 854,317, each of which is incorporated herein by reference).

[0096] Additional methods of sample preparation and techniques for reducing the complexity of nucleic acid samples are described in Dong et al., Genome Research 11, 1418 (2001), U.S. Patents Nos. 6,361,947, 6,391,592, and U.S. Patent Applications Nos. 09 / 916,135, 09 / 920,491 (U.S. Patent Application Publication No. 20030096235), 09 / 910,292 (U.S. Patent Application Publication No. 20030082543), and 10 / 013,598.

[0097] Methods for performing polynucleotide hybridization assays are well developed in the art. The procedures and conditions of the hybridization assays used in the methods of the present invention vary according to the application and are selected according to known general conjugation methods including those mentioned in Maniatis et al. Molecular Cloning: A Laboratory Manual (2nd Ed. Cold Spring Harbor, N.Y., 1989); Berger and Kimmel Methods in Enzymology, Vol. 152, Guide to Molecular Cloning Techniques (Academic Press, Inc., San Diego, Calif., 1987); Young and Davism, P.N.A.S, 80: 1194 (1983). Methods and apparatus for performing iterative and control hybridization reactions are described in U.S. Patent Nos. 5,871,928, 5,874,219, 6,045,996, and 6,386,749, 6,391,623, which are hereby incorporated by reference in their entirety for all purposes.

[0098] The methods of the present invention may also involve detection of the hybridization signal between ligands after (and / or during) hybridization. Reference is made to U.S. Patent Nos. 5,143,854, 5,578,832; 5,631,734; 5,834,758; 5,936,324; 5,981,956; 6,025,601; 6,141,096; 6,185,030; 6,201,639; 6,218,803; and 6,225,625, U.S. Patent Application No. 10 / 389,194, and PCT Application PCT / US99 / 06097 (published as WO99 / 47964), which are also hereby incorporated by reference in their entirety for all purposes.

[0099] Methods and apparatuses for signal detection and processing of intensity data are disclosed, for example, in U.S. Patent Nos. 5,143,854, 5,547,839, 5,578,832, 5,631,734, 5,800,992, 5,834,758, 5,856,092, 5,902,723, 5,936,324, 5,981,956, 6,025,601, 6,090,555, 6,141,096, 6,185,030, 6,201,639, 6,218,803, and 6,225,625; U.S. Patent Application Nos. 10 / 389,194, 60 / 493,495; and PCT Application PCT / US99 / 06097 (published as WO99 / 47964), which are hereby incorporated by reference in their entirety for all purposes.

[0100] Array analysis Molecular profiling according to the present invention includes methods for identifying the genotype of one or more biomarkers by determining whether an individual has one or more nucleotide variants (or amino acid variants) in one or more genes or gene products. Identifying the genotype of one or more genes according to the methods of the present invention in some embodiments can provide additional basis for selecting a treatment.

[0101] The biomarkers of the present invention can be analyzed by any method useful for determining changes in nucleic acids or the proteins they encode. According to one embodiment, one of ordinary skill in the art can analyze one or more genes for variants including deletion mutants, insertion mutants, frameshift mutants, nonsense mutants, missense mutants, and splicing mutants.

[0102] Nucleic acids used for the analysis of one or more genes can be isolated from cells in a sample according to standard methodologies (Sambrook et al., 1989). The nucleic acid can be, for example, genomic DNA, or fractionated or total cellular RNA, or exosomal or miRNA obtained from the cell surface. When using RNA, it may be desirable to convert the RNA to complementary DNA. In one embodiment, the RNA is total cellular RNA; in another embodiment, the RNA is polyA RNA; in another embodiment, the RNA is exosomal RNA. Usually, the nucleic acid is amplified. Depending on the assay format for analyzing one or more genes, the specific nucleic acid of interest is identified directly in the sample using amplification or is identified using a second known nucleic acid after amplification. The identified product is then detected. In certain applications, detection can be performed by visual means (e.g., ethidium bromide staining of a gel). Alternatively, detection can include indirect identification of the product by chemiluminescence, radioactive scintigraphy of a radiolabel, or a fluorescent label, or even a system using an electrical or thermal impulse signal (Affymax Technology; Bellus, 1994).

[0103] It is known that various types of defects occur in the biomarkers of the present invention. The changes include, but are not limited to, deletions, insertions, point mutations, and duplications. Point mutations may be silent or may result in a stop codon, a frameshift mutation, or an amino acid substitution. Mutations may occur inside and outside the coding regions of one or more genes and can be analyzed according to the methods of the present invention. The target site of the nucleic acid of interest may include regions where the sequence varies. Examples include polymorphisms that exist in different forms such as single nucleotide changes, nucleotide repeats, multiple base deletions (where two or more nucleotides are deleted from the consensus sequence), multiple base insertions (where two or more nucleotides are inserted into the consensus sequence), microsatellite repeats (short nucleotide repeats typically having 5 to 1000 repeat units), dinucleotide repeats, trinucleotide repeats, sequence rearrangements (including translocations and duplications), chimeric sequences (where two sequences from different gene origins are fused together), etc., but are not limited thereto. Among sequence polymorphisms, the most frequent polymorphism in the human genome is a single base change, also called a single nucleotide polymorphism (SNP). SNPs are present in large numbers, stably, and widely throughout the genome.

[0104] Molecular profiling includes methods for identifying the haplotypes of one or more genes. A haplotype is a set of genetic determinants located on a single chromosome and typically includes a specific combination of alleles (all of the alternative sequences of a gene) in a region of the chromosome. That is, a haplotype is the phase sequence information of an individual chromosome. In very many cases, the phased SNPs on the chromosome define the haplotype. The combination of haplotypes on a chromosome can determine the genetic profile of a cell. It is the haplotype that determines the association between specific gene markers and disease mutations. Haplotype identification can be performed by any method known in the art. General methods for scoring SNPs include hybridization microarrays or direct gel sequencing as outlined in Landgren et al., Genome Research, 8:769-776, 1998. For example, only one copy of one or more genes can be isolated from an individual and the nucleotides at each of the variant positions determined. Alternatively, allele-specific PCR or similar methods can be used to amplify only one copy of one or more genes in an individual and determine the SNPs at the variant positions of the present invention. The Clark's method known in the art can also be used for haplotype identification. High-throughput molecular haplotype identification methods are also disclosed in Tost et al., Nucleic Acids Res., 30(19):e96 (2002), which is hereby incorporated by reference.

[0105] Thus, as will be apparent to those of ordinary skill in the fields of genetics and haplotype identification, additional variants that are in linkage disequilibrium with the variants and / or haplotypes of the present invention can be identified by haplotype identification methods known in the art. Additional variants that are in linkage disequilibrium with the variants and / or haplotypes of the present invention may also be useful in the various applications described below.

[0106] For genotyping and haplotyping, both genomic DNA and mRNA / cDNA can be used, and both are generally referred to herein as "genes".

[0107] Many techniques for detecting nucleotide variants are known in the art and can be used for the methods of the present invention. These techniques can be protein-based or nucleic acid-based. In either case, the techniques used must be sensitive enough to accurately detect minor nucleotide or amino acid changes. In very many cases, probes labeled with detectable markers are utilized. Any suitable marker known in the art can be used, including but not limited to radioisotopes, fluorescent compounds, biotin detectable using streptavidin, enzymes (e.g., alkaline phosphatase), enzyme substrates, ligands, and antibodies. See Jablonski et al., Nucleic Acids Res., 14:6115-6128 (1986); Nguyen et al., Biotechniques, 13:116-123 (1992); Rigby et al., J. Mol. Biol., 113:237-251 (1977).

[0108] In nucleic acid-based detection methods, a target DNA sample, i.e., a sample containing genomic DNA, cDNA, mRNA, and / or miRNA corresponding to one or more genes, must be collected from the individual to be tested. Any tissue or cell sample containing genomic DNA, miRNA, mRNA, and / or cDNA (or a portion thereof) corresponding to one or more genes can be used. For this purpose, tissue samples containing cell nuclei and thus genomic DNA can be collected from the individual. Blood samples can also be useful, except that only white blood cells and other lymphocytes have cell nuclei, while red blood cells have no nuclei and contain only mRNA or miRNA. Nevertheless, both miRNA and mRNA can be analyzed for the presence of nucleotide variants in their sequences or can serve as templates for cDNA synthesis, and are thus also useful. Tissue or cell samples can be analyzed directly with little processing. Alternatively, the nucleic acid containing the target sequence can be extracted, purified, and / or amplified before being subjected to various detection procedures discussed below. In addition to tissue or cell samples, cDNA or genomic DNA derived from a cDNA or genomic DNA library constructed using tissue or cell samples collected from the individual to be tested is also useful.

[0109] Sequence analysis To determine the presence or absence of a specific nucleotide variant, sequencing of the target genomic DNA or cDNA, particularly the region encompassing the nucleotide variant locus to be detected. Various sequencing techniques, including the Sanger method and the Gilbert chemical method, are generally known and widely used in the art. Pyrosequencing monitors DNA synthesis in real time using a luminescence detection system. Pyrosequencing has been shown to be effective in the analysis of genetic polymorphisms such as single nucleotide polymorphisms and can also be used in the present invention. See Nordstrom et al., Biotechnol. Appl. Biochem., 31(2):107-112 (2000); Ahmadian et al., Anal. Biochem., 280:103-110 (2000).

[0110] Nucleic acid variants can be detected by appropriate detection procedures. Non-limiting examples of methods for detection, quantification, sequencing, etc. include mass detection of mass-modified unit replication sequences (e.g., matrix-assisted laser desorption ionization (MALDI) mass spectrometry and electrospray (ES) mass spectrometry), primer extension methods (e.g., iPLEX™; Sequenom, Inc.), microsequencing methods (e.g., improved methods of primer extension methodology), ligase sequencing methods (e.g., U.S. Pat. Nos. 5,679,524 and 5,952,174, and WO 01 / 27326), mismatch sequencing methods (e.g., U.S. Pat. Nos. 5,851,770; 5,958,692; 6,110,684; and 6,183,958), direct DNA sequencing, restriction fragment length polymorphism (RFLP analysis), allele-specific oligonucleotide (ASO) analysis, methylation-specific PCR (MSPCR), pyrosequencing analysis, acycloprime analysis, reverse dot blot, GeneChip microarray, dynamic allele-specific hybridization (DASH), peptide nucleic acid (PNA) probes and locked nucleic acid (LNA) probes, TaqMan, molecular beacons, intercalating dyes, FRET primers, AlphaScreen, SNPstream, genebit analysis (GBA), multiplex minisequencing, SNaPshot, GOOD assay, microarray minisequencing, array-type primer extension (APEX), microarray primer extension (e.g., microarray sequencing method), tag array, coded microspheres, template-dependent incorporation (TDI), fluorescence polarization, colorimetric oligonucleotide ligation assay (OLA), sequence-coded OLA, microarray ligation, ligase chain reaction, padlock probes, Invader assay, hybridization methods (e.g., hybridization using at least one probe, hybridization using at least one fluorescently labeled probe, etc.), conventional dot blot analysis, single-strand conformational polymorphism analysis (SSCP, e.g., U.S. Pat. Nos. 5,891,625 and 6,013,499; Orita et al., Proc. Natl. Acad. Sci. U.S.A.(Proc. Natl. Acad. Sci. USA 86: 27776-2770 (1989)), denaturing gradient gel electrophoresis (DGGE), heteroduplex analysis, mismatch cleavage detection, and the techniques, cloning and sequencing, electrophoresis, use of hybridization probes and quantitative real-time polymerase chain reaction (QRT-PCR), digital PCR, nanopore sequencing, chips, and combinations thereof described in Sheffield et al., Proc. Natl. Acad. Sci. USA 49: 699-706 (1991), White et al., Genomics 12: 301-306 (1992), Grompe et al., Proc. Natl. Acad. Sci. USA 86: 5855-5892 (1989), and Grompe, Nature Genetics 5: 111-117 (1993). Detection and quantification of alleles or paralogs can be performed using the "closed tube" method described in U.S. Patent Application No. 11 / 950,395, filed December 4, 2007. In some embodiments, the amount of nucleic acid species is measured by mass spectrometry, primer extension, sequencing (e.g., any suitable method, e.g., nanopore or pyrosequencing), quantitative PCR (Q-PCR or QRT-PCR), digital PCR, combinations thereof, etc.

[0111] As used herein, the term "sequence analysis" refers to determining a nucleotide sequence, e.g., the nucleotide sequence of an amplification product. The entire sequence or a partial sequence of a polynucleotide, e.g., DNA or mRNA, can be determined, and the determined nucleotide sequence can be referred to as a "read" or "sequence read". For example, in some embodiments, a linear amplification product can be directly analyzed without further amplification (e.g., by using a single molecule sequencing methodology). In certain embodiments, a linear amplification product can be subjected to further amplification and then analyzed (e.g., by using a sequencing-by-ligation or pyrosequencing methodology). Reads can be subjected to various types of sequence analysis. Any suitable sequencing method can be utilized to detect and measure the amount of a nucleotide sequence species, an amplified nucleic acid species, or a detectable product made from the foregoing. Examples of certain sequencing methods are described below.

[0112] An array analysis device or array analysis component includes a device that can be used by a person skilled in the art to determine a nucleotide sequence (e.g., a linear amplification product and / or an exponential amplification product) resulting from the processes described herein, and one or more components used with such a device. Examples of sequencing platforms include, but are not limited to, the 454 platform (Roche) (Margulies, M. et al. 2005 Nature 437, 376-380), the Illumina Genomic Analyzer (or Solexa platform) or SOLID system (Applied Biosystems) or Helicos True Single Molecule DNA sequencing technology (Harris TD et al. 2008 Science, 320, 106-109), the single molecule real-time (SMRT™) technology of Pacific Biosciences, and nanopore sequencing (Soni G V and Meller A. 2007 Clin Chem 53: 1996-2001). Such platforms enable the sequencing of many nucleic acid molecules isolated from a sample in a highly multiplexed, parallel manner (Dear Brief Funct Genomic Proteomic 2003; 1: 397-416). Each of these platforms enables the sequencing of single molecules of cloned or unamplified nucleic acid fragments. Certain platforms involve, for example, sequencing by ligation of dye-labeled probes (including cyclic ligation and cleavage), pyrosequencing, and single molecule sequencing. Nucleotide sequence species, amplified nucleic acid species, and detectable products made therefrom can be analyzed by such sequencing platforms.

[0113] Array determination by ligation is a nucleic acid sequencing method that depends on the sensitivity of DNA ligase to base pair mismatches. DNA ligase joins the ends of DNA that base pair correctly. By combining the ability of DNA ligase to join only the ends of DNA that base pair correctly with a pooled mixture of fluorescently labeled oligonucleotides or primers, sequencing by fluorescence detection becomes possible. By including primers that contain cleavable linkages that can be cleaved after label identification, longer sequence reads can be obtained. Cleavage at the linker position removes the label, regenerates the 5' phosphate on the end of the ligated primer, and prepares the primer for further rounds of ligation. In some embodiments, the primer can be labeled with two or more fluorescent labels, such as at least 1, 2, 3, 4, or 5 fluorescent labels.

[0114] Array determination by ligation generally involves the following steps. A clonal bead population can be prepared in an emulsion microreactor containing a target nucleic acid template sequence, amplification reaction components, beads, and primers. After amplification, the template is denatured and bead enrichment is performed to separate the beads with extended templates from unwanted beads (e.g., beads with non-extended templates). The template on the selected beads is 3'-modified to enable covalent attachment to a slide, and the modified beads can be deposited on a glass slide. The deposition chamber provides the ability to divide the slide into 1, 4, or 8 chambers during the process of loading the beads. For sequence analysis, primers hybridize to adapter sequences. A set of probes labeled with 4-color dyes compete in the ligation to the sequencing primer. The specificity of probe ligation is achieved by examining every 4th and 5th base during a series of ligations. With 5 - 7 rounds of ligation, detection, and cleavage, the color at every 5th position is recorded along with the number of rounds determined by the type of library used. After each round of ligation, a new complementary primer offset by 1 base in the 5' direction is placed for a further series of ligations. The primer reset and ligation rounds (5 - 7 ligation cycles per round) are repeated 5 times continuously to generate a 25 - 35 base pair sequence for a single tag. In the case of mate pair sequencing, this process is repeated for the second tag.

[0115] Pyrosequencing is a nucleic acid sequencing method based on synthesis sequencing that relies on the detection of pyrophosphate released upon nucleotide incorporation. Generally, synthesis sequencing involves synthesizing a DNA strand complementary to the strand whose sequence is sought, one nucleotide at a time. The target nucleic acid is immobilized on a solid support, hybridized with a sequencing primer, and can be incubated with DNA polymerase, ATP sulfurylase, luciferase, apyrase, adenosine 5′ phosphosulfate, and luciferin. Nucleotide solutions are sequentially added and removed. The accurate incorporation of a nucleotide releases pyrophosphate, which reacts with ATP sulfurylase in the presence of adenosine 5′ phosphosulfate to generate ATP, stimulating the luciferin reaction, which generates a chemiluminescent signal by which sequencing becomes possible. The amount of light generated is proportional to the number of bases added. Thereby, the sequence downstream of the sequencing primer can be determined. Exemplary systems for pyrosequencing involve the steps of ligating an adapter nucleic acid to the nucleic acid under investigation and hybridizing the resulting nucleic acid to beads; amplifying the nucleotide sequence in an emulsion; sorting the beads using a picoliter multiwell solid support; and sequencing the amplified nucleotide sequence by the pyrosequencing methodology (e.g., Nakano et al., “Single-molecule PCR using water-in-oil emulsion;” Journal of Biotechnology 102: 117-124 (2003)).

[0116] Certain single-molecule sequencing modalities are based on the principle of synthetic sequencing, and utilize single-pair fluorescence resonance energy transfer (single-pair FRET) as the mechanism by which photons are emitted as a result of successful nucleotide incorporation. The emitted photons are often detected using a sensitized or highly sensitive cooled charge-coupled device, together with an internal total reflection microscope (TIRM). Photons are emitted only if the introduced reaction solution contains the correct nucleotide for incorporation into the growing nucleic acid strand synthesized as a result of the sequencing process. In single-molecule sequencing based on FRET, energy is transferred between two fluorescent dyes, which are in some cases the polymethine cyanine dyes Cy3 and Cy5, via long-range dipole-dipole interactions. The donor is excited at its specific excitation wavelength, and the excited-state energy is transferred non-radiatively to the acceptor dye, which is then excited. The acceptor dye finally returns to the ground state by radiative emission of a photon. The two dyes used in the energy transfer process represent the "single-pair" in single-pair FRET. Cy3 is often used as the donor fluorophore and is often incorporated as the first labeled nucleotide. Cy5 is often used as the acceptor fluorophore and is used as a nucleotide label for successive nucleotide addition after incorporation of the first Cy3-labeled nucleotide. The fluorophores are generally within 10 nanometers of each other to allow efficient energy transfer.

[0117] Examples of systems that can be used based on single molecule sequencing generally involve the steps of hybridizing a primer to a target nucleic acid sequence to form a complex; binding the complex to a solid phase; repeatedly extending the primer with nucleotides tagged with fluorescent molecules; and taking an image of the fluorescence resonance energy transfer signal after each interaction (e.g., U.S. Patent No. 7,169,314; Braslavsky et al., PNAS 100(7): 3960 - 3964 (2003)). Using such systems, the amplification products (linear amplification products or exponential amplification products) produced by the processes described herein can be directly sequenced. In some embodiments, the amplification products can be hybridized to a primer that contains a sequence complementary to an immobilized capture sequence present on a solid support, such as beads or a glass slide. Hybridization of the primer - amplification product complex to the immobilized capture sequence immobilizes the amplification product on the solid support for single - pair FRET - based synthetic sequencing. The primer is often fluorescent so that an initial reference image of the surface of the slide on which the nucleic acid is immobilized can be made. The initial reference image is useful for determining the positions where true nucleotide incorporation is occurring. Fluorescent signals detected at array positions that are not initially identified in the "primer only" reference image are discarded as non - specific fluorescence. After immobilization of the primer - amplification product complex, the bound nucleic acid is often sequenced in parallel by repeated steps of a) polymerase extension in the presence of one fluorescently labeled nucleotide, b) detection of fluorescence using an appropriate microscope, such as TIRM, c) removal of the fluorescent nucleotide, and d) restoration to step a) with a different fluorescently labeled nucleotide.

[0118] In some embodiments, nucleotide sequencing may be by methods and processes of solid-phase single nucleotide sequencing. The solid-phase single nucleotide sequencing method involves contacting a target nucleic acid with a solid support under conditions where a sample nucleic acid of a single molecule hybridizes to a solid support of a single molecule. Such conditions may include providing a solid support molecule and a single molecule of target nucleic acid in a "microreactor". Such conditions may also include providing a mixture in which the target nucleic acid molecule can hybridize to a solid-phase nucleic acid on the solid support. Single nucleotide sequencing methods useful in the embodiments described herein are described in U.S. Provisional Patent Application No. 61 / 021,871, filed on January 17, 2008.

[0119] In certain embodiments, nanopore sequencing detection methods include: (a) contacting a base nucleic acid (e.g., a ligated probe molecule), which is a target nucleic acid for sequencing, with a detector under conditions such that the sequence-specific detector hybridizes specifically to a substantially complementary sequence of the base nucleic acid; (b) detecting a signal from the detector; and (c) determining the sequence of the base nucleic acid according to the detected signal. In certain embodiments, when the base nucleic acid passes through a pore, if the detector impedes the nanopore structure, the detector hybridized to the base nucleic acid dissociates from the base nucleic acid (e.g., dissociates continuously), and the detector dissociated from the base nucleic acid is detected. In some embodiments, the detector dissociated from the base nucleic acid emits a detectable signal, and the detector hybridized to the base nucleic acid emits a different detectable signal or does not emit a detectable signal. In certain embodiments, nucleotides in a nucleic acid (e.g., a ligated probe molecule) are replaced by specific nucleotide sequences ( "nucleotide representatives") corresponding to specific nucleotides, thereby generating an extended nucleic acid (e.g., U.S. Patent No. 6,723,513), and the detector hybridizes to the nucleotide representatives in the extended nucleic acid that functions as the base nucleic acid. In such embodiments, the nucleotide representatives can be arranged in a two-component or higher-order arrangement (e.g., Soni and Meller, Clinical Chemistry 53(11): 1996-2001 (2007)). In some embodiments, the nucleic acid is not extended, does not generate an extended nucleic acid, acts directly as the base nucleic acid (e.g., the ligated probe molecule functions as the non-extended base nucleic acid), and the detector is contacted directly with the base nucleic acid. For example, a first detector can hybridize to a first sub-sequence, and a second detector can hybridize to a second sub-sequence. In this case, the first detector and the second detector each have a detectable label that can be distinguished from each other, and the signals from the first detector and the second detector can be distinguished from each other when they dissociate from the base nucleic acid.In certain embodiments, the detector comprises regions (e.g., two regions) that hybridize to the underlying nucleic acid, which may be from about 3 to about 100 nucleotides in length (e.g., about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 50, 55, 60, 65, 70, 75, 80, 85, 90, or 95 nucleotides in length). The detector may also comprise one or more regions of nucleotides that do not hybridize to the underlying nucleic acid. In some embodiments, the detector is a molecular beacon. The detector often comprises one or more detectable labels selected independently of those described herein. Each detectable label can be detected by any convenient detection process that can detect the signal (e.g., magnetic, electrical, chemical, optical, etc.) produced by each label. For example, a signal from one or more distinguishable quantum dots linked to the detector can be detected using a CD camera.

[0120] In certain array analysis embodiments, reads can be used to construct larger nucleotide sequences, which can be facilitated by identifying overlapping sequences among different reads and using the identified sequences in the reads. Such array analysis methods and software for constructing larger sequences from reads are known to those of skill in the art (e.g., Venter et al., Science 291: 1304-1351 (2001)). In certain array analysis embodiments, specific reads, partial nucleotide sequence constructs, and complete nucleotide sequence constructs can be compared among nucleotide sequences within a sample nucleic acid (i.e., internal comparison), or can be compared to a reference sequence (i.e., reference comparison). Internal comparisons can be performed when sample nucleic acids are prepared from multiple samples or from a single sample source that contains sequence variations. Reference comparisons are optionally performed when the reference nucleotide sequence is known and the purpose is to determine whether the sample nucleic acid contains a nucleotide sequence that is substantially similar or identical to the reference nucleotide sequence, or a nucleotide sequence that is different from the reference nucleotide sequence. Array analysis can be facilitated by use of the apparatus and components for array analysis described above.

[0121] The primer extension polymorphism detection method is also referred to herein as the "microsequencing" method and is typically performed by hybridizing a complementary oligonucleotide to a nucleic acid carrying a polymorphic site. In these methods, the oligonucleotide typically hybridizes adjacent to the polymorphic site. The term "adjacent" when used in connection with the "microsequencing" method refers to the 3' end of the extension oligonucleotide that is present in the nucleic acid when the extension oligonucleotide hybridizes to the nucleic acid, optionally 1 nucleotide from the 5' end of the polymorphic site, often 2 or 3 nucleotides from the 5' end of the polymorphic site, and sometimes 4, 5, 6, 7, 8, 9, or 10 nucleotides. The extension oligonucleotide is then extended by one or more nucleotides, often by 1, 2, or 3 nucleotides, and the number and / or type of nucleotides added to the extension oligonucleotide determines which one or more polymorphic variants are present. Oligonucleotide extension methods are disclosed, for example, in U.S. Patent Nos. 4,656,127; 4,851,331; 5,679,524; 5,834,189; 5,876,934; 5,908,755; 5,912,118; 5,976,802; 5,981,186; 6,004,744; 6,013,431; 6,017,702; 6,046,005; 6,087,095; 6,210,891; and WO 01 / 20039. The extension products can be detected in any manner, such as by fluorescence methods (see, for example, Chen & Kwok, Nucleic Acids Research 25: 347-353 (1997), and Chen et al., Proc. Natl. Acad. Sci. USA 94 / 20: 10756-10761 (1997)), or by mass spectrometry (e.g., MALDI-TOF mass spectrometry) and other methods described herein.Oligonucleotide extension methods using mass spectrometry are described, for example, in U.S. Patent Nos. 5,547,835; 5,605,798; 5,691,141; 5,849,542; 5,869,242; 5,928,906; 6,043,031; 6,194,144; and 6,258,538. Microsequencing detection methods often incorporate an amplification process that proceeds through an extension step. The amplification process typically amplifies a region derived from a nucleic acid sample containing a polymorphic site. Amplification can be carried out using the methods described above or, for example, in polymerase chain reaction (PCR), using a pair of oligonucleotide primers where one oligonucleotide primer is typically complementary to a region 3' to the polymorphism and the other is typically complementary to a region 5' to the polymorphism. The PCR primer pair can be used in methods disclosed, for example, in U.S. Patent Nos. 4,683,195; 4,683,202, 4,965,188; 5,656,493; 5,998,143; 6,140,054; WO 01 / 27327; and WO 01 / 27329. The PCR primer pair can also be used in any commercially available machine that performs PCR, such as any of the GeneAmp™ systems available from Applied Biosystems.

[0122] Other suitable sequencing methods include multiplex polony sequencing (described in Shendure et al., Accurate Multiplex Polony Sequencing of an Evolved Bacterial Genome, Sciencexpress, Aug. 4, 2005, pg 1, available at www.sciencexpress.org / 4 Aug. 2005 / Page1 / 10.1126 / science.1117389, incorporated herein by reference), which uses immobilized microbeads and sequencing in microfabricated picoliter reactors (described in Margulies et al., Genome Sequencing in Microfabricated High-Density Picolitre Reactors, Nature, August 2005, available at www.nature.com / nature (published online 31 Jul. 2005, doi:10.1038 / nature03959), incorporated herein by reference).

[0123] In some embodiments, whole genome sequencing can also be utilized to identify alleles of RNA transcripts. Examples of whole genome sequencing methods include, but are not limited to, the nanopore-based sequencing methods, sequencing by synthesis, and sequencing by ligation described above.

[0124] In situ hybridization In situ hybridization assays are well known and are generally described in Angerer et al., Methods Enzymol. 152:649-660 (1987). In in situ hybridization assays, for example, cells derived from a biopsy specimen are immobilized on a solid support, typically a slide glass. When probing for DNA, the cells are denatured by heat or alkali. The cells are then contacted with a hybridization solution at an appropriate temperature to allow annealing of the labeled specific probe. The probe is preferably labeled with a radioisotope or a fluorescent reporter. Fluorescence in situ hybridization (FISH) uses fluorescent probes that bind only to portions of sequences that exhibit a high degree of sequence similarity.

[0125] FISH is a cytogenetic technique used to detect a specific polynucleotide sequence within a cell and to identify its location. For example, FISH can be used to detect a DNA sequence on a chromosome. FISH can also be used to detect and identify a specific RNA, such as mRNA, within a tissue sample. FISH uses fluorescent probes that bind to specific nucleotide sequences that exhibit a high degree of sequence similarity. A fluorescence microscope can be used to find out whether and where the fluorescent probe is bound. In addition to detecting specific nucleotide sequences, such as translocations, fusions, deletions, duplications, and other chromosomal abnormalities, FISH can also help to clarify the specific gene copy number and / or the spatial-temporal pattern of gene expression within cells and tissues.

[0126] Comparative genomic hybridization (CGH) uses the dynamics of in situ hybridization to compare the copy numbers of different DNA or RNA sequences derived from a sample, or to compare the copy numbers of different DNA or RNA in one sample with the copy numbers of substantially identical sequences in another sample. In many useful applications of CGH, the DNA or RNA is isolated from the target cells or cell population. The comparison may be qualitative or quantitative. Procedures are described that enable determination of the absolute copy number of DNA sequences across the entire genome of a cell or cell population when the absolute copy number of one or several sequences is known or determined. Different sequences are distinguished from one another by the different positions of their binding sites when hybridized to a reference genome, which is usually metaphase chromosomes but in some cases is interphase nuclei. The copy number information is derived from comparison of the intensities of hybridization signals between different positions on the reference genome. The methods, techniques, and applications of CGH are known as described in U.S. Patent No. 6,335,167 and U.S. Patent Application No. 60 / 804,818, and relevant portions thereof are incorporated herein by reference.

[0127] Other sequence analysis methods Nucleic acid variants can also be detected using standard electrophoresis techniques. The detection step may in some cases be preceded by an amplification step, but amplification is not necessary in the embodiments described herein. Examples of methods for detecting and quantifying nucleic acids using electrophoresis techniques can be found in the art. Non-limiting examples include the step of electrophoresing a sample (e.g., a mixed nucleic acid sample isolated from maternal serum, or an amplified nucleic acid species, for example) in an agarose or polyacrylamide gel. The gel may be labeled (e.g., stained) with ethidium bromide staining (see Sambrook and Russell, Molecular Cloning: A Laboratory Manual 3d ed., 2001). The presence of a band of the same size as the standard control is an indicator of the presence of the target nucleic acid sequence, and the amount thereof can then be compared to the control based on the intensity of the band, thus detecting and quantifying the target sequence of interest. In some embodiments, the target nucleic acid species can be detected and quantified using a restriction enzyme that can distinguish between maternal and paternal alleles. In certain embodiments, an oligonucleotide probe specific for the sequence of interest is used to detect the presence of the target sequence of interest. Using the oligonucleotide, the amount of the target nucleic acid molecule can also be indicated based on the intensity of the signal imparted by the probe, compared to a standard control.

[0128] Specific nucleic acids in a mixture or mixed population containing other nucleic acid species can be detected using array-specific probe hybridization. Under sufficiently stringent hybridization conditions, the probe will specifically hybridize only to substantially complementary sequences. The stringency of the hybridization conditions can be relaxed to allow for various amounts of sequence mismatches. Several hybridization formats are known in the art, including, but not limited to, liquid phase, solid phase, or mixed phase hybridization assays. The following papers provide overviews of various hybridization assay formats: Singer et al., Biotechniques 4:230, 1986; Haase et al., Methods in Virology, pp. 189-226, 1984; Wilkinson, In situ Hybridization, Wilkinson ed., IRL Press, Oxford University Press, Oxford; and Hames and Higgins eds., Nucleic Acid Hybridization: A Practical Approach, IRL Press, 1987.

[0129] Hybridization complexes can be detected by techniques known in the art. Nucleic acid probes that can specifically hybridize to a target nucleic acid (e.g., mRNA or DNA) can be labeled by any suitable method, and the labeled probe can be used to detect the presence of hybridized nucleic acid. One commonly used detection method is 3 H, 125 I, 35 S, 14 C, 32 P, or 33Autoradiography using a probe labeled with P or the like. The choice of radioisotope depends on the ease of synthesis, stability, and research preferences due to the half-life of the isotope selected. Other labels include fluorophores, chemiluminescent agents, and compounds (e.g., biotin and digoxigenin) that bind to anti-ligands or antibodies labeled with an enzyme. In some embodiments, the probe can be directly conjugated with a label such as a fluorophore, chemiluminescent agent, or enzyme. The choice of label depends on the required sensitivity, ease of conjugation with the probe, stability requirements, and available equipment.

[0130] Alternatively, restriction fragment length polymorphism (RFLP) and AFLP methods may be used for molecular profiling. When nucleotide variants in the target DNA corresponding to one or more genes result in the loss or creation of a restriction enzyme recognition site, digestion of the target DNA with a specific restriction enzyme will result in an altered restriction fragment length pattern. Thus, the presence of a specific nucleotide variant will be indicated by the detected RFLP or AFLP.

[0131] Another useful approach is the single-stranded conformational polymorphism assay (SSCA), which is based on the mobility change of single-stranded target DNA spanning the nucleotide variant of interest. A single nucleotide change in the target sequence can result in different intramolecular base pairing patterns, and thus different secondary structures of single-stranded DNA, which can be detected by non-denaturing gels. See Orita et al., Proc. Natl. Acad. Sci. USA, 86:2776-2770 (1989). Denaturing gel-based techniques such as constant denaturant gel electrophoresis (CDGE) and denaturing gradient gel electrophoresis (DGGE) detect differences in the migration rate of mutant sequences compared to wild-type sequences in denaturing gels. See Miller et al., Biotechniques, 5:1016-24 (1999); Sheffield et al., Am. J. Hum, Genet., 49:699-706 (1991); Wartell et al., Nucleic Acids Res., 18:2699-2705 (1990); and Sheffield et al., Proc. Natl. Acad. Sci. USA, 86:232-236 (1989). In addition, double-stranded conformational analysis (DSCA) may also be useful in the present invention. See Arguello et al., Nat. Genet., 18:192-194 (1998).

[0132] The presence or absence of nucleotide variants at specific loci in one or more genes of an individual can also be detected using the Amplification Refractory Mutation System (ARMS) technique. See European Patent No. 0,332,435; Newton et al., Nucleic Acids Res., 17:2503-2515 (1989); Fox et al., Br. J. Cancer, 77:1267-1274 (1998); Robertson et al., Eur. Respir. J., 12:477-482 (1998). In the ARMS method, a primer is synthesized that matches the nucleotide sequence immediately 5' upstream of the locus, except that the 3' terminal nucleotide corresponds to the nucleotide at the locus being tested. For example, the 3' terminal nucleotide may be the same as the nucleotide at the mutant locus. The primer may be of any suitable length as long as it hybridizes to the target DNA under stringent conditions only when its 3' terminal nucleotide matches the nucleotide at the locus being tested. Preferably the primer has at least 12 nucleotides, more preferably about 18-50 nucleotides. If the individual being tested has a mutation at the locus and the nucleotide therein matches the 3' terminal nucleotide of the primer, the primer can be further extended when hybridized to the target DNA template and the primer can initiate a PCR amplification reaction with another suitable PCR primer. In contrast, if the nucleotide at the locus is wild-type, primer extension cannot be achieved. Various forms of the ARMS technique developed over the past few years can be used. See, for example, Gibson et al., Clin. Chem. 43:1336-1341 (1997).

[0133] Similar to the ARMS technique is minisequencing or single nucleotide primer extension, which is based on the incorporation of a single nucleotide. An oligonucleotide primer that is complementary to the nucleotide sequence immediately 5' to the locus being tested is hybridized to the target DNA, mRNA, or miRNA in the presence of a labeled dideoxyribonucleotide. The labeled nucleotide is incorporated or ligated to the primer only if the dideoxyribonucleotide is complementary to the nucleotide of the detected variant locus. Thus, based on the detection label attached to the incorporated dideoxyribonucleotide, the identity of the nucleotide at the variant locus can be determined. See Syvanen et al., Genomics, 8:684-692 (1990); Shumaker et al., Hum. Mutat., 7:346-354 (1996); Chen et al., Genome Res., 10:549-547 (2000).

[0134] Another set of techniques useful in the present invention is the so-called "oligonucleotide ligation assay" (OLA), in which discrimination between a wild-type locus and a mutation is based on the ability of two oligonucleotides to anneal in proximity to each other on a target DNA molecule such that the two oligonucleotides are ligated together by DNA ligase. See Landergren et al., Science, 241:1077-1080 (1988); Chen et al, Genome Res., 8:549-556 (1998); Iannone et al., Cytometry, 39:131-140 (2000). Thus, for example, to detect a single nucleotide mutation at a particular locus in one or more genes, two oligonucleotides can be synthesized, one having a sequence immediately 5' upstream from the locus and whose 3' terminal nucleotide is identical to the nucleotide at the variant locus of that particular gene, and the other having a nucleotide sequence that matches the sequence immediately 3' downstream from the locus of that gene. The oligonucleotides can be labeled for the purpose of detection. When hybridized to the target gene under stringent conditions, the two oligonucleotides are subjected to ligation in the presence of an appropriate ligase. Ligation of the two oligonucleotides indicates that the target DNA has a nucleotide variant at the locus being detected.

[0135] The detection of minor genetic changes can also be accomplished by various hybridization-based approaches. Allele-specific oligonucleotides are most useful. See Conner et al., Proc. Natl. Acad. Sci. USA, 80:278-282 (1983); Saiki et al, Proc. Natl. Acad. Sci. USA, 86:6230-6234 (1989). Oligonucleotide probes (allele-specific) that specifically hybridize to an allele having a specific gene variant at a particular locus but do not hybridize to other alleles can be designed by methods known in the art. The probe can have a length of, for example, 10 to about 50 nucleotide bases. The target DNA and the oligonucleotide probe can be brought into contact with each other under sufficiently stringent conditions such that nucleotide variants can be discriminated from the wild-type gene based on the presence or absence of hybridization. The probe can be labeled to provide a detection signal. Alternatively, the allele-specific oligonucleotide probe can be used as a PCR amplification primer in "allele-specific PCR," and the presence or absence of a specific nucleotide variant is indicated by the presence or absence of a PCR product of the predicted length.

[0136] Other useful techniques based on hybridization anneal two single-stranded nucleic acids together even if there are mismatches due to nucleotide substitutions, insertions, or deletions. Subsequently, various techniques can be used to detect the mismatches. For example, the annealed duplex can be subjected to electrophoresis. Mismatched duplexes can be detected based on their electrophoretic mobilities, which differ from those of perfectly matched duplexes. See Cariello, Human Genetics, 42:726 (1988). Alternatively, in an RNase protection assay, an RNA probe spanning the nucleotide variant to be detected and having a detection marker can be prepared. See Giunta et al., Diagn. Mol. Path., 5:265-270 (1996); Finkelstein et al., Genomics, 7:167-172 (1990); Kinszler et al., Science 251:1366-1370 (1991). The RNA probe can be hybridized to the target DNA or mRNA to form a heteroduplex, which is then subjected to ribonuclease RNase A digestion. RNase A digests the RNA probe in the heteroduplex only at the site of the mismatch. The digestion can be determined on a denaturing electrophoresis gel based on size changes. In addition, mismatches can also be detected by chemical cleavage methods known in the art. See, for example, Roberts et al., Nucleic Acids Res., 25:3377-3378 (1997).

[0137] In the mutS assay, a probe can be prepared that is identical to the gene sequence surrounding the locus where the presence or absence of a mutation is to be detected, except that a predetermined nucleotide is used at the variant locus. After annealing the probe to the target DNA to form a duplex, the mutS protein of Escherichia coli is contacted with the duplex. Since the mutS protein binds only to heteroduplex sequences containing nucleotide mismatches, binding of the mutS protein will indicate the presence of a mutation. See Modrich et al., Ann. Rev. Genet., 25:229-253 (1991).

[0138] Based on the above basic techniques that may be useful for detecting mutations or nucleotide variants in the present invention, a variety of improved and variant methods have been developed in the art. For example, "Sunrise probes" or "molecular beacons" use fluorescence resonance energy transfer (FRET) properties to provide high sensitivity. See Wolf et al., Proc. Nat. Acad. Sci. USA, 85:8790-8794 (1988). Typically, a probe spanning the nucleotide locus to be detected is designed to have a hairpin structure, with one end labeled with a quenching fluorophore and the other end labeled with a reporter fluorophore. In its native state, since one fluorophore is close to the other fluorophore, the fluorescence from the reporter fluorophore is quenched by the quenching fluorophore. When the probe hybridizes to the target DNA, the 5' end is separated from the 3' end, thus regenerating the fluorescence signal. See Nazarenko et al., Nucleic Acids Res., 25:2516-2521 (1997); Rychlik et al., Nucleic Acids Res., 17:8543-8551 (1989); Sharkey et al., Bio / Technology 12:506-509 (1994); Tyagi et al., Nat. Biotechnol., 14:303-308 (1996); Tyagi et al., Nat. Biotechnol., 16:49-53 (1998). To suppress the accumulation of primer dimers, the homoduplex-assisted non-dimer system (HANDS) can be used in combination with the molecular beacon method. See Brownie et al., Nucleic Acids Res., 25:3235-3241 (1997).

[0139] The pigment-labeled oligonucleotide ligation assay is a FRET-based method that combines the OLA assay and PCR. See Chen et al., Genome Res. 8:549-556 (1998). TaqMan is another FRET-based method for detecting nucleotide variants. The TaqMan probe may be an oligonucleotide having a nucleotide sequence of a gene spanning the variant locus of interest and designed to differentially hybridize with different alleles. The two ends of the probe are labeled with a quenching fluorophore and a reporter fluorophore, respectively. The TaqMan probe is incorporated into a PCR reaction for amplifying a target gene region containing the locus of interest using Taq polymerase. Since Taq polymerase exhibits 5'-3' exonuclease activity but does not have 3'-5' exonuclease activity, when the TaqMan probe is annealed to the template of the target DNA, the 5' end of the TaqMan probe is degraded by Taq polymerase during the PCR reaction, thus separating the reporter fluorophore from the quenching fluorophore and releasing a fluorescent signal. See Holland et al., Proc. Natl. Acad. Sci. USA, 88:7276-7280 (1991); Kalinina et al., Nucleic Acids Res., 25:1999-2004 (1997); Whitcombe et al., Clin. Chem., 44:918-923 (1998).

[0140] In addition, detection in the present invention can also use a chemiluminescence-based technique. For example, an oligonucleotide probe can be designed to hybridize to either the wild-type or variant locus, but not to both. The probe is labeled with a highly chemiluminescent acridinium ester. Hydrolysis of the acridinium ester impairs chemiluminescence. Hybridization of the probe to the target DNA prevents hydrolysis of the acridinium ester. Therefore, by measuring the change in chemiluminescence, the presence or absence of a specific mutation in the target DNA is determined. See Nelson et al., Nucleic Acids Res., 24:4998-5003 (1996).

[0141] Detection of genetic changes in the gene according to the present invention may be based on the "Base Excision Sequence Scanning" (BESS) technique. The BESS method is a PCR-based mutation scanning method. BESS T-Scan and BESS G-Tracker, which are similar to the T and G ladders of dideoxy sequencing, are produced. By comparing the sequences of normal DNA and mutant DNA, mutations are detected. See, for example, Hawkins et al., Electrophoresis, 20:1171-1176 (1999).

[0142] Mass spectrometry can be used for molecular profiling according to the present invention. See Graber et al., Curr. Opin. Biotechnol., 9:14-18 (1998). For example, in the primer oligonucleotide extension (PROBE™) method, the target nucleic acid is immobilized on a solid support. A primer is annealed to the target immediately 5' upstream of the locus to be analyzed. Primer extension is performed in the presence of a selected mixture of deoxyribonucleotides and dideoxyribonucleotides. Subsequently, the resulting mixture of newly extended primers is analyzed by MALDI-TOF. See, for example, Monforte et al., Nat. Med., 3:360-362 (1997).

[0143] In addition, microchip or microarray technology is also applicable to the detection method of the present invention. In essence, in a microchip, a large number of different oligonucleotide probes are immobilized on an array on a substrate or carrier, such as a silicon chip or a slide glass. The target nucleic acid sequence to be analyzed can be brought into contact with the immobilized oligonucleotide probes on the microchip. See Lipshutz et al., Biotechniques, 19:442-447 (1995); Chee et al., Science, 274:610-614 (1996); Kozal et al., Nat. Med. 2:753-759 (1996); Hacia et al., Nat. Genet., 14:441-447 (1996); Saiki et al., Proc. Natl. Acad. Sci. USA, 86:6230-6234 (1989); Gingeras et al., Genome Res., 8:435-448 (1998). Alternatively, a plurality of target nucleic acid sequences to be examined are immobilized on a substrate, and a series of probes are brought into contact with the immobilized target sequences. See Drmanac et al., Nat. Biotechnol., 16:54-58 (1998). Incorporating one or more of the above techniques for detecting mutations, a number of microchip technologies have been developed. Microchip technology combined with computer analysis tools enables large-scale rapid screening. The adaptation of microchip technology to the present invention will be apparent to those skilled in the art who are aware of this disclosure.See, for example, U.S. Patent No. 5,925,525 to Fodor et al.; Wilgenbus et al., J. Mol. Med., 77:761-786 (1999); Graber et al., Curr. Opin. Biotechnol., 9:14-18 (1998); Hacia et al., Nat. Genet., 14:441-447 (1996); Shoemaker et al., Nat. Genet., 14:450-456 (1996); DeRisi et al., Nat. Genet., 14:457-460 (1996); Chee et al., Nat. Genet., 14:610-614 (1996); Lockhart et al., Nat. Genet., 14:675-680 (1996); Drobyshev et al., Gene, 188:45-52 (1997).

[0144] As is apparent from the above overview of suitable detection techniques, depending on the detection technique used, it may or may not be necessary to amplify the target DNA molecule, i.e., the gene, its cDNA, mRNA, miRNA, or a portion thereof, in order to increase the number of target DNA molecules. For example, most techniques based on PCR use a combination of amplification of a portion of the target and detection of mutations. PCR amplification is well known in the art and is disclosed in U.S. Patent Nos. 4,683,195 and 4,800,159, both of which are incorporated herein by reference. In detection techniques not based on PCR, amplification can be achieved, if necessary, by, for example, in vivo plasmid propagation or by purifying the target DNA from a large amount of tissue or cell sample. Generally, Sambrook et al., Molecular Cloning: A Laboratory Manual, 2 ndSee, e.g., ed., Cold Spring Harbor Laboratory, Cold Spring Harbor, N.Y., 1989. However, even with rare samples, many sensitive techniques have been developed that can detect minor genetic changes, such as single nucleotide substitutions, without the need to amplify the target DNA in the sample. For example, techniques have been developed that amplify a signal relative to the target DNA by using, for example, branched DNA or dendrimers that can hybridize to the target DNA. Branched DNA or dendrimer DNA provides multiple hybridization sites to which a hybridization probe binds and amplifies the detection signal. See Detmer et al., J. Clin. Microbiol., 34:901-907 (1996); Collins et al., Nucleic Acids Res., 25:2979-2984 (1997); Horn et al., Nucleic Acids Res., 25:4835-4841 (1997); Horn et al., Nucleic Acids Res., 25:4842-4849 (1997); Nilsen et al., J. Theor. Biol., 187:273-284 (1997).

[0145] The Invader™ assay is another technique for detecting single nucleotide changes that can be used for molecular profiling according to the present invention. The Invader™ assay uses a novel linear signal amplification technology that improves the long turnaround times required for typical PCR DNA sequencing-based analysis. See Cooksey et al., Antimicrobial Agents and Chemotherapy 44:1296-1301 (2000). This assay is based on the cleavage of a unique secondary structure formed between two overlapping oligonucleotides that hybridize to a target sequence of interest so as to form a “flap”. Each “flap” then generates thousands of signals per hour. Thus, the results of this technique can be easily interpreted and the method does not require exponential amplification of the DNA target. The Invader™ system utilizes two short DNA probes that hybridize to the DNA target. The structure formed by the hybridization event is recognized by a specific cleavase enzyme, which cleaves one of the probes, releasing a short DNA “flap”. Each released “flap” then binds to a fluorescently labeled probe, forming another cleavage structure. When the cleavase enzyme cleaves the labeled probe, the probe emits a detectable fluorescent signal. See, for example, Lyamichev et al., Nat. Biotechnol., 17:292-296 (1999).

[0146] The rolling circle method is another way to avoid exponential amplification. Lizardi et al., Nature Genetics, 19:225-232 (1998) (incorporated herein by reference). For example, Sniper™, a commercial embodiment of this method, is a high-sensitivity, high-throughput SNP scoring system designed for accurate fluorescence detection of specific variants. For each nucleotide variant, two linear allele-specific probes are designed. These two allele-specific probes are identical except for the 3' base that is altered to complement the variant site. In the first step of the assay, after denaturing the target DNA, it is hybridized to a pair of single-stranded, allele-specific, open circular oligonucleotide probes. If the 3' base exactly complements the target DNA, ligation of the probes occurs preferentially. Subsequent detection of the circularized oligonucleotide probe is by rolling circle amplification, and the amplified probe product is detected by fluorescence. See Clark and Pickering, Life Science News 6, 2000, Amersham Pharmacia Biotech (2000).

[0147] Some other techniques for avoiding amplification simultaneously include, for example, surface enhanced resonance Raman scattering (SERRS), fluorescence correlation spectroscopy, and single molecule electrophoresis. In SERRS, a chromophore-nucleic acid conjugate is adsorbed onto colloidal silver and irradiated with laser light at the resonance frequency of the chromophore. See Graham et al., Anal. Chem., 69:4703-4707 (1997). Fluorescence correlation spectroscopy is based on the spatial-temporal correlation between a fluctuating light signal in an electric field and a trapped single molecule. See Eigen et al., Proc. Natl. Acad. Sci. USA, 91:5740-5747 (1994). In single molecule electrophoresis, the electrophoretic velocity of a fluorescently tagged nucleic acid is measured by measuring the time required for the molecule to travel a predetermined distance between two laser beams. See Castro et al., Anal. Chem., 67:3181-3186 (1995).

[0148] In addition, allele-specific oligonucleotides (ASOs) can also be used in in situ hybridization using tissues or cells as samples. Oligonucleotide probes that can differentially hybridize with wild-type gene sequences, or gene sequences having mutations, can be labeled with radioisotopes, fluorescence, or other detectable markers. In situ hybridization techniques are well known in the art, and the adaptation of the present invention for detecting the presence or absence of nucleotide variants in one or more genes of a particular individual will be apparent to those skilled in the art who are informed of the present disclosure.

[0149] In particular, protein-based detection techniques are also useful for molecular profiling when nucleotide variants cause amino acid substitutions, deletions, insertions, or frameshifts that affect the primary, secondary, or tertiary structure of the protein. To detect amino acid changes, protein sequencing techniques can be used. For example, a protein corresponding to a gene or a fragment thereof can be synthesized by recombinant expression using a DNA fragment isolated from an individual to be tested. Preferably, a cDNA fragment of 100 to 150 base pairs or less that includes the polymorphic locus to be determined is used. Thereafter, the amino acid sequence of the peptide can be determined by conventional protein sequencing methods. Alternatively, HPLC-microscopy tandem mass spectrometry techniques can also be used to determine amino acid sequence changes. In this technique, the protein is subjected to proteolytic digestion, and the resulting peptide mixture is separated by reverse-phase chromatography. Then tandem mass spectrometry is performed and the data collected therefrom is analyzed. See Gatlin et al., Anal. Chem., 72:757-763 (2000).

[0150] Other protein-based detection molecular profiling techniques include immunoaffinity assays based on antibodies that selectively immunoreact with the protein encoded by the mutant gene according to the present invention. Methods for producing such antibodies are known in the art. Antibodies can be used to immunoprecipitate a specific protein from a liquid sample or to immunoblot a protein separated, for example, by polyacrylamide gel. Immunocytochemical methods can also be used to detect specific protein polymorphisms in tissues or cells. Other well-known techniques based on antibodies can also be used, including enzyme-linked immunosorbent assay (ELISA), radioimmunoassay (RIA), immunoradiometric assay (IRMA), and immunoenzymometric assay (IEMA), including sandwich assays using monoclonal or polyclonal antibodies. See, for example, U.S. Patent Nos. 4,376,110 and 4,486,530, both of which are incorporated herein by reference.

[0151] Therefore, the presence or absence of one or more gene nucleotide variants or amino acid variants in an individual can be determined using any of the above detection methods.

[0152] Typically, once the presence or absence of one or more gene nucleotide variants or amino acid variants is determined, a physician or genetic counselor or patient or other researcher can be informed of the results. Specifically, the results can be released in a transmissible form that can be communicated or transmitted to other researchers or physicians or genetic counselors or patients. Such forms can be various and can be tangible or intangible. The results regarding the presence or absence of the nucleotide variants of the present invention in the tested individual can be embodied in an explanatory description, diagram, photograph, chart, image, or any other visual form. For example, an image of gel electrophoresis of a PCR product can be used to explain the results. A diagram showing where in the individual's gene the variant occurs is also useful for showing the test results. The description and visual forms can be recorded on a tangible medium, such as paper, a computer-readable medium, such as a floppy disk, compact disk, etc., or an intangible medium, such as the Internet or an intranet in the form of an email or website in electronic media. In addition, the results regarding the presence or absence of nucleotide variants or amino acid variants in the tested individual can be recorded in an audio form and transmitted through any suitable medium, such as an analog or digital cable line, fiber optic cable, etc., by telephone, facsimile, wireless mobile phone, Internet phone, etc.

[0153] Accordingly, information and data regarding test results can be created anywhere in the world and transmitted to different locations. For example, if a genotyping assay is performed overseas, the information and data regarding the test results can be created and released in the above-mentioned transmissible form. The test results in transmissible form can thus be imported into the United States. Accordingly, the present invention also encompasses a method for creating transmissible form information regarding the genotypes of two or more samples suspected of being cancerous derived from an individual. The method includes: (1) determining the genotype of DNA from the sample according to the method of the present invention; and (2) embodying the results of the determining step in a transmissible form. The transmissible form is the product of the creating method.

[0154] Data and Analysis The implementation of the present invention may use conventional biological methods, software, and systems. The computer software product of the present invention typically includes a computer-readable medium having computer-executable instructions for implementing the logical steps of the method of the present invention. Suitable computer-readable media include floppy disks, CD-ROM / DVD / DVD-ROM, hard disk drives, flash memory, ROM / RAM, magnetic tapes, etc. The computer-executable instructions can be written in a suitable computer language or a combination of several languages. Basic computational biology methods are described, for example, in Setubal and Meidanis et al., Introduction to Computational Biology Methods (PWS Publishing Company, Boston, 1997); Salzberg, Searles, Kasif, (Ed.), Computational Methods in Molecular Biology, (Elsevier, Amsterdam, 1998); Rashidi and Buehler, Bioinformatics Basics: Application in Biological Science and Medicine (CRC Press, London, 2000), and Ouelette and Bzevanis Bioinformatics: A Practical Guide for Analysis of Gene and Proteins (Wiley & Sons, Inc., 2.sup.nd ed., 2001). See U.S. Patent No. 6,420,108.

[0155] The present invention can also use various computer program products and software for various purposes such as probe design, data management, analysis, and instrument operation. See U.S. Patent Nos. 5,593,839, 5,795,716, 5,733,729, 5,974,164, 6,066,454, 6,090,555, 6,185,561, 6,188,783, 6,223,127, 6,229,911, and 6,308,170.

[0156] Furthermore, the present invention relates to aspects including methods for providing genetic information on a network such as the Internet, as shown in U.S. Patent Application Nos. 10 / 197,621, 10 / 063,559 (U.S. Publication No. 20020183936), 10 / 065,856, 10 / 065,868, 10 / 328,818, 10 / 328,872, 10 / 423,403, and 60 / 482,389. For example, one or more molecular profiling techniques can be performed in one location, such as a city, state, country, or continent, and the results can be transmitted to another city, state, country, or continent. Then, treatment options can be made in all or part of the second location. The methods of the present invention include the transmission of information between different locations.

[0157] Molecular Profiling for Treatment Selection The method of the present invention provides a candidate treatment option to a subject in need thereof. Using molecular profiling, one or more candidate therapeutic agents can be identified for an individual suffering from a condition in which one or more of the biomarkers disclosed herein are therapeutic targets. For example, the method can identify one or more chemotherapeutic treatments for cancer. In one aspect, the present invention comprises the steps of performing immunohistochemical (IHC) analysis on a sample from a subject to determine an IHC expression profile for at least five proteins; performing microarray analysis on the sample to determine a microarray expression profile for at least ten genes; performing fluorescence in situ hybridization (FISH) analysis on the sample to determine a FISH mutation profile for at least one gene; performing DNA sequencing on the sample to determine a sequencing mutation profile for at least one gene; and comparing the IHC expression profile, the microarray expression profile, the FISH mutation profile, and the sequencing mutation profile to a rule database comprising mapping for treatments for which biological activity against i) diseased cells overexpressing or underexpressed one or more proteins included in the IHC expression profile, ii) diseased cells overexpressing or underexpressed one or more genes included in the microarray expression profile, iii) diseased cells having no mutation or having one or more mutations in one or more genes included in the FISH mutation profile, and / or iv) diseased cells having no mutation or having one or more mutations in one or more genes included in the sequencing mutation profile has been determined; and identifying the treatment if comparison to the rule database indicates that the treatment should have biological activity against the diseased cells and comparison to the rule database does not indicate a contraindication to the treatment for the treatment of the diseased cells. The disease may be cancer. The molecular profiling steps can be performed in any order.In some embodiments, not all of the molecular profiling steps are performed. By way of non-limiting example, as described herein, if the quality of the sample does not meet a threshold, microarray analysis is not performed. In another example, sequencing is performed only if the FISH analysis meets a threshold. Any relevant biomarker can be evaluated using one or more of the molecular profiling described herein or known in the art. The marker only needs to have some direct or indirect relationship to a useful treatment.

[0158] Molecular profiling involves profiling at least one gene (or gene product) for each assay technique performed. Various numbers of genes can be assayed with various techniques. Any marker disclosed herein that is directly or indirectly related to a targeted therapy can be evaluated based on a gene, such as a DNA sequence, and / or a gene product, such as mRNA or protein. Such nucleic acids and / or polypeptides can be profiled as prescribed with respect to presence or absence, level or amount, mutations, sequence, haplotype, rearrangement, copy number, etc. In some embodiments, a single gene and / or one or more corresponding gene products are assayed by two or more molecular profiling techniques. A gene or gene product (also referred to herein as a "marker" or "biomarker"), such as mRNA or protein, is evaluated using applicable techniques (e.g., for evaluating DNA, RNA, protein) including, but not limited to, FISH, microarray, IHC, sequencing, or immunoassays. Thus, any of the markers disclosed herein can be assayed by a single molecular profiling technique or by multiple methods disclosed herein (e.g., profiling a single marker by one or more of IHC, FISH, sequencing, microarray, etc.). In some embodiments, at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or at least about 100 genes or gene products are profiled by at least one technique, multiple techniques, or each of FISH, microarray, IHC, and sequencing.In some embodiments, at least about 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, 11,000, 12,000, 13,000, 14,000, 15,000, 16,000, 17,000, 18,000, 19,000, 20,000, 21,000, 22,000, 23,000, 24,000, 25,000, 26,000, 27,000, 28,000, 29,000, 30,000, 31,000, 32,000, 33,000, 34,000, 35,000, 36,000, 37,000, 38,000, 39,000, 40,000, 41,000, 42,000, 43,000, 44,000, 45,000, 46,000, 47,000, 48,000, 49,000, or at least about 50,000 genes or gene products are profiled by each technique. The number of markers to assay can depend on the technique used. For example, microarrays and next-generation sequencing are useful for high-throughput analysis.

[0159] In some embodiments, a sample from a subject in need thereof is profiled for one or more of the following, using methods including but not limited to IHC expression profiling, microarray expression profiling, FISH mutation profiling, and / or sequencing mutation profiling (such as by PCR, RT-PCR, pyrosequencing): ABCC1, ABCG2, ACE2, ADA, ADH1C, ADH4, AGT, androgen receptor, AR, AREG, ASNS, BCL2, BCRP, BDCA1, BIRC5, B-RAF, BRCA1, BRCA2, CA2, caveolin, CD20, CD25, CD33, CD52, CDA, CDK2, CDW52, CES2, CK 14, CK 17, CK 5 / 6, c-KIT, c-Myc, COX-2, cyclin D1, DCK, DHFR, DNMT1, DNMT3A, DNMT3B, E-cadherin, ECGF1, EGFR, EPHA2, epiregulin, ER, ERBR2, ERCC1, ERCC3, EREG, ESR1, FLT1, folate receptor, FOLR1, FOLR2, FSHB, FSHPRH1, FSHR, FYN, GART, GNRH1, GNRHR1, GSTP1, HCK, HDAC1, Her2 / Neu, HGF, HIF1A, HIG1, HSP90, HSP90AA1, HSPCA, IL13RA1, IL2RA, KDR, KIT, K-RAS, LCK, LTB, lymphotoxin beta receptor, LYN, MGMT, MLH1, MRP1, MS4A1, MSH2, Myc, NFKB1, NFKB2, NFKBIA, ODC1, OGFR, p53, p95, PARP-1, PDGFC, PDGFR, PDGFRA, PDGFRB, PGP, PGR, PI3K, POLA, POLA1, PPARG, PPARGC1, PR, PTEN, PTGS2, RAF1, RARA, RRM1, RRM2, RRM2B, RXRB, RXRG, SPARC, SPARC MC, SPARCPC, SRC, SSTR1, SSTR2, SSTR3, SSTR4, SSTR5, survivin, TK1, TLE3, TNF, TOP1, TOP2A, TOP2B, TOPO1, TOPO2B, topoisomerase II, TS, TXN, TXNRD1, TYMS, VDR, VEGF, VEGFA, VEGFC, VHL, YES1, and ZAP70.

[0160] In some embodiments, additional molecular profiling methods are performed. These can include, without limitation, PCR, RT-PCR, Q-PCR, SAGE, MPSS, immunoassays, and other techniques for evaluating biological systems described herein or known to those of skill in the art. The selection of genes and gene products to be assayed can be updated as new treatments and new drug targets are identified. Once the expression or mutation of a biomarker has been associated with a treatment option, it can be evaluated by molecular profiling. Such molecular profiling is not limited to the techniques disclosed herein, and one of skill in the art will understand that it can include any methodology conventional for evaluating nucleic acids or proteins, sequence information, or both. The methods of the invention can also utilize any improvements to current methods, or new molecular profiling techniques developed in the future. In some embodiments, a gene or gene product is evaluated by a single molecular profiling technique. In other embodiments, genes and / or gene products are evaluated by multiple molecular profiling techniques. By way of non-limiting example, gene sequences can be analyzed by one or more of FISH and pyrosequencing analysis, mRNA gene products can be assayed by one or more of RT-PCR and microarray, and protein gene products can be assayed by one or more of IHC and immunoassays. One of skill in the art will understand that any combination of biomarkers and molecular profiling techniques that provide a benefit in the treatment of disease is contemplated by the present invention.

[0161] Genes and gene products that are known to play a role in cancer and can be assayed by any of the molecular profiling techniques of the present invention include, but are not limited to, 2AR, disintegrin, activator of thyroid and retinoic acid receptor (ACTR), ADAM 11, adipogenesis inhibitory factor (ADIF), α6 integrin subunit, αV integrin subunit, α-catenin, amplification in breast cancer 1 (AIB1), amplification in breast cancer 3 (AIB3), amplification in breast cancer 4 (AIB4), amyloid precursor protein secretase (APPS), AP-2γ, APPS, ATP-binding cassette transporter (ABCT), placenta-specific (ABCP), ATP-binding cassette subfamily C member (ABCC1), BAG-1, basigin (BSG), BCEI, B cell differentiation factor (BCDF), B cell leukemia 2 (BCL-2), B cell stimulating factor-2 (BSF-2), BCL-1, BCL-2-binding X protein (BAX), BCRP, β1 integrin subunit, β3 integrin subunit, β5 integrin subunit, β2 interferon, β-catenin, β-catenin, bone sialoprotein (BSP), breast cancer estrogen-induced sequence (BCEI), breast cancer resistance protein (BCRP), breast cancer type 1 (BRCA1), breast cancer type 2 (BRCA2), breast cancer amplified sequence 2 (BCAS2), cadherin, epithelial cadherin-11, cadherin-binding protein, calcitonin receptor (CTR), calcium placenta protein (CAPL), calcyclin, CALLA, CAM5, CAPL, carcinoembryonic antigen (CEA), catenin, α1, cathepsin B, cathepsin D, cathepsin K, cathepsin L2, cathepsin O, cathepsin O1, cathepsin V, CD10, CD146, CD147, CD24, CD29, CD44, CD51, CD54, CD61, CD66e, CD82, CD87, CD9, CEA, cellular retinol-binding protein 1 (CRBP1), c-ERBB-2, CK7, CK8, CK18, CK19, CK20, claudin-7, c-MET, collagenase, fibroblast, collagenase, stroma, collagenase 3, common acute lymphoblastic leukemia antigen (CALLA), connexin 26 (Cx26), connexin 43 (Cx43), cortactin, COX-2, CTLA-8, CTR, CTSD, cyclin D1,Cyclooxygenase-2, Cytokeratin 18, Cytokeratin 19, Cytokeratin 8, Cytotoxic T lymphocyte-associated serine esterase 8 (CTLA-8), Differentiation inhibitory activity (DIA), DNA amplified in breast cancer 1 (DAM1), DNA topoisomerase IIα, DR-NM23, E-cadherin, EMMPRIN, EMS1, Endothelial cell growth factor (ECGR), Platelet-derived (PD-ECGF), Enkephalinase, Epidermal growth factor receptor (EGFR), Epicitin, Epithelial membrane antigen (EMA), ER-α, ERBB2, ERBB4, ER-β, ERF-1, Erythrocyte enhancing activity (EPA), ESR1, Estrogen receptor-α, Estrogen receptor-β, ETS-1, Extracellular matrix metalloproteinase inducer (EMMPRIN), Fibronectin receptor, β polypeptide (FNRB), Fibronectin receptor β subunit (FNRB), FLK-1, GA15.3, GA733.2, Galectin-3, γ-catenin, Gap junction protein (26 kDa), Gap junction protein (43 kDa), Gap junction protein α-1 (GJA1), Gap junction protein β-2 (GJB2), GCP1, Gelatinase A, Gelatinase B, Gelatinase (72 kDa), Gelatinase (92 kDa), Gliostatin, Glucocorticoid receptor interacting protein 1 (GRIP1), Glutathione S-transferase p, GM-CSF, Granulocyte chemotactic protein 1 (GCP1), Granulocyte macrophage colony-stimulating factor, Growth factor receptor-bound-7 (GRB-7), GSTp, HAP, Heat shock cognate protein 70 (HSC70), Heat-stable antigen, Hepatocyte growth factor (HGF), Hepatocyte growth factor receptor (HGFR), Hepatocyte stimulating factor III (HSF III), HER-2, HER2 / NEU, HERMES antigen, HET, HHM, Humoral hypercalcemia of malignancy (HHM), ICERE-1, INT-1, Intercellular adhesion molecule-1 (ICAM-1), Interferon γ-inducing factor (IGIF), Interleukin-1α (IL-1A), Interleukin-1β (IL-1B), Interleukin-11 (IL-11), Interleukin-17 (IL-17), Interleukin-18 (IL-18), Interleukin-6 (IL-6), Interleukin-8 (IL-8)Estrogen receptor expression and inverse correlation-1 (ICERE-1), KAI1, KDR, keratin 8, keratin 18, keratin 19, KISS-1, leukemia inhibitory factor (LIF), LIF, lost in inflammatory breast cancer (LIBC), LOT ("lost in transformation"), lymphocyte homing receptor, macrophage colony-stimulating factor, MAGE-3, mammaglobin, maspin, MC56, M-CSF, MDC, MDNCF, MDR, melanoma cell adhesion molecule (MCAM), membrane metalloendopeptidase (MME), membrane-bound neutral endopeptidase (NEP), cysteine-rich protein (MDC), metastasin (MTS-1), MLN64, MMP1, MMP2, MMP3, MMP7, MMP9, MMP11, MMP13, MMP14, MMP15, MMP16, MMP17, moesin, monocyte arginine-serpin, monocyte-derived neutrophil chemotactic factor, monocyte-derived plasminogen activator inhibitor, MTS-1, MUC-1, MUC18, mucin-like carcinoma-associated antigen (MCA), mucin, MUC-1, multidrug resistance protein 1 (MDR, MDR1), multidrug resistance-associated protein-1 (MRP, MRP-1), N-cadherin, NEP, NEU, neutral endopeptidase, neutrophil activation peptide 1 (NAP1), NM23-H1, NM23-H2, NME1, NME2, nuclear receptor coactivator-1 (NCoA-1), nuclear receptor coactivator-2 (NCoA-2), nuclear receptor coactivator-3 (NCoA-3), nucleoside diphosphate kinase A (NDPKA), nucleoside diphosphate kinase B (NDPKB), oncostatin M (OSM), ornithine decarboxylase (ODC), osteoclast differentiation factor (ODF), osteoclast differentiation factor receptor (ODFR), osteonectin (OSN, ON), osteopontin (OPN), oxytocin receptor (OXTR), p27 / kip1, p300 / CBP co-integrator associated protein (p / CIP), p53, p9Ka, PAI-1, PAI-2, parathyroid adenomatosis 1 (PRAD1), parathyroid hormone-like hormone (PTHLH), parathyroid hormone-related peptide (PTHrP), P-cadherin, PD-ECGF, PDGF, peanut-reactive urinary mucin (PUM), P-glycoprotein (P-GP), PGP-1, PHGS-2, PHS-2, PIP, placoglobin,Plasminogen activator inhibitor (type 1), plasminogen activator inhibitor (type 2), plasminogen activator (tissue-type), plasminogen activator (urokinase-type), platelet glycoprotein IIIa (GP3A), PLAU, pleomorphic adenoma gene-like 1 (PLAGL1), polymorphic epithelial mucin (PEM), PRAD1, progesterone receptor (PgR), progesterone resistance, prostaglandin endoperoxide synthase-2, prostaglandin G / H synthase-2, prostaglandin H synthase-2, pS2, PS6K, sorcin, PTHLH, PTHrP, RAD51, RAD52, RAD54, RAP46, receptor-binding coactivator 3 (RAC3), estrogen receptor activity suppressor (REA), S100A4, S100A6, S100A7, S6K, SART-1, scaffold attachment factor B (SAF-B), scatter factor (SF), secreted phosphoprotein-1 (SPP-1), secreted protein, acidic and cysteine-rich (SPARC), stanniocalcin, steroid receptor coactivator-1 (SRC-1), steroid receptor coactivator-2 (SRC-2), steroid receptor coactivator-3 (SRC-3), steroid receptor RNA activator (SRA), stromelysin-1, stromelysin-3, tenascin-C (TN-C), testicular-specific protease 50, thrombospondin I, thrombospondin II, thymidine phosphorylase (TP), thyroid hormone receptor activator molecule 1 (TRAM-1), tight junction protein 1 (TJP1), TIMP1, TIMP2, TIMP3, TIMP4, tissue-type plasminogen activator, TN-C, TP53, tPA, transcriptional intermediary factor 2 (TRANSCRIPTIONAL INTERMEDIARY FACTOR 2, TIF2), trefoil factor 1 (TFF1), TSG101, TSP-1, TSP1, TSP-2, TSP2, TSP50, tumor cell collagenase stimulatory factor (TCSF), tumor-associated epithelial mucin, uPA, uPAR, urokinase, urokinase-type plasminogen activator, urokinase-type plasminogen activator receptor (uPAR), uvomorulin, vascular endothelial growth factor, vascular endothelial growth factor receptor-2 (VEGFR2)It includes vascular endothelial growth factor-A, vascular permeability factor, VEGFR2, very late antigen T cell beta (VLA-β), vimentin, vitronectin receptor alpha polypeptide (VNRA), vitronectin receptor, von Willebrand factor, VPF, VWF, WNT-1, ZAC, ZO-1, and occludin-1.

[0162] Gene products used for IHC expression profiling include, without limitation, one or more of SPARC, PGP, Her2 / neu, ER, PR, c-kit, AR, CD52, PDGFR, TOP2A, TS, ERCC1, RRM1, BCRP, TOPO1, PTEN, MGMT, and MRP1. IHC profiling of EGFR can also be performed. IHC is also used to detect or test various gene products including, without limitation, one or more of the following: EGFR, SPARC, C-kit, ER, PR, androgen receptor, PGP, RRM1, TOPO1, BRCP1, MRP1, MGMT, PDGFR, DCK, ERCC1, thymidylate synthase, Her2 / neu, or TOPO2A. In some embodiments, IHC is used to detect one or more of the following proteins including, without limitation, the following: ADA, AR, ASNA, BCL2, BRCA2, CD33, CDW52, CES2, DNMT1, EGFR, ERBB2, ERCC3, ESR1, FOLR2, GART, GSTP1, HDAC1, HIF1A, HSPCA, IL2RA, KIT, MLH1, MS4A1, MASH2, NFKB2, NFKBIA, OGFR, PDGFC, PDGFRA, PDGFRB, PGR, POLA, PTEN, PTGS2, RAF1, RARA, RXRB, SPARC, SSTR1, TK1, TNF, TOP1, TOP2A, TOP2B, TXNRD1, TYMS, VDR, VEGF, VHL, or ZAP70.

[0163] Microarray expression profiling can be used to simultaneously measure the expression of one or more genes or gene products, including but not limited to: ABCC1, ABCG2, ADA, AR, ASNS, BCL2, BIRC5, BRCA1, BRCA2, CD33, CD52, CDA, CES2, DCK, DHFR, DNMT1, DNMT3A, DNMT3B, ECGF1, EGFR, EPHA2, ERBB2, ERCC1, ERCC3, ESR1, FLT1, FOLR2, FYN, GART, GNRH1, GSTP1, HCK, HDAC1, HIF1A, HSP90AA1, IL2RA, HSP90AA1, KDR, KIT, LCK, LYN, MGMT, MLH1, MS4A1, MSH2, NFKB1, NFKB2, OGFR, PDGFC, PDGFRA, PDGFRB, PGR, POLA1, PTEN, PTGS2, RAF1, RARA, RRM1, RRM2, RRM2B, RXRB, RXRG, SPARC, SRC, SSTR1, SSTR2, SSTR3, SSTR4, SSTR5, TK1, TNF, TOP1, TOP2A, TOP2B, TXNRD1, TYMS, VDR, VEGFA, VHL, YES1, and ZAP70. In some embodiments, the genes used in microarray expression profiling include one or more of EGFR, SPARC, C-kit, ER, PR, androgen receptor, PGP, RRM1, TOPO1, BRCP1, MRP1, MGMT, PDGFR, DCK, ERCC1, thymidylate synthase, Her2 / neu, TOPO2A, ADA, AR, ASNA, BCL2, BRCA2, CD33, CDW52, CES2, DNMT1, EGFR, ERBB2, ERCC3, ESR1, FOLR2, GART, GSTP1, HDAC1, HIF1A, HSPCA, IL2RA, KIT, MLH1, MS4A1, MASH2, NFKB2, NFKBIA, OGFR, PDGFC, PDGFRA, PDGFRB, PGR, POLA, PTEN, PTGS2, RAF1, RARA, RXRB, SPARC, SSTR1, TK1, TNF, TOP1, TOP2A, TOP2B, TXNRD1, TYMS, VDR, VEGF, VHL, or ZAP70.Microarray expression profiling can be performed using a low-density microarray, an expression microarray, a comparative genomic hybridization (CGH) microarray, a single nucleotide polymorphism (SNP) microarray, a proteomics array, an antibody array, or other arrays disclosed herein or known to those of skill in the art. In some embodiments, a high-throughput expression array is used. Such systems include, without limitation, systems commercially available from Agilent or Illumina, which are described in more detail herein.

[0164] One or more of EGFR and HER2 can be profiled using FISH mutation profiling. In some embodiments, FISH is used to detect or assay one or more of the following genes, including but not limited to: EGFR, SPARC, C-kit, ER, PR, androgen receptor, PGP, RRM1, TOPO1, BRCP1, MRP1, MGMT, PDGFR, DCK, ERCC1, thymidylate synthase, HER2, or TOPO2A. In some embodiments, FISH is used to detect or assay various biomarkers including, without limitation, one or more of the following: ADA, AR, ASNA, BCL2, BRCA2, CD33, CDW52, CES2, DNMT1, EGFR, ERBB2, ERCC3, ESR1, FOLR2, GART, GSTP1, HDAC1, HIF1A, HSPCA, IL2RA, KIT, MLH1, MS4A1, MASH2, NFKB2, NFKBIA, OGFR, PDGFC, PDGFRA, PDGFRB, PGR, POLA, PTEN, PTGS2, RAF1, RARA, RXRB, SPARC, SSTR1, TK1, TNF, TOP1, TOP2A, TOP2B, TXNRD1, TYMS, VDR, VEGF, VHL, or ZAP70.

[0165] In some embodiments, the genes used for sequencing mutation profiling include one or more of KRAS, BRAF, c-KIT, and EGFR. The sequencing analysis may also include assessing mutations in one or more of ABCC1, ABCG2, ADA, AR, ASNS, BCL2, BIRC5, BRCA1, BRCA2, CD33, CD52, CDA, CES2, DCK, DHFR, DNMT1, DNMT3A, DNMT3B, ECGF1, EGFR, EPHA2, ERBB2, ERCC1, ERCC3, ESR1, FLT1, FOLR2, FYN, GART, GNRH1, GSTP1, HCK, HDAC1, HIF1A, HSP90AA1, IL2RA, HSP90AA1, KDR, KIT, LCK, LYN, MGMT, MLH1, MS4A1, MSH2, NFKB1, NFKB2, OGFR, PDGFC, PDGFRA, PDGFRB, PGR, POLA1, PTEN, PTGS2, RAF1, RARA, RRM1, RRM2, RRM2B, RXRB, RXRG, SPARC, SRC, SSTR1, SSTR2, SSTR3, SSTR4, SSTR5, TK1, TNF, TOP1, TOP2A, TOP2B, TXNRD1, TYMS, VDR, VEGFA, VHL, YES1, and ZAP70.

[0166] In related aspects, the present invention provides a method for identifying a candidate treatment for a subject in need thereof by using molecular profiling of a set of known biomarkers. For example, the method can identify chemotherapeutic agents for an individual having cancer. The method includes the steps of collecting a sample from the subject; performing immunohistochemical (IHC) analysis on the sample to determine an IHC expression profile for at least 5 of SPARC, PGP, Her2 / neu, ER, PR, c-kit, AR, CD52, PDGFR, TOP2A, TS, ERCC1, RRM1, BCRP, TOPO1, PTEN, MGMT, and MRP1; performing microarray analysis on the sample to determine a microarray expression profile for at least 5 of ABCC1, ABCG2, ADA, AR, ASNS, BCL2, BIRC5, BRCA1, BRCA2, CD33, CD52, CDA, CES2, DCK, DHFR, DNMT1, DNMT3A, DNMT3B, ECGF1, EGFR, EPHA2, ERBB2, ERCC1, ERCC3, ESR1, FLT1, FOLR2, FYN, GART, GNRH1, GSTP1, HCK, HDAC1, HIF1A, HSP90AA1, IL2RA, HSP90AA1, KDR, KIT, LCK, LYN, MGMT, MLH1, MS4A1, MSH2, NFKB1, NFKB2, OGFR, PDGFC, PDGFRA, PDGFRB, PGR, POLA1, PTEN, PTGS2, RAF1, RARA, RRM1, RRM2, RRM2B, RXRB, RXRG, SPARC, SRC, SSTR1, SSTR2, SSTR3, SSTR4, SSTR5, TK1, TNF, TOP1, TOP2A, TOP2B, TXNRD1, TYMS, VDR, VEGFA, VHL, YES1, and ZAP70; performing fluorescence in situ hybridization (FISH) analysis on the sample to determine a FISH mutation profile for at least 1 of EGFR and HER2; performing DNA sequencing on the sample to determine a sequencing mutation profile for at least 1 of KRAS, BRAF, c-KIT, and EGFR;and i) diseased cells that overexpress or underexpress one or more proteins included in the IHC expression profile, ii) diseased cells that overexpress or underexpress one or more genes included in the microarray expression profile, iii) diseased cells that do not have a mutation in one or more genes included in the FISH mutation profile or have one or more mutations, and / or iv) diseased cells that do not have a mutation in one or more genes included in the sequencing mutation profile or have one or more mutations, comparing the IHC expression profile, the microarray expression profile, the FISH mutation profile, and the sequencing mutation profile to a rule database that includes mapping for treatments for which biological activity has been determined; and identifying the treatment if comparison to the rule database indicates that the treatment should have biological activity against the disease and comparison to the rule database does not indicate a treatment contraindication for treatment of the disease. The disease may be cancer. The molecular profiling steps can be performed in any order. In some embodiments, not all of the molecular profiling steps are performed. By way of non-limiting example, as described herein, microarray analysis is not performed if the quality of the sample does not meet a threshold. In some embodiments, IHC expression profiling is performed for at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% of the gene products described above. In some embodiments, microarray expression profiling is performed for at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% of the genes described above.;

[0167] In related aspects, the present invention provides a method for identifying a candidate treatment for a subject in need thereof by using molecular profiling of a predefined set of known biomarkers. For example, the method can identify chemotherapeutic agents for an individual having cancer. The method includes the steps of obtaining a sample from the subject, wherein the sample comprises formalin-fixed paraffin-embedded (FFPE) tissue or fresh frozen tissue and wherein the sample comprises cancer cells; performing immunohistochemistry (IHC) analysis on the sample to determine an IHC expression profile for at least SPARC, PGP, Her2 / neu, ER, PR, c-kit, AR, CD52, PDGFR, TOP2A, TS, ERCC1, RRM1, BCRP, TOPO1, PTEN, MGMT, and MRP1; performing microarray analysis on the sample to determine a microarray expression profile for at least ABCC1, ABCG2, ADA, AR, ASNS, BCL2, BIRC5, BRCA1, BRCA2, CD33, CD52, CDA, CES2, DCK, DHFR, DNMT1, DNMT3A, DNMT3B, ECGF1, EGFR, EPHA2, ERBB2, ERCC1, ERCC3, ESR1, FLT1, FOLR2, FYN, GART, GNRH1, GSTP1, HCK, HDAC1, HIF1A, HSP90AA1, IL2RA, HSP90AA1, KDR, KIT, LCK, LYN, MGMT, MLH1, MS4A1, MSH2, NFKB1, NFKB2, OGFR, PDGFC, PDGFRA, PDGFRB, PGR, POLA1, PTEN, PTGS2, RAF1, RARA, RRM1, RRM2, RRM2B, RXRB, RXRG, SPARC, SRC, SSTR1, SSTR2, SSTR3, SSTR4, SSTR5, TK1, TNF, TOP1, TOP2A, TOP2B, TXNRD1, TYMS, VDR, VEGFA, VHL, YES1, and ZAP70; performing fluorescence in situ hybridization (FISH) analysis on the sample to determine a FISH mutation profile for at least EGFR and HER2; and performing DNA sequencing on the sample to determine a sequencing mutation profile for at least KRAS, BRAF, c-KIT, and EGFR.Compare the IHC expression profile, microarray expression profile, FISH mutation profile, and sequencing mutation profile against a rules database, where the rules database includes mappings for treatments for which the biological activity against i) diseased cells that overexpress or underexpress one or more proteins included in the IHC expression profile; ii) diseased cells that overexpress or underexpress one or more genes included in the microarray expression profile; iii) diseased cells that do not have a mutation or have one or more mutations in one or more genes included in the FISH mutation profile; and / or iv) diseased cells that do not have a mutation or have one or more mutations in one or more genes included in the sequencing mutation profile has been determined; and identify a treatment if: comparison against the rules database indicates that the treatment should have biological activity against the disease; and comparison against the rules database does not indicate a contraindication to the treatment for the disease. The disease may be cancer. The molecular profiling steps can be performed in any order. In some embodiments, not all of the molecular profiling steps are performed. By way of non-limiting example, as described herein, microarray analysis is not performed if the quality of the sample does not meet a threshold. In some embodiments, the biological material is mRNA, and the quality control tests include the A260 / A280 ratio and / or the Ct value of RT-PCR using a housekeeping gene, such as RPL13a. In some embodiments, if the A260 / A280 ratio is lower than 1.5 or the RPL13a Ct value is higher than 30, the mRNA fails the quality control test. In this case, microarray analysis may not be performed. Alternatively, the microarray results may be attenuated, for example, given a lower priority compared to the results of other molecular profiling techniques.

[0168] In some embodiments, molecular profiling is always performed on specific genes or gene products, and profiling of other genes or gene products is optional. For example, IHC expression profiling can be performed with respect to at least SPARC, TOP2A, and / or PTEN. Similarly, microarray expression profiling can be performed with respect to at least CD52. In other embodiments, additional genes are used in addition to the above genes to identify a treatment. For example, the group of genes used for IHC expression profiling can further include DCK, EGFR, BRCA1, CK 14, CK 17, CK 5 / 6, E-cadherin, p95, PARP-1, SPARC, and TLE3. In some embodiments, the group of genes used for IHC expression profiling further includes Cox-2 and / or Ki-67. In some embodiments, HSPCA is assayed by microarray analysis. In some embodiments, FISH mutations are performed with respect to c-Myc and TOP2A. In some embodiments, sequencing is performed with respect to PI3K.

[0169] The method of the present invention can be carried out in any setting where differential expression or mutation analysis is associated with the efficacy of various treatments. In some embodiments, the method is used to identify a candidate treatment for a subject having cancer. Under these conditions, the sample used for molecular profiling preferably contains cancer cells. The percentage of cancer in the sample can be determined by methods known to those skilled in the art, such as using pathological techniques. Cancer cells can also be enriched from the sample, for example, using microdissection techniques. The sample may be required to have a specific threshold of cancer cells before it is used for molecular profiling. The threshold may be at least about 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 95% cancer cells. The threshold may depend on the analytical method. For example, techniques that reveal expression in individual cells may require a lower threshold than techniques that use samples extracted from a mixture of different cells. In some embodiments, the diseased sample is compared to a normal sample taken from the same patient, such as adjacent non-cancerous tissue.

[0170] Treatment selection Using the systems and methods of the present invention, any treatment whose predicted effectiveness can be associated with molecular profiling results can be selected. The present invention includes the use of molecular profiling results to suggest relevance to treatment response. In one aspect, appropriate biomarkers for molecular profiling are selected based on the tumor type of the subject. Using the biomarkers thus suggested, an initial set list of biomarkers can be changed. In other aspects, molecular profiling is independent of the source material. In some aspects, rules are used to provide chemotherapeutic treatments proposed based on molecular profiling test results. In one aspect, the rules are created from a summary of clinical oncology literature with peer review. Expert opinion rules can also be used, but are optional. In one aspect, a hierarchy derived from the evidence scoring system used by the US Preventive Services Task Force is used to evaluate clinical citations with respect to their relevance to the methods of the present invention. "Best evidence" can be used as the basis for the rules. The simplest rules are constructed in the form "if the biomarker is positive, there is one treatment option; otherwise, there are two treatment options". Treatment options include treatment without a particular drug, treatment with a particular drug, or treatment with a combination of drugs. In some aspects, more complex rules are constructed that include the interaction of two or more biomarkers. In such cases, the more complex interactions are typically supported by clinical studies that analyze the interactions between the biomarkers included in the rules. Finally, a report can be created that describes the relevance of chemotherapeutic response to biomarkers and a summary of the best evidence indicating the treatment selected. Ultimately, the treating physician determines the best course of treatment.

[0171] As a non-limiting example, molecular profiling can reveal that the EGFR gene is amplified or overexpressed, and by doing so, indicate a choice of treatment that can block EGFR activity, such as the monoclonal antibody inhibitors cetuximab and panitumumab, or small molecule kinase inhibitors such as gefitinib, erlotinib, and lapatinib, which are effective in patients with activating mutations in EGFR. Other anti-EGFR monoclonal antibodies in clinical development include zalutumumab, nimotuzumab, and matuzumab. The candidate treatment selected may depend on the setting revealed by molecular profiling. For example, when it is found that EGFR has an activating mutation, a kinase inhibitor is often prescribed. Continuing with the exemplary scenario, molecular profiling may also reveal that some or all of these treatments are likely to be of low efficacy. For example, patients taking gefitinib or erlotinib may ultimately develop drug-resistant mutations in EGFR. Thus, the presence of drug-resistant mutations represents a contraindication to the choice of small molecule kinase inhibitors. Extending this example, one of ordinary skill in the art will understand that molecular profiling can lead to the selection of other candidate treatments that act on genes or gene products whose differential expression is revealed by molecular profiling. Similarly, when molecular profiling reveals a particular nucleic acid variant, a candidate agent that has been found to be effective against diseased cells carrying such a variant can be selected.

[0172] Cancer therapies that can be identified as candidate therapies by the method of the present invention include, but are not limited to, 13-cis-retinoic acid, 2-CdA, 2-chlorodeoxyadenosine, 5-azacitidine, 5-fluorouracil, 5-FU, 6-mercaptopurine, 6-MP, 6-TG, 6-thioguanine, Abraxane, Accutane®, actinomycin-D, Adriamycin®, Adrucil®, Afinitor®, Agrylin®, Ala-Cort®, aldesleukin, alemtuzumab, ALIMTA,alitretinoin, Alkaban-AQ®, Alkeran®, all-trans retinoic acid, α-interferon, altretamine, amethopterin, amifostine, aminoglutethimide, anagrelide, Anandron®, anastrozole, arabinosylcytosine, Ara-C, Aranesp®, Aredia®, Arimidex®, Aromasin®, Arranon®, arsenic trioxide, asparaginase, ATRA, Avastin®, azacitidine, BCG, BCNU, bendamustine, bevacizumab, bexarotene, BEXXAR®, bicalutamide, BiCNU, Blenoxane®, bleomycin, bortezomib, busulfan, Busulfex®, C225, leucovorin calcium, Campath®, Camptosar®, camptothecin-11, capecitabine, Carac™, carboplatin, carmustine, carmustine wafer, Casodex®, CC-5013, CCI-779, CCNU, CDDP, CeeNU, Cerubidine®, cetuximab, chlorambucil, cisplatin, citrovorum factor, cladribine, cortisone, Cosmegen®, CPT-11, cyclophosphamide, Cytadren®, cytarabine, cytarabine liposome, Cytosar-U®, Cytoxan®, dacarbazine, Dacogen, dactinomycin, darbepoetin alfa, dasatinib, daunomycin daunorubicin, daunorubicin hydrochloride,Daunorubicin liposome, DaunoXome®, Decadron, Decitabine, Delta-Cortef®, Deltasone®, Denileukin, Diftitox, DepoCyt™, Dexamethasone, Dexamethasone acetate, Dexamethasone sodium phosphate, Dexasone, Dexrazoxane, DHAD, DIC, Diodex, Docetaxel, Doxil®, Doxorubicin, Doxorubicin liposome, Droxia™, DTIC, DTIC-Dome®, Duralone®, Efudex®, Eligard™, Ellence™, Eloxatin™, Elspar®, Emcyt®, Epirubicin, Epoetin alfa, Erbitux, Erlotinib, Erwinia L-asparaginase, Estramustine, Ethyol, Etopophos®, Etoposide, Etoposide phosphate, Eulexin®, Everolimus, Evista®, Exemestane, Fareston®, Faslodex®, Femara®, Filgrastim, Floxuridine, Fludara®, Fludarabine, Fluoroplex®, Fluorouracil, Fluorouracil (cream), Fluoxymesterone, Flutamide, Folic acid, FUDR®, Fulvestrant, G-CSF, Gefitinib, Gemcitabine, Gemtuzumab ozogamicin, Gemzar, Gleevec™, Gliadel®, Wafer, GM-CSF, Goserelin, Granulocyte colony-stimulating factor, Granulocyte-macrophage colony-stimulating factor, Halotestin®, Herceptin®, Hexadrol, Hexalen®, Hexamethylmelamine, HMM, Hycamtin®, Hydrea®, Hydrocort Acetate®, Hydrocortisone, Hydrocortisone phosphate sodium, Hydrocortisone succinate sodium, Hydrocortone Phosphate, Hydroxyurea, Ibritumomab, Ibritumomab tiuxetan, Idamycin®,Idarubicin, Ifex (registered trademark), IFN-α, Ifosfamide, IL-11, IL-2, Imatinib Mesylate, Imidazole Carboxamide, Interferon α, Interferon α-2b (PEG conjugate), Interleukin-2, Interleukin-11, Intron A (registered trademark) (Interferon α-2b), Iressa (registered trademark), Irinotecan, Isotretinoin, Ixabepilone, Ixempra (trademark), Kidrolase (t), Lanacort (registered trademark), Lapatinib, L-Asparaginase, LCR, Lenalidomide, Letrozole, Leucovorin, Leukeran, Leukine (trademark), Leuprolide, Leukocristine, Leustatin (trademark), Liposomal Ara-C Liquid Pred (registered trademark), Lomustine, L-PAM, L-Sarcolysin, Lupron (registered trademark), Lupron Depot (registered trademark), Matulane (registered trademark), Maxidex, Mechlorethamine, Mechlorethamine Hydrochloride, Medralone (registered trademark), Medrol (registered trademark), Megace (registered trademark), Megestrol, Megestrol Acetate, Melphalan, Mercaptopurine, Mesna, Mesnex (trademark), Methotrexate, Methotrexate Sodium, Methylprednisolone, Meticorten (registered trademark), Mitomycin, Mitomycin-C, Mitoxantrone, M-Prednisol (registered trademark), MTC, MTX, Mustargen (registered trademark), Mustine, Mutamycin (registered trademark), Myleran (registered trademark), Mylocel (trademark), Mylotarg (registered trademark), Navelbine (registered trademark), Nelarabine, Neosar (registered trademark), Neulasta (trademark), Neumega (registered trademark), Neupogen (registered trademark), Nexavar (registered trademark), Nilandron (registered trademark), Nilutamide, Nipent (registered trademark), Nitrogen Mustard, Novaldex (registered trademark), Novantrone (registered trademark), Octreotide, Octreotide Acetate, Oncospar (registered trademark), Oncovin (registered trademark), Ontak (registered trademark), Onxal (trademark), Oprelvekin, Orapred (registered trademark), Orasone (registered trademark)Oxaliplatin, Paclitaxel, Protein-Bound Paclitaxel, Pamidronate, Panitumumab, Panretin®, Paraplatin®, Pediapred®, PEG Interferon, Pegaspargase, Pegfilgrastim, PEG-INTRON™, PEG-L-Asparaginase, Pemetrexed, Pentostatin, Phenylalanine Mustard, Platinol®, Platinol-AQ®, Prednisolone, Prednisone, Prelone®, Procarbazine, PROCRIT®, Proleukin®, Prolifeprospan 20 with Carmustine Implant, Purinethol®, Raloxifene, Revlimid®, Rheumatrex®, Rituxan®, Rituximab, Roferon-A® (Interferon α-2a), Rubex®, Rubidomycin Hydrochloride, Sandostatin®, Sandostatin LAR®, Sargramostim, Solu-Cortef®, Solu-Medrol®, Sorafenib, SPRYCEL™, STI-571, Streptozocin, SU11248, Sunitinib, Sutent®, Tamoxifen, Tarceva®, Targretin®, Taxol®, Taxotere®, Temodar®, Temozolomide, Temsirolimus, Teniposide, TESPA, Thalidomide, Thalomid®, TheraCys®, Thioguanine, Thioguanine Tabloid®, Thiophosphoramide, Thioplex®, Thiotepa, TICE®, Toposar®, Topotecan, Toremifene, Torisel®, Tositumomab, Trastuzumab, Treanda®, Tretinoin, Trexall™, Trisenox®, TSPA, TYKERB®, VCR, Vectibix™, Velban®, Velcade®, VePesid®Vesanoid (registered trademark), Viadur (trademark), Vidaza (registered trademark), vinblastine, vinblastine sulfate, Vincasar Pfs (registered trademark), vincristine, vinorelbine, vinorelbine tartrate, VLB, VM-26, vorinostat, VP-16, Vumon (registered trademark), Xeloda (registered trademark), Zanosar (registered trademark), Zevalin (trademark), Zinecard (registered trademark), Zoladex (registered trademark), zoledronic acid, Zolinza, Zometa (registered trademark), and any combination thereof are included.

[0173] In some embodiments, a database is created that maps treatment and molecular profiling results. Treatment information can include the predicted efficacy of a therapeutic agent against cells having certain characteristics that can be measured by molecular profiling. Molecular profiling can include differential expression or mutations in specific genes, proteins, or other biomolecules of interest. Through mapping, the results of molecular profiling can be compared against the database to select a treatment. The database can include both positive and negative mappings between treatment and molecular profiling results. In some embodiments, the mapping is created by reviewing the literature regarding the relationship between biological agents and therapeutic agents. For example, scientific papers, patent publications or published patent applications, scientific publications, etc. can be reviewed for potential mapping. The mapping can include in vivo results, such as animal or clinical trials, or in vitro experiments, such as cell culture results. For example, any mapping found, such as the cytotoxic effect of a therapeutic agent against cells expressing a certain gene or protein, can be incorporated into the database. In this manner, the database can be continuously updated. It will be understood that the methods of the present invention are similarly updated.

[0174] The rules for mapping can include various supplementary information. In some aspects, the database includes prioritization criteria. For example, a treatment with higher predictive efficacy in a given setting is considered more preferable than a treatment predicted to have lower efficacy. A mapping derived from a particular setting, such as a clinical trial, may be prioritized over a mapping derived from another setting, such as a cell culture experiment. A treatment with stronger literature support may be prioritized over a treatment supported by more preliminary results. A treatment generally applicable to the type of disease in question, such as cancer of a particular tissue origin, may be prioritized over a treatment not adapted to that particular disease. The mapping can include both positive and negative correlations between treatments and molecular profiling results. In a non-limiting example, one mapping may propose the use of a kinase inhibitor such as erlotinib for tumors having activating mutations in EGFR, whereas another mapping may propose the contrary if EGFR also has drug-resistant mutations. Similarly, a treatment may be shown to be effective in cells overexpressing a particular gene or protein, but may be shown to be ineffective if the gene or protein is underexpressed.

[0175] The selection of a candidate treatment for an individual can be based on molecular profiling results by any one or more of the methods described. Alternatively, the selection of a candidate treatment for an individual can be based on molecular profiling results by two or more of the methods described. For example, the selection of a treatment for an individual can be based on molecular profiling results by FISH only, IHC only, or microarray analysis only. In other embodiments, the selection of a treatment for an individual can be based on molecular profiling results by IHC, FISH, and microarray analysis; IHC and FISH; IHC and microarray analysis, or FISH and microarray analysis. The selection of a treatment for an individual can also be based on molecular profiling results by other methods of sequencing or mutation detection. The molecular profiling results may include mutation analysis in combination with one or more methods such as IHC, immunoassays, and / or microarray analysis. Various combinations and sequential results can be used. For example, treatments can be prioritized according to the results obtained by molecular profiling. In one embodiment, the prioritization is based on the following algorithm: 1) If IHC / FISH and microarray indicate the same target, this is given first priority; 2) A single IHC positive result is given the next priority; or 3) A single microarray positive result is given the final priority. Sequencing can also be used to guide the selection. In some embodiments, sequencing reveals drug resistance mutations such that the resulting drug is not selected even if techniques including IHC, microarray, and / or FISH indicate differential expression of the target molecule. Any such contraindications, e.g., differential expression or mutation of another gene or gene product, may override the selection of treatment.

[0176] An exemplary list of microarray expression results and predicted treatments is shown in Table 1. Molecular profiling is performed to determine whether a particular gene or gene product in a sample is differentially expressed compared to a control. The control may be any control suitable for that setting, including, without limitation, the expression level of a control gene such as a housekeeping gene, the expression of the same gene in healthy tissue from the same or other individuals, statistical criteria, level of detection, etc. One of ordinary skill in the art will understand that the expression status can be determined using any applicable molecular profiling technique; for example, the results of microarray analysis, PCR, Q-PCR, RT-PCR, immunoassays, SAGE, IHC, FISH, or sequencing. The expression status of a gene or gene product is used to select drugs predicted to be effective or ineffective. For example, Table 1 shows that overexpression of the ADA gene or protein presents pentostatin as a possible treatment. On the other hand, underexpression of the ADA gene or protein implies resistance to cytarabine, suggesting that cytarabine is not the optimal treatment.

[0177] (Table 1) Molecular Profiling Results and Predicted Treatments TIFF2025106337000001.tif226170TIFF2025106337000002.tif228170TIFF2025106337000003.tif233170TIFF2025106337000004.tif67170

[0178] Table 2 shows a more comprehensive summary of the rules regarding treatment selection. For each biomarker in the table, the type of assay and the assay results are shown. The summary of the effectiveness of various therapeutic agents given the assay results can be derived from the medical literature or other medical knowledge bases. Using the results, one can lead to the selection of specific therapeutic agents that are recommended or not recommended. In some embodiments, this table is continuously updated as new literature reports and treatments become available. In this way, the molecular profiling of the present invention will evolve and improve over time. The rules in Table 2 can be stored in a database. For example, if molecular profiling results such as differential expression or mutation of a certain gene or gene product are obtained, the results can be compared against the database to lead to treatment selection. The set of rules in the database can be updated as new treatments and new treatment data become available. In some embodiments, the rules database is continuously updated. In some embodiments, the rules database is periodically updated. The rules database can be updated at least once a day, two days, three days, four days, five days, six days, one week, ten days, two weeks, three weeks, four weeks, one month, six weeks, two months, three months, four months, five months, six months, seven months, eight months, nine months, ten months, eleven months, twelve months, one year, eighteen months, two years, or at least every three years. The molecular profiling results can be compared with the rules database using any relevant correlation or comparison approach. In one embodiment, molecular profiling identifies that a certain gene or gene product is differentially expressed. Query the rules database and select the item regarding the gene or gene product. Extract the treatment selection information selected from the rules database and use this to select a treatment. For example, the information for recommending or not recommending a specific treatment may depend on whether the gene or gene product is overexpressed or underexpressed. In some cases, multiple rules and treatments can be derived depending on the results of the molecular profiling. In some embodiments, the treatment options are ranked in a list and presented to the end user. In some embodiments, the treatment options are presented without ranking information.In any case, an individual, such as a treating physician or a similar caregiver, can select from among the available treatment options.

[0179] (Table 2) Summary of rules regarding treatment selection TIFF2025106337000005.tif239170TIFF2025106337000006.tif241164TIFF2025106337000007.tif241168TIFF2025106337000008.tif241168TIFF2025106337000009.tif241168TIFF2025106337000010.tif238170TIFF2025106337000011.tif241168TIFF2025106337000012.tif238170TIFF2025106337000013.tif241168TIFF2025106337000014.tif239170TIFF2025106337000015.tif238170TIFF2025106337000016.tif239170TIFF2025106337000017.tif239170TIFF2025106337000018.tif240167TIFF2025106337000019.tif240167TIFF2025106337000020.tif239170TIFF2025106337000021.tif241163TIFF2025106337000022.tif241167TIFF2025106337000023.tif241164TIFF2025106337000024.tif238170TIFF2025106337000025.tif238170TIFF2025106337000026.tif238170TIFF2025106337000027.tif241164TIFF2025106337000028.tif241167TIFF2025106337000029.tif239170TIFF2025106337000030.tif239170TIFF2025106337000031.tif239167TIFF2025106337000032.tif239163TIFF2025106337000033.tif241168TIFF2025106337000034.tif238170TIFF2025106337000035.tif241168TIFF2025106337000036.tif241168TIFF2025106337000037.tif241164TIFF2025106337000038.tif241168TIFF2025106337000039.tif241168TIFF2025106337000040.tif239170TIFF2025106337000041.tif241164TIFF2025106337000042.tif239170TIFF2025106337000043.tif241168TIFF2025106337000044.tif241164TIFF2025106337000045.tif239170TIFF2025106337000046.tif238170TIFF2025106337000047.tif241165TIFF2025106337000048.tif241152.

[0180] By using the methods described herein to provide individualized treatment options, the survival period of a subject can be extended. In some embodiments, the subject has previously been treated with one or more therapeutic agents for treating a disease such as cancer. The cancer may be refractory to one of these agents, for example by acquiring a drug-resistant mutation. In some embodiments, the cancer is metastatic. In some embodiments, the subject has not previously been treated with one or more candidate therapeutic agents identified by the method. Candidate treatments can be selected using molecular profiling, regardless of the stage, anatomical location, or anatomical origin of the cancer cells.

[0181] Progression-free survival (PFS) means, for example, the likelihood that an individual or group of individuals suffering from a disease such as cancer remains free of disease progression after the start of a treatment process. This can refer to the percentage of individuals within a group who are likely to maintain a stable state of the disease (e.g., show no signs of progression) after a specific duration. Progression-free survival is an indicator of the effectiveness of a specific treatment. Similarly, disease-free survival (DFS) means the likelihood that an individual or group of individuals suffering from cancer remains free of the disease after the start of a specific treatment. This can refer to the percentage of individuals within a group who are likely to be free of the disease after a specific duration. Disease-free survival is an indicator of the effectiveness of a specific treatment. Based on the PFS or DFS obtained in a similar patient group, treatment strategies can be compared. When described in the context of cancer survival, disease-free survival is often used together with the term overall survival.

[0182] The progression-free survival (PFS) (Period B) using a therapy selected by molecular profiling can be compared with the PFS (Period A) of the most recent therapy when the patient started progression, so that a candidate therapy selected by molecular profiling according to the present invention can be compared with a therapy selected other than by molecular profiling. See Figure 32. In one setting, a PFS(B) / PFS(A) ratio ≥ 1.3 was used to show that a therapy selected by molecular profiling provides a benefit to the patient (Robert Temple, Clinical measurement in drug evaluation. Edited by Wu Ningano and G.T. Thicker John Wiley and Sons Ltd. 1995; Von Hoff, D.D. Clin Can Res. 4: 1079, 1999: Dhani et al. Clin Cancer Res. 15: 118-123, 2009). Other ways to compare a therapy selected by molecular profiling with a therapy selected other than by molecular profiling include determining the response rate (RECIST) and the percentage of patients who have not progressed or died at the 4-month time point. The term "about" when used in connection with a numerical value of PFS means a variation of + / - 10 percent (10%) with respect to the numerical value. The PFS with a therapy selected by molecular profiling can be extended by at least 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or at least 90% compared with a therapy selected other than by molecular profiling. In some embodiments, the PFS with a therapy selected by molecular profiling can be extended by at least 100%, 150%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, or at least about 1000% compared with a therapy selected other than by molecular profiling. In yet other embodiments, the PFS ratio (PFS in a therapy selected by molecular profiling or a novel therapy / PFS in a previous therapy or treatment) is at least about 1.3.In yet other embodiments, the PFS ratio is at least about 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, or 2.0. In yet other embodiments, the PFS ratio is at least about 3, 4, 5, 6, 7, 8, 9, or 10.

[0183] Similarly, DFS in patients selected with or without using molecular profiling can be compared. In some embodiments, DFS with a treatment selected by molecular profiling is extended by at least 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or at least 90% compared to a non-molecular profiling selected treatment. In some embodiments, DFS with a treatment selected by molecular profiling can be extended by at least 100%, 150%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, or at least about 1000% compared to a non-molecular profiling selected treatment. In yet other embodiments, the DFS ratio (DFS in a treatment selected by molecular profiling or a novel treatment / DFS in a previous treatment or therapy) is at least about 1.3. In yet other embodiments, the DFS ratio is at least about 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, or 2.0. In yet other embodiments, the DFS ratio is at least about 3, 4, 5, 6, 7, 8, 9, or 10.

[0184] In some embodiments, although molecular profiling provides valuable patient benefits, the candidate therapies of the present invention do not increase the PFS ratio or DFS ratio in patients. For example, in certain cases, no preferred treatment has been identified for a patient. In such cases, when nothing has been identified currently, molecular profiling provides a way to identify candidate therapies. Molecular profiling can extend PFS, DFS, or lifespan by at least 1 week, 2 weeks, 3 weeks, 4 weeks, 1 month, 5 weeks, 6 weeks, 7 weeks, 8 weeks, 2 months, 9 weeks, 10 weeks, 11 weeks, 12 weeks, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months, 9 months, 10 months, 11 months, 12 months, 13 months, 14 months, 15 months, 16 months, 17 months, 18 months, 19 months, 20 months, 21 months, 22 months, 23 months, 24 months, or 2 years. Molecular profiling can extend PFS, DFS, or lifespan by at least two and a half years, 3 years, 4 years, 5 years, or more. In some embodiments, the methods of the present invention improve the outcome such that the patient is in a remission state.

[0185] The effectiveness of the treatment can also be monitored by other means. Complete remission (CR) includes the complete disappearance of the disease: no disease is evident on examination, scan, or other tests. Partial remission (PR) refers to the presence of some disease in the body, but with a 30% or greater decrease in the size or number of lesions. Stable disease (SD) refers to a disease in which the size and number of lesions remain relatively unchanged. Generally, a decrease of less than 50% in size or a slight increase is described as stable disease. Progressive disease (PD) means that the disease has increased in size or number during treatment. In some embodiments, the molecular profiling according to the present invention results in complete remission or partial remission. In some embodiments, the methods of the present invention result in stable disease. In some embodiments, the present invention can achieve stable disease when non-molecular profiling leads to progressive disease.

[0186] Computer system The conventional data networking, application development, and other functional aspects of this system (and the components of the individual operating components of this system) cannot be described in detail herein, but are part of the present invention. Further, the connecting lines shown in the various drawings included herein are intended to represent exemplary functional relationships and / or physical couplings between various elements. It should be noted that many alternative or additional functional relationships or physical couplings may exist in the actual system.

[0187] The various system components discussed herein may include one or more of the following: a host server or other computer system including a processor for processing digital data; a memory coupled to the processor for storing digital data; an input digitizer coupled to the processor for inputting digital data; an application program stored in the memory and accessible by the processor for instructing the processor to process digital data; a display device coupled to the processor and the memory for displaying information obtained from the digital data processed by the processor; and a plurality of databases. The various databases used herein may include: patient data such as family history, demographic and environmental data, biosample data, data on previous treatments and protocols, clinical data of patients, molecular profiling data of biosamples, data on therapeutic and / or investigational drugs, gene libraries, disease libraries, drug libraries, patient tracking data, file management data, financial management data, billing data, and / or similar data useful in the operation of the system. As will be understood by those skilled in the art, the user computer may include an operating system (e.g., Windows NT, 95 / 98 / 2000, OS2, UNIX, Linux, Solaris, MacOS, etc.), as well as various conventional support software and drivers typically associated with a computer. The computer may include any suitable personal computer, network computer, workstation, minicomputer, mainframe, etc. The user computer may be in a home or medical / business environment having access to a network. In an exemplary aspect, the access is via a network or the Internet through a commercially available web browser software package.

[0188] As used herein, the term "network" includes any telecommunications means incorporating both its hardware components and software components. Communication between interested parties can be achieved by any suitable communication channel, such as, for example, a telephone line network, an extranet, an intranet, the Internet, a point of interaction device, a personal digital assistant (e.g., Palm Pilot®, Blackberry®), a cell phone, a kiosk, etc.), online communication, satellite communication, offline communication, wireless communication, transponder communication, a local area network (LAN), a wide area network (WAN), networked or connected devices, a keyboard, a mouse, and / or any suitable communication or data input format. Further, although the system is frequently described herein as being implemented with the TCP / IP communication protocol, the system can also be implemented using IPX, Appletalk, IP-6, NetBIOS, OSI, or any of several existing or future protocols. When a network has the nature of a public network such as the Internet, it can be beneficial to presume that the network is not secure and is open to eavesdroppers. Specific information regarding protocols, standards, and application software used in connection with the Internet is generally known to those of ordinary skill in the art and, thus, need not be described herein. For example, reference is made to DILIP NAIK, INTERNET STANDARDS AND PROTOCOLS (1998); JAVA 2 COMPLETE, various authors, (Sybex 1999); DEBORAH RAY AND ERIC RAY, MASTERING HTML 4.0 (1997); and LOSHIN, TCP / IP CLEARLY EXPLAINED (1997), and DAVID GOURLEY AND BRIAN TOTTY, HTTP, THE DEFINITIVE GUIDE (2002), the contents of which are hereby incorporated by reference into this specification.

[0189] Various system components can be appropriately connected to a network, either independently, individually, or collectively, via a data link, which is typically connected and used through a standard modem communication, cable modem, dish network, ISDN, digital subscriber line (DSL), or various wireless communication methods, such as, for example, a connection to an Internet service provider (ISP) through a local loop. See, for example, GILBERT HELD, UNDERSTANDING DATA COMMUNICATIONS (1996), which is incorporated herein by reference. Note that the network can also be implemented as other types of networks, such as an interactive television (ITV) network. Further, the present system contemplates the use, sale, or distribution of any article, service, or information on any network having similar functionality as described herein.

[0190] As used herein, "transmission" can include sending electronic data from one system component to another via a network connection. Further, as used herein, "data" can include comprehensive information such as commands, queries, files, stored data, etc. in digital or any other form.

[0191] The present system contemplates uses related to web services, utility computing, pervasive and personal computing, security and identity solutions, autonomous computing, commodity computing, mobile and wireless solutions, open source, biometrics, grid computing, and / or mesh computing.

[0192] Any database discussed herein may include a relational structure, a hierarchical structure, a graphic structure, or an object-oriented structure, and / or any other database configuration. Common database products that may be used to implement a database include DB2 by IBM (White Plains, NY), various database products available from Oracle Corporation (Redwood Shores, CA), Microsoft Access or Microsoft SQL Server by Microsoft Corporation (Redmond, Washington), or any other suitable database product. Further, the database may be organized in any suitable manner, such as, for example, as a data table or a lookup table. Each record may be a single file, a series of files, a linked series of data fields, or any other data structure. The association of specific data may be achieved by any desired data association technique, such as an association technique known or practiced in the art. For example, the association may be achieved either manually or automatically. Automatic association techniques may include, for example, database searching, database merging, GREP, AGREP, SQL, using key fields in a table for fast searching, sequential searching across all tables and files, sorting the records in a file according to a known order to simplify searching, and / or the like. The association step may be achieved, for example, in a preselected database or database sector using, for example, a "key field" by means of a database merge function.

[0193] More specifically, the "key field" divides the database according to the high-level class of objects defined by the key field. For example, a particular type of data can be designated as a key field in multiple related data tables, and then the data tables can be linked based on the type of data in the key field. The data corresponding to the key field in each of the linked data tables is preferably the same or of the same type. However, data tables having similar data even if not identical in the key field can also be linked, for example, by using AGREP. According to one aspect, any suitable data storage technique can be utilized to store data without a standard format. The data set can be stored using any suitable technique, including, for example, storing individual files using the ISO / IEC 7816-4 file structure; implementing a domain in which a dedicated file that publishes one or more basic files containing one or more data sets is selected; using a data set stored in individual files using a hierarchical filing system; storing a data set as a record in a single file (including compression, SQL-accessible, hashed via one or more keys, numbers, alphabets in the first tuple, etc.); binary large object (BLOB); storing as ungrouped data elements encoded using ISO / IEC 7816-6 data elements; storing as ungrouped data elements encoded using ISO / IEC Abstract Syntax Notation (ASN.1) as in ISO / IEC 8824 and 8825; and / or other proprietary techniques that may include fractal compression methods, image compression methods, etc.

[0194] In one exemplary aspect, the ability to store diverse information in different formats is facilitated by storing the information as a BLOB. Accordingly, any binary information can be stored in a storage area associated with the dataset. The BLOB method can store the dataset as ungrouped data elements formatted as two-component blocks via a fixed memory offset, using any of best practices regarding fixed storage allocation, circular queue techniques, or memory management (such as least recently used paged memory, etc.). By using the BLOB method, the ability to store various datasets having different formats facilitates the storage of data by multiple and unrelated owners of the dataset. For example, a first dataset that can be stored can be provided by a first party, a second dataset that can be stored can be provided by an unrelated second party, and a third dataset that can be further stored can be provided by a third party unrelated to the first and second parties. Each of these three exemplary datasets can include different information stored using different data storage formats and / or techniques. Further, each dataset can include a subset of data that can also be different from other subsets.

[0195] As noted above, in various aspects, data can be stored regardless of the common format. However, in one exemplary aspect, a dataset (e.g., a BLOB) can be annotated in a standard manner when provided for manipulating the data. The annotation can include a short header, a postscript, or other suitable indicators associated with each dataset configured to convey information useful for managing the various datasets. For example, the annotation may be referred to herein as a "condition header", "header", "postscript", or "status", and can include an indication of the status of the dataset or an identifier correlated with a particular issuer or owner of the data. The next byte of data can be used, for example, to indicate the identity of the issuer or owner of the data, the user, the transaction / member account identifier, etc. Each of these conditional annotations is further discussed herein.

[0196] Dataset annotations can also be used for other types of status information and various other purposes. For example, dataset annotations can include security information that establishes access levels. The access levels can be configured to permit access to the dataset only by, for example, a particular individual, an employee level, a company, or other entity, or to permit access to a particular dataset based on a transaction, the issuer or owner of the data, a user, etc. Further, the security information can restrict / permit only certain actions such as access to the dataset, modification of the dataset, and / or deletion of the dataset. In one example, dataset annotations indicate that only the owner or user of the dataset is permitted to delete the dataset, that various identified users may be permitted to access the dataset for reading, and that others are completely excluded from access to the dataset. However, other access restriction parameters that permit various entities to access the dataset at various permission levels as needed can also be used. Data including headers or footers can be received by a stand-alone type of interactive device configured to add to, delete, modify, or supplement the data according to the header or footer.

[0197] For security reasons, any database, system, device, server, or other component of this system can consist of any combination of them in a single location or multiple locations, and those skilled in the art will also understand that each database or system can include any of various appropriate security functions such as firewalls, access codes, encryption, decryption, compression, decompression, and / or the like.

[0198] The arithmetic unit of the web client may further include an Internet browser connected to the Internet or an intranet using standard dial-up, cable, DSL, or any other Internet protocol known in the art. To prevent unauthorized access from users of other networks, transactions going out from the web client may pass through a firewall. Further, to enhance security even more, additional firewalls may be deployed between various components of the CMS.

[0199] The firewall may include any hardware and / or software appropriately configured to protect the CMS components and / or the enterprise's computing resources from users of other networks. Further, the firewall may be configured to limit or restrict access from the web client connected via the web server to various systems and components behind the firewall. The firewall may exist in various configurations including, among others, stateful inspection, proxy-based, and packet filtering. The firewall may be integrated within the web server or any other CMS component, or may exist as a further separate entity.

[0200] The computers discussed in this specification can provide a suitable website or other Internet-based graphical user interface accessible by a user. In one aspect, Microsoft Internet Information Server (IIS), Microsoft Transaction Server (MTS), and Microsoft SQL Server are used with the Microsoft operating system, Microsoft NT web server software, Microsoft SQL Server database system, and Microsoft Commerce Server. Additionally, components such as Access or Microsoft SQL Server, Oracle, Sybase, Informix MySQL, Interbase, etc. can be used to provide an Active Data Object (ADO)-compliant database management system.

[0201] Any of the communications, inputs, storage, databases, or displays discussed in this specification may be facilitated via a website having a web page. The term "web page" as used herein is not intended to limit the types of documents and applications that may be used to interact with a user. For example, a typical website may include, in addition to standard HTML documents, various formats, Java applets, JavaScript, Active Server Pages (ASP), Common Gateway Interface scripts (CGI), Extensible Markup Language (XML), Dynamic HTML, Cascading Style Sheets (CSS), helper applications, plug-ins, etc. A server may include a web service that receives requests from a web server, the requests including a URL (http: / / yahoo.com / stockquotes / ge) and an IP address (123.56.789.234). The web server searches for an appropriate web page and transmits the data or application of that web page to the IP address. A web service is an application that can interact with other applications over a communication means such as the Internet. Web services are typically based on standards or protocols such as XML, XSLT, SOAP, WSDL, and UDDI. The web service methodology is well known in the art and is covered in many standard textbooks. See, for example, ALEX NGHIEM, IT WEB SERVICES: A ROADMAP FOR THE ENTERPRISE (2003), which is incorporated herein by reference.

[0202] The web-based clinical database for the system and method of the present invention preferably has the ability to upload and store clinical data files in native format and is searchable for any clinical parameter. This database is also extensible and can enter clinical annotations from any study for easy integration with other studies using an EAV data model (metadata). In addition, the web-based clinical database can be adaptable and may be XML and XSLT that can dynamically add customized questions to the user. Further, this database includes the ability to export to CDISC ODM.

[0203] Implementers will also understand that there are several ways to display data within browser-based documents. The data can be displayed as standard text or within fixed lists, scrollable lists, drop-down lists, editable text areas, fixed text areas, pop-up windows, etc. Similarly, there are several ways available to modify data on a web page, such as free text input using, for example, a keyboard, selection of menu items, check boxes, option boxes, etc.

[0204] The present system and method may be described herein with respect to functional block components, screen shots, option selections, and various processing stages. It should be understood that such functional blocks may be implemented by a number of hardware and / or software components configured to perform a particular function. For example, the present system may use various integrated circuit components such as memory elements, processing elements, logic elements, look-up tables, etc. that may execute various functions under the control of one or more microprocessors or other control devices. Similarly, the software elements of the present system may be executed by any programming or scripting language such as C, C++, Macromedia Cold Fusion, Microsoft Active Server Pages, Java, COBOL, assembler, PERL, Visual Basic, SQL Stored Procedures, Extensible Markup Language (XML), etc., and various algorithms may be executed by any combination of data structures, objects, processes, routines, or other programming elements. Further, it should be noted that the present system may use a number of conventional techniques for data transmission, signal transmission, data processing, network control, etc. Still further, the present system may be used by client-side scripting languages such as JavaScript, VBScript, etc. to detect or prevent security problems.For a basic introduction to encryption and network security, please refer to any of the following references: (1) "Applied Cryptography: Protocols, Algorithms, And Source Code In C", by Bruce Schneier, published by John Wiley & Sons (second edition, 1995); (2) "Java Cryptography", by Jonathan Knudson, published by O'Reilly & Associates (1998); (3) "Cryptography & Network Security: Principles & Practice", by William Stallings, published by Prentice Hall; all of which are hereby incorporated by reference into this specification.

[0205] As used herein, the terms "end user", "consumer", "customer", "client", "treating physician", "hospital", or "business" may be used interchangeably with each other, each meaning any person, entity, machine, hardware, software, or business. Each party interacts with the system and is equipped with a computing device to facilitate online data access and data entry. Customers have computing devices in the form of personal computers, although other types of computing devices may also be used, including laptop, notebook, handheld computers, set-top boxes, mobile phones, push-button phones, etc. The owner / operator of the system and method of the present invention has a computing device executed in the form of a computer server, although other executions by systems including computing centers shown as mainframe computers, minicomputers, PC servers, networks of computers located in the same or different geographical locations, etc. are also contemplated. Further, the system contemplates the use, sale, or distribution of any article, service, or information on any network having similar functionality as described herein.

[0206] In one exemplary embodiment, each client customer may be issued an "account" or "account number." As used herein, an account or account number may be any device, code, number, character, symbol, digital certificate, smart chip, digital signal, analog signal, biometric, or other identifier / marker (e.g., an authorization / access code, personal identification number (PIN), Internet code, other identification code, and / or one or more of the like) appropriately configured to allow a consumer to access, interact with, or communicate with the system. The account number may optionally be located on or associated with a charge card, credit card, debit card, prepaid card, embossed card, smart card, magnetic stripe card, barcode card, transponder, radio frequency card, or related account. The system may include or be associated with any of the foregoing cards or devices, or a fob having a transponder and an RFID reader in RF communication with the fob. The system may include fob embodiments, but the invention is not so limited. In fact, the system may include any device having a transponder configured to communicate with an RFID reader via RF communication. Exemplary devices may include, for example, a keyholder, tag, card, mobile phone, wristwatch, or any such form that may be presented for inquiry. Further, the systems, computing devices, or devices discussed herein may include "pervasive computing devices," which may include conventional non-computerized devices having a computing device incorporated therein. The account number may be distributed and stored in any form of plastic device, electronic device, magnetic device, radio frequency device, wireless device, acoustic device, and / or optical device capable of transmitting or downloading data from itself to a second device.

[0207] As will be understood by those skilled in the art, the present system can be embodied as a customization of an existing system, an add-on product, upgraded software, a stand-alone system, a distributed system, a method, a data processing system, a device for data processing, and / or a computer program product. Accordingly, the present system can take the form of an all-software aspect, an all-hardware aspect, or an aspect that combines both software and hardware aspects. Further, the present system may take the form of a computer program product on a computer-readable storage medium having computer-readable program code means embodied in the storage medium. Any suitable computer-readable storage medium including a hard disk, a CD-ROM, an optical storage device, a magnetic storage device, and / or the like can be utilized.

[0208] The present system and method are described herein with reference to screen shots, block diagrams, and flowchart diagrams of methods, apparatus (e.g., systems), and computer program products in various aspects. It will be understood that each functional block in the block diagrams and flowchart diagrams, as well as combinations of functional blocks in the block diagrams and flowchart diagrams, can be implemented by computer program instructions.

[0209] Reference will now be made to FIGS. 2-25, but the process flows and screen shots shown are merely examples and are not intended to limit the scope of the invention described herein. For example, the steps recited in any of the method or process descriptions can be executed in any order and are not limited to the order of display. It will be understood that the following description appropriately refers not only to the steps and user interface elements shown in FIGS. 2-25, but also to the various system components described above with reference to FIG. 1.

[0210] These computer program instructions may be loaded onto a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed on the computer or other programmable data processing apparatus create means for implementing the functions specified in one or more flowchart blocks. These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instruction means for implementing the functions specified in one or more flowchart blocks. The computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus and to produce a process implemented by the computer, the instructions executed on the computer or other programmable apparatus providing steps for implementing the functions specified in one or more flowchart blocks.

[0211] Accordingly, the functional blocks of the block diagrams and flowchart diagrams support combinations of means for performing a particular function, combinations of steps for performing a particular function, and program instruction means for performing a particular function. It will also be understood that each functional block of the block diagrams and flowchart diagrams, as well as combinations of functional blocks in the block diagrams and flowchart diagrams, can be executed by either a special purpose hardware-based computer system for performing a particular function or step, or a suitable combination of special purpose hardware and computer instructions. Further, the diagrams of the process flow and its description may refer to user windows, web pages, websites, web forms, prompts, etc. Implementers will understand that the illustrated steps described herein may be included in a number of configurations including the use of windows, web pages, web forms, pop-up windows, prompts, etc. It should be further understood that the multiple steps illustrated and described may be integrated into a single web page and / or window, but are extended for simplicity. In other cases, the steps illustrated and described as a single process step may be divided among multiple web pages and / or windows, but are integrated for simplicity.

[0212] Advantages, other advantages, and solutions to problems have been described herein with respect to specific embodiments. However, none of the advantages, advantages, solutions, and any element that may bring about or clarify any of these advantages, advantages, or solutions should be construed as a critical, necessary, or essential feature or element of any or all of the claims of the present invention. Accordingly, the scope of the present invention should not be limited by anything other than the claims, and in the claims, a reference to an element in the singular is not intended to mean "one and only one" unless expressly stated so, but is intended to mean "one or more." All structural, chemical, and functional equivalents to the elements of the above exemplary embodiments known to those skilled in the art are expressly incorporated herein by reference and are intended to be encompassed by the claims of the present invention. Further, the apparatus or method is not required to address every problem the present invention seeks to solve by being encompassed by the claims of the present invention. Further, elements, components, or method steps in the present disclosure are not intended to be dedicated to the public whether or not such elements, components, or method steps are expressly recited in the claims. Elements of the claims herein are not to be construed according to the provisions of 35 U.S.C. § 112, paragraph 6, unless the element is expressly recited using the phrase "means for." As used herein, "comprising," "comprises," or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Further, elements described herein are not necessary for the practice of the present invention unless expressly described as "essential" or "critical."

[0213] Figure 1 shows a block diagram of an exemplary embodiment of a system 10 for determining personalized medical interventions for a particular medical condition using molecular profiling of a patient's biological specimen. System 10 includes a user interface 12, a host server 14 including a processor 16 for processing data, a memory 18 coupled to the processor, an application program 20 stored in the memory 18 and accessible by the processor 16 for instructing the processor 16 to process data, a plurality of internal databases 22, and an external database 24, as well as an interface to a wired or wireless communication network 26 (e.g., the Internet, etc.). System 10 may also include an input digitizer 28 coupled to the processor 16 for inputting digital data from data received from the user interface 12.

[0214] The user interface 12 includes an input device 30 and a display 32 for inputting data into the system 10 and for displaying information obtained from data processed by the processor 16. The user interface 12 may also include a printer 34 for printing information obtained from data processed by the processor 16, such as a patient report including test results regarding a target and a proposed drug therapy based on the test results.

[0215] The internal databases 22 may include, but are not limited to, information and tracking of a patient's biological sample / specimen, clinical data, patient data, patient tracking, file management, research protocols, test results of a patient by molecular profiling, as well as billing information and tracking. The external database 24 may include, but is not limited to, a drug library, a gene library, a disease library, as well as public and private databases such as UniGene, OMIM, GO, TIGR, GenBank, KEGG, and Biocarta.

[0216] Molecular profiling method In accordance with system 10, various methods can be used. FIG. 2 shows a flowchart of an exemplary embodiment of a method 50 for determining personalized medical interventions for a particular medical condition that utilizes molecular profiling of a patient's biological specimen that is not disease specific. In step 52, at least one test is performed on at least one target from a biological sample of an affected patient to determine a medical intervention for a particular medical condition using molecular profiling that is not dependent on disease series diagnosis (i.e., not limited to a single disease). A target is defined as any molecular finding that can be obtained from a molecular test. For example, a target can include one or more genes, one or more gene expression proteins, one or more molecular mechanisms, and / or combinations thereof. For example, the expression level of a target can be measured by analysis of the mRNA level of the target or gene, or the protein level of the gene. Tests for finding such targets can include, but are not limited to, fluorescence in situ hybridization (FISH), in situ hybridization (ISH), and other molecular tests known to those of skill in the art. PCR-based methods such as real-time PCR or quantitative PCR can also be used. Additionally, in the methods disclosed herein, microarray analysis such as comparative genomic hybridization (CGH) microarray, single nucleotide polymorphism (SNP) microarray, proteomics array, or antibody array analysis can also be used. In some embodiments, microarray analysis includes identifying whether a gene is upregulated or downregulated with a significance of p<0.001 compared to a reference. Testing or analysis of a target can also include immunohistochemical (IHC) analysis. In some embodiments, IHC analysis includes determining whether 30% or more of the sample is stained, whether the staining intensity is +2 or greater, or both.

[0217] Furthermore, the methods disclosed herein also include profiling two or more targets. For example, the expression of multiple genes can be identified. Further, the identification of multiple targets in a sample may be by one method or by various means. For example, the expression of a first gene can be measured by one method, and the expression level of a second gene can be measured by a different method. Alternatively, the same method can be used to detect the expression levels of the first gene and the second gene. For example, the first method may be IHC, and the second method may be by microarray analysis such as detecting gene expression of a gene.

[0218] In some embodiments, molecular profiling may also include identifying genetic variants such as target mutations, polymorphisms (such as SNPs), deletions, or insertions. For example, the identification of SNPs in a gene can be determined by microarray analysis, real-time PCR, or sequencing. One or more target variants can also be identified using other methods disclosed herein.

[0219] Accordingly, one or more of the following can be performed: IHC analysis in step 54, microanalysis in step 56, and other molecular tests known to those skilled in the art in step 58.

[0220] A biological sample is obtained from an affected patient by taking a tumor biopsy specimen, by performing minimally invasive surgery, or, if a recent tumor is not available, by taking a sample of the patient's blood, or a cell extract, nuclear extract, cell lysate, or biological product or substance of biological origin, such as excretions, blood, serum, plasma, urine, sputum, tears, feces, saliva, membrane extracts, etc., including but not limited to any other biological fluid sample.

[0221] In step 60, a determination is made as to whether one or more of the targets tested in step 52 exhibit a change in expression as compared to the normal reference for that particular target. In one exemplary method of the present invention, IHC analysis can be performed in step 54, and in step 64, a determination is made as to whether any target by IHC analysis exhibits a change in expression by determining whether 30% or more of the cells of the biological sample for a particular target are stained +2 or higher. Those skilled in the art will understand that the staining results can vary depending on the technician performing the test and the type of target being tested, and that +1 or higher staining may in some cases indicate a change in expression. In another exemplary aspect of the present invention, microarray analysis can be performed in step 56, and in step 66, a determination is made as to whether any target by microarray analysis exhibits a change in expression by determining whether the fold change in expression of a particular target relative to the normal tissue of the origin reference is significant with p < 0.001, and identifying which targets are upregulated or downregulated. A change in expression can also be demonstrated by the absence of one or more genes, gene expression proteins, molecular mechanisms, or other molecular findings.

[0222] After determining in step 60 which targets exhibit a change in expression, in step 70, at least one non-disease-specific agent that interacts with each target having a change in expression is identified. The agent can be any drug or compound having a therapeutic effect. A non-disease-specific agent is a therapeutic drug or compound that was previously not associated with the treatment of the diagnosed disease of the patient and that can interact with the target derived from the biological sample of the patient that exhibits a change in expression. A portion of the non-disease-specific agents that have been found to interact with specific targets found in various cancer patients is shown in Table 3 below.

[0223] (Table 3) TIFF2025106337000049.tif72151

[0224] Finally, at step 80, a patient profile report can be provided that includes the patient's test results regarding various targets and any treatment methods proposed based on those results. An exemplary patient profile report 100 is shown in FIGS. 3A - 3D. The patient profile report 100 shown in FIG. 3A identifies the tested target 102, such tested target 104 that showed a significant change in expression, and the non-disease-specific agent 106 proposed to interact with the target. The patient profile report 100 shown in FIG. 3B identifies the result 108 of immunohistochemical analysis regarding a specific gene expression protein 110, and whether the gene expression protein is a molecular target 112 by determining whether 30% or more of the tumor cells had a staining of +2 or more. The report 100 also identifies the immunohistochemical tests 114 that were not performed. The patient profile report 100 shown in FIG. 3C identifies the gene 116 analyzed by microarray analysis and whether the gene was underexpressed or overexpressed compared to a reference 118. Finally, the patient profile report 100 shown in FIG. 3D identifies the patient's clinical history 120 and the specimen 122 submitted by the patient. The molecular profiling technique can be performed anywhere, for example overseas, and the results can be transmitted via a network to appropriate parties such as, for example, patients, physicians, research institutes, or other remotely located interested parties.

[0225] Figure 4 shows a flowchart of an exemplary embodiment of a method 200 for identifying a drug therapy / drug that can interact with a target. At step 202, a molecular target is identified that shows a change in expression in some affected individuals. Next, at step 204, a drug therapy / drug is administered to the affected individuals. After administration of the drug therapy / drug, any change in the molecular target identified at step 202 is identified at step 206 to determine whether the drug therapy / drug administered at step 204 interacts with the molecular target identified at step 202. If it is determined that the drug therapy / drug administered at step 204 interacts with the molecular target identified at step 202, instead of approving the drug therapy / drug for a particular disease, the drug therapy / drug can be approved for treating patients showing a change in expression of the identified molecular target.

[0226] Figures 5 to 14 are flowcharts and diagrams showing various parts of an information-based personalized medicine discovery system and method according to the present invention. Figure 5 is a diagram showing an exemplary clinical decision support system of the information-based personalized medicine discovery system and method of the present invention. Data obtained through clinical research and clinical care, such as clinical trial data, biomedical / molecular imaging data, genomics / proteomics / chemical library / literature / expert curation, biospecimen tracking / LIMS, family history / environmental records, and clinical data, are collected and stored as databases and data marts in a data warehouse. Figure 6 is a diagram showing the flow of information through a clinical decision support system of the information-based personalized medicine discovery system and method of the present invention using web services. A user interacts with the system by inputting data into the system via form-based input / upload of a dataset, formulating queries and executing data analysis jobs, and obtaining and evaluating the display of output data. The data warehouse in the web-based system is where data is extracted, transformed, and loaded from various database systems. The data warehouse is also where common formatting, mapping, and transformation occur. This web-based system also includes data marts created based on data views of interest.

[0227] A flowchart of an exemplary clinical decision support system of the information-based personalized medicine discovery system and method of the present invention is shown in Figure 7. The clinical information management system includes a research institute information management system, and the medical information contained in the data warehouse and databases includes medical information libraries such as drug libraries, gene libraries, and disease libraries in addition to literature text mining. The information management system for a specific patient and both the medical information database and the data warehouse are integrated at a data linkage center where diagnostic information and treatment options can be obtained. A financial management system can also be incorporated into the clinical decision support system of the information-based personalized medicine discovery system and method of the present invention.

[0228] FIG. 8 is a diagram showing an exemplary biological specimen tracking and management system that can be utilized as part of the information-based personalized medicine discovery system and method of the present invention. FIG. 8 shows two host medical centers that transfer specimens to a tissue / blood bank. The specimens can pass through laboratory analysis prior to shipment. The research can be performed on the samples via microarray, genotyping, and proteomics analysis. This information can be redistributed to the tissue / blood bank. FIG. 9 is a flowchart of an exemplary biological specimen tracking and management system that can be utilized with the information-based personalized medicine discovery system and method of the present invention. The host medical center obtains a sample from a patient and then sends the patient sample to a molecular profiling laboratory that can also perform isolation and analysis of RNA and DNA.

[0229] FIG. 10 shows a diagram of a method for maintaining a clinical standard vocabulary for use with the information-based personalized medicine discovery system and method of the present invention. FIG. 10 shows how the findings and patient information of a physician related to a patient of one physician can be made accessible to another physician so that the other physician can utilize the data in making diagnostic and treatment decisions regarding that patient.

[0230] FIG. 11 shows a schematic diagram of an exemplary microarray gene expression database that can be used as part of the information-based personalized medicine discovery system and method of the present invention. The microarray gene expression database includes both an external database and an internal database accessible via a web-based system. The external database can include, but is not limited to, UniGene, GO, TIGR, GenBank, KEGG. The internal database can include, but is not limited to, tissue tracking, LIMS, clinical data, and patient tracking. FIG. 12 shows a diagram of an exemplary microarray gene expression database data warehouse that can be used as part of the information-based personalized medicine discovery system and method of the present invention. Institute data, clinical data, and patient data can all be housed within the microarray gene expression database data warehouse, and the data can then be accessed by public / private disclosure and utilized by data analysis tools.

[0231] Another schematic diagram showing the flow of information through the information-based personalized medicine discovery system and method of the present invention is shown in FIG. 13. Similar to FIG. 7, this schematic diagram includes clinical information management, medical and literature information management, and financial management of the information-based personalized medicine discovery system and method of the present invention. FIG. 14 is a schematic diagram showing an exemplary network of the information-based personalized medicine discovery system and method of the present invention. In order to provide patients with proposed treatments or drugs based on various identified targets, patients, medical practitioners, host medical centers, and research institutes all share and exchange various information.

[0232] FIGS. 15-25 are printed outputs of computer screens related to various parts of the information-based personalized medicine discovery system and method based on the information shown in FIGS. 5-14. FIGS. 15 and 16 show computer screens for entering physician information and insurance company information on behalf of the client. FIGS. 17-19 show computer screens where information can be entered to order analyses and tests on patient samples.

[0233] Figure 20 is a computer screen showing the results of microarray analysis of specific genes tested using patient samples. This information and the computer screen are similar to the information detailed in the patient profile report shown in Figure 3C. Figure 22 is a computer screen showing the results of immunohistochemical tests of a specific patient for various genes. This information is similar to the information included in the patient profile report shown in Figure 3B.

[0234] Figure 21 is a computer screen showing options for finding specific patients, publishing patient reports regarding test orders and / or results, and tracking current cases / patients.

[0235] Figure 23 is a computer screen outlining some of the steps for creating a patient profile report as shown in Figures 3A through 3D. Figure 24 shows a computer screen for ordering immunohistochemical tests on patient samples, and Figure 25 shows a computer screen for entering information regarding the primary tumor site for microarray analysis. Those skilled in the art will understand that any number and type of computer screens can be utilized for the purpose of entering the information necessary to utilize the information-based personalized medicine discovery system and method of the present invention, and for obtaining the information resulting from utilizing the information-based personalized medicine discovery system and method of the present invention.

[0236] Figures 26 to 31 show a table indicating the frequency of significant changes in the expression of specific genes and gene expression proteins according to tumor type, that is, the number of times a gene and / or gene expression protein was flagged as a target according to tumor type as being significantly overexpressed or underexpressed (see also Examples 1 to 3). These tables show the total number of times a gene and / or gene expression protein was overexpressed or underexpressed in a specific tumor type, and whether the change in expression was determined by immunohistochemical analysis (Figures 26, 28), or by microarray analysis (Figures 27, 30). These tables also identify, using immunohistochemistry, the total number of times overexpression of any gene expression protein occurred in a specific tumor type, and, using microarray analysis, the total number of times overexpression or underexpression of any gene occurred in a specific tumor type.

[0237] Accordingly, the present invention provides a method and system for analyzing diseased tissue using IHC tests and gene microarray tests according to the IHC and microarray tests described above. The patient may be in an advanced stage of the disease. Biomarker patterns or biomarker signature sets can be determined in several tumor types, diseased tissue types, or diseased cells including adipose, adrenal cortex, adrenal gland, adrenal-medulla, appendix, bladder, blood vessels, bone, osteochondral, brain, breast, cartilage, cervix, colon, sigmoid colon, dendritic cells, skeletal muscle, endometrium, esophagus, fallopian tube, fibroblasts, gallbladder, kidney, larynx, liver, lung, lymph nodes, melanocytes, mesothelium, myoepithelial cells, osteoblasts, ovaries, pancreas, parotid gland, prostate, salivary gland, nasal tissue, skeletal muscle, skin, small intestine, smooth muscle, stomach, synovium, intra-articular tissue, tendon, testes, thymus, thyroid, uterus, and uterine body.

[0238] The method of the present invention can be used to select the treatment of any cancer, including but not limited to breast cancer, pancreatic cancer, colon and / or rectal cancer, leukemia, skin cancer, bone cancer, prostate cancer, liver cancer, lung cancer, brain cancer, laryngeal cancer, gallbladder cancer, parathyroid cancer, thyroid cancer, adrenal cancer, nerve tissue cancer, head and neck cancer, stomach cancer, bronchial cancer, kidney cancer, basal cell carcinoma, ulcerative and papillary squamous cell carcinoma, metastatic skin cancer, osteosarcoma, Ewing's sarcoma, reticulosarcoma, myeloma, giant cell tumor, small cell lung tumor, islet cell carcinoma, primary brain tumor, acute and chronic lymphocytic and granulocytic tumors, hairy cell tumor, adenoma, hyperplasia, medullary carcinoma, pheochromocytoma, mucosal neuroma, enteric ganglion cell tumors, hyperplastic corneal nerve tumors, marfanoid habitus tumor, Wilms tumor, seminoma, ovarian tumor, leiomyoma, cervical dysplasia and intraepithelial carcinoma, neuroblastoma, retinoblastoma, soft tissue sarcoma, malignant carcinoid, local skin lesions, fungating polyps, rhabdomyosarcoma, Kaposi's sarcoma, osteogenic and other sarcomas, malignant hypercalcemia, renal cell tumor, polycythemia vera, adenocarcinoma, glioblastoma multiforme, leukemia, lymphoma, malignant melanoma, and epidermoid carcinoma.

[0239] Biomarker patterns or biomarker signature sets can also be determined in some tumor types, affected tissue types, or affected cells, including but not limited to cancer of the paranasal sinuses, middle ear, and inner ear, adrenal gland, appendix, hematopoietic system, bone and joints, spinal cord, breast, cerebellum, cervix, connective soft tissue, uterine body, esophagus, eye, nose, eyeball, fallopian tube, extrahepatic bile duct, other oral cavity, intrahepatic bile duct, kidney, appendix - colon, larynx, lips, liver, lung and bronchus, lymph nodes, cerebrum, spinal cord, nasal cartilage, retina, eye, excluding specific inabilities, oropharynx, other endocrine glands, other female genitalia, ovary, pancreas, penis and scrotum, pituitary gland, pleura, prostate, rectum renal pelvis, ureter, peritoneum, salivary gland, skin, small intestine, stomach, testis, thymus, thyroid, tongue, unknown, bladder, uterus, specific inabilities, vagina and labia, and vulva, and specific inabilities.

[0240] Thus, using a biomarker pattern or biomarker signature set, a therapeutic agent or treatment protocol that can interact with the biomarker pattern or signature set can be determined. For example, for advanced breast cancer, immunohistochemical analysis can be used to determine one or more gene expression proteins that are overexpressed. Thus, a biomarker pattern or biomarker signature set for advanced breast cancer can be identified, and a therapeutic agent or treatment protocol that can interact with the biomarker pattern or signature set can be identified.

[0241] Such examples of biomarker patterns or biomarker signature sets for advanced breast cancer are just one example of the numerous biomarker patterns or biomarker signatures for several advanced diseases or cancers that can be identified from the tables shown in FIGS. 26-31. In addition, by utilizing the above-described method steps of the present invention as shown in FIGS. 1-2 and FIGS. 5-14, several non-disease-specific therapies or treatment protocols can be identified for treating patients having these biomarker patterns or biomarker signature sets.

[0242] The biomarker patterns and / or biomarker signature sets disclosed in the tables shown in FIGS. 26 and 28 and the tables shown in FIGS. 27 and 30 can be used for several purposes including, but not limited to, detection of a specific cancer / disease, treatment of a specific cancer / disease, and identification of a new drug therapy or protocol for a specific cancer / disease. The biomarker patterns and / or biomarker signature sets disclosed in the tables shown in FIGS. 26 and 28 and the tables shown in FIGS. 27 and 30 can also show a drug resistance expression profile for a specific tumor type or cancer type. The biomarker patterns and / or biomarker signature sets disclosed in the tables shown in FIGS. 26 and 28 and the tables shown in FIGS. 27 and 30 show an advanced drug resistance profile.

[0243] A biomarker pattern and / or biomarker signature set can include at least one biomarker. In still other embodiments, a biomarker pattern or signature set can include at least 2, 3, 4, 5, 6, 7, 8, 9, or 10 biomarkers. In some embodiments, a biomarker signature set or biomarker pattern can include at least 15, 20, 30, 40, 50, or 60 biomarkers. In some embodiments, a biomarker signature set or biomarker pattern can include at least 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, 15,000, 20,000, 25,000, 30,000, 35,000, 40,000, 45,000, or 50,000 biomarkers. Analysis of one or more biomarkers can be by one or more methods. For example, two biomarkers can be analyzed using a microarray. Alternatively, one biomarker can be analyzed by IHC and another biomarker can be analyzed by microarray. Any such combination of methods and biomarkers is contemplated herein.

[0244] One or more biomarkers may be selected from the group consisting of, but not limited to, Her2 / Neu, ER, PR, c-kit, EGFR, MLH1, MSH2, CD20, p53, Cyclin D1, bcl2, COX-2, androgen receptor, CD52, PDGFR, AR, CD25, VEGF, HSP90, PTEN, RRM1, SPARC, survivin, TOP2A, BCL2, HIF1A, AR, ESR1, PDGFRA, KIT, PDGFRB, CDW52, ZAP70, PGR, SPARC, GART, GSTP1, NFKBIA, MSH2, TXNRD1, HDAC1, PDGFC, PTEN, CD33, TYMS, RXRB, ADA, TNF, ERCC3, RAF1, VEGF, TOP1, TOP2A, BRCA2, TK1, FOLR2, TOP2B, MLH1, IL2RA, DNMT1, HSP...

Claims

【Claim 1】 A method for identifying a candidate treatment for a subject in need of such treatment, comprising the following steps: (a) performing immunohistochemical (IHC) analysis on a sample from the subject to determine an IHC expression profile for at least five proteins; (b) performing microarray analysis on the sample to determine a microarray expression profile for at least ten genes; (c) performing fluorescence in situ hybridization (FISH) analysis on the sample to determine a FISH mutation profile for at least one gene; (d) performing DNA sequencing on the sample to determine a sequencing mutation profile for at least one gene; and (e) i. cancer cells that overexpress or underexpress one or more proteins included in the IHC expression profile, ii. cancer cells that overexpress or underexpress one or more genes included in the microarray expression profile, iii. cancer cells that have no mutations or one or more mutations in one or more genes included in the FISH mutation profile, and / or iv. cancer cells that have no mutations or one or more mutations in one or more genes included in the sequencing mutation profile comparing the IHC expression profile, the microarray expression profile, the FISH mutation profile, and the sequencing mutation profile against a regulatory database that includes mapping for treatments for which the biological activity against the cancer cells described in (e) above is known; and (f) i. where the comparison in step (e) indicates that the treatment should have biological activity against the cancer, and ii. where the comparison in step (e) does not indicate a contraindication to the treatment for the treatment of the cancer, identifying a candidate treatment.

Citation Information

Patent Citations

  • Method for determining susceptibility of antharcycline-based anticancer agent, and device therefor

    JP2008122258A

  • Methods of Screening for Cell Proliferative or Neoplastic Disorders

    JP2008504018A

  • Inhibitors of progastrin-induced repression of ICAT for treating and / or preventing colorectal cancer, adenomatous polyposis or metastasis displaying progastrin-secreting cells and cells in which the beta-catenin / tef-4-mediated transcriptional pathway

    WO2007135542A2

  • System and method for determining individualized medical intervention for a disease state

    WO2007137187A2

  • Assay for metastatic colorectal cancer

    WO2008057305A2