Next-generation molecular profiling
A machine learning model using molecular profiling and a voting methodology improves treatment prediction accuracy for colorectal cancer by identifying biomarker signatures, addressing inefficiencies in traditional cancer treatment approaches.
Patent Information
- Application Number
- JP2024049383
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-01-07
- Filing Date
- 2024-03-26
- Publication Date
- 2026-01-14
- Estimated Expiration
- 2039-12-02
AI Technical Summary
Traditional cancer treatment approaches are one-size-fits-all, leading to inefficiencies and side effects due to lack of personalized molecular profiling, particularly in colorectal cancer where treatments like FOLFOX have limited efficacy and significant side effects.
A machine learning model utilizing comprehensive molecular profiling data and a 'voting' methodology to predict treatment efficacy by combining multiple classifier models, identifying biomarker signatures that correlate with response to treatments like FOLFOX.
Enhances treatment prediction accuracy by identifying biomarker signatures that indicate likely benefit from FOLFOX or alternative regimens, reducing side effects and improving patient outcomes.
Smart Images

Figure 0007798947000022 
Figure 0007798947000023 
Figure 0007798947000024
Abstract
Description
[Technical Field]
[0001] Priority claims This application claims the benefit of U.S. Provisional Patent Application No. 62 / 744,082, filed November 30, 2018; U.S. Provisional Patent Application No. 62 / 788,689, filed January 4, 2019; and U.S. Provisional Patent Application No. 62 / 789,495, filed January 7, 2019, the entire contents of which are incorporated herein by reference.
[0002] Technical Field The present disclosure relates to the fields of data structures, data processing and machine learning and their use in precision medicine, for example, the use of molecular profiling to guide personalized treatment recommendations for victims of various diseases and disorders, including cancer. [Background technology]
[0003] background Drug therapy for cancer patients has long been a challenge. Traditionally, when a patient was diagnosed with cancer, the treating physician typically selected from a predetermined list of treatment options that traditionally matched the patient's observable clinical factors, such as the type and stage of the cancer. As a result, cancer patients generally received the same treatment as other patients with the same type and stage of cancer. Because patients with the same type and stage of cancer often respond differently to the same treatment, the efficacy of such treatments is determined by trial and error. Moreover, when patients do not immediately respond to any such "one-size-fits-all" treatment, or when previously successful treatments stop working, physicians' treatment choices will often be based on anecdotal evidence at best.
[0004] Until the late 2000s, limited molecular testing was available to assist physicians in making more informed choices from a list of conventional treatments corresponding to a patient's cancer type, also known as "cancer lineage." For example, if a breast cancer patient's physician was presented with a list of conventional treatment options, including Herceptin®, they could have tested the patient's tumor for overexpression of the gene HER2 / neu. HER2 / neu was then known to be associated with breast cancer and responsiveness to Herceptin®. Approximately one-third of breast cancer patients whose tumors were known to overexpress the HER2 / neu gene showed an initial response to treatment with Herceptin®, but the majority of these patients began to progress within one year. See, e.g., Bartsch, R. et al., Trastuzumab in the management of early and advanced stage breast cancer, Biologies. 2007 Mar;1(1):19-31. This type of molecular test helped explain why known treatments for a particular type of cancer are more effective than others in treating some patients with that type of cancer, but the test did not identify or rule out any further treatment options for the patient.
[0005] Frustrated with a one-size-fits-all approach to treating cancer patients and faced with the reality that many patients' tumors will progress and eventually exhaust all conventional therapies, oncologist Daniel Von Hoff sought to identify additional, non-traditional treatment options for his patients. Recognizing the limitations of making treatment decisions based on clinical observation and the limitations of lineage-specific molecular testing, and believing that these limitations could result in effective treatment options being overlooked, Von Hoff and his colleagues developed a system and method for determining personalized treatment regimens for cancer based on a comprehensive assessment of the tumor's molecular characteristics. Their approach to "molecular profiling" used a variety of testing technologies to collect molecular information from a patient's tumor, creating a unique molecular profile regardless of cancer type. Physicians could then use the results of that molecular profile to assist in the selection of candidate treatments for patients, regardless of the stage, anatomical location, or anatomical origin of the cancer cells. See Von Hoff DD, et al., Pilot study using molecular profiling of patients' tumors to find potential targets and select treatments for their refractory cancers. J Clin Oncol. 2010 Nov 20;28(33):4877-83 (Non-Patent Document 2). Such molecular profiling approaches can suggest potential benefits of therapies that would otherwise be overlooked by treating physicians, as well as non-potential benefits of certain therapies, thereby avoiding the time, expense, disease progression, and side effects associated with ineffective treatment. Molecular profiling can be particularly beneficial in the "salvage therapy" setting when patients have failed to respond to or developed resistance to multiple treatment regimens. In addition, such approaches can be used to guide decision-making for first-line and other standard treatment regimens.
[0006] Colorectal cancer (CRC) is the second most common cancer in women and the third most common cancer in men. In 2015, there were 835,000 deaths attributed to CRC worldwide (see Global Burden of Disease Cancer Collaboration, JAMA Oncol. 2017;3(4):524). While surgery is the primary treatment, systemic therapy including 5-fluorouracil / leucovorin in combination with oxaliplatin (FOLFOX) or irinotecan (FOLFIRI) has been shown to be effective in some patients, especially those with metastatic colorectal cancer (Mohelnikova-Duchonova et al., World J Gastroenterol. 2014 Aug 14;20(30):10316-10330).
[0007] FOLFOX has become the standard of care for CRC in the metastatic and adjuvant settings, but only about half of patients respond to treatment. In addition, 20–100% of patients receiving FOLFOX experience at least one of the following: hair loss, pain or peeling of the palms and soles, rash, diarrhea, nausea, vomiting, constipation, loss of appetite, difficulty swallowing, sore mouth, heartburn, infections with low white blood cell counts, anemia, bruising or bleeding, headache, fatigue, numbness, tingling or pain in the extremities, difficulty breathing, cough, and fever; and 4–20% experience chest pain, abnormal heartbeat, fainting, infusion site reactions, hives, weight gain, weight loss, abdominal pain, bruising (black stools, vomit, or urine stains). experience at least one of the following side effects: bleeding (including coughing up blood, coughing up blood, vaginal or testicular bleeding, and bleeding in the brain), taste changes, blood clots, liver damage, yellowing of the eyes and skin, allergic reactions, changes in voice, confusion, dizziness, weakness, blurred vision, sensitivity to light, tics or twitches, difficulty with motor skills (walking, use of hands, mouth opening, speaking, balance / hearing, smell, eating, sleeping, urination), and hearing loss; up to 3% experience serious side effects including at least one of heart damage and development of another treatment-induced cancer.
[0008] A machine learning model can be configured to analyze labeled training data and draw inferences from the training data. When a machine learning model is trained, a set of unlabeled data can be provided to the machine learning model as input. The machine learning model can process the input data, e.g., molecular profiling data, and perform predictions about the input based on inferences learned during training. The present disclosure provides a "voting" methodology for combining multiple classifier models to achieve more accurate classification than can be achieved by using a single model.
[0009] Comprehensive molecular profiling provides abundant data about the molecular status of patient samples.The inventors have carried out such profiling on more than 100,000 tumor patients from virtually all cancer types, and have followed up the patient outcomes and response to treatment in thousands of these patients.For example, our molecular profiling data can be compared with the patient benefit or lack of benefit to treatment, and processed using machine learning algorithms, such as "voting" methodology, to identify additional biomarker signatures that predict the effectiveness of various treatments.Here, this "next-generation profiling" (NGP) method is applied to identify the biomarker signatures that predict the benefit of FOLFOX treatment regimen in colorectal cancer patients. [Prior art documents] [Non-patent literature]
[0010] [Non-Patent Document 1] Bartsch, R. et al., Trastuzumab in the management of early and advanced stage breast cancer, Biologies. 2007 Mar; 1(1): 19-31 [Non-patent document 2] Von Hoff DD, et al., Pilot study using molecular profiling of patients' tumors to find potential targets and select treatments for their refractory cancers. J Clin Oncol. 2010 Nov 20;28(33):4877-83 [Non-patent document 3] Global Burden of Disease Cancer Collaboration, JAMA Oncol. 2017;3(4):524 [Non-patent document 4] Mohelnikova-Duchonova et al., World J Gastroenterol. 2014 Aug 14; 20(30): 10316-10330 Summary of the Invention
[0011] overview Comprehensive molecular profiling provides a wealth of data about the molecular state of a patient sample. Such data can be compared with patient response to treatment to identify biomarker signatures that predict response or non-response to such treatment. This approach has been applied to identify biomarker signatures that correlate with benefit or lack of benefit of the FOLFOX treatment regimen in colorectal cancer patients.
[0012] Described herein is a method for training a machine learning model to predict the efficacy of a treatment for a disease or disorder in a subject having a particular set of biomarkers.
[0013] Provided herein is a data processing apparatus for generating input data structures for use in training a machine learning model to predict the effectiveness of a treatment for a disease or disorder in a subject, the data processing apparatus including one or more processors and one or more storage devices that store instructions that, when executed by the one or more processors, cause the one or more processors to perform operations, the operations including obtaining, by the data processing apparatus, one or more biomarker data structures and one or more outcome data structures; extracting, by the data processing apparatus, first data representing one or more biomarkers associated with the subject from the one or more biomarker data structures, second data representing the disease or disorder and the treatment from the one or more outcome data structures, and third data representing an outcome of the treatment for the disease or disorder. generating, by the data processing device, a data structure for input to a machine learning model based on first data representing one or more biomarkers and second data representing a disease or disorder and a treatment; providing, by the data processing device, the generated data structure as input to the machine learning model; obtaining, by the data processing device, an output generated by the machine learning model based on processing of the machine learning model of the generated data structure; determining, by the data processing device, a difference between third data representing an outcome of the treatment for the disease or disorder and the output generated by the machine learning model; and adjusting, by the data processing device, one or more parameters of the machine learning model based on the difference between the third data representing the outcome of the treatment for the disease or disorder and the output generated by the machine learning model.
[0014] In some embodiments, the set of one or more biomarkers comprises one or more biomarkers set forth in any one of Tables 2-8. In some embodiments, the set of one or more biomarkers comprises each of the biomarkers in Tables 2-8. In some embodiments, the set of one or more biomarkers comprises at least one of the biomarkers in Tables 2-8, and optionally, the set of one or more biomarkers comprises a biomarker in Table 5, Table 6, Table 7, Table 8, or any combination thereof.
[0015] Also provided herein is a data processing apparatus for generating an input data structure for use in training a machine learning model to predict a subject's therapeutic responsiveness to a particular treatment, the apparatus comprising: one or more processors; and one or more storage devices that store instructions that, when executed by the one or more processors, cause the one or more processors to perform operations, the operations including: obtaining, by the data processing apparatus, from a first distributed data source, a first data structure that structures data representing a set of one or more biomarkers associated with a subject, the first data structure including a key-value that identifies the subject; storing, by the data processing apparatus, the first data structure in one or more memory devices; obtaining, by the data processing apparatus, from a second distributed data source, a second data structure that structures data representing outcome data of the subject having the one or more biomarkers, the outcome data including a key-value that identifies the subject. the second data structure includes data identifying the subject, the treatment, and an indication of the effectiveness of the treatment, and the second data structure also includes a key-value identifying the subject; storing, by the data processing device, the second data structure in one or more memory devices; using, by the data processing device, the first data structure and the second data structure stored in the memory device to generate a labeled training data structure including (i) data representing the set of one or more biomarkers, the disease or disorder, and the treatment, and (ii) a label providing an indication of the effectiveness of the treatment for the disease or disorder (the generating, by the data processing device, using the first data structure and the second data structure includes correlating, by the data processing device, the first data structure structuring the data representing the set of one or more biomarkers associated with the subject based on the key-value identifying the subject, with the second data structure representing outcome data for the subject having the one or more biomarkers);and a step of training, by the data processing device, a machine learning model using the generated labeled training data structure (wherein training the machine learning model using the generated labeled training data structure includes providing, by the data processing device, the generated labeled training data structure to the machine learning model as an input to the machine learning model);
[0016] In some embodiments, the operations further include obtaining, by the data processing device, from the machine learning model an output generated by the machine learning model based on the machine learning model's processing of the generated labeled training data structure; and determining, by the data processing device, a difference between the output generated by the machine learning model and a label that provides an indication of the effectiveness of a treatment for the disease or disorder.
[0017] In some embodiments, the operations further include adjusting, by the data processing device, one or more parameters of the machine learning model based on the determined difference between the output generated by the machine learning model and the label that provides an indication of the effectiveness of the treatment for the disease or disorder.
[0018] In some embodiments, the set of one or more biomarkers comprises one or more biomarkers set forth in any one of Tables 2-8. In some embodiments, the set of one or more biomarkers comprises each of the biomarkers in Tables 2-8. In some embodiments, the set of one or more biomarkers comprises at least one of the biomarkers in Tables 2-8, and optionally, the set of one or more biomarkers comprises the biomarkers in Table 5, Table 6, Table 7, Table 8, or any combination thereof.
[0019] Relatedly, provided herein is a method including steps corresponding to each of the operations of the data processing apparatus. Still further, provided herein is a system including one or more computers and one or more storage media storing instructions that, when executed by the one or more computers, cause the one or more computers to perform each of the operations described with reference to the data processing apparatus. Still further, provided herein is a non-transitory computer-readable medium storing software executable by one or more computers and including instructions that, when so executed, cause the one or more computers to perform the operations described with reference to the data processing apparatus.
[0020] In another aspect, provided herein is a method for entity classification, the method including, for each particular machine learning model of a plurality of machine learning models, i) providing input data representing a type of entity to be classified to the particular machine learning model trained to determine a prediction or classification; ii) obtaining output data representing the entity classifications of a plurality of candidate entity classes into initial entity classes produced by the particular machine learning model based on processing of the input data by the particular machine learning model; providing the obtained output data for each of the plurality of machine learning models to a voting unit, the provided output data including data representing the initial entity classes determined by each of the plurality of machine learning models; and determining, by the voting unit, an actual entity class for the entity based on the provided output data.
[0021] In some aspects, the actual entity class for an entity is determined by applying a majority rule to the provided output data.
[0022] In some embodiments, determining, by the voting unit, the actual entity class for the entity based on the provided output data includes: determining, by the voting unit, a number of occurrences of each initial entity class among the plurality of candidate entity classes; and selecting, by the voting unit, the initial entity class having the greatest number of occurrences among the plurality of candidate entity classes.
[0023] In some embodiments, each machine learning model of the plurality of machine learning models comprises a random forest classification algorithm, a support vector machine, a logistic regression, a k-nearest neighbor model, an artificial neural network, a naive Bayes model, a quadratic discriminant analysis, or a Gaussian process model.
[0024] In some embodiments, each machine learning model of the plurality of machine learning models comprises a random forest classification algorithm.
[0025] In some embodiments, the multiple machine learning models include multiple representations of the same type of classification algorithm.
[0026] In some embodiments, the input data represents (i) entity attributes and (ii) types of treatments for diseases or disorders.
[0027] In some embodiments, the plurality of candidate entity classes comprises a reactive class or a non-responsive class.
[0028] In some embodiments, the entity attributes include one or more biomarkers for the entity.
[0029] In some embodiments, the one or more biomarkers comprise a panel of genes that is less than all known genes of the entity.
[0030] In some embodiments, the one or more biomarkers comprise a panel of genes that includes all known genes for the entity.
[0031] In some embodiments, the input data further comprises data representing a type of disease or disorder.
[0032] Relatedly, provided herein is a system including one or more computers and one or more storage media storing instructions that, when executed by the one or more computers, cause the one or more computers to perform each of the operations described with reference to the method for entity classification. Still further, provided herein is a non-transitory computer-readable medium storing software including instructions executable by one or more computers and, when so executed, causing the one or more computers to perform the operations described with reference to the method for entity classification.
[0033] In yet another aspect, provided herein is a method comprising obtaining a biological sample comprising cells from a cancer in a subject; and performing an assay to evaluate at least one biomarker in the biological sample, wherein the biomarker is: (a) Group 1, including 1, 2, 3, 4, 5, or all 6 of MYC, EP300, U2AF1, ASXL1, MAML2, and CNTRL; (b) group 2, including 1, 2, 3, 4, 5, 6, 7, or all 8 of MYC, EP300, U2AF1, ASXL1, MAML2, CNTRL, WRN, and CDX2; (c) group 3, including 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or all 14 of BCL9, PBX1, PRRX1, INHBA, YWHAE, GNAS, LHFPL6, FCRL4, HOXA11, AURKA, BIRC3, IKZF1, CASP8, and EP300; (d) group 4, including 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or all 13 of PBX1, BCL9, INHBA, PRRX1, YWHAE, GNAS, LHFPL6, FCRL4, AURKA, IKZF1, CASP8, PTEN, and EP300; (e) group 5, including all 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 of BCL9, PBX1, PRRX1, INHBA, GNAS, YWHAE, LHFPL6, FCRL4, PTEN, HOXA11, AURKA, and BIRC3; (f) group 6, including one, two, three, four, or all five of BCL9, PBX1, PRRX1, INHBA, and YWHAE; (g) group 7, including all 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 of BCL9, PBX1, GNAS, LHFPL6, CASP8, ASXL1, FH, CRKL, MLF1, TRRAP, AKT3, ACKR3, MSI2, PCM1, and MNX1; (h)BX1, GNAS, AURKA, CASP8, ASXL1, CRKL, MLF1, GAS7, MN1, SOX10, TCL1A, LMO1, BRD3, SMARCA4, PER1, PAX7, SBDS, SEPT5, PDGFB, AKT2, TERT, KEAP1, ETV6, TOP1, TLX3, COX6C, NFIB, ARFRP1, ARID1A, MAP2K4, NFKBIA, WWTR1, ZNF217, IL2, NSD3, CR Group 8, including all 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, or 45 of EB1, BRIP1, SDC4, EWSR1, FLT3, FLT1, FAS, CCNE1, RUNX1T1, and EZR; and (i) Group 9, including 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or all 11 of BCL9, PBX1, PRRX1, INHBA, YWHAE, GNAS, LHFPL6, FCRL4, BIRC3, AURKA, and HOXA11; The method includes at least one of the following:
[0034] In some embodiments, the biological sample comprises formalin-fixed paraffin-embedded (FFPE) tissue, fixed tissue, core needle biopsy, fine needle aspirate, unstained slide, fresh frozen (FF) tissue, formalin sample, tissue contained in a solution that preserves nucleic acid or protein molecules, fresh sample, malignant fluid, bodily fluid, tumor sample, tissue sample, or any combination thereof.
[0035] In some embodiments, the biological sample comprises cells from a solid tumor.
[0036] In some embodiments, the biological sample comprises a bodily fluid.
[0037] In some embodiments, the bodily fluid comprises malignant fluid, pleural fluid, peritoneal fluid, or any combination thereof.
[0038] In some embodiments, the bodily fluid comprises peripheral blood, serum, plasma, ascites, urine, cerebrospinal fluid (CSF), sputum, saliva, bone marrow, synovial fluid, aqueous humor, amniotic fluid, earwax, breast milk, bronchoalveolar lavage fluid, semen, prostatic fluid, Cowper's fluid, pre-ejaculate fluid, female ejaculate, sweat, feces, tears, cyst fluid, pleural fluid, peritoneal fluid, pericardial fluid, lymph, chyme, chyle, bile, interstitial fluid, menstrual fluid, pus, sebum, vomit, vaginal fluid, mucosal secretions, stool water, pancreatic juice, nasal washings, bronchopulmonary aspirate, blastocyst fluid, or umbilical cord blood.
[0039] In some embodiments, the evaluation comprises determining the presence, level, or status of a protein or nucleic acid for each biomarker, and optionally, the nucleic acid comprises deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or a combination thereof. In some embodiments, (a) the presence, level, or status of a protein is determined using immunohistochemistry (IHC), flow cytometry, immunoassay, antibody or functional fragment thereof, aptamer, or any combination thereof; and / or (b) the presence, level, or status of a nucleic acid is determined using polymerase chain reaction (PCR), in situ hybridization, amplification, hybridization, microarray, nucleic acid sequencing, dye terminator sequencing, pyrosequencing, next-generation sequencing (NGS; high-throughput sequencing), or any combination thereof.
[0040] In some embodiments, the state of a nucleic acid comprises a sequence, mutation, polymorphism, deletion, insertion, substitution, translocation, fusion, truncation, duplication, amplification, repeat, copy number, copy number variation (CNV; copy number alteration; CNA), or any combination thereof.
[0041] In some embodiments, the state of the nucleic acid comprises copy number.
[0042] In some embodiments, the method includes performing an assay to determine the copy number of all members of Group 1 (i.e., MYC, EP300, U2AF1, ASXL1, MAML2, and CNTRL) or genomic regions adjacent thereto.
[0043] In some embodiments, the method includes performing an assay to determine the copy number of all members of Group 2 (i.e., MYC, EP300, U2AF1, ASXL1, MAML2, CNTRL, WRN, and CDX2) or genomic regions adjacent thereto.
[0044] In some embodiments, the method includes performing an assay to determine the copy number of all members of Group 3 (i.e., BCL9, PBX1, PRRX1, INHBA, YWHAE, GNAS, LHFPL6, FCRL4, HOXA11, AURKA, BIRC3, IKZF1, CASP8 and EP300) or genomic regions adjacent thereto.
[0045] In some embodiments, the method includes performing an assay to determine the copy number of all members of Group 4 (i.e., PBX1, BCL9, INHBA, PRRX1, YWHAE, GNAS, LHFPL6, FCRL4, AURKA, IKZF1, CASP8, PTEN, and EP300) or genomic regions adjacent thereto.
[0046] In some embodiments, the method includes performing an assay to determine the copy number of all members of Group 5 (i.e., BCL9, PBX1, PRRX1, INHBA, GNAS, YWHAE, LHFPL6, FCRL4, PTEN, HOXA11, AURKA, and BIRC3) or genomic regions adjacent thereto.
[0047] In some embodiments, the method comprises performing an assay to determine the copy number of all members of Group 6 (i.e., BCL9, PBX1, PRRX1, INHBA, and YWHAE) or genomic regions adjacent thereto.
[0048] In some embodiments, the method includes performing an assay to determine the copy number of all members of Group 7 (i.e., BCL9, PBX1, GNAS, LHFPL6, CASP8, ASXL1, FH, CRKL, MLF1, TRRAP, AKT3, ACKR3, MSI2, PCM1 and MNX1) or genomic regions adjacent thereto.
[0049] In some embodiments, the method comprises performing an assay to determine the copy number of all members of Group 8 (i.e., BX1, GNAS, AURKA, CASP8, ASXL1, CRKL, MLF1, GAS7, MN1, SOX10, TCL1A, LMO1, BRD3, SMARCA4, PER1, PAX7, SBDS, SEPT5, PDGFB, AKT2, TERT, KEAP1, ETV6, TOP1, TLX3, COX6C, NFIB, ARFRP1, ARID1A, MAP2K4, NFKBIA, WWTR1, ZNF217, IL2, NSD3, CREB1, BRIP1, SDC4, EWSR1, FLT3, FLT1, FAS, CCNE1, RUNX1T1 and EZR) or genomic regions adjacent thereto.
[0050] In some embodiments, the method includes performing an assay to determine the copy number of all members of Group 9 (i.e., BCL9, PBX1, PRRX1, INHBA, YWHAE, GNAS, LHFPL6, FCRL4, BIRC3, AURKA, and HOXA11) or genomic regions adjacent thereto.
[0051] In some embodiments, the method comprises performing an assay to determine the copy number of (a) at least one or all members of Group 1 and Group 2 or genomic regions adjacent thereto; (b) at least one or all members of Group 3 or genomic regions adjacent thereto; or (c) at least one or all members of Group 2, Group 6, Group 7, Group 8, and Group 9 or genomic regions adjacent thereto.
[0052] In some embodiments, the method further comprises comparing the copy number of the biomarker to a reference copy number (e.g., diploid) to identify biomarkers with copy number variation (CNV).
[0053] In some embodiments, the method further comprises generating a molecular profile that identifies the gene or a region adjacent to the gene that has the CNV.
[0054] In some embodiments, the presence or level of PTEN protein is determined, and optionally, the presence or level of PTEN protein is determined using immunohistochemistry (IHC).
[0055] In some embodiments, the method further comprises determining the level of proteins including TOPO1 and one or more mismatch repair proteins (e.g., MLH1, MSH2, MSH6 and PMS2), and optionally, the presence or level of PTEN protein is determined using immunohistochemistry (IHC).
[0056] In some embodiments, the method further comprises comparing the level of the one or more proteins to a reference level for that protein.
[0057] In some embodiments, the method further comprises generating a molecular profile that identifies proteins having levels that differ from the reference level, e.g., that differ significantly from the reference level.
[0058] In some embodiments, the method further comprises selecting a treatment of likely benefit based on the evaluated biomarkers, optionally wherein the treatment comprises 5-fluorouracil / leucovorin in combination with oxaliplatin (FOLFOX) or an alternative treatment thereto, and optionally wherein the alternative treatment comprises 5-fluorouracil / leucovorin in combination with irinotecan (FOLFIRI).
[0059] In some embodiments, selecting a treatment of likely benefit is based on (a) the copy number determined for the group; and / or (b) the molecular profile determined as described above.
[0060] In some embodiments, selecting a treatment of likely benefit based on the copy number determined for the group comprises use of a voting module.
[0061] In some aspects, the voting module is a voting module provided herein.
[0062] In some embodiments, the voting module includes the use of at least one random forest model.
[0063] In some embodiments, the use of the voting module comprises applying a machine learning classification model to the copy numbers obtained for each of group 2, group 6, group 7, group 8, and group 9 (see above), optionally wherein each machine learning classification model is a random forest model, optionally wherein the random forest model is a random forest model described in Table 10 below.
[0064] In some embodiments, the subject has not been previously treated with a therapy of potential benefit.
[0065] In some embodiments, the cancer comprises metastatic cancer, recurrent cancer, or a combination thereof.
[0066] In some embodiments, the subject has not previously undergone cancer treatment.
[0067] In some embodiments, the method further comprises administering to the subject a treatment of potential benefit.
[0068] In some embodiments, progression-free survival (PFS), disease-free survival (DFS) or life span is extended by administration of said treatment.
[0069] In some embodiments, the cancer is acute lymphoblastic leukemia; acute myeloid leukemia; adrenocortical carcinoma; AIDS-related cancer; AIDS-related lymphoma; anal cancer; appendix cancer; astrocytoma; atypical teratoid / rhabdoid tumor; basal cell carcinoma; bladder cancer; brain stem glioma; brain tumor, brain stem glioma, central nervous system atypical teratoid / rhabdoid tumor, central nervous system embryonal tumor, astrocytoma, craniopharyngioma, ependymoblastoma, ependymoma, medulloblastoma, medulloepithelioma, intermediate pineal parenchymal tumor, supratentorial primitive neuroectodermal tumor and pineoblastoma; breast cancer; bronchoma Tumor; Burkitt's lymphoma; Cancer of unknown primary (CUP); Carcinoid tumor; Carcinoma of unknown primary; Central nervous system atypical teratoid / rhabdoid tumor; Central nervous system embryonal tumor; Cervical cancer; Childhood cancer; Chordoma; Chronic lymphocytic leukemia; Chronic myeloid leukemia; Chronic myeloproliferative disorder; Colon cancer; Colorectal cancer; Craniopharyngioma; Cutaneous T-cell lymphoma; Endocrine pancreatic islet cell tumor; Endometrial cancer; Ependymoblastoma; Ependymoma; Esophageal cancer; Nasal neuroblastoma; Ewing's sarcoma; Extracranial germ cell tumor; Extragonadal germ cell tumor; Extrahepatic bile duct cancer; Gallbladder cancer; Gastric cancer (stomach) cancer); gastrointestinal carcinoid tumor; gastrointestinal stromal cell tumor; gastrointestinal stromal tumor (GIST); gestational trophoblastic tumor; glioma; hairy cell leukemia; head and neck cancer; cardiac cancer; Hodgkin's lymphoma; hypopharyngeal cancer; intraocular melanoma; pancreatic islet tumor; Kaposi's sarcoma; kidney cancer; Langerhans cell histiocytosis; laryngeal cancer; lip cancer; liver cancer; malignant fibrous histiocytoma bone cancer; medulloblastoma; medulloepithelioma; melanoma; Merkel cell carcinoma; Merkel cell skin cancer; mesothelioma; metastatic squamous neck cancer of unknown primary; oral cancer; multiple endocrine neoplasia syndrome; multiple myeloma; multiple myeloma / plasma cell neoplasm; mycosis fungoides; myelodysplastic syndrome; myeloproliferative neoplasm; nasal cavity cancer; nasopharyngeal carcinoma; neuroblastoma; non-Hodgkin's lymphoma; non-melanoma skin cancer; non-small cell lung cancer; oral cancer cancer); oral cavity cancer; oropharyngeal cancer; osteosarcoma; other brain and spinal cord tumors; ovarian cancer; ovarian epithelial cancer; ovarian germ cell tumor; ovarian low malignant potential tumor; pancreatic cancer; papillomatosis; sinonasal cancer; parathyroid cancer; pelvic cancer; penile cancer; pharyngeal cancer; intermediate pineal parenchymal tumor; pineoblastoma; pituitary tumor; plasma cell neoplasm / multiple myeloma; pleuropulmonary blastoma; primary central nervous system (CNS) lymphoma;Primary hepatocellular carcinoma; prostate cancer; rectal cancer; renal cancer; renal cell (kidney) cancer; renal cell carcinoma; airway cancer; retinoblastoma; rhabdomyosarcoma; salivary gland cancer; Sézary syndrome; small cell lung cancer; small intestine cancer; soft tissue sarcoma; squamous cell carcinoma; cervical squamous cell carcinoma; stomach (gastric) cancer; supratentorial primitive neuroectodermal tumor; T-cell lymphoma; testicular cancer; throat cancer; thymic carcinoma; thymoma; thyroid cancer; transitional cell carcinoma; transitional cell carcinoma of the renal pelvis and ureter; trophoblastic tumor; ureteral cancer; urethral cancer; uterine cancer; uterine sarcoma; vaginal cancer; vulvar cancer; Waldenstrom's macroglobulinemia; or Wilms' tumor.
[0070] In some embodiments, the cancer is acute myeloid leukemia (AML), breast cancer, cholangiocarcinoma, colorectal adenocarcinoma, extrahepatic bile duct adenocarcinoma, female genital malignancies, gastric adenocarcinoma, gastroesophageal adenocarcinoma, gastrointestinal stromal tumor (GIST), glioblastoma, head and neck squamous cell carcinoma, leukemia, hepatocellular carcinoma, low-grade glioma, lung bronchoalveolar carcinoma (BAC), non-small cell lung cancer (NSCLC), small cell lung cancer (SCLC), lymphoma, male reproductive tract cancer (MRCC), or malignant tumors of the genital tract. Includes organ malignancies, malignant solitary fibrous tumor of the pleura (MSFT), melanoma, multiple myeloma, neuroendocrine tumors, nodular diffuse large B-cell lymphoma, non-epithelial ovarian cancer (non-EOC), ovarian surface epithelial carcinoma, pancreatic adenocarcinoma, pituitary carcinoma, oligodendroglioma, prostate adenocarcinoma, retroperitoneal or peritoneal carcinoma, retroperitoneal or peritoneal sarcoma, small intestine malignancies, soft tissue tumors, thymic carcinoma, thyroid carcinoma or uveal melanoma.
[0071] In some embodiments, the cancer comprises colorectal cancer.
[0072] Further provided herein is a method of selecting a treatment for a subject having colorectal cancer, comprising the steps of obtaining a biological sample comprising cells from the colorectal cancer; performing next-generation sequencing on genomic DNA from the biological sample to identify (a) group 2 comprising one, two, three, four, five, six, seven, or all eight of: MYC, EP300, U2AF1, ASXL1, MAML2, CNTRL, WRN, and CDX2; (b) group 3 comprising one, two, three, four, five, six, seven, or all five of: BCL9, PBX1, PRRX1, INHBA, and YWHAE; (c) Group 6, including 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or all 15 of BCL9, PBX1, GNAS, LHFPL6, CASP8, ASXL1, FH, CRKL, MLF1, TRRAP, AKT3, ACKR3, MSI2, PCM1, and MNX1; (d) Group 7, including 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or all 15 of BX1, GNAS, AURKA, CASP8, ASXL1, CRKL, MLF1, GAS7, MN1, SOX10, TCL1A, LMO1, BRD3, SMARCA4, PER1, PA 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 X7, SBDS, SEPT5, PDGFB, AKT2, TERT, KEAP1, ETV6, TOP1, TLX3, COX6C, NFIB, ARFRP1, ARID1A, MAP2K4, NFKBIA, WWTR1, ZNF217, IL2, NSD3, CREB1, BRIP1, SDC4, EWSR1, FLT3, FLT1, FAS, CCNE1, RUNX1T1 and EZR , 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44 or 45 of Group 8, and (e) Group 9, including one, two, three, four, five, six, seven, eight, nine, ten or eleven of BCL9, PBX1, PRRX1, INHBA, YWHAE, GNAS, LHFPL6, FCRL4, BIRC3, AURKA and HOXA11;The method includes applying a machine learning classification model to the copy numbers obtained for each of Groups 2, 6, 7, 8, and 9 (optionally, each machine learning classification model is a random forest model, and optionally, the random forest model is a random forest model described in Table 10); obtaining from each machine learning classification model an indication of whether the subject is likely to benefit from 5-fluorouracil / leucovorin in combination with oxaliplatin (FOLFOX); and selecting FOLFOX if a majority of the machine learning classification models indicate that the subject is likely to benefit from the treatment, or selecting an alternative treatment to FOLFOX (optionally, the alternative treatment is 5-fluorouracil / leucovorin in combination with irinotecan (FOLFIRI)) if a majority of the machine learning classification models indicate that the subject is unlikely to benefit from FOLFOX. In some embodiments, the method further includes administering the selected treatment to the subject.
[0073] Still further provided herein is a method for generating a molecular profiling report, comprising generating a report summarizing the results of performing the method. In some embodiments, the report includes (a) a treatment of likely benefit determined as disclosed above; or (b) a selected treatment determined as disclosed above. In some embodiments, the report is computer-generated; is a printed report or computer file; or is accessible via a web portal.
[0074] Relatedly, provided herein is a system for identifying a treatment for cancer in a subject, the system including: (a) at least one host server; (b) at least one user interface for accessing the at least one host server to access and input data; (c) at least one processor for processing the input data; (d) at least one memory coupled to the processor for storing the processed data and instructions for (1) accessing the results of analyzing the biological sample as described above; and (2) determining a treatment of likely benefit as described above or a selected treatment as described above; and (e) at least one display for displaying the cancer treatment (the treatment is FOLFOX or an alternative thereto, e.g., FOLFIRI).
[0075] In some embodiments, at least one display includes a report including the results of analyzing the biological sample and treatments that have potential benefit in treating the cancer or have been selected for treating the cancer.
[0076] Additionally provided herein is a method of providing recommendations for cancer treatment to provide longer progression-free survival, longer disease-free survival, longer overall survival, or extended lifespan, comprising the steps of: obtaining a biological sample comprising nucleic acids and / or proteins from an individual diagnosed with cancer; performing molecular testing on the biological sample to determine one or more molecular features selected from the group consisting of: nucleic acid sequences of a set of target genes or portions thereof; the presence of copy number variation of the set of target genes; the presence of gene fusions or other genomic alterations; one or more levels of a set of proteins and / or transcripts; and / or the epigenetic status of the set of target genes, e.g., as described herein, thereby generating a molecular profile for the cancer; comparing the molecular profile of the cancer to a reference molecular profile for that type of cancer; generating a list of molecular features that exhibit differences, e.g., significant differences, when compared to the reference molecular profile; and generating a list of one or more treatment recommendations for the individual based on the list of molecular features that exhibit differences when compared to the reference sequence profile of the target genes.
[0077] In some embodiments, the molecular test is at least one of next generation sequencing, Sanger sequencing, ISH, fragment analysis, PCR, IHC, and immunoassay.
[0078] In some embodiments, the biological sample comprises a cell, a tissue sample, a blood sample, or a combination thereof.
[0079] In some embodiments, the molecular test detects at least one of a mutation, polymorphism, deletion, insertion, substitution, translocation, fusion, truncation, duplication, amplification, or repeat.
[0080] In some embodiments, the nucleic acid sequence comprises a deoxyribonucleic acid sequence.
[0081] In some embodiments, the nucleic acid sequence comprises a ribonucleic acid sequence.
[0082] Unless otherwise specified, all scientific and technical terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention belongs. Methods and materials for use in the present invention are described herein; however, other suitable methods and materials known in the art can also be used. Materials, methods, and examples are illustrative only and are not intended to be limiting. All publications, patent applications, patents, sequences, database entries, and other references mentioned herein are incorporated herein by reference in their entirety. In case of conflict, the present specification, including definitions, will control.
[0083] [The present invention 1001] 1. A data processing apparatus for generating an input data structure for use in training a machine learning model for predicting the efficacy of a treatment for a disease or disorder of interest, the apparatus comprising: the data processing apparatus including one or more processors and one or more storage devices that store instructions that, when executed by the one or more processors, cause the one or more processors to perform operations; The operation is obtaining, with said data processing device, one or more biomarker data structures and one or more outcome data structures; extracting, with the data processing device, first data representing one or more biomarkers associated with the subject from the one or more biomarker data structures, second data representing a disease or disorder and a treatment from the one or more outcome data structures, and third data representing an outcome of a treatment for the disease or disorder; generating, by the data processing device, a data structure for input into a machine learning model based on first data representative of the one or more biomarkers and second data representative of the disease or disorder and treatment; providing, by the data processing device, the generated data structure as an input to the machine learning model; obtaining, by the data processing device, an output generated by the machine learning model based on processing of the generated data structure by the machine learning model; determining, by the data processing device, a difference between third data representing an outcome of treatment for the disease or disorder and an output generated by the machine learning model; and adjusting, by the data processing device, one or more parameters of the machine learning model based on a difference between third data representing an outcome of treatment for the disease or disorder and an output generated by the machine learning model. The data processing device comprising: [The present invention 1002] 1001. A data processing apparatus of the present invention, wherein the set of one or more biomarkers comprises one or more biomarkers listed in any one of Tables 2-8. [The present invention 1003] A data processing device of the present invention 1001, wherein the set of one or more biomarkers includes each of the biomarkers of the present invention 1002. [The present invention 1004] The data processing device of the present invention 1001, wherein the set of one or more biomarkers comprises at least one of the biomarkers of the present invention 1002, and optionally the set of one or more biomarkers comprises the markers in Table 5, Table 6, Table 7, Table 8 or any combination thereof. [The present invention 1005] 1. A data processing apparatus for generating an input data structure for use in training a machine learning model for predicting a subject's therapeutic responsiveness to a particular treatment, the apparatus comprising: the data processing apparatus including one or more processors and one or more storage devices that store instructions that, when executed by the one or more processors, cause the one or more processors to perform operations; The operation is obtaining, by the data processing device, from a first distributed data source, a first data structure that structures data representing a set of one or more biomarkers associated with a subject, the first data structure including a key-value that identifies the subject; storing, by the data processing apparatus, the first data structure in one or more memory devices; obtaining, by the data processing device, from a second distributed data source, a second data structure that structures data representing outcome data for subjects having the one or more biomarkers, the outcome data including data identifying a disease or disorder, a treatment, and an indication of efficacy of the treatment, and the second data structure also including key-values that identify the subjects; storing, by the data processing apparatus, the second data structure in one or more memory devices; generating, by the data processing device, a labeled training data structure using the first data structure and the second data structure stored in the memory device, the labeled training data structure comprising (i) data representative of a set of one or more biomarkers, the disease or disorder, and a treatment, and (ii) a label providing an indication of the effectiveness of a treatment for the disease or disorder, wherein the generating, by the data processing device, using the first data structure and the second data structure includes correlating, by the data processing device, a first data structure that structures data representative of the set of one or more biomarkers associated with the subject based on a key-value that identifies the subject, with a second data structure that represents outcome data for subjects having the one or more biomarkers; and training, by the data processing device, a machine learning model using the generated labeled training data structure, wherein training the machine learning model using the generated labeled training data structure comprises providing, by the data processing device, the generated labeled training data structure to the machine learning model as an input to the machine learning model. The data processing device comprising: [The present invention 1006] The operation is obtaining, by the data processing device, from the machine learning model, an output generated by the machine learning model based on the machine learning model's processing of the generated labeled training data structure; and determining, by the data processing device, the difference between the output generated by the machine learning model and a label that provides an indication of the effectiveness of a treatment for a disease or disorder. The data processing device of the present invention 1005 further includes: [The present invention 1007] The operation is adjusting, by the data processing device, one or more parameters of the machine learning model based on the determined difference between the output generated by the machine learning model and the label that provides an indication of the effectiveness of the treatment for the disease or disorder. The data processing device of the present invention 1006 further includes: [The present invention 1008] 1005. A data processing device of the present invention, wherein the set of one or more biomarkers comprises one or more biomarkers set forth in any one of Tables 2-8, and optionally, the set of one or more biomarkers comprises markers in Table 5, Table 6, Table 7, Table 8, or any combination thereof. [The present invention 1009] A data processing apparatus of the present invention 1005, wherein the set of one or more biomarkers includes each of the biomarkers of the present invention 1008. [The present invention 1010] The data processing device of the present invention 1005, wherein the set of one or more biomarkers includes one of the biomarkers of the present invention 1008. [The present invention 1011] A method including steps corresponding to any of the operations 1001 to 1010 of the present invention. [The present invention 1012] A system including one or more computers and one or more data storage media storing instructions that, when executed by the one or more computers, cause the one or more computers to perform each of the operations of the present invention 1001 to 1010. [The present invention 1013] Instructions executable by one or more computers that, when so executed, cause the one or more computers to perform any of the operations of the present invention 1001-1010. A non-transitory computer readable medium storing software including: [The present invention 1014] 1. A method for classification of entities, comprising: With respect to each particular machine learning model of the plurality of machine learning models, providing input data representing the type of entity to be classified to a specific machine learning model trained to determine a prediction or classification; obtaining output data representing entity classifications of a plurality of candidate entity classes into initial entity classes generated by the particular machine learning model based on processing of input data by the particular machine learning model; providing output data obtained for each of the plurality of machine learning models to a voting unit, the provided output data including data representing an initial entity class determined by each of the plurality of machine learning models; and determining, by the voting unit, a real entity class for the entity based on the provided output data. The method comprising: [The present invention 1015] The method of the present invention 1014, wherein the actual entity class for the entity is determined by applying a majority rule to the provided output data. [The present invention 1016] determining, by the voting unit, a real entity class for the entity based on the provided output data; determining, by the voting unit, a number of occurrences of each initial entity class of a plurality of candidate entity classes; and selecting, by the voting unit, an initial entity class having a maximum number of occurrences from among the plurality of candidate entity classes; The method of the present invention 1014 or 1015, comprising: [The present invention 1017] Any of the methods of claims 1014 to 1016, wherein each machine learning model of the plurality of machine learning models comprises a random forest classification algorithm, a support vector machine, a logistic regression, a k-nearest neighbor model, an artificial neural network, a naive Bayes model, a quadratic discriminant analysis, or a Gaussian process model. [The present invention 1018] The method of any of claims 1014 to 1016, wherein each machine learning model of the plurality of machine learning models comprises a random forest classification algorithm. [The present invention 1019] The method of any of claims 1014 to 1018, wherein the plurality of machine learning models comprises multiple representations of the same type of classification algorithm. [The present invention 1020] The method of any of claims 1014-1018, wherein the input data represents (i) entity attributes and (ii) types of treatment for a disease or disorder. [The present invention 1021] The method of the present invention 1020, wherein the plurality of candidate entity classes includes a reactive class or a non-reactive class. [The present invention 1022] The method of any one of claims 1020 to 1021, wherein the entity attributes include one or more biomarkers for the entity. [The present invention 1023] The method of claim 1022, wherein the one or more biomarkers comprise a panel of genes that is less than all known genes of the entity. [The present invention 1024] The method of claim 1022, wherein the one or more biomarkers comprises a panel of genes including all known genes for the entity. [The present invention 1025] The method of any one of claims 1020 to 1024, wherein the input data further comprises data representing the type of disease or disorder. [The present invention 1026] A system comprising one or more computers and one or more data storage media storing instructions that, when executed by the one or more computers, cause the one or more computers to perform each of the operations of the present invention 1014-1025. [The present invention 1027] Instructions executable by one or more computers that, when so executed, cause the one or more computers to perform any of the operations of the present invention 1014-1025. A non-transitory computer readable medium storing software including: [The present invention 1028] Obtaining a biological sample containing cells derived from a cancer in a subject; and performing an assay to evaluate at least one biomarker in said biological sample. A method comprising: The biomarker is (a) Group 1, including 1, 2, 3, 4, 5, or all 6 of MYC, EP300, U2AF1, ASXL1, MAML2, and CNTRL; (b) group 2, including 1, 2, 3, 4, 5, 6, 7, or all 8 of MYC, EP300, U2AF1, ASXL1, MAML2, CNTRL, WRN, and CDX2; (c) group 3, including 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or all 14 of BCL9, PBX1, PRRX1, INHBA, YWHAE, GNAS, LHFPL6, FCRL4, HOXA11, AURKA, BIRC3, IKZF1, CASP8, and EP300; (d) group 4, including 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or all 13 of PBX1, BCL9, INHBA, PRRX1, YWHAE, GNAS, LHFPL6, FCRL4, AURKA, IKZF1, CASP8, PTEN, and EP300; (e) group 5, including all 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 of BCL9, PBX1, PRRX1, INHBA, GNAS, YWHAE, LHFPL6, FCRL4, PTEN, HOXA11, AURKA, and BIRC3; (f) group 6, including one, two, three, four, or all five of BCL9, PBX1, PRRX1, INHBA, and YWHAE; (g) group 7, including all 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 of BCL9, PBX1, GNAS, LHFPL6, CASP8, ASXL1, FH, CRKL, MLF1, TRRAP, AKT3, ACKR3, MSI2, PCM1, and MNX1; (h)BX1, GNAS, AURKA, CASP8, ASXL1, CRKL, MLF1, GAS7, MN1, SOX10, TCL1A, LMO1, BRD3, SMARCA4, PER1, PAX7, SBDS, SEPT5, PDGFB, AKT2, TERT, KEAP1, ETV6, TOP1, TLX3, COX6C, NFIB, ARFRP1, ARID1A, MAP2K4, NFKBIA, WWTR1, ZNF217, IL2, NSD3, CR Group 8, including all 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, or 45 of EB1, BRIP1, SDC4, EWSR1, FLT3, FLT1, FAS, CCNE1, RUNX1T1, and EZR; and (i) Group 9, including 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or all 11 of BCL9, PBX1, PRRX1, INHBA, YWHAE, GNAS, LHFPL6, FCRL4, BIRC3, AURKA, and HOXA11; The method includes at least one of the following: [The present invention 1029] The method of the present invention 1028, wherein the biological sample comprises formalin-fixed paraffin-embedded (FFPE) tissue, fixed tissue, core needle biopsy, fine needle aspirate, unstained slide, fresh frozen (FF) tissue, formalin sample, tissue contained in a solution that preserves nucleic acid or protein molecules, fresh sample, malignant fluid, body fluid, tumor sample, tissue sample, or any combination thereof. [The present invention 1030] The method of any one of claims 1028 to 1029, wherein the biological sample comprises cells from a solid tumor. [The present invention 1031] The method of any one of claims 1028 to 1029, wherein the biological sample comprises a body fluid. [The present invention 1032] 1032. The method of any of claims 1028 to 1031, wherein the body fluid comprises malignant fluid, pleural fluid, peritoneal fluid, or any combination thereof. [The present invention 1033] 1032. The method of any of claims 1028 to 1032, wherein the body fluid comprises peripheral blood, serum, plasma, ascites, urine, cerebrospinal fluid (CSF), sputum, saliva, bone marrow, synovial fluid, aqueous humor, amniotic fluid, earwax, breast milk, bronchoalveolar lavage fluid, semen, prostatic fluid, Cowper's fluid, pre-ejaculate fluid, female ejaculate, sweat, feces, tears, cyst fluid, pleural fluid, peritoneal fluid, pericardial fluid, lymph, chyme, chyle, bile, interstitial fluid, menstrual secretions, pus, sebum, vomit, vaginal fluid, mucosal secretions, stool water, pancreatic juice, nasal washings, bronchopulmonary aspirate, blastocyst fluid, or umbilical cord blood. [The present invention 1034] Any of the methods of claims 1028 to 1033, wherein the evaluation comprises determining the presence, level or status of a protein or nucleic acid for each biomarker, and optionally the nucleic acid comprises deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or a combination thereof. [This invention 1035] (a) the presence, level, or status of the protein is determined using immunohistochemistry (IHC), flow cytometry, immunoassays, antibodies or functional fragments thereof, aptamers, or any combination thereof; and / or (b) The method of the present invention 1034, wherein the presence, level or state of the nucleic acid is determined using polymerase chain reaction (PCR), in situ hybridization, amplification, hybridization, microarray, nucleic acid sequencing, dye terminator sequencing, pyrosequencing, next generation sequencing (NGS; high-throughput sequencing) or any combination thereof. [The present invention 1036] The method of the present invention 1035, wherein the nucleic acid state comprises a sequence, mutation, polymorphism, deletion, insertion, substitution, translocation, fusion, truncation, duplication, amplification, repeat, copy number, copy number variation (CNV; copy number alteration; CNA), or any combination thereof. [This invention 1037] The method of claim 1036, wherein the state of the nucleic acid comprises copy number. [The present invention 1038] 1037. The method of claim 1037, comprising the step of performing an assay to determine the copy number of all members of Group 1 (i.e., MYC, EP300, U2AF1, ASXL1, MAML2 and CNTRL) or genomic regions adjacent thereto. [This invention 1039] 1037. The method of claim 1037, comprising the step of performing an assay to determine the copy number of all members of Group 2 (i.e., MYC, EP300, U2AF1, ASXL1, MAML2, CNTRL, WRN, and CDX2) or genomic regions adjacent thereto. [The present invention 1040] 1037. The method of claim 1037, comprising the step of performing an assay to determine the copy number of all members of Group 3 (i.e., BCL9, PBX1, PRRX1, INHBA, YWHAE, GNAS, LHFPL6, FCRL4, HOXA11, AURKA, BIRC3, IKZF1, CASP8 and EP300) or genomic regions adjacent thereto. [The present invention 1041] 1037. The method of claim 1037, comprising the step of performing an assay to determine the copy number of all members of Group 4 (i.e., PBX1, BCL9, INHBA, PRRX1, YWHAE, GNAS, LHFPL6, FCRL4, AURKA, IKZF1, CASP8, PTEN and EP300) or genomic regions adjacent thereto. [The present invention 1042] 1037. The method of claim 1037, comprising the step of performing an assay to determine the copy number of all members of Group 5 (i.e., BCL9, PBX1, PRRX1, INHBA, GNAS, YWHAE, LHFPL6, FCRL4, PTEN, HOXA11, AURKA and BIRC3) or genomic regions adjacent thereto. [This invention 1043] 1037. The method of claim 1037, comprising the step of performing an assay to determine the copy number of all members of Group 6 (i.e., BCL9, PBX1, PRRX1, INHBA, and YWHAE) or genomic regions adjacent thereto. [This invention 1044] 1037. The method of claim 1037, comprising the step of performing an assay to determine the copy number of all members of Group 7 (i.e., BCL9, PBX1, GNAS, LHFPL6, CASP8, ASXL1, FH, CRKL, MLF1, TRRAP, AKT3, ACKR3, MSI2, PCM1 and MNX1) or genomic regions adjacent thereto. [This invention 1045] 1037. The method of claim 1037, comprising performing an assay to determine the copy number of all members of Group 8 (i.e., BX1, GNAS, AURKA, CASP8, ASXL1, CRKL, MLF1, GAS7, MN1, SOX10, TCL1A, LMO1, BRD3, SMARCA4, PER1, PAX7, SBDS, SEPT5, PDGFB, AKT2, TERT, KEAP1, ETV6, TOP1, TLX3, COX6C, NFIB, ARFRP1, ARID1A, MAP2K4, NFKBIA, WWTR1, ZNF217, IL2, NSD3, CREB1, BRIP1, SDC4, EWSR1, FLT3, FLT1, FAS, CCNE1, RUNX1T1 and EZR) or genomic regions adjacent thereto. [The present invention 1046] 1037. The method of claim 1037, comprising the step of performing an assay to determine the copy number of all members of Group 9 (i.e., BCL9, PBX1, PRRX1, INHBA, YWHAE, GNAS, LHFPL6, FCRL4, BIRC3, AURKA, and HOXA11) or genomic regions adjacent thereto. [This invention 1047] (a) at least one or all members of Group 1 and Group 2 or genomic regions adjacent thereto; (b) at least one or all members of Group 3 or genomic regions adjacent thereto; or (c) at least one or all members of Group 2, Group 6, Group 7, Group 8, and Group 9, or genomic regions adjacent thereto 1037. The method of claim 1037, comprising the step of performing an assay to determine the copy number of the gene. [This invention 1048] Any of the methods of claims 1037 to 1047, further comprising the step of comparing the copy number of the biomarker with a reference copy number (e.g., diploid) to identify biomarkers with copy number variation (CNV). [This invention 1049] The method of claim 1048, further comprising generating a molecular profile that identifies the gene or a region adjacent thereto that has the CNV. [The present invention 1050] 1049. The method of any of claims 1028 to 1049, wherein the presence or level of PTEN protein is determined, and optionally, the presence or level of PTEN protein is determined using immunohistochemistry (IHC). [This invention 1051] Any of the methods of claims 1028 to 1050, further comprising determining the level of proteins including TOPO1 and one or more mismatch repair proteins (e.g., MLH1, MSH2, MSH6 and PMS2), and optionally, wherein the presence or level of PTEN protein is determined using immunohistochemistry (IHC). [This invention 1052] 1052. The method of any one of claims 1050 to 1051, further comprising the step of comparing the level of the protein or proteins to a reference level for each of said protein or proteins. [This invention 1053] The method of claim 1052, further comprising generating a molecular profile that identifies proteins having levels that differ from a reference level, e.g., that differ significantly from said reference level. [This invention 1054] Any of the methods of claims 1028 to 1053, further comprising selecting a treatment of likely benefit based on the evaluated biomarkers, optionally wherein the treatment comprises 5-fluorouracil / leucovorin in combination with oxaliplatin (FOLFOX) or an alternative treatment thereto, and optionally wherein the alternative treatment comprises 5-fluorouracil / leucovorin in combination with irinotecan (FOLFIRI). [This invention 1055] The process of selecting treatments with promising benefits is (a) the determined copy number of any of 1037 to 1047 of the present invention; and / or (b) Molecular profile of 1049 or 1053 of the present invention The method of the present invention 1054 is based on the above. [The present invention 1056] The method of any one of claims 1037 to 1047, wherein the step of selecting a treatment of likely benefit based on the determined copy number comprises use of a voting module. [This invention 1057] The method of invention 1056, wherein the voting module is any one of inventions 1014 to 1025. [This invention 1058] The method of any one of claims 1056 to 1057, wherein the voting module includes the use of at least one random forest model. [This invention 1059] Any of the methods of claims 1056 to 1058, wherein using a voting module comprises applying a machine learning classification model to the copy numbers obtained for each of group 2, group 6, group 7, group 8, and group 9, optionally wherein each machine learning classification model is a random forest model, and optionally wherein the random forest model is one described in Table 10. [The present invention 1060] The method of any of claims 1054 to 1059, wherein the subject has not been previously treated with a therapy of potential benefit. [The present invention 1061] The method of any one of claims 1028 to 1060, wherein the cancer comprises metastatic cancer, recurrent cancer, or a combination thereof. [This invention 1062] The method of any of claims 1028 to 1061, wherein the subject has not previously received treatment for cancer. [The present invention 1063] The method of any of claims 1054 to 1062, further comprising administering to the subject a treatment of likely benefit. [The present invention 1064] The method of claim 1063, wherein progression-free survival (PFS), disease-free survival (DFS) or life span is prolonged by administration of said treatment. [This invention 1065] Cancers include acute lymphoblastic leukemia; acute myeloid leukemia; adrenocortical carcinoma; AIDS-related cancer; AIDS-related lymphoma; anal cancer; appendix cancer; astrocytoma; atypical teratoid / rhabdoid tumor; basal cell carcinoma; bladder cancer; brain stem glioma; brain tumor, brain stem glioma, central nervous system atypical teratoid / rhabdoid tumor, central nervous system embryonal tumor, astrocytoma, craniopharyngioma, ependymomas, ependymomas, medulloblastomas, medulloepithelioma, intermediate pineal parenchymal tumor, supratentorial primitive neuroectodermal tumor and pineoblastoma; breast cancer; bronchial tumor; Burkitt's disease Lymphoma; Cancer of Unknown Primary Source (CUP); Carcinoid Tumor; Carcinoma of Unknown Primary Source; Central Nervous System Atypical Teratoid / Rhabdoid Tumor; Central Nervous System Embryonal Tumor; Cervical Cancer; Childhood Cancer; Chordoma; Chronic Lymphocytic Leukemia; Chronic Myeloid Leukemia; Chronic Myeloproliferative Disorder; Colon Cancer; Colorectal Cancer; Craniopharyngioma; Cutaneous T-Cell Lymphoma; Endocrine Pancreatic Islet Cell Tumor; Endometrial Cancer; Ependymoblastoma; Ependymoma; Esophageal Cancer; Nasal Neuroblastoma; Ewing Sarcoma; Extracranial Germ Cell Tumor; Extragonadal Germ Cell Tumor; Extrahepatic Bile Duct Cancer; Gallbladder Cancer; Gastric Cancer (stomach) cancer); gastrointestinal carcinoid tumor; gastrointestinal stromal cell tumor; gastrointestinal stromal tumor (GIST); gestational trophoblastic tumor; glioma; hairy cell leukemia; head and neck cancer; cardiac cancer; Hodgkin's lymphoma; hypopharyngeal cancer; intraocular melanoma; pancreatic islet tumor; Kaposi's sarcoma; kidney cancer; Langerhans cell histiocytosis; laryngeal cancer; lip cancer; liver cancer; malignant fibrous histiocytoma bone cancer; medulloblastoma; medulloepithelioma; melanoma; Merkel cell carcinoma; Merkel cell skin cancer; mesothelioma; metastatic squamous neck cancer of unknown primary; oral cancer; multiple endocrine neoplasia syndrome; multiple myeloma; multiple myeloma / plasma cell neoplasm; mycosis fungoides; myelodysplastic syndrome; myeloproliferative neoplasm; nasal cavity cancer; nasopharyngeal carcinoma; neuroblastoma; non-Hodgkin's lymphoma; non-melanoma skin cancer; non-small cell lung cancer; oral cancer cancer); oral cavity cancer; oropharyngeal cancer; osteosarcoma; other brain and spinal cord tumors; ovarian cancer; ovarian epithelial cancer; ovarian germ cell tumor; ovarian low malignant potential tumor; pancreatic cancer; papillomatosis; sinonasal cancer; parathyroid cancer; pelvic cancer; penile cancer; pharyngeal cancer; intermediate pineal parenchymal tumor; pineoblastoma; pituitary tumor; plasma cell neoplasm / multiple myeloma; pleuropulmonary blastoma; primary central nervous system (CNS) lymphoma; primary hepatocellular carcinoma;Any of the methods of claims 1028-1064, including prostate cancer; rectal cancer; kidney cancer; renal cell (kidney) cancer; renal cell carcinoma; airway cancer; retinoblastoma; rhabdomyosarcoma; salivary gland cancer; Sézary syndrome; small cell lung cancer; small intestine cancer; soft tissue sarcoma; squamous cell carcinoma; cervical squamous cell carcinoma; stomach (gastric) cancer; supratentorial primitive neuroectodermal tumor; T-cell lymphoma; testicular cancer; throat cancer; thymic carcinoma; thymoma; thyroid cancer; transitional cell carcinoma; transitional cell carcinoma of the renal pelvis and ureter; trophoblastic tumor; ureteral cancer; urethral cancer; uterine cancer; uterine sarcoma; vaginal cancer; vulvar cancer; Waldenstrom's macroglobulinemia; or Wilms' tumor. [The present invention 1066] Cancers include acute myeloid leukemia (AML), breast cancer, bile duct cancer, colorectal adenocarcinoma, extrahepatic bile duct adenocarcinoma, female genital malignancies, gastric adenocarcinoma, gastroesophageal adenocarcinoma, gastrointestinal stromal tumor (GIST), glioblastoma, head and neck squamous cell carcinoma, leukemia, hepatocellular carcinoma, low-grade glioma, lung bronchoalveolar carcinoma (BAC), non-small cell lung cancer (NSCLC), small cell lung cancer (SCLC), lymphoma, male genital malignancies, and solitary fibrous pleural malignancies. 10. The method of any of claims 1028 to 1064, including tumors (MSFT), melanoma, multiple myeloma, neuroendocrine tumors, nodular diffuse large B-cell lymphoma, non-epithelial ovarian cancer (non-EOC), ovarian surface epithelial carcinoma, pancreatic adenocarcinoma, pituitary carcinoma, oligodendroglioma, prostate adenocarcinoma, retroperitoneal or peritoneal carcinoma, retroperitoneal or peritoneal sarcoma, small intestine malignant tumor, soft tissue tumor, thymic carcinoma, thyroid carcinoma, or uveal melanoma. [This invention 1067] 1065. The method of any one of claims 1028 to 1064, wherein the cancer comprises colorectal cancer. [The present invention 1068] 1. A method of selecting a treatment for a subject with colorectal cancer, comprising: obtaining a biological sample comprising cells derived from colorectal cancer; performing next-generation sequencing on genomic DNA from the biological sample; (a) Group 2, including 1, 2, 3, 4, 5, 6, 7, or all 8 of MYC, EP300, U2AF1, ASXL1, MAML2, CNTRL, WRN, and CDX2; (b) group 6, including one, two, three, four, or all five of BCL9, PBX1, PRRX1, INHBA, and YWHAE; (c) group 7, including all 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 of BCL9, PBX1, GNAS, LHFPL6, CASP8, ASXL1, FH, CRKL, MLF1, TRRAP, AKT3, ACKR3, MSI2, PCM1, and MNX1; (d)BX1, GNAS, AURKA, CASP8, ASXL1, CRKL, MLF1, GAS7, MN1, SOX10, TCL1A, LMO1, BRD3, SMARCA4, PER1, PAX7, SBDS, SEPT5, PDGFB, AKT2, TERT, KEAP1, ETV6, TOP1, TLX3, COX6C, NFIB, ARFRP1, ARID1A, MAP2K4, NFKBIA, WWTR1, ZNF217, IL2, NSD3, CR Group 8, including all 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, or 45 of EB1, BRIP1, SDC4, EWSR1, FLT3, FLT1, FAS, CCNE1, RUNX1T1, and EZR; and (e) Group 9, including 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or all 11 of BCL9, PBX1, PRRX1, INHBA, YWHAE, GNAS, LHFPL6, FCRL4, BIRC3, AURKA, and HOXA11. determining the copy number for each of the genes or their adjacent genomic regions; applying a machine learning classification model to the copy numbers obtained for each of group 2, group 6, group 7, group 8, and group 9, optionally wherein each machine learning classification model is a random forest model, optionally wherein the random forest model is one described in Table 10; Obtaining from each machine learning classification model an indication of whether the subject is likely to benefit from treatment with 5-fluorouracil / leucovorin in combination with oxaliplatin (FOLFOX); and selecting FOLFOX if a majority of the machine learning classification models indicate that the subject is likely to benefit from the treatment, and selecting an alternative treatment to FOLFOX if a majority of the machine learning classification models indicate that the subject is unlikely to benefit from FOLFOX, optionally wherein the alternative treatment is 5-fluorouracil / leucovorin in combination with irinotecan (FOLFIRI). The method comprising: [This invention 1069] The method of claim 1068, further comprising administering the selected treatment to the subject. [The present invention 1070] A method for generating a molecular profiling report, comprising the step of creating a report summarizing the results of carrying out any of the methods of the present inventions 1028 to 1069. [This invention 1071] The report, (a) a treatment of any of the potential benefits of inventions 1054 to 1059; or (b) the selected treatment of invention 1068 or 1069 The method of the present invention 1070, comprising: [This invention 1072] The method of invention 1070 or 1071, wherein the report is computer-generated; a printed report or computer file; or accessible via a web portal. [This invention 1073] 1. A system for identifying a treatment for cancer in a subject, comprising: (a) at least one host server; (b) at least one user interface for accessing said at least one host server to access and input data; (c) at least one processor for processing input data; (d) the processed data; (1) Access the results of analyzing any one of the biological samples of the present invention 1028 to 1069, and (2) Determine the treatment of any one of inventions 1054 to 1059 with promising benefits or the selected treatment of inventions 1068 or 1069 With instructions for at least one memory coupled to the processor for storing (e) at least one display for displaying a cancer treatment that is FOLFOX or an alternative thereto, e.g., FOLFIRI; The system comprising: [This invention 1074] The system of the present invention 1073, wherein at least one display includes a report containing the results of analyzing the biological sample and treatments that have promising benefits in treating the cancer or have been selected for treating the cancer. Other features and advantages of the present invention will become apparent from the following detailed description and drawings, and from the appended claims. [Brief explanation of the drawings]
[0084] [Figure 1A] FIG. 1 is a block diagram of an example of a prior art system for training a machine learning model. [Figure 1B] FIG. 1 is a block diagram of a system for generating a training data structure for training a machine learning model to predict the efficacy of a treatment for a disease or disorder in a subject having a particular set of biomarkers. [Figure 1C] FIG. 1 is a block diagram of a system for using a trained machine learning model to predict the efficacy of a treatment for a disease or disorder in a subject having a particular set of biomarkers. [Figure 1D] 1 is a flowchart of a process for generating training data for training a machine learning model to predict the efficacy of a treatment for a disease or disorder in a subject having a particular set of biomarkers. [Figure 1E]1 is a flowchart of a process for using a trained machine learning model to predict the efficacy of a treatment for a disease or disorder in a subject having a particular set of biomarkers. [Figure 1F] FIG. 1 is a block diagram of a system for predicting the efficacy of a treatment for a disease or disorder in a subject having a particular set of biomarkers by interpreting outputs generated by multiple machine learning models using a voting unit. [Figure 1G] FIG. 6 is a block diagram of system components that can be used to implement the systems of FIGS. [Figure 1H] FIG. 1 shows a block diagram of an exemplary embodiment of a system for determining personalized medical interventions for cancer that utilizes molecular profiling of patient biospecimens. [Figure 2A] A method for determining personalized medicine interventions for cancer that utilizes molecular profiling of patient biospecimens. [Figure 2B] A method for identifying a signature or molecular profile that can be used to predict benefit from a therapy. [Figure 2C] 10 is a flowchart of an exemplary embodiment of an alternative version of (B). [Figure 3A] Paired hazard ratio graphs showing model performance using CNV profiling of eight markers in the setting of treatment with FOLFOX. CNA = copy number alteration. The eight markers were MYC, EP300, U2AF1, ASXL1, MAML2, CNTRL, WRN, and CDX2. [Figure 3B] Paired hazard ratio graph showing model performance using CNV profiling of eight markers in the setting of treatment with FOLFIRI. CNA = copy number alteration. The eight markers were MYC, EP300, U2AF1, ASXL1, MAML2, CNTRL, WRN, and CDX2. [Figure 3C]
[0023] Figure 1 is a paired hazard ratio graph showing model performance using CNV profiling of six markers in the setting of treatment with FOLFOX. The six markers were MYC, EP300, U2AF1, ASXL1, MAML2, and CNTRL. [Figure 3D] 1 is a paired hazard ratio graph showing model performance using CNV profiling of six markers in the setting of treatment with FOLFIRI. The six markers were MYC, EP300, U2AF1, ASXL1, MAML2, and CNTRL. [Figure 3E] 3A-B show exemplary random forest decision trees for the eight marker signatures shown in FIGS. [Figure 4A] We demonstrate the development of a biosignature to predict benefit from FOLFOX regimens in patients with metastatic colorectal cancer. [Figure 4B] We demonstrate the development of a biosignature to predict benefit from FOLFOX regimens in patients with metastatic colorectal cancer. [Figure 4C] We demonstrate the development of a biosignature to predict benefit from FOLFOX regimens in patients with metastatic colorectal cancer. [Figure 4D] We demonstrate the development of a biosignature to predict benefit from FOLFOX regimens in patients with metastatic colorectal cancer. [Figure 4E] We demonstrate the development of a biosignature to predict benefit from FOLFOX regimens in patients with metastatic colorectal cancer. [Figure 4F] We demonstrate the development of a biosignature to predict benefit from FOLFOX regimens in patients with metastatic colorectal cancer. [Figure 4G] We demonstrate the development of a biosignature to predict benefit from FOLFOX regimens in patients with metastatic colorectal cancer. [Figure 4H] We demonstrate the development of a biosignature to predict benefit from FOLFOX regimens in patients with metastatic colorectal cancer. [Figure 4I]We demonstrate the development of a biosignature to predict benefit from FOLFOX regimens in patients with metastatic colorectal cancer. [Figure 4J] We demonstrate the development of a biosignature to predict benefit from FOLFOX regimens in patients with metastatic colorectal cancer. [Figure 4K] We demonstrate the development of a biosignature to predict benefit from FOLFOX regimens in patients with metastatic colorectal cancer. [Figure 4L] We demonstrate the development of a biosignature to predict benefit from FOLFOX regimens in patients with metastatic colorectal cancer. [Figure 4M] We demonstrate the development of a biosignature to predict benefit from FOLFOX regimens in patients with metastatic colorectal cancer. [Figure 4N] We demonstrate the development of a biosignature to predict benefit from FOLFOX regimens in patients with metastatic colorectal cancer. [Figure 4O] We demonstrate the development of a biosignature to predict benefit from FOLFOX regimens in patients with metastatic colorectal cancer. [Figure 5A] We demonstrate the development of a biosignature to predict benefit from FOLFOX regimens in patients with colorectal cancer. [Figure 5B] We demonstrate the development of a biosignature to predict benefit from FOLFOX regimens in patients with colorectal cancer. [Figure 5C] We demonstrate the development of a biosignature to predict benefit from FOLFOX regimens in patients with colorectal cancer. DETAILED DESCRIPTION OF THE INVENTION
[0085] Detailed Description Described herein are methods and systems for identifying therapeutic agents for use in personalized treatment by using molecular profiling, including systems, methods, devices and computer programs for training machine learning models and then using the trained machine learning models to predict the effectiveness of treatment for a subject's disease or disorder.In some embodiments, the system can include one or more computer programs on one or more computers at one or more locations, configured for use in, for example, the methods described herein.
[0086] Aspects of the present disclosure relate to a system for generating a set of one or more training data structures that can be used to train a machine learning model to provide various classifications, such as characterizing the phenotype of a biological sample. Phenotypic characterization can include providing a diagnosis, prognosis, theranostic, or other related classification. For example, the classification can be a classification that predicts the disease state of a subject having a particular set of biomarkers or the effectiveness of a treatment for a disease or disorder. Once trained, the trained machine learning model can be used to process input data provided by the system and make a prediction based on the processed input data. The input data can include a set of features associated with the subject, such data representing one or more subject biomarkers, and the data representing a disease or disorder. In some embodiments, the input data can further include features representing a proposed type of treatment, and a prediction can be made that describes the subject's likely response to the treatment. The prediction can include data output by the machine learning model based on processing of a particular set of features provided to the machine learning model as input. The data can optionally include data representing one or more subject biomarkers, data representing a disease or disorder, and data representing a proposed type of treatment.
[0087] An innovative aspect of the present disclosure involves the extraction of specific data from the incoming data stream for use in generating a training data structure. Of critical importance is the selection of a specific set of one or more biomarkers to include in the training data structure. This is because the presence, absence, or state of a specific biomarker can indicate a desired classification. For example, specific biomarkers can be selected to determine whether a treatment for a disease or disorder is effective or ineffective. By way of illustration, in this disclosure, applicants present a specific set of biomarkers that, when used in training a machine learning model, produces a trained model that can predict treatment efficacy more accurately than a different set of biomarkers. See Examples 2-4.
[0088] The system is configured to obtain output data generated by the trained machine learning model based on processing the data. In various embodiments, the data includes biological data representing one or more biomarkers, data representing a disease or disorder, and data representing a treatment type. The system can then predict the effectiveness of a treatment for a subject with a specific set of biomarkers. In some embodiments, the disease or disorder may include a type of cancer, and the treatment for the subject may include one or more therapeutic agents, such as small molecule drugs, biologics, and various combinations thereof. In this setting, the output of the trained machine learning model, generated based on processing the input data including the set of biomarkers, the disease or disorder, and the treatment type, includes data representing the level of responsiveness of the subject to a treatment for the disease or disorder.
[0089] In some embodiments, the output data generated by the trained machine learning model may include a probability of a desired classification. Illustratively, such a probability may be the probability that a subject will respond favorably to a treatment for a disease or disorder. In other embodiments, the output data may include any output data generated by the trained machine learning model based on processing the trained machine learning model of input data. In some aspects, the input data includes a set of biomarkers, data representing a disease or disorder, and data representing a treatment type.
[0090] In some embodiments, the training data structure generated by the present disclosure may include multiple training data structures, each including a field representing a feature vector corresponding to a particular training sample. The feature vector includes a set of features derived from and representative of the training sample. The training sample may include, for example, one or more biomarkers of a subject, a disease or disorder of the subject, and a proposed treatment for the disease or disorder. The training data structures are flexible because each training data structure may be assigned a weight representing each feature of the feature vector. Thus, each training data structure of the multiple training data structures can be specifically configured to enable a specific inference to be made by the machine learning model during training.
[0091] Consider a non-limiting example in which a model is trained to make predictions about the likely benefit of a particular treatment for a disease or disorder. As a result, the novel training data structure generated in accordance with the present specification is designed to improve the performance of a machine learning model, since it can be used to train the machine learning model to predict the effectiveness of a treatment for a disease or disorder in subjects with a particular set of biomarkers. Illustratively, a machine learning model that was unable to make predictions about the effectiveness of a treatment for a disease or disorder in subjects with a particular set of biomarkers before being trained using the training data structure, system, and operations described herein can learn to make predictions about the effectiveness of a treatment for the subject's disease or disorder by being trained using the training data structure, system, and operations described herein. Thus, this process takes an otherwise general-purpose machine learning model and transforms it into a specialized computer for performing the specific task of predicting the effectiveness of a treatment for a disease or disorder in subjects with a particular set of biomarkers.
[0092] FIG. 1A is a block diagram of an example prior art system 100 for training a machine learning model 110. In some embodiments, the machine learning model may be, for example, a support vector machine. Alternatively, the machine learning model may include a neural network model, a linear regression model, a random forest model, a logistic regression model, a naive Bayes model, a quadratic discriminant analysis model, a k-nearest neighbor model, a support vector machine, etc. The machine learning model training system 100 may be implemented as a computer program on one or more computers at one or more locations, on which the systems, components, and techniques described below can be implemented. The machine learning model training system 100 trains the machine learning model 110 using training data items from a database (or dataset) 120 of training data items. The training data items may include multiple feature vectors. Each training vector may include multiple values, each corresponding to a particular feature of the training sample that the training vector represents. The training features are sometimes referred to as independent variables. Additionally, the system 100 maintains a respective weight for each feature included in the feature vector.
[0093] The machine learning model 110 is configured to receive input training data items 122 and process the input training data items 122 to generate output 118. The input training data items may include multiple features (or independent variables "X") and training labels (or dependent variables "Y"). The machine learning model may be trained using the training items and, when trained, is capable of predicting X=f(Y).
[0094] To enable the machine learning model 110 to generate accurate outputs for received data items, the machine learning model training system 100 may train the machine learning model 110 to adjust the values of the parameters of the machine learning model 110, e.g., to determine trained values of the parameters from initial values. These parameters derived from the training process may include weights that can be used during the prediction stage using the fully trained machine learning model 110.
[0095] When training the machine learning model 110, the machine learning model training system 100 uses training data items stored in a database (dataset) 120 of labeled training data items. The database 120 stores a set of training data items, with each training data item in the set of training data items associated with a respective label. Generally, the label for a training data item identifies the correct classification (or prediction) for the training data item, i.e., the classification that should be identified as the classification of the training data item by the output value generated by the machine learning model 110. Referring to FIG. 1A, a training data item 122 may be associated with a training label 122a.
[0096] The machine learning model training system 100 trains the machine learning model 110 to optimize an objective function. Optimizing the objective function may include, for example, minimizing a loss function 130. In general, the loss function 130 is a function that depends on (i) the output 118 produced by the machine learning model 110 by processing a given training data item 122, and (ii) the label 122 a for the training data item 122, i.e., the target output that the machine learning model 110 should have produced by processing the training data item 122.
[0097] A conventional machine learning model training system 100 can train a machine learning model 110 to minimize a (cumulative) loss function 130 by performing multiple iterations of conventional machine learning model training techniques, such as hinge loss, stochastic gradient methods, and stochastic gradient descent with backpropagation, on training data items from a database 120 to iteratively adjust the values of parameters of the machine learning model 110. The fully trained machine learning model 110 can then be deployed as a predictive model that can be used to make predictions based on unlabeled input data.
[0098] FIG. 1B is a block diagram of a system 200 for generating a training data structure for training a machine learning model to predict the efficacy of a treatment for a disease or disorder in a subject having a particular set of biomarkers.
[0099] The system 200 includes two or more distributed computers 210, 310, a network 230, and an application server 240. The application server 240 includes an extraction unit 242, a memory unit 244, a vector generation unit 250, and a machine learning model 270. The machine learning model 270 may include one or more of a vector support machine, a neural network model, a linear regression model, a random forest model, a logistic regression model, a naive Bayes model, a quadratic discriminant analysis model, a k-nearest neighbor model, a support vector machine, etc. Each distributed computer 210, 310 may include a smartphone, a tablet computer, a laptop computer, a desktop computer, etc. Alternatively, the distributed computers 210, 310 may include a server computer that receives data input by one or more terminals 205, 305, respectively. The terminal computers 205, 305 may include any user device, including a smartphone, a tablet computer, a laptop computer, a desktop computer, etc. The network 230 may include one or more networks 230, such as a LAN, a WAN, a wired Ethernet network, a wireless network, a cellular network, the Internet, or any combination thereof.
[0100] Application server 240 is configured to obtain or otherwise receive data records 220, 222, 224, 320 provided by one or more distributed computers, such as first distributed computer 210 and second distributed computer 310, using network 230. In some embodiments, each distributed computer 210, 310 may provide different types of data records 220, 222, 224, 320. For example, first distributed computer 210 may provide biomarker data records 220, 222, 224 representing subject biomarkers, and second distributed computer 310 may provide outcome data 320 representing subject outcome data obtained from outcome database 312.
[0101] Biomarker data records 220, 222, and 224 may include any type of biomarker data describing a subject's biometric attributes. Illustratively, the example in FIG. 1B shows biomarker data records as including data records representing DNA biomarkers 220, protein biomarkers 222, and RNA biomarkers 224. Each of these biomarker data records may include a data structure having fields structuring information 220a, 222a, and 224a describing a subject's biomarkers, e.g., a subject's DNA biomarker 220a, protein biomarker 222a, or RNA biomarker 224a. However, the present disclosure need not be so limited. For example, biomarker data records 220, 222, and 224 may include next-generation sequencing data, such as DNA alterations. Such next-generation sequencing data may include single variants, insertions and deletions, substitutions, translocations, fusions, truncations, duplications, amplifications, losses, copy number, repeats, total genetic dose, microsatellite instability, and the like. Alternatively or additionally, biomarker data records 220, 222, 224 may also include in situ hybridization data, such as DNA copies. Such in situ hybridization data may include gene copies, gene translocations, etc. Alternatively or additionally, biomarker data records 220, 222, 224 may include RNA data, such as gene expression or gene fusions, including whole transcriptome sequencing. Alternatively or additionally, biomarker data records 220, 222, 224 may include protein expression data, such as obtained using immunohistochemistry (IHC). Alternatively or additionally, biomarker data records 220, 222, 224 may include ADAPT data, such as complex numbers.
[0102] In some embodiments, the set of one or more biomarkers includes one or more biomarkers listed in any one of Tables 2-8. However, the disclosure need not be so limited, and other types of biomarkers may be used instead. For example, biomarker data may be obtained by whole-exome sequencing, whole-transcriptome sequencing, or a combination thereof.
[0103] The outcome data record 320 may describe the outcome of a treatment for a subject. For example, the outcome data record 320 obtained from the outcome database 312 may include one or more data structures with fields structuring subject data attributes, such as a disease or disorder 320a, a treatment the subject received for the disease or disorder 320a, a treatment result 320a, or a combination of both. In addition, the outcome data record 320 may also include fields structuring data attributes describing details of the treatment and the subject's response to the treatment. An example disease or disorder may include, for example, certain types of cancer. The type of treatment may include, for example, the type of drug, biologic, or other treatment the subject received for the disease or disorder included in the outcome data record 320. The treatment result may include data describing the subject's outcome of the treatment regimen, such as benefit, moderate benefit, or no benefit. In some embodiments, the treatment result may include the type of cancerous tumor at the end of treatment, such as the amount the tumor shrank, the overall size of the tumor after treatment, etc. Alternatively or additionally, the treatment results may include counts or ratios of white blood cells, red blood cells, etc. Treatment details may include dosages, e.g., amount of medication taken, drug regimen, number of missed doses, etc. Thus, while the example of FIG. 1B shows that outcome data may include a disease or disorder, a treatment, and a treatment result, the outcome data may include other types of information as described herein. Moreover, the outcome data need not be limited to human “patients.” Instead, the outcome data records 220, 222, 224 and biometric data record 320 may be associated with any desired subject, including any non-human organism.
[0104] In some embodiments, each of the data records 220, 222, 224, 320 may include keyed data that allows the data records from each distributed computer to be correlated by the application server 240. The keyed data may include, for example, data representing a subject identifier. The subject identifier may include any form of data that identifies a subject that can associate the subject's biomarkers with the subject's outcome data.
[0105] The first distributed computer 210 may provide 208 the biomarker data records 220, 222, 224 to the application server 240. The second distributed computer 310 may provide 210 the outcome data records 320 to the application server 240. The application server 240 may provide the biomarker data records 220 and the outcome data records 220, 222, 224 to the extraction unit 242.
[0106] The extraction unit 242 may process the received biomarker data 220, 222, 224 and outcome data records 320 to extract data 220a-1, 222a-1, 224a-1, 320a-1, 320a-2, 320a-3 that can be used to train a machine learning model. For example, the extraction unit 242 may obtain data structured by fields of the data structure of the biometric data records 220, 222, 224, data structured by fields of the data structure of the outcome data record 320, or a combination thereof. The extraction unit 242 may perform one or more information extraction algorithms, such as keyed data extraction, pattern matching, natural language processing, etc., to identify and obtain the data 220a-1, 222a-1, 224a-1, 320a-1, 320a-2, 320a-3 from the biometric data records 220, 222, 224 and outcome data record 320, respectively. The extraction unit 242 may provide the extracted data to a memory unit 244. The extracted data may be stored in the memory unit 244, such as a flash memory (as opposed to a hard disk), to improve data access time and reduce latency in accessing the extracted data to improve system performance. In some embodiments, the extracted data may be stored in the memory unit 244 as an in-memory data grid.
[0107] More specifically, the extraction unit 242 may be configured to filter the portions of the biomarker data records 220, 222, 224 and outcome data records 320 used to generate the input data structure 260 for processing by the machine learning model 270 from the portions of the outcome data records 320 used as labels for the generated input data structure 260. Such filtering includes the extraction unit 242 separating the biomarker data and a first portion of the outcome data including the disease or disorder, treatment, treatment details, or a combination thereof from the treatment results. The application server 240 can then generate the input data structure 260 using the biomarker data 220a-1, 222a-1, 224a-1, 320a-1, 320a-2 and the first portion of the outcome data including the disease or disorder 320a-1, treatment 320a-2, treatment details (not shown in FIG. 1B ), or a combination thereof. Additionally, the application server 240 may use a second portion of the outcome data describing the treatment results 320a-3 as a label for the generated data structure.
[0108] The application server 240 may process the extracted data stored in the memory unit 244 and correlate the biomarker data 220a-1, 222a-1, 224a-1 extracted from the biomarker data records 220, 222, 224 with a first portion of the outcome data 320a-1, 320a-2. The purpose of this correlation is to cluster the biomarker data with the outcome data such that the subject's outcome data is clustered with the subject's biomarker data. In some embodiments, the correlation of the biomarker data with the first portion of the outcome data may be based on keyed data associated with each of the biomarker data records 220, 222, 224 and the outcome data record 320. For example, the keyed data may include a subject identifier.
[0109] Application server 240 provides the extracted biomarker data 220a-1, 222a-1, 224a-1 and the extracted first portions of outcome data 320a-1, 320a-2 as input to vector generation unit 250. Vector generation unit 250 is used to generate a data structure based on the extracted biomarker data 220a-1, 222a-1, 224a-1 and the extracted first portions of outcome data 320a-1, 320a-2. The generated data structure is feature vector 260 that includes multiple values that numerically represent the extracted first portions of the extracted biomarker data 220a-1, 222a-1, 224a-1 and the outcome data 320a-1, 320a-2. Feature vector 260 may include a field for each type of biomarker and each type of outcome data. For example, feature vector 260 may include one or more fields corresponding to (i) one or more types of next-generation sequencing data, such as single variants, insertions and deletions, substitutions, translocations, fusions, truncations, duplications, amplifications, losses, copy number, repeats, total genetic mutation burden, microsatellite instability; (ii) one or more types of in situ hybridization data, such as DNA copies, gene copies, gene translocations; (iii) one or more types of RNA data, such as gene expression or gene fusions; (iv) one or more types of protein data, such as obtained using immunohistochemistry; (v) one or more types of ADAPT data, such as complex numbers; and (vi) one or more types of outcome data, such as disease or disorder, treatment type, details of each treatment type, etc.
[0110] Vector generation unit 250 is configured to assign a weight to each field of feature vector 260 that indicates the extent to which the extracted first portions of extracted biomarker data 220a-1, 222a-1, 224a-1 and outcome data 320a-1, 320a-2 contain the data represented by each field. In one embodiment, for example, vector generation unit 250 may assign a "1" to each field of the feature vector that corresponds to a feature found in the extracted first portions of extracted biomarker data 220a-1, 222a-1, 224a-1 and outcome data 320a-1, 320a-2. In such an embodiment, vector generation unit 250 may also assign a "0" to each field of the feature vector that corresponds to a feature not found in the extracted first portions of extracted biomarker data 220a-1, 222a-1, 224a-1 and outcome data 320a-1, 320a-2, for example. The output of the vector generation unit 250 may include a data structure, such as a feature vector 260, that can be used to train a machine learning model 270.
[0111] The application server 240 can label the training feature vector 260. Specifically, the application server can use the extracted second portion of the patient outcome data 320a-3 to label the generated feature vector 260 with a treatment outcome 320a-3. The label of the training feature vector 260 generated based on the treatment outcome 320a-3 can provide an indication of the effectiveness of the treatment 320a-2 for the disease or disorder 320a-1 of interest as determined by a particular set of biomarkers 220a-1, 222a-1, 224a-1 (each described in the training data structure 260).
[0112] The application server 240 may train the machine learning model 270 by providing the feature vector 260 as input to the machine learning model 270. The machine learning model 270 may process the generated feature vector 260 and generate an output 272. The application server 240 may use a loss function 280 to determine the amount of error between the output 272 of the machine learning model 280 and the values specified by the training labels (generated based on a second portion of the extracted patient outcome data that describes the treatment results 320a-3). The output 282 of the loss function 280 may be used to adjust the parameters of the machine learning model 282.
[0113] In some embodiments, adjusting the parameters of the machine learning model 270 may include manual tuning of the machine learning model parameters. Alternatively, in some embodiments, the parameters of the machine learning model 270 may be tuned automatically by one or more algorithms executed by the application server 242.
[0114] The application server 240 may perform multiple iterations of the process described above with reference to FIG. 1B for each outcome data record 320 stored in the outcome database corresponding to the subject's set of biomarker data. This may include hundreds of iterations, thousands of iterations, tens of thousands of iterations, hundreds of thousands of iterations, millions of iterations, or more iterations until each outcome data record 320 having a corresponding set of biomarker data for the subject stored in the outcome database 312 is exhausted, until the machine learning model 270 is trained to within a particular error range, or a combination thereof. The machine learning model 270 is trained within a particular error range, for example, when the machine learning model 270 can predict the efficacy of a treatment for a subject having biomarker data based on the set of unlabeled biomarker data, disease or disorder data, and treatment data. Efficacy may include, for example, a probability, a general indicator of whether a treatment will be successful or unsuccessful, etc.
[0115] FIG. 1C is a block diagram of a system for using a trained machine learning model to predict the efficacy of a treatment for a disease or disorder in a subject having a particular set of biomarkers.
[0116] The machine learning model 370 includes a machine learning model trained using the process described with reference to the system of FIG. 1B above. The trained machine learning model 370 can predict a level of effectiveness of a treatment in treating a disease or disorder in a subject having biomarkers based on an input feature vector representing one or more sets of biomarkers, a disease or disorder, and a treatment. In some embodiments, the "treatment" may include medication, treatment details (e.g., dosage, regimen, missed doses, etc.), or any combination thereof.
[0117] The application server 240 hosting the machine learning model 370 is configured to receive unlabeled biomarker data records 320, 322, 324. The biomarker data records 320, 322, 324 include one or more data structures having fields that structure data representing one or more particular biomarkers, such as DNA biomarkers 320a, protein biomarkers 322a, RNA biomarkers 324a, or any combination thereof. As mentioned above, the received biomarker data records may include types of biomarkers not represented by FIG. 1C, such as (i) one or more types of next-generation sequencing data, e.g., single variants, insertions and deletions, substitutions, translocations, fusions, truncations, duplications, amplifications, losses, copy number, repeats, total gene mutation dose, microsatellite instability; (ii) one or more types of in situ hybridization data, e.g., DNA copies, gene copies, gene translocations; (iii) one or more types of RNA data, e.g., gene expression or gene fusions; (iv) one or more types of protein data, such as obtained using immunohistochemistry; or (v) one or more types of ADAPT data, e.g., complex numbers.
[0118] The application server 240 hosting the machine learning model 370 is also configured to receive data representing suggested treatment data 422a for a disease or disorder described by the disease or disorder data 420a for subjects having the biomarkers represented by the received biomarker data records 320, 322, 324. The suggested treatment data 422a for the disease or disorder 422a is also unlabeled and is merely a suggestion for treating a subject having the biomarkers represented by the biomarker data records 320, 322, 324.
[0119] In some embodiments, disease or disorder data 420a and proposed treatment 422a are provided (305) by terminal 405 via network 230, and biomarker data is obtained from second distributed computer 310. Biomarker data may be derived from laboratory equipment used to perform various assays. In other embodiments, disease or disorder data 420a, proposed treatment 422a, and biomarker data 320, 322, 324 may each be received from terminal 405. For example, terminal 405 may be a user device of a physician, an employee working for a physician, or a physician's representative, or other person who inputs data representing a disease or disorder, data representing a proposed treatment, and data representing one or more biomarkers of a subject having the disease or disorder. In some embodiments, treatment data 422 may include a data structure structuring fields of data representing a proposed treatment described by medication name. In other embodiments, treatment data 422 may include a data structure structuring fields of data representing more complex treatment data, such as dosage, dosing regimen, allowed number of missed doses, etc.
[0120] The application server 240 receives the biomarker data records 320, 322, 324, the disease or disorder data 420, and the treatment data 422. The application server 240 provides the biomarker data records 320, 322, 324, the disease or disorder data 420, and the treatment data 422 to the extraction unit 242, which is configured to extract (i) specific biomarker data, e.g., DNA biomarker data 320a-1, protein expression data 322a-1, 324a-1, (ii) disease or disorder data 420a-1, and (iii) suggested treatment data 420a-1 from fields of the biomarker data records 320, 322, 324 and the outcome data records 420, 422. In some embodiments, the extracted data is stored in the memory unit 244 as a buffer, cache, etc., and then provided as input to the vector generation unit 250 when the vector generation unit 250 has the bandwidth to receive the input for processing. In other embodiments, the extracted data is provided directly to vector generation unit 250 for processing. For example, in some embodiments, multiple vector generation units 250 may be used to allow parallel processing of inputs to reduce latency.
[0121] The vector generation unit 250 can generate a data structure, such as a feature vector 360, that includes multiple fields, including one or more fields for each type of biomarker data and one or more fields for each type of outcome data. For example, each field in the feature vector 360 can correspond to (i) each type of extracted biomarker data that can be extracted from the biomarker data records 320, 322, 324, such as each type of next-generation sequencing data, each type of in situ hybridization data, each type of RNA data, each type of immunohistochemistry data, and each type of ADAPT data, and (ii) each type of outcome data that can be extracted from the outcome data records 420, 422, such as each type of disease or disorder, each type of treatment, and each type of treatment details.
[0122] Vector generation unit 250 is configured to assign a weight to each field of feature vector 360 that indicates the extent to which the extracted biomarker data 320 a-1, 322 a-1, 324 a-1, extracted disease or disorder 420 a-1, and extracted treatment 422 a-1 include the data represented by each field. In one embodiment, for example, vector generation unit 250 may assign a “1” to each field of feature vector 360 that corresponds to features found in the extracted biomarker data 320 a-1, 322 a-1, 324 a-1, extracted disease or disorder 420 a-1, and extracted treatment 422 a-1. In such an embodiment, vector generation unit 250 may also assign a “0” to each field of the feature vector that corresponds to features not found in the extracted biomarker data 320 a-1, 322 a-1, 324 a-1, extracted disease or disorder 420 a-1, and extracted treatment 422 a-1, for example. The output of the vector generation unit 250 may include a data structure such as a feature vector 360 that can be provided as input to a trained machine learning model 370.
[0123] The trained machine learning model 370 processes the generated feature vector 360 based on the adjusted parameters determined during the training phase and described with reference to FIG. 1B . The output 272 of the trained machine learning model provides an indication of the effectiveness of treatment 422a-1 for disease or disorder 420a-1 in subjects having biomarkers 320a-1, 322a-1, 324a-1. In some embodiments, the output 272 may include a probability indicating the effectiveness of treatment 422a-1 for disease or disorder 420a-1 in subjects having biomarkers 320a-1, 322a-1, 324a-1. In such embodiments, the output 272 may be provided 311 to the terminal 405 using the network 230. The terminal 405 may then generate an output on the user interface 420 indicating a predicted level of effectiveness of treatment for the disease or disorder in persons having the biomarkers represented by the feature vector 360.
[0124] In other embodiments, the output 272 may be provided to a prediction unit 380 configured to decipher the meaning of the output 272. For example, the prediction unit 380 may be configured to map the output 272 to one or more categories of validity. The output of the prediction unit 328 may then be used as part of a message 390 that is provided 311 to the terminal 305 using the network 230 for review by the subject, the subject's guardian, a nurse, a doctor, etc.
[0125] 1D is a flowchart of a process 400 for generating training data for training a machine learning model to predict the effectiveness of a treatment for a disease or disorder in a subject having a particular set of biomarkers. In one aspect, process 400 may include obtaining from a first distributed data source a first data structure including fields structuring data representing a set of one or more biomarkers associated with the subject (410), storing the first data structure in one or more memory devices (420), obtaining from a second distributed data source a second data structure including fields structuring data representing outcome data for the subject having the one or more biomarkers (430), storing the second data structure in one or more memory devices (440), generating a labeled training data structure based on the first data structure and the second data structure (450), the labeled training data structure including data representing (i) the one or more biomarkers, (ii) the disease or disorder, (iii) a treatment, and (iv) the effectiveness of the treatment for the disease or disorder, and training a machine learning model using the generated labeled training data (460).
[0126] 1E is a flowchart of a process 500 using a trained machine learning model to predict the effectiveness of a treatment for a disease or disorder in a subject having a particular set of biomarkers. In one aspect, process 500 may include obtaining a data structure representing a set of one or more biomarkers associated with a subject (510), obtaining data representing the subject's disease or disorder type (520), obtaining data representing a treatment type for the subject (530), generating a data structure for input to a machine learning model representing (i) the one or more biomarkers, (ii) the disease or disorder, and (iii) the treatment type (540), providing the generated data structure as input to a machine learning model trained using labeled training data representing the one or more obtained biomarkers, one or more treatment types, and one or more diseases or disorders (550), obtaining output generated by the machine learning model based on machine learning model processing of the provided data structure (560), and determining a predicted outcome for the treatment of the subject's disease or disorder based on the obtained output generated by the machine learning model (570).
[0127] Provided herein is a method for improving classification performance using multiple machine learning models. Traditionally, a single model is selected to perform a desired prediction / classification. For example, during the training phase, various model parameters or model types, such as random forests, support vector machines, logistic regression, k-nearest neighbors, artificial neural networks, naive Bayes, quadratic discriminant analysis, or Gaussian process models, may be compared to identify a model with optimal desired performance. The applicants have recognized that selecting a single model may not provide optimal performance in every setting. Instead, multiple models can be trained to perform prediction / classification, and classification can be performed using joint prediction. In this scenario, each model is allowed to "vote," and the classification that receives the majority of votes is deemed the winner.
[0128] The voting strategy disclosed herein can be applied to any machine learning classification, including both model building (e.g., using training data) and applications for classifying naive samples. Such settings include, but are not limited to, data from the fields of biology, finance, communications, media, and entertainment. In some preferred embodiments, the data is high-dimensional "big data." In some embodiments, the data includes biological data, including biological data obtained by molecular profiling as described herein. See, e.g., Example 1. Molecular profiling data can include, but is not limited to, high-dimensional next-generation sequencing data for a particular biomarker panel (see, e.g., Example 1) or whole-exome and / or whole-transcriptome data. The classification can be any classification useful, for example, for characterizing a phenotype. For example, the classification can provide a diagnosis (e.g., diseased or healthy), a prognosis (e.g., predicting a good or bad outcome), or theranostic (e.g., predicting or monitoring therapeutic efficacy or lack thereof). Examples of the application of the voting strategy are provided herein in Examples 2-4.
[0129] 1F is a block diagram of a system 600 that interprets outputs generated by multiple machine learning models using a voting unit. System 600 is similar to system 300 of FIG. 1C. However, instead of a single machine learning model 370, system 600 includes multiple machine learning models 370-0, 370-1... 370-x (where x is any non-zero integer greater than 1). In addition, system 600 also includes a voting unit 480.
[0130] As a non-limiting example, system 600 can be used to predict the effectiveness of a treatment for a disease or disorder in a subject having a particular set of biomarkers. See Examples 2-4.
[0131] Each machine learning model 370-0, 370-1, 370-x may include a machine learning model trained to classify a particular type of input data 320-0, 320-1... 320-x (where x is any non-zero integer greater than 1 and equal to the number x of machine learning models). In some embodiments, each of the machine learning models 370-0, 370-1, 370-x may be the same type. For example, each of the machine learning models 370-0, 370-1, 370-x may be a random forest classification algorithm trained using, for example, different parameters. In other embodiments, the machine learning models 370-0, 370-1, 370-x may be different types. For example, there may be one or more random forest model classifiers, one or more neural networks, one or more k-nearest neighbor classifiers, other types of machine learning models, or any combination thereof.
[0132] Input data, such as input data 0 (320-0), input data 1 (320-1), and input data x (320-x), can be obtained by the application server 240. In some embodiments, the input data 320-0, 320-1, and 320-x are obtained from one or more distributed computers 310, 405 over the network 230. Illustratively, one or more of the input data items 320-0, 320-1, and 320-x can be generated by correlating data from multiple different data sources 210, 405. In such embodiments, (i) first data describing a subject's biomarkers can be obtained from the first distributed computer 310, and (ii) second data describing a disease or disorder and associated treatment can be obtained from the second computer 405. The application server 240 can correlate the first data and the second data to generate an input data structure, such as input data structure 320-0. This process is described in further detail in FIG. 1C. Input data items 320-0, 320-1, 320-x may be provided sequentially, one at a time, as respective inputs, to, for example, a vector generation unit. The vector generation unit may generate input vectors 360-0, 360-1, 360-x corresponding to respective input data 320-0, 320-1, 320-x. While some embodiments may generate vectors 360-0, 360-1, 360-x sequentially, the disclosure need not be so limited.
[0133] Alternatively, in some embodiments, vector generation unit 250 may be configured to operate multiple parallel vector generation units that can parallelize the vector generation process. In such embodiments, vector generation unit 250 may concurrently receive input data 320-0, 320-1, 320-x, concurrently process input data 320-0, 320-1, 320-x, and concurrently generate respective vectors 360-0, 360-1, 360-x, each corresponding to one of input data 320-0, 320-1, 320-x.
[0134] In some embodiments, vectors 360-0, 360-1, and 360-x may each be generated based on a corresponding piece of input data, such as input data 320-0, 320-1, and 320-x. That is, vector 360-0 is generated based on and represents input data 320-0. Similarly, vector 360-1 is generated based on and represents input data 320-1. Similarly, vector 360-x is generated based on and represents input data 320-x.
[0135] In some embodiments, each input data structure 320-0, 320-1, 320-x can include data representing a subject's biomarkers, data describing a disease or disorder associated with the subject, data describing a proposed treatment for the subject, or any combination thereof. The data representing a subject's biomarkers can include data describing a particular subset or panel of genes from the subject. Alternatively, in some embodiments, the data representing a subject's biomarkers can include data representing the complete set of known genes for the subject. The complete set of known genes for the subject can include all of the subject's genes. In some embodiments, each of machine learning models 370-0, 370-1, 370-x is the same type of machine learning model, e.g., a neural network trained to classify input data vectors as corresponding to subjects likely to respond or not likely to respond to a treatment identified as associated by the vector processed by the machine learning model. In such an embodiment, each of the machine learning models 370-0, 370-1, 370-x is the same type of machine learning model, but each of the machine learning models 370-0, 370-1, 370-x may be trained in a different manner. The machine learning models 370-1, 370-1, 370-x can generate output data 272-0, 272-1, 272-x, respectively, that represent whether a subject associated with the input vector 360-0, 360-1, 360-x is likely to respond or not respond to the treatment associated with the input vector 360-0, 360-1, 360-x. In this example, the input data sets and their corresponding input vectors are the same. For example, each set of input data has the same biomarkers, the same disease or disorder, the same treatment, or any combination.Nevertheless, given the various training methods used to train each machine learning model 370-0, 370-1, 370-x, different outputs 272-0, 272-1, 272-x may be generated based on each machine learning model 370-0, 370-1, 370-x processing input vectors 360-0, 361-1, 361-x, as shown in FIG. 1F.
[0136] Alternatively, each of machine learning models 370-0, 370-1, 370-x can be a different type of machine learning model trained or otherwise configured to classify input data as representing subjects likely to respond or likely not respond to a treatment for a disease or disorder. For example, first machine learning model 370-1 can include a neural network, machine learning model 370-1 can include a random forest classification algorithm, and machine learning model 370-x can include a k-nearest neighbor algorithm. In this example, each of these different types of machine learning models 370-0, 370-1, 370-x can be trained or otherwise configured to receive and process an input vector and determine whether the input vector is associated with a subject likely to respond or likely not respond to a treatment also associated with the input vector. In this example, the input datasets and their corresponding input vectors can be the same. For example, each set of input data has the same biomarkers, the same disease or disorder, the same treatment, or any combination. Thus, machine learning model 370-0 can be a neural network trained to process input vector 360-0 and generate output data 272-0 indicating whether a subject associated with input vector 360-0 is likely to respond or not respond to a treatment also associated with input vector 360-0. Additionally, machine learning model 370-1 can be a random forest classification algorithm trained to process input vector 360-1, which in this example is the same as input vector 360-0, and generate output data 272-1 indicating whether a subject associated with input vector 360-1 is likely to respond or not respond to a treatment also associated with input vector 360-1. This input vector analysis method can continue with each of the x inputs, x input vectors, and x machine learning models.Continuing with this example with reference to FIG. 1F, machine learning model 370-x can be a k-nearest neighbor algorithm trained to process input vector 360-x, which in this example is the same as input vectors 360-0 and 360-1, and generate output data 272-x that indicates whether the subject associated with input vector 360-x is likely to respond or not respond to the treatment also associated with input vector 360-x.
[0137] Alternatively, each of the machine learning models 370-0, 370-1, and 370-x can be the same type of machine learning model, or different types of machine learning models configured to receive different inputs. For example, the input to the first machine learning model 370-0 can include a vector 360-0 containing data representing a first subset or panel of genes of a subject, and then, based on the machine learning model 370-0 processing of the vector 360-0, predict whether the subject is likely to respond or not to a treatment. Additionally, in this example, the input to the second machine learning model 370-1 can include a vector 360-1 containing data representing a second subset or panel of genes of the subject, different from the first subset or panel of genes. The second machine learning model can then generate second output data 272-1 indicating whether the subject associated with input vector 360-1 is likely to respond or not to a treatment associated with input vector 360-2. This input vector analysis method can continue with each of the x inputs, x input vectors, and x machine learning models. The input to the xth machine learning model 370-x can include a vector 360-x containing data representing an xth subset or xth panel of genes in a subject that is different from (i) at least one, (ii) two or more, or (iii) each of the other x−1 input data vectors 370-0 through 370-x−1. In some embodiments, at least one of the x input data vectors includes data representing a complete set of genes from the subject. The xth machine learning model 370-x can then generate second output data 272-x, which indicates whether the subject associated with the input vector 360-x is likely to respond or not respond to the treatment associated with the input vector 360-x.
[0138] The above-described embodiments of system 400 are not intended to be limiting, but instead are merely examples of configurations of machine learning models 370-0, 370-1, 370-x and their respective inputs that can be used when using the present disclosure. When referring to these examples, the subject can be any human, non-human animal, plant, or other subject. As described above, input feature vectors can be generated based on and represent input data. Thus, each input vector can represent data including one or more biomarkers, a disease or disorder, and a treatment, the level of effectiveness of the treatment in treating the disease or disorder in a subject having the biomarkers. A "treatment" can include data describing any therapeutic agent, e.g., a small molecule drug or biologic, details of the treatment (e.g., dosage, regimen, missed doses, etc.), or any combination thereof.
[0139] In the embodiment of FIG. 1F , output data 272-0, 272-1, 272-x can be analyzed using voting unit 480. For example, output data 272-0, 272-1, 272-x can be input to voting unit 480. In some embodiments, output data 272-0, 272-1, 272-x can be data indicative of whether a subject associated with an input vector processed by the machine learning model is likely to respond or not respond to a treatment associated with the vector processed by the machine learning model. The data indicative of the subject associated with the input vector and generated by each machine learning model can include a “0” or a “1.” A “0” generated by machine learning model 370-0 based on machine learning model 370-0's processing of input vector 360-0 can indicate that the subject associated with input vector 360-0 is likely not to respond to a treatment associated with input vector 360-0. Similarly, a "1" generated by machine learning model 360-0 based on machine learning model 370-0's processing of input vector 360-0 may indicate that the subject associated with input vector 360-0 is likely to respond to the treatment associated with input vector 360-0. While this example uses a "0" as "non-responder" and a "1" as "responder," the disclosure is not so limited. Instead, any values may be generated as output data to represent the "responder" and "non-responder" classes. For example, in some embodiments, a "1" may be used to represent the "non-responder" class, and a "0" may be used to represent the "responder" class. In still other embodiments, output data 272-0, 272-1, 272-x may include a probability indicating the likelihood that the subject associated with the input vector processed by the machine learning model will be associated with the "responder" or "non-responder" class. In such embodiments, for example, the generated probability may be applied to a threshold, and if the threshold is met, the subject associated with the input vector processed by the machine learning model may be determined to be in the "responder" class.
[0140] The voting unit 480 can evaluate the received output data 270-0, 272-1, 272-x and determine whether the subject associated with the processed input vector 360-0, 360-1, 360-x is likely to respond or not respond to the treatment associated with the processed input vector 360-0, 360-1, 360-x. The voting unit 480 can then determine, based on the set of received output data 270-0, 272-1, 272-x, whether the subject associated with the input vector 360-0, 360-1, 360-x is likely to respond to the treatment associated with the input vector 360-0, 360-2, 360-x. In some embodiments, the voting unit 480 can apply a "majority rule." Applying the majority rule, the voting unit 480 can tally the outputs 272-0, 272-1, and 272-x indicating that the subject will respond and the outputs 272-0, 272-1, and 272-x indicating that the subject will not respond. The class with the majority of predictions or votes (e.g., the response or non-response class) is then selected as the appropriate classification for the subject associated with the input vector 360-0, 360-1, and 360-x. This selected class can be referred to as the actual entity class, and each of the predictions or votes output by the machine learning models 370-0, 370-1, and 370-x can be referred to as the initial entity class.
[0141] Thus, in some embodiments, determining the majority of predictions or votes can be achieved by the voting unit 480 tallying the number of occurrences of predictions or votes for each initial entity class. For example, the system 600 can determine the number of times each initial entity class is predicted or voted for by the machine learning models 370-0, 370-1, 370-x, and then select the entity class associated with the greatest number of occurrences of predictions or votes.
[0142] In some embodiments, the voting unit 480 can complete a more nuanced analysis. For example, in some embodiments, the voting unit 480 can store a confidence score for each machine learning model 370-0, 370-1, 370-x. This confidence score for each machine learning model 370-0, 370-1, 370-x can initially be set to a default value, such as 0 or 1. Thereafter, after each round of processing of input vectors, the voting unit 480 or another module of the application server 240 can adjust the confidence score of each machine learning model 370-0, 370-1, 370-x based on whether the machine learning model accurately predicted the target classification selected by the voting unit 480 during the previous iteration. Thus, the stored confidence score for each machine learning model can provide an indication of the historical accuracy of each machine learning model.
[0143] In a more subtle approach, the voting unit 480 can adjust the output data 272-0, 272-0, 272-x generated by each machine learning model 370-0, 370-1, 370-x, respectively, based on a confidence score calculated for the machine learning model. Thus, a confidence score indicating that a machine learning model is historically accurate can be used to boost the value of the output data generated by the machine learning model. Similarly, a confidence score indicating that a machine learning model is historically inaccurate can be used to decrease the value of the output data generated by the machine learning model. Such boosting or decreasing the value of the output data generated by the machine learning model can be achieved, for example, by using the confidence score as a multiplier, less than 1 for a decrease and greater than 1 for a boost. Other operations can also be used to adjust the value of the output data, for example, by subtracting the confidence score from the value of the output data to decrease the value of the output data or by adding the confidence score to the value of the output data to boost the value of the output data. The use of confidence scores to boost or reduce the value of output data generated by a machine learning model is particularly useful when the machine learning model is configured to output probabilities that apply to one or more thresholds for determining whether a subject will respond or not respond to a treatment, because the confidence scores for adjusting the output of the machine learning model can be used to move the generated output value above or below a class threshold, thereby altering the predictions made by the machine learning model based on its historical accuracy.
[0144] The use of the voting unit 480 to evaluate the output of multiple machine learning models can result in greater accuracy in predicting the efficacy of a treatment for a particular set of biomarkers of interest, as the consensus among multiple machine learning models can be evaluated instead of the output of only a single machine learning model.
[0145] FIG. 1G is a block diagram of system components that can be used to implement the systems of FIGS.
[0146] Computing device 600 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Computing device 650 is intended to represent various forms of mobile devices, such as personal digital assistants, mobile phones, smartphones, and other similar computing devices. Additionally, computing device 600 or 650 may include a Universal Serial Bus (USB) flash drive. A USB flash drive may store an operating system and other applications. A USB flash drive may include input / output components, such as a wireless transmitter or a USB connector, that may be inserted into a USB port of another computing device. The components, their connections and relationships, and their functions shown herein are exemplary only and are not intended to limit the embodiments of the invention(s) described and / or claimed herein.
[0147] Computing device 600 includes a processor 602, memory 604, a storage device 608, a high-speed interface 608 connecting to memory 604 and a high-speed expansion port 610, and a low-speed interface 612 connecting to a low-speed bus 614 and storage device 608. Each of components 602, 604, 608, 608, 610, and 612 may be interconnected using various buses, implemented on a common motherboard, or attached in any other suitable manner. Processor 602 may process instructions for execution within computing device 600, including instructions stored in memory 604 or storage device 608 for displaying graphical information for a GUI on an external input / output device, such as a display 616 coupled to high-speed interface 608. In other embodiments, multiple processors and / or multiple buses may be used, along with multiple memories and memory types, as appropriate. Multiple computing devices 600, each providing a portion of the required operations, may also be connected, for example, as a server bank, a group of blade servers, or a multiprocessor system.
[0148] The memory 604 stores information within the computing device 600. In one embodiment, the memory 604 is one or more volatile memory units. In another embodiment, the memory 604 is one or more non-volatile memory units. The memory 604 can also be another form of computer-readable medium, such as a magnetic or optical disk.
[0149] The storage device 608 can provide mass storage for the computing device 600. In one embodiment, the storage device 608 can be or include a computer-readable medium, such as a floppy disk device, a hard disk device, an optical disk device or tape device, a flash memory or other similar solid-state memory device, or an array of devices including devices in a storage area network or other configuration. A computer program product can be tangibly embodied in an information carrier. The computer program product can also include instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as the memory 604, the storage device 608, or the on-processor memory 602.
[0150] High-speed controller 608 manages bandwidth-intensive operations for computing device 600, while low-speed controller 612 manages low-bandwidth-intensive operations. This allocation of functionality is merely exemplary. In one embodiment, high-speed controller 608 is coupled to memory 604, display 616 (e.g., via a graphics processor or accelerator), and high-speed expansion port 610, which can accept various expansion cards (not shown). In an embodiment, low-speed controller 612 is coupled to storage device 608 and low-speed expansion port 614. The low-speed expansion port, which can include various communication ports, e.g., USB, Bluetooth, Ethernet, wireless Ethernet, can be coupled, e.g., via a network adapter, to one or more input / output devices, e.g., a keyboard, a pointing device, a microphone / speaker pair, a scanner, or a networking device, e.g., a switch or router. Computing device 600 can be implemented in several different forms, as shown. For example, it can be implemented as a standard server 620, or multiplexed as a group of such servers. It can also be implemented as part of a rack server system 624. Additionally, it may be implemented as a personal computer, such as a laptop computer 622. Alternatively, components from computing device 600 may be combined with other components in a mobile device (not shown), such as device 650. Each such device may include one or more computing devices 600, 650, or the entire system may be made up of multiple computing devices 600, 650 in communication with each other.
[0151] The computing device 600, as shown in the figure, can be implemented in several different forms. For example, it can be implemented as a standard server 620, or multiple such servers. It can also be implemented as part of a rack server system 624. Additionally, it can be implemented as a personal computer, such as a laptop computer 622. Alternatively, components from the computing device 600 can be combined with other components in a mobile device (not shown), such as device 650. Each such device can include one or more computing devices 600, 650, or the entire system can be made up of multiple computing devices 600, 650 in communication with each other.
[0152] Computing device 650 includes, among other things, a processor 652, a memory 664, and input / output devices such as a display 654, a communication interface 666, and a transceiver 668. Device 650 may also include a storage device, such as a microdrive or other device, to provide additional storage. Each of components 650, 652, 664, 654, 666, and 668 are interconnected using various buses, and some of the components may be mounted on a common motherboard or in any other suitable manner.
[0153] The processor 652 can execute instructions within the computing device 650, including instructions stored in the memory 664. The processor can be implemented as a chipset of chips including separate and multiple analog and digital processors. Additionally, the processor can be implemented using any of several architectures. For example, the processor 610 can be a Complex Instruction Set Computer (CISC) processor, a Reduced Instruction Set Computer (RISC) processor, or a Minimal Instruction Set Computer (MISC) processor. The processor can, for example, provide coordination of other components of the device 650, such as control of a user interface, applications executed by the device 650, and wireless communication by the device 650.
[0154] Processor 652 can communicate with a user via a control interface 658 and a display interface 656 coupled to a display 654. Display 654 can be, for example, a Thin-Film-Transistor Liquid Crystal Display (TFT) display or an Organic Light Emitting Diode (OLED) display or other suitable display technology. Display interface 656 can include appropriate circuitry for driving display 654 to present graphical and other information to a user. Control interface 658 can receive commands from a user and convert them for submission to processor 652. Additionally, an external interface 662 can be provided in communication with processor 652 to enable near-field communication between device 650 and other devices. External interface 662 can, for example, provide for wired communication in some embodiments and wireless communication in other embodiments, or multiple interfaces can be used.
[0155] Memory 664 stores information within computing device 650. Memory 664 may be embodied as one or more of a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. Expansion memory 674 may also be provided and connected to device 650 via expansion interface 672, which may include, for example, a Single In Line Memory Module (SIMM) card interface. Such expansion memory 674 may provide extra storage space for device 650 or may store applications or other information for device 650. Specifically, expansion memory 674 may include instructions for performing or supplementing the above-described processes or may include security information. Thus, for example, expansion memory 674 may be provided as a security module for device 650 and may be programmed with instructions that allow secure use of device 650. Additionally, secure applications may be provided via a SIMM card along with additional information, such as by placing identification information on the SIMM card in an unhackable manner.
[0156] The memory may include, for example, flash memory and / or NVRAM memory, as described in more detail below. In one embodiment, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as memory 664, expansion memory 674, or on-processor memory 652, which may be received, for example, via transceiver 668 or external interface 662.
[0157] Device 650 can communicate wirelessly via communication interface 666, which can include digital signal processing circuitry if necessary. Communication interface 666 can provide communications under various modes or protocols, such as GSM voice calls, SMS, EMS, or MMS messaging, CDMA, TDMA, PDC, WCDMA, CDMA2000, or GPRS, among others. Such communications can be implemented, for example, via radio frequency transceiver 668. Additionally, short-range communications can be implemented, for example, using Bluetooth, Wi-Fi, or other such transceivers (not shown). Additionally, a Global Positioning System (GPS) receiver module 670 can provide further navigation-related and location-related wireless data to device 650, which can be used appropriately by applications running on device 650.
[0158] Device 650 can also communicate audibly using an audio codec 660 that can receive voice information from a user and convert it into usable digital information. Audio codec 660 can also generate audible sounds for the user, such as through a speaker in the handset of device 650. Such sounds can include sounds from a telephone call, recordings such as voice messages, music files, etc., or sounds generated by applications running on device 650.
[0159] The computing device 650, as shown, can be implemented in several different forms, such as a mobile phone 680, or as part of a smartphone 682, personal digital assistant, or other similar mobile device.
[0160] Various embodiments of the systems and methods described herein can be implemented in digital electronic circuitry, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations of such embodiments. These various embodiments can include implementation as one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which can be special purpose or general purpose, coupled to receive data and instructions from, and send data and instructions to, a storage system, at least one input device, and at least one output device.
[0161] These computer programs (also known as programs, software, software applications, or code) contain machine instructions for a programmable processor and may be implemented in high-level procedural and / or object-oriented programming languages and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device used to provide machine instructions and / or data to a programmable processor, e.g., magnetic disks, optical disks, memory, programmable logic devices (PLDs), including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0162] To provide for user interaction, the systems and techniques described herein can be implemented on a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user, and a keyboard and pointing device, such as a mouse or trackball, through which the user can provide input to the computer. Other types of devices can also be used to provide for user interaction. For example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; input from the user can be received in any form, including acoustic, speech, or tactile input.
[0163] The systems and techniques described herein can be implemented as a computing system that includes back-end components, e.g., data servers, or middleware components, e.g., application servers, or front-end components, e.g., client computers having a graphical user interface or web browser through which a user can interact with embodiments of the systems and techniques described herein, or any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communications network. Examples of communications networks include a local area network ("LAN"), a wide area network ("WAN"), and the Internet.
[0164] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0165] Computer Systems The implementation of the method may also use computer-related software and systems. A computer software product as described herein typically includes a computer-readable medium having computer-executable instructions for performing the logical steps of the method as described herein. Suitable computer-readable media include floppy disks, CD-ROMs / DVDs / DVD-ROMs, hard disk drives, flash memory, ROM / RAM, magnetic tape, etc. The computer-executable instructions may be written in a suitable computer language or a combination of several languages. Basic computational biology methods are described, for example, in Setubal and Meidanis at al., Introduction to Computational Biology Methods (PWS Publishing Company, Boston, 1997); Salzberg, Searles, Kasif, (Eds.), Computational Methods in Molecular Biology (Elsevier, Amsterdam, 1998); Rashidi and Buehler, Bioinformatics Basics: Application in Biological Science and Medicine (CRC Press, London, 2000); and Ouelette and Bzevanis, Bioinformatics: A Practical Guide for Analysis of Genes and Proteins (Wiley & Sons, Inc., 2nd ed., 2001). See U.S. Patent No. 6,420,108.
[0166] The present methods may also utilize various computer program products and software for a variety of purposes, such as probe design, data management, analysis, and instrument operation. See U.S. Patent Nos. 5,593,839, 5,795,716, 5,733,729, 5,974,164, 6,066,454, 6,090,555, 6,185,561, 6,188,783, 6,223,127, 6,229,911, and 6,308,170.
[0167] Additionally, the present method relates to embodiments that include methods for providing genetic information via a network, such as the Internet, as described in U.S. Patent Application Nos. 10 / 197,621, 10 / 063,559 (U.S. Patent Application Publication No. 20020183936), 10 / 065,856, 10 / 065,868, 10 / 328,818, 10 / 328,872, 10 / 423,403, and 60 / 482,389. For example, one or more molecular profiling techniques can be performed in one location, such as a city, state, country, or continent, and the results can be transmitted to a different city, state, country, or continent. Treatment selection can then be made, in whole or in part, at the second location. The methods described herein include processes for transferring information between different locations.
[0168] Conventional data networking, application development, and other functional aspects of the system (and components of the system's individual operating components) may not be described in detail herein, but are part of the system as described herein. Additionally, the connecting lines shown in the various figures contained herein are intended to represent example functional relationships and / or physical couplings between the various elements. It should be noted that many alternative or additional functional relationships or physical connections may exist in an actual system.
[0169] The various system components detailed herein may include one or more of the following: a host server or other computing system including a processor for processing digital data; a memory coupled to the processor for storing digital data; an input digitizer coupled to the processor for inputting digital data; an application program stored in the memory and accessible by the processor for instructing the processing of the digital data by the processor; a display device coupled to the processor and memory for displaying information derived from the digital data processed by the processor; and a plurality of databases. The various databases used herein may include patient data, such as family history, age group and environmental data, biosample data, previous treatment and protocol data, patient clinical data, molecular profiling data of biosamples, data regarding therapeutic and / or investigational drugs, gene libraries, disease libraries, drug libraries, patient follow-up data, file management data, financial management data, billing data, and / or similar types of data useful for the operation of the system. As one skilled in the art will appreciate, a user computer may include an operating system (e.g., Windows NT, 95 / 98 / 2000, OS2, UNIX, Linux, Solaris, MacOS, etc.), as well as various conventional support software and drivers that typically accompany a computer. The computer may include any suitable personal computer, network computer, workstation, minicomputer, mainframe, etc. The user computer may be in a home or medical / business environment with network access. In exemplary embodiments, access is via a network or via the Internet via a commercially available web browser software package.
[0170] As used herein, the term "network" is intended to include any electronic communication means incorporating both hardware and software components. Communication between parties may be achieved through any suitable communication channel, such as a telephone network, an extranet, an intranet, the Internet, a point of interaction device, a personal digital assistant (e.g., Palm Pilot®, Blackberry®), a mobile phone, a kiosk, online communication, satellite communication, offline communication, wireless communication, transponder communication, a local area network (LAN), a wide area network (WAN), networked or linked devices, a keyboard, a mouse, and / or any suitable communication or data entry modality. Moreover, although systems are often described herein as being implemented using the TCP / IP communication protocol, systems may also be implemented using IPX, Appletalk, IP-6, NetBIOS, OSI, or any number of existing or future protocols. If the network has the nature of a public network, such as the Internet, it may be advantageous to assume that the network is not secure and is subject to eavesdropping. Specific information relating to protocols, standards, and application software used in connection with the Internet is generally known to those skilled in the art and, therefore, need not be detailed herein. See, for example, Dilip Naik, Internet Standards and Protocols (1998); Java 2 Complete, various authors, (Sybex 1999); Deborah Ray and Eric Ray, Mastering HTML 4.0 (1997); and Loshin, TCP / IP Clearly Explained (1997) and David Gourley and Brian Totty, HTTP, The Definitive Guide (2002), the contents of which are incorporated herein by reference.
[0171] The various system components may be suitably coupled independently, individually, or collectively to a network via data links, including, by way of example, connections to an Internet Service Provider (ISP) via standard modem communications, cable modem, Dish Network, ISDN, DSL (Digital Subscriber Line), or a local loop such as those commonly used in connection with various wireless communication methods. See, for example, Gilbert Held, Understanding Data Communications (1996), incorporated herein by reference. Note that the network may also be implemented as other types of networks, such as an interactive television (ITV) network. Moreover, the system contemplates the use, sale, or distribution of any goods, services, or information over any network having similar functionality as described herein.
[0172] As used herein, "transmitting" may include sending electronic data from one system component to another over a network connection. Additionally, as used herein, "data" may include including information, such as commands, queries, files, data for storage, etc., in digital or any other form.
[0173] The system contemplates use in connection with web services, utility computing, pervasive and personalized computing, security and identity solutions, autonomic computing, commodity computing, mobility and wireless solutions, open source, biometric authentication, grid computing and / or mesh computing.
[0174] Any database detailed herein may include a relational, hierarchical, graphical, or object-oriented structure and / or any other database configuration. Common database products that may be used to implement a database include DB2 from IBM (White Plains, NY), various database products commercially available from Oracle Corporation (Redwood Shores, CA), Microsoft Access or Microsoft SQL Server from Microsoft Corporation (Redmond, Washington), or any other suitable database product. Moreover, a database may be organized in any suitable manner, such as, for example, a data table or a lookup table. Each record may be a single file, a set of files, a linked set of data fields, or any other data structure. The association of specific data may be achieved by any desired data association technique known or implemented in the art. For example, the association may be achieved manually or automatically. Examples of automatic association techniques include, for example, database search, database merge, GREP, AGREP, SQL, using key fields in tables to speed up searches, sequentially searching all tables and files, sorting records in files according to a known order to simplify retrieval, etc. The association process may be accomplished, for example, by a database merge function using pre-selected "key fields" of databases or data sectors.
[0175] More specifically, a "key field" divides a database according to the superclass of the object defined by the key field. For example, a particular type of data may be designated as a key field in multiple related data tables, in which case the data tables may be linked based on the type of data in the key field. The data corresponding to the key field in each of the linked data tables is preferably the same or of the same type. However, data tables having similar but non-identical data in their key fields may also be linked, for example, using AGREP. According to one embodiment, any suitable data storage technology may be used to store data without a standard format. Datasets may be stored using any suitable technique, including, for example, storing individual files using the ISO / IEC 7816-4 file structure; realizing a domain where a dedicated file exposes one or more base files containing one or more datasets is selected; using datasets stored in individual files using a hierarchical filing system; using datasets stored as records in a single file (compressed, SQL accessible, alphabetical with one or more hashed keys, numeric, first tuples, etc.); BLOBs (Binary Large Objects); storage as ungrouped data elements coded using ISO / IEC 7816-6 data elements; storage as ungrouped data elements coded using ISO / IEC Abstract Syntax Notation (ASN.1) as in ISO / IEC 8824 and 8825; and / or using other proprietary techniques, which may include fractal compression schemes, image compression methods, etc.
[0176] In one exemplary embodiment, the ability to store a wide variety of information in different formats is facilitated by storing the information as a blob. Thus, any binary information can be stored in the storage space associated with a dataset. The blob method may store datasets as ungrouped data elements formatted as blocks of binary data at fixed memory offsets using either fixed storage allocation, circular queue techniques, or memory management best practices (e.g., paged memory, least recently used, etc.). Using the blob method, the ability to store various datasets with different formats facilitates the storage of data by multiple, unrelated owners of the datasets. For example, a first dataset that can be stored can be provided by a first party, a second dataset that can be stored can be provided by an unrelated second party, and yet a third dataset that can be stored can be provided by a third party unrelated to the first and second parties. Each of these three exemplary datasets can contain different information stored using different data storage formats and / or techniques. Furthermore, each dataset can contain a subset of data that can also differ from the other subsets.
[0177] As noted above, in various embodiments, data can be stored without regard to a common format. However, in one exemplary embodiment, data sets (e.g., blobs) can be annotated in a standard manner when provided for manipulating the data. The annotations can include short headers, trailers, or other suitable indicators associated with each data set that are configured to convey information useful in managing various data sets. For example, annotations, sometimes referred to herein as "conditional headers," "headers," "trailers," or "status," can include an indication of the status of the data set or can include an identifier correlated to a particular publisher or owner of the data. Subsequent bytes of data can be used to indicate, for example, the identity of the data publisher or owner, a user, a transaction / membership account identifier, or the like. Each of these conditional annotations is described in further detail herein.
[0178] Dataset annotations may also be used for other types of status information and various other purposes. For example, dataset annotations may include security information establishing access levels. Access levels may be configured, for example, so that only certain individuals, employee levels, companies, or other entities are allowed to access the dataset, or may be configured to grant access to specific datasets based on transactions, data issuers or owners, users, etc. Furthermore, security information may restrict / allow only certain actions, such as accessing, modifying, and / or deleting the dataset. In one example, a dataset annotation may indicate that only the dataset owner or user is allowed to delete the dataset, various identified users may be allowed to access the dataset for reading, and all other users are excluded from accessing the dataset. However, other access restriction parameters may be used that allow various entities to access the dataset with various permission levels, as appropriate. Data, including a header or trailer, may be received by a standalone interactive device configured to add, delete, modify, or augment the data according to the header or trailer.
[0179] Those skilled in the art will also understand that for security reasons, any database, system, device, server, or other component of a system may consist of any combination thereof in one location or multiple locations, and each database or system may include any of a variety of appropriate security mechanisms, such as firewalls, access codes, encryption, decryption, compression, decompression, etc.
[0180] The web client's computing unit may further include an internet browser connected to the internet or intranet using standard dial-up, cable, DSL, or any other internet protocol known in the art. Transactions occurring at the web client may pass through firewalls to prevent unauthorized access from users of other networks. Additionally, additional firewalls may be deployed between various components of the CMS to further enhance security.
[0181] A firewall may include any hardware and / or software appropriately configured to protect CMS components and / or enterprise computing resources from users of other networks. Additionally, a firewall may be configured to limit or restrict access to various systems and components behind the firewall in the case of web clients connecting through a web server. Firewalls may exist in a variety of configurations, including stateful inspection, proxy-based firewalls, and packet filtering, among others. A firewall may be integrated into a web server or any other CMS component, or may exist as a separate entity.
[0182] The computers detailed herein may provide a suitable website or other Internet-based graphical user interface accessible by a user. In one embodiment, Microsoft Internet Information Server (IIS), Microsoft Transaction Server (MTS), and Microsoft SQL Server are used in conjunction with the Microsoft operating system, Microsoft NT web server software, the Microsoft SQL Server database system, and Microsoft Commerce Server. Additionally, components such as Access or Microsoft SQL Server, Oracle, Sybase, Informix MySQL, Interbase, etc. may be used to provide an Active Data Object (ADO)-compliant database management system.
[0183] Any of the communication, input, storage, database, or display detailed herein can be facilitated through a website having a web page. The term "web page" as used herein is not meant to limit the types of documents and applications that can be used to interact with a user. For example, a typical website may include, in addition to standard HTML documents, various forms, Java applets, JavaScript, Active Server Pages (ASP), Common Gateway Interface (CGI) scripts, Extensible Markup Language (XML), Dynamic HTML, Cascading Style Sheets (CSS), helper applications, plug-ins, and the like. A server may include a web service that receives a request from a web server, including a URL (http: / / yahoo.com / stockquotes / ge) and an IP address (123.56.789.234). The web server retrieves the appropriate web page and sends the data or application for the web page to the IP address. A web service is an application that can interact with other applications via a communication medium such as the Internet. Web services are typically based on standards or protocols such as XML, XSLT, SOAP, WSDL, and UDDI. Web services methods are well known in the art and are covered in many standard textbooks, see, for example, Alex Nghiem, IT Web Services: A Roadmap for the Enterprise (2003), incorporated herein by reference.
[0184] The web-based clinical database for the subject systems and methods preferably has the ability to upload and store clinical data files in their native format and is searchable by any clinical parameter. The database is also extensible and can input clinical annotations from any study using the EAV data model (metadata) for easy integration with other studies. In addition, the web-based clinical database is flexible and can be XML and XSLT enabled to allow users to dynamically add customized queries. Furthermore, the database includes export capabilities to CDISC ODM.
[0185] Implementers will also appreciate that there are many ways to display data within a browser-based document: data may be displayed as standard text, in fixed lists, scrollable lists, drop-down lists, editable text fields, fixed text fields, pop-up windows, etc. Similarly, there are many ways available to change data within a web page, such as free text entry using the keyboard, selecting menu items, check boxes, option boxes, etc.
[0186] The systems and methods may be described herein in terms of functional block components, screenshots, optional selections, and various processing steps. It should be understood that such functional blocks may be implemented by any number of hardware and / or software components configured to perform the specified functions. For example, the system may use various integrated circuit components, such as memory elements, processing elements, logic elements, lookup tables, etc., that may perform various functions under the control of one or more microprocessors or other control devices. Similarly, the software elements of the system may be implemented in any programming or scripting language, such as C, C++, Macromedia Cold Fusion, Microsoft Active Server Pages, Java, COBOL, Assembler, PERL, Visual Basic, SQL Stored Procedures, or Extensible Markup Language (XML), and various algorithms may be implemented in any combination of data structures, objects, processes, routines, or other programming elements. Furthermore, it should be noted that the system may use any number of conventional techniques for data transmission, signaling, data processing, network control, etc. Still further, the system may also be used to detect or prevent security issues in client-side scripting languages, such as JavaScript and VBScript.For an introduction to the basics of cryptography and network security, see any of the following references, all of which are incorporated herein by reference: (1) "Applied Cryptography: Protocols, Algorithms, And Source Code In C," by Bruce Schneier, published by John Wiley & Sons (second edition, 1995); (2) "Java Cryptography" by Jonathan Knudson, published by O'Reilly & Associates (1998); (3) "Cryptography & Network Security: Principles & Practice" by William Stallings, published by Prentice Hall.
[0187] As used herein, the terms "end user," "consumer," "customer," "client," "treating physician," "hospital," or "business" may be used interchangeably and refer to any person, entity, machine, hardware, software, or business, respectively. Each participant is equipped with a computing device to interact with the system and facilitate online data access and data entry. Customers have computing units in the form of personal computers, although other types of computing units may be used, including laptops, notebooks, handheld computers, set-top boxes, cellular phones, touch-tone phones, and the like. The system and owner / operator of the method have computing units implemented in the form of computer servers, although other embodiments are contemplated, including systems including computing centers depicted as mainframe computers, minicomputers, PC servers, networks of computers in the same or different geographic locations, and the like. Moreover, the system contemplates the use, sale, or distribution of any goods, services, or information over any network having similar functionality as described herein.
[0188] In one exemplary embodiment, each client customer may be issued an "account" or "account number." As used herein, an account or account number may include any device, code, number, letter, symbol, digital certificate, smart chip, digital signal, analog signal, biometric or other identifier / indicia (e.g., one or more of an authentication / access code, personal identification number (PIN), internet code, other identification code, etc.) suitably configured to allow a consumer to access, interact with, or communicate with the system. The account number may optionally be located on or associated with a charge card, credit card, debit card, prepaid card, embossed card, smart card, magnetic stripe card, bar code card, transponder, radio frequency card, or related account. The system may include or interface with any of the foregoing cards or devices, or a fob having a transponder and RFID reader for RF communication with the fob. While the system may include a fob embodiment, the method is not so limited. Indeed, the system may include any device having a transponder configured to communicate with an RFID reader via RF communication. Exemplary devices may include, for example, a key ring, a tag, a card, a mobile phone, a watch, or any such form that can be presented for interrogation. Moreover, the systems, computing units, or devices detailed herein may include "pervasive computing devices," which may include traditional non-computerized devices with embedded computing units. Account numbers may be distributed and stored in any form of plastic, electronic, magnetic, radio frequency, wireless, audio, and / or optical device that can transmit or download data from itself to a second device.
[0189] As will be appreciated by those skilled in the art, the system may be embodied as a customization of an existing system, an add-on product, upgraded software, a standalone system, a distributed system, a method, a data processing system, a device for data processing, and / or a computer program product. Accordingly, the system may take the form of an all-software embodiment, an all-hardware embodiment, or an embodiment combining both software and hardware aspects. Furthermore, the system may take the form of a computer program product on a computer-readable storage medium having computer-readable program code means embodied in the storage medium. Any suitable computer-readable storage medium may be used, including a hard disk, a CD-ROM, an optical storage device, a magnetic storage device, or the like.
[0190] Systems and methods are described herein with reference to screenshots, block diagrams, and flowchart illustrations of methods, apparatus (e.g., systems), and computer program products according to various aspects. It will be understood that each functional block of the block diagrams and flowchart illustrations, and combinations of functional blocks in the block diagrams and flowchart illustrations, respectively, can be implemented by computer program instructions.
[0191] These computer program instructions may be loaded into a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, where the instructions executing on the computer or other programmable data processing apparatus create means for implementing the functions specified in one or more flowchart blocks. These computer program instructions may also be stored in a computer-readable memory that can instruct the computer or other programmable data processing apparatus to function in a particular manner to produce an article of manufacture including instruction means for implementing the functions specified in one or more flowchart blocks, where the instructions stored in the computer-readable memory. Computer program instructions may also be loaded into a computer or other programmable data processing apparatus to cause a series of operating steps executed on the computer or other programmable apparatus to produce a computer-implemented process, where the instructions executing on the computer or other programmable apparatus provide steps for performing the functions specified in one or more flowchart blocks.
[0192] Accordingly, the functional blocks in the block diagrams and flowchart diagrams support combinations of means for performing the specified functions, combinations of steps for performing the specified functions, and program instruction means for performing the specified functions. It will also be understood that each functional block in the block diagrams and flowchart diagrams, and combinations of functional blocks in the block diagrams and flowchart diagrams, can be implemented by a dedicated hardware-based computer system that performs the specified functions or steps, or by an appropriate combination of dedicated hardware and computer instructions. Furthermore, the process flow illustrations and descriptions may refer to user windows, web pages, websites, web forms, prompts, and the like. Practitioners will understand that the illustrated steps described herein may include any number of configurations, including the use of windows, web pages, web forms, pop-up windows, prompts, and the like. Furthermore, it should be understood that multiple steps illustrated and described may be combined into a single web page and / or window, but have been expanded for clarity. In other cases, steps illustrated and described as a single process step may be separated into multiple web pages and / or windows, but have been combined for clarity.
[0193] Molecular Profiling Molecular profiling techniques provide a method for selecting candidate treatments for individuals that can improve the clinical course for individuals with a disease or disorder, such as cancer. Molecular profiling techniques provide clinical benefits for individuals, such as identifying treatment regimens that provide longer progression-free survival (PFS), longer disease-free survival (DFS), longer overall survival (OS), or longer lifespan. The methods and systems described herein relate to molecular profiling of cancer on an individual basis, which can identify optimal treatment regimens. Molecular profiling provides a personalized approach to selecting candidate treatments that are likely to benefit the cancer. The molecular profiling methods described herein can be used to guide treatment in any desired setting, including first-line / standard-of-care settings, or for patients with poor prognosis, such as patients with metastatic disease, or whose cancer has progressed on standard first-line treatment, or whose cancer has progressed on previous chemotherapy or hormonal therapy.
[0194] The systems and methods of the invention may be used to classify patients as likely or unlikely to benefit from or respond to various treatments. Unless otherwise specified, the terms "response" or "non-response" as used herein refer to any appropriate indication that a treatment provides benefit to the patient (a "responder" or "beneficiary") or lacks benefit to the patient (a "non-responder" or "non-beneficiary"). Such indications may be determined using accepted clinical response criteria, such as standard RECIST (Response Evaluation Criteria in Solid Tumors) criteria, or other useful patient response criteria, such as progression-free survival (PFS), time to progression (TTP), disease-free survival (DFS), time to next treatment (TNT, TTNT), tumor shrinkage or disappearance, etc. RECIST is a set of rules published by an international consortium that defines whether a tumor improves ("responds"), remains stable ("stable"), or worsens ("progresses") during treatment in cancer patients. As used herein, unless otherwise specified, a patient "benefit" from a treatment can refer to any appropriate measure of improvement, including a RECIST response or longer PFS / TTP / DFS / TNT / TTNT, and a "lack of benefit" from a treatment can refer to any appropriate measure of disease worsening during treatment. Generally, disease stabilization is considered a benefit, although in certain circumstances, stabilization may be considered a lack of benefit if so specified herein. If there is no acceptable level of prediction of benefit or lack of benefit, the predicted or indicated benefit may be described as "indeterminate." In some cases, benefit is considered indeterminate if it cannot be calculated, for example, due to a lack of necessary data.
[0195] Personalized medicine based on pharmacogenetic insights, such as those provided by molecular profiling as described herein, is increasingly taken for granted by some practitioners and the general public, but it forms the basis of hope for improved cancer treatment. However, molecular profiling as taught herein represents a radical departure from traditional approaches to tumor treatment, in which patients are largely grouped and treated based on findings from light microscopy and disease stage. Traditionally, differential response to a particular therapeutic strategy has been determined only after treatment has been administered, i.e., post hoc. The "standard" approach to disease treatment relies on what is generally true for a given cancer diagnosis, and treatment response is explored through randomized phase III clinical trials, forming the "standard of care" in medical practice. The results of these trials are summarized in consensus statements by guideline organizations such as the National Comprehensive Cancer Network and the American Society of Clinical Oncology. The NCCN Compendium™ contains authoritative, scientifically derived information designed to support decision-making regarding the appropriate use of drugs and biologics in cancer patients. The NCCN Compendium™ is recognized by the Centers for Medicare & Medicaid Services (CMS) and UnitedHealthcare as the authoritative reference source for oncology insurance coverage. The Compendium treatments are those recommended by such guides. Biostatistical methods used to validate the results of clinical trials rely on minimizing interpatient differences and declaring the probability of error that one method is superior to another for patient groups determined solely by light microscopy and disease stage (not by individual differences in tumors). The molecular profiling method described herein exploits such individual differences. This method can provide candidate treatments, which can then be selected by a physician to treat the patient.
[0196] Molecular profiling can be used to provide a comprehensive view of the biological state of a sample. In some embodiments, molecular profiling is used for whole-tumor profiling. Thus, several molecular techniques are used to assess the state of a tumor. Whole-tumor profiling can be used to select candidate treatments for a tumor. Molecular profiling can be used to select candidate therapeutics for any sample at any stage. In some embodiments, the methods described herein are used to profile a newly diagnosed cancer. Candidate treatments indicated by molecular profiling can be used to select a therapy for treating the newly diagnosed cancer. In other embodiments, the methods described herein are used to profile a cancer that has already been treated, for example, with one or more standard therapies. In some embodiments, the cancer is refractory to previous treatments. For example, the cancer can be refractory to standard therapies for cancer. The cancer can be metastatic or other recurrent cancers. The treatment can be on-compendium or off-compendium.
[0197] Molecular profiling can be carried out by any known means for detecting molecules in biological samples.Molecular profiling includes methods such as nucleic acid sequencing, for example, DNA sequencing or RNA sequencing; immunohistochemistry (IHC); in situ hybridization (ISH); fluorescent in situ hybridization (FISH); colorimetric in situ hybridization (CISH); PCR amplification (for example, qPCR or RT-PCR); various types of microarrays (mRNA expression array, low-density array, protein array, etc.); various types of sequencing (Sanger, pyrosequencing, etc.); comparative genomic hybridization (CGH); high-throughput or next-generation sequencing (NGS); Northern blot; Southern blot; immunoassay; and any other suitable technology for testing the presence or amount of biological molecules of interest.In various embodiments, any one or more of these methods can be used in parallel or sequentially with each other to evaluate the target genes disclosed herein.
[0198] Molecular profiling of individual samples is used to select one or more candidate treatments for the disorder in a subject, for example, by identifying targets for drugs that may be effective for a given cancer. For example, the candidate treatments can be treatments known to affect cells that differentially express the genes identified by molecular profiling techniques, experimental drugs, government or regulatory approved drugs, or any combination of such drugs (which may have been studied and approved for a specific indication that is the same or different from the indication of the subject for which the biological sample is collected and molecularly profiled).
[0199] When multiple biomarker targets are identified by evaluating target genes through molecular profiling, one or more decision rules can be applied to prioritize the selection of specific therapeutic agents for individual treatment on an individualized basis. Rules such as those described herein can help prioritize treatments, for example, based on the direct results of molecular profiling, the expected efficacy of the therapeutic agent, previous treatment history with the same or other treatments, expected side effects, availability of the therapeutic agent, cost of the therapeutic agent, drug-drug interactions, and other factors considered by the treating physician. Based on the recommended and prioritized therapeutic agent targets, the physician can determine the course of treatment for a particular individual. Thus, molecular profiling methods and systems such as those described herein can select candidate treatments based on the individual characteristics of diseased cells, e.g., tumor cells, in a subject requiring treatment and other personalized factors, as opposed to relying on traditional one-size-fits-all approaches routinely used to treat individuals suffering from diseases, particularly cancer. In some cases, the recommended treatment is one that is not typically used to treat the disease or disorder afflicting the subject. In some cases, the recommended treatment is used after standard treatments no longer provide sufficient efficacy.
[0200] The treating physician can use the results of molecular profiling method to optimize the treatment regimen for the patient.The candidate treatment identified by the method as described herein can be used to treat the patient, but such treatment is not required for the method.In fact, the analysis of molecular profiling results and the identification of candidate treatment based on such results can be automated and do not require the involvement of a physician.
[0201] Biological entities Nucleic acids include deoxyribonucleotides or ribonucleotides and their polymers in either single-stranded or double-stranded form, or their complements. Nucleic acids can contain known synthetic, natural, and non-natural nucleotide analogs or modified backbone residues or linkages, which have similar binding properties to the reference nucleic acid and are metabolized in a similar manner to the reference nucleotide. Examples of such analogs include, but are not limited to, phosphorothioates, phosphoramidates, methyl phosphonates, chiral-methyl phosphonates, 2-O-methyl ribonucleotides, and peptide-nucleic acids (PNAs). Nucleic acid sequences can encompass the sequences specified, as well as conservatively modified variants thereof (e.g., degenerate codon substitutions) and complementary sequences. Specifically, degenerate codon substitutions can be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and / or deoxyinosine residues (Batzer et al., Nucleic Acid Res. 19:5081 (1991); Ohtsuka et al., J. Biol. Chem. 260:2605-2608 (1985); Rossolini et al., Mol. Cell Probes 8:91-98 (1994)). The terms nucleic acid can be used interchangeably with gene, cDNA, mRNA, oligonucleotide, and polynucleotide.
[0202] A particular nucleic acid sequence can implicitly encompass nucleic acid sequences encoding the specific sequence as well as "splice variants" and truncated forms. Similarly, a particular protein encoded by a nucleic acid can encompass any protein encoded by a splice variant or truncated form of that nucleic acid. A "splice variant," as the name suggests, is the product of alternative splicing of a gene. After transcription, an initial nucleic acid transcript may be spliced such that different (alternative) nucleic acid splice products encode different polypeptides. Mechanisms for producing splice variants vary but include alternative splicing of exons. Alternative polypeptides obtained from the same nucleic acid by read-through transcription are also encompassed by this definition. Any product of a splicing reaction, including recombinant forms of splice products, is included in this definition. Nucleic acids can be truncated at the 5' or 3' end. Polypeptides can be truncated at the N- or C-terminus. Truncated versions of nucleic acid or polypeptide sequences can be natural or can be produced using recombinant techniques.
[0203] The terms "gene variant" and "nucleotide variant" are used interchangeably herein to refer to changes or alterations to a reference human gene or cDNA sequence at a specific locus, including, but not limited to, deletions, insertions, inversions, and substitutions of nucleotide bases in coding and non-coding regions. Deletions can be of a single nucleotide base, a portion or region of the nucleotide sequence of a gene, or the entire gene sequence. Insertions can be of one or more nucleotide bases. Gene variants or nucleotide variants can occur in transcriptional regulatory regions, untranslated regions of mRNA, exons, introns, exon / intron junctions, etc. Gene variants or nucleotide variants can potentially result in stop codons, frameshifts, amino acid deletions, altered gene transcript splice forms, or altered amino acid sequences.
[0204] Alleles or gene alleles generally include naturally occurring genes having a reference sequence or genes containing particular nucleotide variants.
[0205] A haplotype refers to a combination of genetic (nucleotide) variants in a region of genomic DNA on a chromosome or mRNA found in an individual. Thus, a haplotype contains several genetically linked polymorphic variants that are typically inherited together as a unit.
[0206] As used herein, the term "amino acid variant" refers to amino acid changes relative to a reference human protein sequence that result from genetic or nucleotide variants relative to the reference human gene encoding the reference protein. The term "amino acid variant" is intended to encompass not only single amino acid substitutions in the amino acid sequence of the reference protein, but also amino acid deletions, insertions, and other significant changes.
[0207] The term "genotype" as used herein refers to the nature of the nucleotide at a particular nucleotide variant marker (or locus) in either one allele or both alleles of a gene (or a particular chromosomal region). With respect to a particular nucleotide position of a gene of interest, the nucleotide at that locus or its equivalent in one or both alleles forms the genotype of the gene at that locus. The genotype can be homozygous or heterozygous. Thus, "genotyping" refers to determining the genotype, i.e., the nucleotide at a particular gene locus. Genotyping can also be performed by determining the amino acid variant at a particular position of a protein, which can be used to infer the corresponding nucleotide variant.
[0208] The term "locus" refers to a specific position or site in a gene sequence or protein. Thus, there may be one or more consecutive nucleotides at a particular gene locus, or one or more amino acids at a particular locus in a polypeptide. Moreover, a locus may refer to a specific position in a gene where one or more nucleotides have been deleted, inserted, or inverted.
[0209] Unless otherwise specified or understood by those skilled in the art, the terms "polypeptide," "protein," and "peptide" are used interchangeably herein to refer to an amino acid chain in which amino acid residues are linked by covalent peptide bonds. The amino acid chain can be of any length of at least two amino acids, including full-length proteins. Unless otherwise specified, polypeptides, proteins, and peptides also encompass various modified forms thereof, including, but not limited to, glycosylated forms, phosphorylated forms, and the like. A polypeptide, protein, or peptide can also be referred to as a gene product.
[0210] A list of genes and gene products that can be assayed by molecular profiling techniques is provided herein. The list of genes may be provided in conjunction with molecular profiling techniques that detect gene products (e.g., mRNA or protein). Those skilled in the art will understand that this refers to the detection of the gene products of the listed genes. Similarly, the list of gene products may be provided in conjunction with molecular profiling techniques that detect gene sequence or copy number. Those skilled in the art will understand that this refers to the detection of the genes corresponding to the gene products, including, for example, the DNA that encodes the gene products. As will be recognized by those skilled in the art, "biomarker" or "marker" includes genes and / or gene products depending on the context.
[0211] The terms "label" and "detectable label" can refer to any composition detectable by spectroscopic, photochemical, biochemical, immunochemical, electrical, optical, chemical, or similar methods. Such labels include biotin for staining with labeled streptavidin conjugates, magnetic beads (e.g., DYNABEADS™), fluorescent dyes (e.g., fluorescein, Texas Red, rhodamine, green fluorescent protein, etc.), radioactive labels (e.g., 3 H, 125 I, 35 S, 14 C, or 32P), enzymes (e.g., horseradish peroxidase, alkaline phosphatase, and others commonly used in ELISA), and colorimetric labels such as colloidal gold or colored glass or plastic (e.g., polystyrene, polypropylene, latex, etc.) beads. Patents teaching the use of such labels include U.S. Pat. Nos. 3,817,837; 3,850,752; 3,939,350; 3,996,345; 4,277,437; 4,275,149; and 4,366,241. Means for detecting such labels are well known to those skilled in the art. Thus, for example, radioactive labels may be detected using photographic film or a scintillation counter, and fluorescent markers may be detected using a photodetector to detect emitted light. Enzymatic labels are typically detected by providing a substrate to the enzyme and detecting the reaction product produced by the action of the enzyme on the substrate, while colorimetric labels are detected by simply visualizing the colored label. Labels can include, for example, ligands that bind to labeled antibodies, fluorophores, chemiluminescent agents, enzymes, and antibodies that can serve as members of binding pairs specific for the labeled ligand. General information on labels, labeling procedures, and label detection can be found in Polak and Van Noorden, Introduction to Immunocytochemistry, 2nd ed., Springer Verlag, NY (1997); and Haugland Handbook of Fluorescent Probes and Research Chemicals (1996), a combined handbook and catalog published by Molecular Probes, Inc.
[0212] Detectable labels include, but are not limited to, nucleotides (labeled or unlabeled), compomers, sugars, peptides, proteins, antibodies, chemical compounds, conducting polymers, binding moieties such as biotin, mass tags, colorimetric agents, luminescent agents, chemiluminescent agents, light scattering agents, fluorescent tags, radioactive tags, charge tags (electrical or magnetic), volatile tags and hydrophobic tags, biomolecules (e.g., binding pairs antibody / antigen, antibody / antibody, antibody / antibody fragment, antibody / antibody receptor, antibody / protein A or protein G, hapten / antihapten, biotin / avidin, biotin / streptavidin, folate / folate binding protein, vitamin B12 / intrinsic factor, chemically reactive groups / complementary chemically reactive groups (e.g., sulfhydryl / maleimide, sulfhydryl / haloacetyl derivatives, amine / isotriocyanate, amine / succinimidyl ester, and amine / sulfonyl halide), and the like.
[0213] The terms "primer," "probe," and "oligonucleotide" are used interchangeably herein to refer to relatively short nucleic acid fragments or sequences. They can include DNA, RNA, or hybrids thereof, or chemically modified analogs or derivatives thereof. Typically, they are single-stranded. However, they can also be double-stranded, with two complementary strands that can be separated by denaturation. Primers, probes, and oligonucleotides usually range from about 8 to about 200 nucleotides in length, preferably from about 12 to about 100 nucleotides in length, and more preferably from about 18 to about 50 nucleotides in length. They can be labeled with a detectable marker or modified using conventional methods for various molecular biology applications.
[0214] The term "isolated," when used in reference to a nucleic acid (e.g., genomic DNA, cDNA, mRNA, or a fragment thereof), is intended to mean that the nucleic acid molecule is present in a form that is substantially separated from other naturally occurring nucleic acids with which the molecule is normally associated. Because naturally occurring chromosomes (or their viral equivalents) contain long nucleic acid sequences, an isolated nucleic acid can be a nucleic acid molecule that contains only a portion of the nucleic acid sequence in the chromosome, but lacks one or more other portions present on the same chromosome. More specifically, an isolated nucleic acid can contain naturally occurring nucleic acid sequences that flank the nucleic acid in a naturally occurring chromosome (or its viral equivalent). An isolated nucleic acid can be substantially separated from other naturally occurring nucleic acids on different chromosomes of the same organism. An isolated nucleic acid can also be a composition in which a particular nucleic acid molecule is significantly enriched, such that it constitutes at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or at least 99% of the total nucleic acids in the composition.
[0215] An isolated nucleic acid can be a hybrid nucleic acid having a specific nucleic acid molecule covalently linked to one or more nucleic acid molecules, where the one or more nucleic acid molecules are not naturally adjacent to the specific nucleic acid. For example, an isolated nucleic acid can be in a vector. In addition, a specific nucleic acid can have the same nucleotide sequence as a natural nucleic acid, or its modified form or mutein with one or more mutations, such as nucleotide substitutions, deletions / insertions, inversions, etc.
[0216] An isolated nucleic acid can be prepared from a recombinant host cell (wherein the nucleic acid has been recombinantly amplified and / or expressed), or can be a chemically synthesized nucleic acid having a naturally occurring nucleotide sequence or an artificially modified form thereof.
[0217] The term "high stringency hybridization conditions," when used in reference to nucleic acid hybridization, includes hybridization overnight at 42°C in a solution containing 50% formamide, 5x SSC (750 mM NaCl, 75 mM sodium citrate), 50 mM sodium phosphate, pH 7.6, 5x Denhardt's solution, 10% dextran sulfate, and 20 micrograms / ml denatured, fragmented salmon sperm DNA, with the hybridization filter washed in 0.1x SSC at approximately 65°C. The term "moderately stringent hybridization conditions," when used in reference to nucleic acid hybridization, includes overnight hybridization at 37° C. in a solution containing 50% formamide, 5×SSC (750 mM NaCl, 75 mM sodium citrate), 50 mM sodium phosphate, pH 7.6, 5×Denhardt's solution, 10% dextran sulfate, and 20 micrograms / ml denatured, fragmented salmon sperm DNA, with the hybridization filter washed in 1×SSC at approximately 50° C. It should be noted that many other hybridization methods, solutions, and temperatures can be used to achieve similarly stringent hybridization conditions, as will be apparent to those skilled in the art.
[0218] For purposes of comparing two different nucleic acid or polypeptide sequences, one sequence (test sequence) may be described as being a certain percentage identical to another sequence (comparison sequence). The percentage of identity can be determined using the algorithm of Karlin and Altschul, Proc. Natl. Acad. Sci. USA, 90:5873-5877 (1993), which is incorporated into various BLAST programs. The percentage of identity can be determined using the "BLAST 2 Sequences" tool available on the National Center for Biotechnology Information (NCBI) website. See Tatusova and Madden, FEMS Microbiol. Lett., 174(2):247-250 (1999). For DNA-DNA pairwise comparisons, the BLASTN program is used with default parameters (e.g., match: 1; mismatch: -2; open gap: 5 penalty; extension gap: 2 penalty; gap x_dropoff: 50; expectation: 10; and word size: 11, with filter). For protein-protein pairwise comparisons, the BLASTP program can be employed with default parameters (e.g., matrix: BLOSUM62; gap open: 11; gap extension: 1; x_dropoff: 15; expectation: 10.0; and word size: 3, with filter). The percent identity of two sequences is calculated by aligning the test sequence with the comparison sequence using BLAST, determining the number of amino acids or nucleotides in the aligned test sequence that are identical to amino acids or nucleotides at the same positions in the comparison sequence, and dividing the number of identical amino acids or nucleotides by the number of amino acids or nucleotides in the comparison sequence. When BLAST is used to compare two sequences, BLAST aligns the sequences and generates a percent identity over a given aligned region. When two sequences are aligned over their entire length, the percent identity produced by BLAST is the percent identity of these two sequences.If BLAST does not align two sequences over their entire length, the number of identical amino acids or nucleotides in the unaligned regions of the test and comparison sequences is assumed to be zero, and the percent identity is calculated by adding the number of identical amino acids or nucleotides in the aligned regions and dividing that number by the length of the comparison sequences. Various versions of the BLAST program can be used to compare sequences, such as BLAST 2.1.2 or BLAST+ 2.2.22.
[0219] The subject or individual can be any animal that may benefit from the methods described herein, including, for example, humans and non-human mammals, such as primates, rodents, horses, dogs and cats.Subjects include, but are not limited to, eukaryotes, most preferably mammals, such as primates, for example, chimpanzees or humans, cows; dogs; cats; rodents, such as guinea pigs, rats, mice; rabbits; or birds; reptiles; or fish.The subjects specifically intended for treatment using the methods described herein include humans.Subjects may also be referred to herein as individuals or patients.In this method, the subject has colorectal cancer, for example, has been diagnosed with colorectal cancer.Methods for identifying subjects with colorectal cancer, for example, using biopsy, are known in the art. See, for example, Fleming et al., J Gastrointest Oncol. 2012 Sep; 3(3): 153-173; Chang et al., Dis Colon Rectum. 2012; 55(8):831-43.
[0220] Treating a disease or individual using the methods described herein is an approach to achieving beneficial or desired medical results, including clinical results, but not necessarily a cure. For the purposes of the methods described herein, beneficial or desired clinical results include, but are not limited to, alleviation or amelioration of one or more symptoms, whether detectable or undetectable, reduction in the extent of disease, stable disease (i.e., not worsening), prevention of disease spread, delay or slowing of disease progression, amelioration or palliation of disease, and remission (whether partial or complete). Treatment also includes extending survival compared to the expected survival if no treatment or a different treatment were administered. Treatment can include administering either a FOLFOX or FOLFIRI regimen. Biomarkers generally refer to molecules, including but not limited to genes or their products, nucleic acids (e.g., DNA, RNA), proteins / peptides / polypeptides, carbohydrate structures, lipids, glycolipids, which, when detected in tissues or cells, have characteristics that can provide predictive, diagnostic, prognostic and / or theranostic information about susceptibility or resistance to a candidate treatment.
[0221] biological samples As used herein, a sample includes any relevant biological sample that can be used for molecular profiling, such as biopsies or tissues removed during surgical or other procedures, body fluids, autopsy samples, and tissue sections such as frozen sections taken for histological purposes. Such samples include blood and blood fractions or products (e.g., serum, buffy coat, plasma, platelets, red blood cells, etc.), sputum, malignant effusions, buccal tissue, cultured cells (e.g., primary cultures, explants, and transformed cells), feces, urine, other biological fluids or bodily fluids (e.g., prostatic fluid, gastric fluid, intestinal fluid, renal fluid, lung fluid, cerebrospinal fluid, etc.), and others. Samples can include biological materials that are fresh-frozen and formalin-fixed paraffin-embedded (FFPE) blocks, formalin-fixed paraffin-embedded, or in RNA preservative plus formalin fixative. More than one sample of more than one type can be used for each patient. In a preferred embodiment, the sample includes a fixed tumor sample.
[0222] The samples used in the systems and methods of the present invention can be formalin-fixed, paraffin-embedded (FFPE) samples. FFPE samples can be one or more of fixed tissue, unstained slides, bone marrow cores or clots, core needle biopsies, malignant fluids, and fine needle aspirates (FNAs). In one embodiment, the fixed tissue comprises a tumor-containing formalin-fixed, paraffin-embedded (FFPE) block from surgery or biopsy. In another embodiment, the unstained slide comprises an unstained, charged, unbaked slide from a paraffin block. In another embodiment, the bone marrow cores or clots comprise decalcified cores. Formalin-fixed cores and / or clots can be paraffin-embedded. In yet another embodiment, the core needle biopsy comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more, e.g., 3-4, paraffin-embedded biopsy samples. An 18-gauge needle biopsy can be used. The malignant fluid may contain a volume of fresh pleural / peritoneal fluid sufficient to produce a 5 x 5 x 2 mm cell pellet. The fluid may be formalin-fixed in a paraffin block. In some embodiments, core needle biopsies include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more, e.g., 4-6, paraffin-embedded aspirates.
[0223] The sample can be processed according to techniques understood by those skilled in the art. The sample can be, but is not limited to, fresh, frozen, or fixed cells or tissues. In some embodiments, the sample includes formalin-fixed paraffin-embedded (FFPE) tissue, fresh tissue, or fresh-frozen (FF) tissue. The sample can include cultured cells, including primary or immortalized cell lines, derived from a subject sample. The sample can also refer to an extract from a sample derived from a subject. For example, the sample can include DNA, RNA, or protein extracted from tissue or body fluid. Many techniques and commercially available kits are available for such purposes. Fresh samples from individuals can be treated with agents to preserve RNA before further processing, such as cell lysis and extraction. The sample can also include frozen samples collected for other purposes. The sample can be associated with relevant information such as age, sex, and clinical symptoms present in the subject; the origin of the sample; and the method of sample collection and storage. The sample is typically obtained from a subject.
[0224] Biopsy includes the process of removing a tissue sample for diagnostic or prognostic evaluation, as well as the tissue specimen itself. Any biopsy technique known in the art can be applied to the molecular profiling method of the present disclosure. The biopsy technique applied can depend, among other factors, on the type of tissue to be evaluated (e.g., colon, prostate, kidney, bladder, lymph node, liver, bone marrow, blood cells, lung, breast, etc.), the size and type of tumor (e.g., solid or floating, blood or ascites), etc. Representative biopsy techniques include, but are not limited to, excision biopsy, incision biopsy, needle biopsy, surgical biopsy, and bone marrow biopsy. An "excision biopsy" refers to the removal of an entire tumor mass along with a small margin of surrounding normal tissue. An "incision biopsy" refers to the removal of a wedge of tissue containing the cross-sectional diameter of the tumor. Molecular profiling can use a "core needle biopsy" of the tumor mass, or a "fine needle aspiration biopsy," which generally obtains a suspension of cells from within the tumor mass. Biopsy techniques are discussed, for example, in Chapter 70 and throughout Part V of Harrison's Principles of Internal Medicine, Kasper, et al., eds., 16th ed., 2005.
[0225] Unless otherwise specified, the "sample" referred to herein for molecular profiling of a patient may include more than one physical specimen.As a non-limiting example, a "sample" may include multiple sections from a tumor, such as multiple sections of an FFPE block or multiple core needle biopsy sections.As another non-limiting example, a "sample" may include multiple biopsy specimens, such as one or more surgical biopsy specimens, one or more core needle biopsy specimens, one or more fine needle aspiration biopsy specimens, or any useful combination thereof.As yet another non-limiting example, a molecular profile may be generated for a subject using a "sample" including a solid tumor specimen and a body fluid specimen.In some embodiments, a sample is a unit sample, i.e., a single physical specimen.
[0226] Standard molecular biology techniques known in the art and not specifically described are generally described in Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York (1989), and Ausubel et al., Current Protocols in Molecular Biology, John Wiley and Sons, Baltimore, Md. (1989), and Perbal, A Practical Guide to Molecular Cloning, John Wiley & Sons, New York (1988), and Watson et al., Recombinant DNA, Scientific American Books, New York, and Birren et al (eds) Genome Analysis: A Laboratory Manual Series, Vols. 1-4, Cold Spring Harbor Laboratory Press, New York. (1998), and U.S. Patent Nos. 4,666,828; 4,683,202; 4,801,531; 5,192,659 and 5,272,057, which are incorporated herein by reference. Polymerase chain reaction (PCR) can generally be performed as described in PCR Protocols: A Guide to Methods and Applications, Academic Press, San Diego, Calif. (1990).
[0227] Vesicles The sample can include a vesicle. The methods described herein can include examining one or more vesicles, including examining a population of vesicles. As used herein, a vesicle is a membrane vesicle shed from a cell. Vesicles or membrane vesicles include, but are not limited to, circulating microvesicles (cMVs), microvesicles, exosomes, nanovesicles, dexosomes, blebs, blebbies, prostasomes, microparticles, intraluminal vesicles, membrane fragments, intraluminal endosomal vesicles, endosome-like vesicles, exocytosis vehicles, endosome vesicles, endosomal vesicles, apoptotic bodies, multivesicular bodies, secretory vesicles, phospholipid vesicles, liposomal vesicles, argosomes, texasomes, secretomes, tolerosomes, melanosomes, oncosomes, or exocytosed vehicles. Furthermore, while vesicles may be produced by different cellular processes, the methods described herein are not limited to or dependent on any one mechanism, so long as such vesicles are present in a biological sample and can be characterized by the methods disclosed herein. Unless otherwise specified, methods utilizing one type of vesicle can be applied to other types of vesicles. Vesicles comprise spherical structures with a lipid bilayer, similar to a cell membrane, surrounding an internal compartment that can contain soluble components, sometimes referred to as the payload. In some embodiments, the methods described herein utilize exosomes, which are small secreted vesicles approximately 40-100 nm in diameter. For a review of membrane vesicles, including types and characteristics, see Thery et al., Nat Rev Immunol. 2009 Aug;9(8):581-93. Some properties of different types of vesicles include those listed in Table 1:
[0228] Table 1. Vesicle characteristics TIFF0007798947000001.tif127169Abbreviations: Phosphatidylserine (PPS); Electron microscopy (EM)
[0229] Vesicles include excreted membrane-bound particles or "microparticles" derived from either the plasma membrane or endomembranes. Vesicles can be released from cells into the extracellular environment. Cells releasing vesicles include, but are not limited to, cells derived from or derived from the ectoderm, endoderm, or mesoderm. Cells may have undergone genetic, environmental, and / or any other variation or alteration. For example, the cells may be tumor cells. Vesicles can reflect any changes in the source cell, thereby reflecting changes in the derived cell, for example, cells with various genetic mutations. In one mechanism, vesicles are generated intracellularly when segments of the cell membrane spontaneously invaginate and are eventually exocytotic (see, e.g., Keller et al., Immunol. Lett. 107 (2): 102-8 (2006)). Vesicles also include cell-derived structures bound by lipid bilayer membranes, resulting from both the separation of blebbing and the sealing of plasma membrane portions, or from the export of any intracellular membrane-bound vesicular structures containing various membrane-associated proteins of tumor origin, which contain surface-bound molecules obtained from the host circulation that selectively bind to tumor-derived proteins along with molecules contained in the vesicle lumen, including, but not limited to, tumor-derived microRNAs or intracellular proteins. Blebs and blebbing are further described in Charras et al., Nature Reviews Molecular and Cell Biology, Vol. 9, No. 11, pp. 730-736 (2008). Vesicles released from tumor cells into the circulation or body fluids may be referred to as "circulating tumor-derived vesicles." When such vesicles are exosomes, they may be referred to as circulating tumor-derived exosomes (CTEs). In some cases, vesicles can be derived from a specific cellular origin. CTEs, like cell-origin-specific vesicles, typically have one or more unique biomarkers that allow for isolation of the CTEs or cell-origin-specific vesicles, sometimes in a specific manner, e.g., from bodily fluids.For example, cell or tissue specific marker is used to identify cell origin.The example of such cell or tissue specific marker is disclosed herein and can be found in Tissue-specific Gene Expression and Regulation (TiGER) database, available at bioinfo.wilmer.jhu.edu / tiger / ; Liu et al. (2008) TiGER: a database for tissue-specific gene expression and regulation. BMC Bioinformatics. 9:271; TissueDistributionDBs, available at genome.dkfz-heidelberg.de / menu / tissue_db / index.html.
[0230] Vesicles can have diameters greater than about 10 nm, 20 nm, or 30 nm. Vesicles can have diameters greater than 40 nm, 50 nm, 100 nm, 200 nm, 500 nm, 1000 nm, or greater than 10,000 nm. Vesicles can have diameters of about 30-1000 nm, about 30-800 nm, about 30-200 nm, or about 30-100 nm. In some embodiments, vesicles have diameters of 10,000 nm, 1000 nm, 800 nm, 500 nm, 200 nm, 100 nm, 50 nm, 40 nm, 30 nm, less than 20 nm, or less than 10 nm. As used herein, the term "about" in connection with a numerical value means that a 10% variation above and below the numerical value is within the range ascribed to the particular value. Typical sizes for various types of vesicles are listed in Table 1. Vesicles can be examined to measure the diameter of a single vesicle or any number of vesicles. For example, the diameter range of a vesicle population or the average diameter of a vesicle population can be determined.The diameter of vesicle can be examined by methods known in the art, for example, imaging techniques such as electron microscopy.In some embodiments, the diameter of one or more vesicles is determined by optical particle detection.See, for example, U.S. Patent No. 7,751,053, entitled "Optical Detection and Analysis of Particles," issued on July 6, 2010; and U.S. Patent No. 7,399,600, entitled "Optical Detection and Analysis of Particles," issued on July 15, 2010.
[0231] In some embodiments, vesicles are assayed directly from a biological sample without prior isolation, purification, or enrichment. For example, the amount of vesicles in a sample can itself provide a biosignature for diagnostic, prognostic, or theranostic determinations. Alternatively, vesicles in a sample may be isolated, captured, purified, or enriched from the sample prior to analysis. As described above, isolation, capture, or purification, as used herein, includes partial isolation, partial capture, or partial purification away from other components in the sample. Vesicle isolation can be performed using various techniques, such as those described herein or known in the art, including, but not limited to, size exclusion chromatography, density gradient centrifugation, differential centrifugation, nanomembrane ultrafiltration, immunosorbent capture, affinity purification, affinity capture, immunoassay, immunoprecipitation, microfluidic separation, flow cytometry, or a combination thereof.
[0232] Vesicles can be examined and their characteristics compared to a standard to provide phenotypic characterization. In some embodiments, surface antigens on vesicles are examined. Vesicles or vesicle populations bearing a particular marker can be referred to as positive (biomarker+) vesicles or vesicle populations. For example, a DLL4+ population refers to a vesicle population that binds DLL4. Conversely, a DLL4- population does not bind DLL4. Surface antigens can provide an indication of the anatomical and / or cellular origin of vesicles and other phenotypic information, such as tumor status. For example, vesicles found in a patient sample can be examined for surface antigens indicative of colorectal origin and the presence of cancer, thereby identifying vesicles associated with colorectal cancer cells. Surface antigens can include any biological entity that provides information detectable on the vesicle membrane surface, including, but not limited to, surface proteins, lipids, carbohydrates, and other membrane components. For example, positive detection of vesicles obtained from the colon expressing a tumor antigen can indicate that the patient has colorectal cancer. Thus, methods such as those described herein can be used to characterize any disease or condition associated with an anatomical or cellular origin, for example, by examining disease-specific and cell-specific biomarkers in one or more vesicles obtained from a subject.
[0233] In various embodiments, one or more vesicle payloads are examined to provide phenotypic characterization. Vesicle-bearing payloads include any informative biological entity that can be detected as encapsulated within the vesicle, including, but not limited to, proteins and nucleic acids, e.g., genomes or cDNA, mRNA, or functional fragments thereof, as well as microRNAs (miRs). Additionally, methods such as those described herein are directed to detecting vesicle surface antigens (in addition to or exclusively with the vesicle payload) to provide phenotypic characterization. For example, vesicles can be characterized by using a binding agent (e.g., an antibody or aptamer) specific for the vesicle surface antigen, and the bound vesicles can be further examined to identify one or more payload components disclosed herein. As described herein, the level of a surface antigen of interest or a vesicle bearing a payload of interest can be compared to a reference to characterize a phenotype. For example, overexpression of a cancer-associated surface antigen or vesicle payload, e.g., tumor-associated mRNA or microRNA, in a sample compared to a reference can indicate the presence of cancer in the sample. The biomarkers examined may be present or absent, increased or decreased, based on the selection of a desired target sample and comparison of the target sample with a desired reference sample. Non-limiting examples of target samples include disease; treated / untreated; different time points, for example, in longitudinal studies; non-limiting examples of reference samples include non-disease; normal; different time points; and sensitive or resistant to a candidate treatment.
[0234] In certain aspects, molecular profiling as described herein involves the analysis of microvesicles, such as circulating microvesicles.
[0235] microRNA Various biomarker molecules can be examined in biological samples or vesicles obtained from such biological samples. MicroRNAs comprise one class of biomarkers that can be examined via the methods described herein. MicroRNAs, also referred to herein as miRNAs or miRs, are short RNA strands approximately 21-23 nucleotides in length. miRNAs are encoded by genes that are transcribed from DNA but not translated into protein, and thus comprise non-coding RNAs. miRs are processed from a primary transcript known as a pri-miRNA into a short stem-loop structure called a pre-miRNA and finally into the resulting single-stranded miRNA. The pre-miRNA typically forms a structure that folds back on itself in a self-complementary region. These structures are then processed by the nuclease Dicer in animals or DCL1 in plants. Mature miRNA molecules are partially complementary to one or more messenger RNA (mRNA) molecules and can function to regulate protein translation. Identified sequences of miRNAs can be accessed from publicly available databases such as www.microRNA.org, www.mirbase.org, or www.mirz.unibas.ch / cgi / miRNA.cgi.
[0236] miRNAs are generally assigned numbers according to the naming convention "mir-[number]." miRNA numbers are assigned according to their order of discovery relative to previously identified miRNA species. For example, if the last published miRNA was mir-121, the next discovered miRNA would be named mir-122, and so on. If a miRNA is discovered to be homologous to a known miRNA from a different organism, its name can be given an optional organism identifier in the format [organism identifier]-mir-[number]. Identifiers include hsa for Homo sapiens and mmu for Mus musculus. For example, the human homolog of mir-121 might be referred to as hsa-mir-121, while the mouse homolog could be referred to as mmu-mir-121.
[0237] Mature microRNAs are usually named with the prefix "miR", while genes or precursor miRNAs are named with the prefix "mir". For example, mir-121 is the precursor for miR-121. When different miRNA genes or precursors are processed into the same mature miRNA, the genes / precursors can be described by numbered suffixes. For example, mir-121-1 and mir-121-2 can refer to different genes or precursors that are processed into miR-121. Lettered suffixes are used to indicate closely related mature sequences. For example, mir-121a and mir-121b can be processed into closely related miRNAs, miR-121a and miR-121b, respectively. In the context of the present disclosure, any microRNA (miRNA or miR) designated herein with the prefix mir-* or miR-* is understood to encompass both the precursor and / or mature species, unless otherwise specified.
[0238] It is sometimes observed that two mature miRNA sequences originate from the same precursor. When one sequence is more abundant than the other, the "*" suffix can be used to designate the less common variant. For example, miR-121 is the main product, while miR-121* is the less common variant found in the opposite arm of the precursor. When the main variant is not identified, miRs can be identified by the suffix "5p" for variants from the 5' arm of the precursor and "3p" for variants from the 3' arm. For example, miR-121-5p originates from the 5' arm of the precursor, while miR-121-3p originates from the 3' arm. Less commonly, 5p and 3p variants are referred to as sense ("s") and antisense ("as") forms, respectively. For example, miR-121-5p may be referred to as miR-121-s, while miR-121-3p may be referred to as miR-121-as.
[0239] The above naming conventions have evolved over time and are general guidelines rather than absolute prescriptions. For example, the let and lin families of miRNAs continue to be referred to by their nicknames. The mir / miR convention for precursor / mature forms is also a guideline, and context should be taken into account when determining which form is being referred to. Further details on miR nomenclature can be found at www.mirbase.org or in Ambros et al., A uniform system for microRNA annotation, RNA 9:277-279 (2003).
[0240] Plant miRNAs follow a different naming convention as described in Meyers et al., Plant Cell. 2008 20(12):3186-3190.
[0241] Several miRNAs are involved in gene regulation and are part of an expanding class of noncoding RNAs now recognized as a major layer of gene control. In some cases, miRNAs can disrupt translation by binding to regulatory sites embedded in the 3'-UTR of target mRNAs, resulting in translational repression. Target recognition involves complementary base pairing between the target site and the seed region of the miRNA (positions 2-8 at the 5' end of the miRNA). However, the exact degree of seed complementarity is not precisely determined and can be modified by 3' pairing. In other cases, miRNAs function like small interfering RNAs (siRNAs), binding to perfectly complementary mRNA sequences and disrupting target transcripts.
[0242] Characterization of some miRNAs has shown that they affect a variety of processes, including early development, cell proliferation and cell death, apoptosis, and fat metabolism. For example, some miRNAs, such as lin-4, let-7, mir-14, mir-23, and bantam, have been shown to play important roles in cell differentiation and tissue development. Others are also thought to play important roles due to their differential spatial and temporal expression patterns.
[0243] The miRNA database available at miRBase (www.mirbase.org) contains a searchable database of published miRNA sequences and annotations. Further information regarding miRBase can be found in the following documents, each of which is incorporated herein by reference in its entirety: Griffiths-Jones et al., miRBase: tools for microRNA genomics. NAR 2008 36(Database Issue):D154-D158; Griffiths-Jones et al., miRBase: microRNA sequences, targets and gene nomenclature. NAR 2006 34(Database Issue):D140-D144; and Griffiths-Jones, S. The microRNA Registry. NAR 2004 32(Database Issue):D109-D111. Representative miRNAs included in Release 16 of miRBase were made available in September 2010.
[0244] As described herein, microRNA is known to be involved in cancer and other diseases, and can be examined to characterize the phenotype in samples.See, for example, Ferracin et al., Micromarkers: miRNAs in cancer diagnosis and prognosis, Exp Rev Mol Diag, April 2010, Vol.10, No.3, Pages 297-308; Fabbri, miRNAs as molecular biomarkers of cancer, Exp Rev Mol Diag, May 2010, Vol.10, No.4, Pages 435-444.
[0245] In certain aspects, molecular profiling as described herein includes the analysis of microRNAs.
[0246] Techniques for isolating and characterizing vesicles and miRs are known to those of skill in the art. In addition to the methodologies presented herein, additional methods are described in U.S. Pat. No. 7,888,035, issued February 15, 2011, entitled "METHODS FOR ASSESSING RNA PATTERNS"; and U.S. Pat. No. 7,897,356, issued March 1, 2011, entitled "METHODS AND SYSTEMS OF USING EXOSOMES FOR DETERMINING PHENOTYPES"; and International Patent Publications WO / 2011 / 066589, issued November 30, 2010, entitled "METHODS AND SYSTEMS FOR ISOLATING, STORING, AND ANALYZING VESICLES"; WO / 2011 / 088226, issued January 13, 2011, entitled "DETECTION OF GASTROINTESTINAL DISORDERS"; and BIOMARKERS FOR and WO / 2011 / 127219, published on April 6, 2011, entitled "CIRCULATING BIOMARKERS FOR DISEASE," each of which is incorporated herein by reference in its entirety.
[0247] Circulating biomarkers Circulating biomarkers include biomarkers that can be detected in bodily fluids, such as blood, plasma, and serum. Examples of circulating cancer biomarkers include cardiac troponin T (cTnT), prostate-specific antigen (PSA) for prostate cancer, and CA125 for ovarian cancer. Circulating biomarkers according to the present disclosure include any suitable biomarker that can be detected in bodily fluids, including, but not limited to, proteins, nucleic acids (e.g., DNA, mRNA, and microRNA), lipids, carbohydrates, and metabolites. Circulating biomarkers can include biomarkers that are not associated with cells, such as biomarkers that are membrane-bound, biomarkers embedded in membrane fragments, biomarkers that are part of biological complexes, or biomarkers that are free in solution. In one embodiment, the circulating biomarker is a biomarker associated with one or more vesicles present in a subject's biological fluid.
[0248] Circulating biomarkers have been identified for use in characterizing various phenotypes, including cancer detection. For example, Ahmed N, et al., Proteomic-based identification of haptoglobin-1 precursor as a novel circulating biomarker of ovarian cancer. Br. J. Cancer 2004; Mathelin _et al., Circulating proteinic biomarkers and breast cancer, Gynecol Obstet Fertil. 2006 Jul-Aug;34(7-8):638-46. Epub 2006 Jul 28; Ye et al., Recent technical strategies to identify diagnostic biomarkers for ovarian cancer. Expert Rev Proteomics. 2007 Feb;4(1):121-31; Carney, Circulating oncoproteins HER2 / neu, EGFR and CAIX (MN) as novel cancer biomarkers. Expert Rev Mol Diagn. 2007 May;7(3):309-19; Gagnon, Discovery and application of protein biomarkers for ovarian cancer, Curr Opin Obstet Gynecol. 2008 Feb;20(1):9-13; Pasterkamp et al., Immune regulatory cells: circulating biomarker factories in cardiovascular disease. Clin Sci (Lond). 2008 Aug;115(4):129-31; Fabbri, miRNAs as biomarker moleculars of cancer, Exp Rev Mol Diag, May 2010, Vol. 10, No.4, Pages 435-444; PCT Patent Publication WO / 2007 / 088537; U.S. Patent Nos. 7,745,150 and 7,655,479; U.S. Patent Application Publication Nos. 20110008808, 20100330683, 20100248290, 20100222230, 20100203566, 20100173788, 20090291932, and 2 See US Pat. Nos. 0090239246, 20090226937, 20090111121, 20090004687, 20080261258, 20080213907, 20060003465, 20050124071, and 20040096915, each of which is incorporated by reference in its entirety. In certain aspects, molecular profiling as described herein includes analysis of circulating biomarkers.
[0249] Gene expression profiling The methods and systems described herein include expression profiling, which involves examining the differential expression of one or more target genes disclosed herein. Differential expression can include overexpression and / or underexpression of a biological product, such as a gene, mRNA, or protein, compared to a control (or standard). The control can include cells similar to the sample but without the disease (e.g., an expression profile obtained from a sample from a healthy individual). The control can be a previously determined level indicative of the effectiveness of a drug target associated with a particular disease and a particular drug target. The control can be derived from the same patient, e.g., a normal adjacent part of the same organ as the diseased cells, or it can be obtained from healthy tissue from another patient, or it can be a previously determined threshold value indicating whether a disease responds or does not respond to a particular drug target. The control can also be a control found in the same sample, such as a housekeeping gene or its product (e.g., mRNA or protein). For example, the control nucleic acid can be one known to show no difference depending on the cancerous or non-cancerous state of the cell. The expression level of a control nucleic acid can be used to normalize the signal level in the test population and the reference population. Illustrative control genes include, but are not limited to, β-actin, glyceraldehyde-3-phosphate dehydrogenase, and ribosomal protein P1. Multiple controls or types of controls can be used. The cause of differential expression can vary. For example, gene copy number can increase in cells, resulting in increased expression of the gene. Alternatively, gene transcription can be altered by, for example, chromatin remodeling, differential methylation, differential expression or activity of transcription factors, etc. Translation can also be altered by, for example, differential expression of factors that degrade mRNA, translate mRNA, or silence translation, such as microRNA or siRNA. In some embodiments, differential expression includes differential activity. For example, a protein can have a mutation that increases the activity of the protein, such as constitutive activation, contributing to a pathological condition.Molecular profiling, which reveals changes in activity, can be used to guide treatment selection.
[0250] Gene expression profiling methods include polynucleotide hybridization analysis-based methods and polynucleotide sequencing-based methods.The commonly used methods known in the art for quantifying mRNA expression in a sample include Northern blotting and in situ hybridization (Parker & Barnes (1999) Methods in Molecular Biology 106:247-283); RNase protection assay (Hod (1992) Biotechniques 13:852-854); and reverse transcription polymerase chain reaction (RT-PCR) (Weis et al. (1992) Trends in Genetics 8:263-264).Alternatively, antibodies can be employed that can recognize specific duplexes, including DNA duplexes, RNA duplexes, and DNA-RNA hybrid duplexes or DNA-protein duplexes. Representative methods for sequencing-based gene expression analysis include Serial Analysis of Gene Expression (SAGE), gene expression analysis by massively parallel signature sequencing (MPSS), and / or next-generation sequencing.
[0251] RT-PCR Reverse transcription polymerase chain reaction (RT-PCR) is a variation of polymerase chain reaction (PCR). With this technique, RNA strand is reverse transcribed into its DNA complement (i.e., complementary DNA, or cDNA) using an enzyme called reverse transcriptase, and the resulting cDNA is amplified using PCR. Real-time polymerase chain reaction is another PCR variation, also called quantitative PCR, Q-PCR, qRT-PCR, or sometimes RT-PCR. Either reverse transcription PCR or real-time PCR can be used for molecular profiling according to the present disclosure, and RT-PCR can be referred to unless otherwise specified or as understood by those skilled in the art.
[0252] RT-PCR can be used to determine the RNA level, for example, mRNA or miRNA level, of biomarkers as described herein. RT-PCR can be used to compare such RNA levels of biomarkers as described herein in different sample populations, in normal tissues and tumor tissues, with or without drug treatment, to characterize patterns of gene expression, to distinguish closely related RNAs, and to analyze RNA structure.
[0253] The first step is to isolate RNA, such as mRNA, from sample.Starting material can be the total RNA isolated from human tumor or tumor cell line and corresponding normal tissue or cell line.Therefore, RNA can be isolated from sample, such as tumor cell or tumor cell line, and compared with the DNA pooled from healthy donors.If the origin of mRNA is primary tumor, mRNA can be extracted from, for example, frozen tissue sample or paraffin-embedded and fixed (for example, formalin-fixed) archived tissue sample.
[0254] General methods for mRNA extraction are well known in the art and are disclosed in standard molecular biology textbooks, including Ausubel et al. (1997) Current Protocols of Molecular Biology, John Wiley and Sons. Methods for extracting RNA from paraffin-embedded tissues are disclosed, for example, in Rupp & Locker (1987) Lab Invest. 56:A67 and De Andres et al., BioTechniques 18:42044 (1995). In particular, RNA isolation can be performed using a purification kit, buffer set, and protease from a commercial manufacturer such as Qiagen, according to the manufacturer's instructions (QIAGEN Inc., Valencia, CA). For example, total RNA from cultured cells can be isolated using Qiagen RNeasy mini-columns. Numerous RNA isolation kits are commercially available and can be used for the methods described herein.
[0255] Alternatively, the first step is the isolation of miRNA from target samples.The starting material is typically total RNA isolated from human tumors or tumor cell lines and corresponding normal tissues or cell lines.Therefore, RNA can be isolated from various primary tumors or tumor cell lines, along with pooled DNA from healthy donors.When the source of miRNA is primary tumor, miRNA can be extracted, for example, from frozen tissue samples or archival tissue samples that have been paraffin-embedded and fixed (e.g., formalin-fixed).
[0256] General methods for miRNA extraction are well known in the art and are disclosed in standard molecular biology textbooks, including Ausubel et al. (1997) Current Protocols of Molecular Biology, John Wiley and Sons. Methods for extracting RNA from paraffin-embedded tissues are disclosed, for example, in Rupp & Locker (1987) Lab Invest. 56:A67 and De Andres et al., BioTechniques 18:42044 (1995). In particular, RNA isolation can be performed using a purification kit, buffer set, and protease from a commercial manufacturer such as Qiagen, according to the manufacturer's instructions. For example, total RNA from cultured cells can be isolated using Qiagen RNeasy mini-columns. Many miRNA isolation kits are commercially available and can be used for the methods described herein.
[0257] Whether the RNA contains mRNA, miRNA, or other types of RNA, gene expression profiling by RT-PCR can involve reverse transcription of the RNA template into cDNA, followed by amplification in a PCR reaction. Commonly used reverse transcriptases include, but are not limited to, avian myeloblastosis virus reverse transcriptase (AMV-RT) and Moloney murine leukemia virus reverse transcriptase (MMLV-RT). The reverse transcription step is typically primed using specific primers, random hexamers, or oligo-dT primers, depending on the context and goal of expression profiling. For example, extracted RNA can be reverse transcribed using a GeneAmp RNA PCR kit (Perkin Elmer, Calif., USA) according to the manufacturer's instructions. The resulting cDNA can then be used as a template in a subsequent PCR reaction.
[0258] Although a variety of thermostable DNA-dependent DNA polymerases can be used, the PCR process typically employs Taq DNA polymerase, which possesses 5'-3' nuclease activity but lacks 3'-5' proofreading endonuclease activity. TaqMan PCR typically uses the 5'-nuclease activity of Taq or Tth polymerase to hydrolyze hybridization probes bound to the target amplicon, although any enzyme with equivalent 5' nuclease activity can be used. Two oligonucleotide primers are used to generate the amplicon typical of a PCR reaction. A third oligonucleotide, or probe, is designed to detect the nucleotide sequence located between the two PCR primers. The probe is non-extendable by Taq DNA polymerase enzyme and is labeled with a reporter and a quencher fluorescent dye. When the two dyes are positioned in close proximity to each other on the probe, any laser-induced emission from the reporter dye is quenched by the quencher dye. During the amplification reaction, Taq DNA polymerase enzyme cleaves the probe in a template-dependent manner. The resulting probe fragments dissociate in solution, and the signal from the released reporter dye is freed from the quenching effect of the second fluorophore. Because one molecule of reporter dye is liberated for each new molecule synthesized, detection of the unquenched reporter dye provides the basis for quantitative interpretation of the data.
[0259] TaqMan™ RT-PCR can be performed using commercially available equipment, such as the ABI PRISM 7700™ Sequence Detection System™ (Perkin-Elmer-Applied Biosystems, Foster City, Calif., USA) or the LightCycler (Roche Molecular Biochemicals, Mannheim, Germany). In a specific embodiment, the 5' nuclease procedure is performed in a real-time quantitative PCR device such as the ABI PRISM 7700 Sequence Detection System. The system consists of a thermocycler, a laser, a charge-coupled device (CCD), a camera, and a computer. The system amplifies samples in a 96-well format using a thermocycler. During amplification, laser-induced fluorescent signals are collected in real time for all 96 wells through a fiber optic cable and detected by the CCD. The system includes software for operating the instrument and analyzing the data.
[0260] TaqMan data are initially expressed as Ct, or threshold cycle. As mentioned above, fluorescence values are recorded during each cycle and represent the amount of product amplified to that point in the amplification reaction. The point at which the fluorescent signal is first recorded as statistically significant is the threshold cycle (Ct).
[0261] To minimize errors and the effects of sample-to-sample variation, RT-PCR is usually performed using an internal standard. An ideal internal standard is expressed at a consistent level among different tissues and is not affected by experimental treatments. The RNAs most frequently used to normalize patterns of gene expression are mRNAs for the housekeeping genes glyceraldehyde-3-phosphate-dehydrogenase (GAPDH) and β-actin.
[0262] Real-time quantitative PCR (also known as quantitative real-time polymerase chain reaction, QRT-PCR, or Q-PCR) is a more recent variation of the RT-PCR technique. Q-PCR can measure the accumulation of PCR products through dual-labeled fluorogenic probes (i.e., TaqMan probes). Real-time PCR is compatible with both quantitative competitive PCR, in which an internal competitor for each target sequence is used for normalization, and quantitative comparative PCR, which uses normalization genes contained in the sample or housekeeping genes for RT-PCR. See, for example, Held et al. (1996) Genome Research 6:986-994.
[0263] Protein-based detection techniques are also useful for molecular profiling, especially when nucleotide variants cause amino acid substitutions, deletions, insertions, or frameshifts that affect the primary, secondary, or tertiary structure of a protein. Protein sequencing techniques can be used to detect amino acid variations. For example, a protein or fragment corresponding to a gene can be synthesized by recombinant expression using DNA fragments isolated from a test individual. Preferably, a cDNA fragment of 100 to 150 base pairs or less encompassing the polymorphic locus to be determined is used. The amino acid sequence of the peptide can then be determined by conventional protein sequencing methods. Alternatively, HPLC-microscopy tandem mass spectrometry techniques can be used to determine amino acid sequence variations. In this technique, the protein is subjected to proteolytic digestion, and the resulting peptide mixture is separated by reverse-phase chromatography. Tandem mass spectrometry is then performed, and the collected data is analyzed. See Gatlin et al., Anal. Chem., 72:757-763 (2000).
[0264] Microarray Biomarkers as described herein can also be identified, confirmed, and / or measured using microarray technology. Thus, expression profile biomarkers can be measured in cancer samples using microarray technology. In this method, the polynucleotide sequences of interest are plated or arrayed on a microchip substrate. The arrayed sequences are then hybridized with specific DNA probes from cells or tissues of interest. The mRNA source can be the total RNA isolated from a sample, for example, human tumor or tumor cell line and corresponding normal tissue or cell line. Thus, RNA can be isolated from a variety of primary tumors or tumor cell lines. When the mRNA source is a primary tumor, mRNA can be extracted, for example, from frozen tissue samples or paraffin-embedded and fixed (e.g., formalin-fixed) preserved tissue samples, which are routinely prepared and stored in daily clinical practice.
[0265] The expression profile of biomarkers can be measured in either fresh or paraffin-embedded tumor tissue or body fluid using microarray technology.In this method, the polynucleotide sequence of interest is plated or arrayed on a microchip substrate.Then, the arrayed sequence is hybridized with specific DNA probes from cells or tissues of interest.Similar to RT-PCR method, the source of miRNA is typically the total RNA isolated from human tumor or tumor cell line and corresponding normal tissue or cell line, including body fluids such as serum, urine, tear and exosome.Therefore, RNA can be isolated from various sources.When the source of miRNA is primary tumor, miRNA can be extracted from, for example, frozen tissue samples, which are routinely prepared and stored in daily clinical practice.
[0266] The cDNA microarray technique, also known as biochip, DNA chip or gene array, can determine the gene expression level in biological samples.The cDNA or oligonucleotide representing each given gene is immobilized and tagged on a substrate, such as a small chip, bead or nylon membrane, and serves as a probe to indicate whether they are expressed in the biological sample of interest.The simultaneous expression of thousands of genes can be monitored simultaneously.
[0267] In a specific embodiment of microarray technology, PCR-amplified inserts of cDNA clones are applied to a substrate in the form of a high-density array.In one aspect, at least 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 1,500, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 15,000, 20,000, 25,000, 30,000, 35,000, 40,000, 45,000 or at least 50,000 nucleotide sequences are applied to a substrate.Each sequence can correspond to a different gene, or multiple sequences can be arrayed per gene.The microarrayed genes immobilized on the microchip are suitable for hybridization under stringent conditions. Fluorescently labeled cDNA probes may be generated through the incorporation of fluorescent nucleotides by reverse transcription of RNA extracted from tissues of interest. Labeled cDNA probes applied to a chip specifically hybridize to each DNA spot on the array. After stringent washing to remove nonspecifically bound probes, the chip is scanned by confocal laser microscopy or another detection method, such as a CCD camera. Quantitation of hybridization of each arrayed element allows for the determination of corresponding mRNA abundance. cDNA probes generated from two RNA sources and separately labeled with dual-color fluorescence are hybridized to the array in pairs. Thus, the relative abundance of transcripts from the two sources corresponding to each specific gene is determined simultaneously. Miniaturized-scale hybridization allows for convenient and rapid evaluation of expression patterns for a large number of genes. Such methods have been shown to have the sensitivity necessary to detect rare transcripts expressed at a few copies per cell and to reproducibly detect at least approximately two-fold differences in expression levels (Schena et al. (1996) Proc. Natl. Acad. Sci. USA 93(2):106-149).Microarray analysis can be performed by commercially available equipment following manufacturer protocols, including, but not limited to, Affymetrix GeneChip technology (Affymetrix, Santa Clara, CA), Agilent (Agilent Technologies, Inc., Santa Clara, CA), or Illumina (Illumina, Inc., San Diego, CA) microarray technology.
[0268] The development of microarray methods for large-scale analysis of gene expression allows for the systematic search for molecular markers for cancer classification and outcome prediction in diverse tumor types.
[0269] In some embodiments, the Agilent Whole Human Genome Microarray Kit (Agilent Technologies, Inc., Santa Clara, CA) is used. The system is capable of analyzing over 41,000 unique human genes and transcripts, all of which are represented by public domain annotations. The system is used according to the manufacturer's instructions.
[0270] In some embodiments, the Illumina Whole Genome DASL assay (Illumina Inc., San Diego, CA) is used. This system provides a method for simultaneously profiling over 24,000 transcripts in a high-throughput manner from minimal RNA input from both fresh-frozen (FF) and formalin-fixed, paraffin-embedded (FFPE) tissue sources.
[0271] Microarray expression analysis involves identifying whether a gene or gene product is up-regulated or down-regulated compared to a reference. This identification can be performed using a statistical test to determine the statistical significance of any observed differential expression. In some embodiments, statistical significance is determined using a parametric statistical test. Parametric statistical tests can include, for example, fractional factorial design, analysis of variance (ANOVA), t-test, least squares, Pearson correlation, linear regression, nonlinear regression, multiple linear regression, or multiple nonlinear regression. Alternatively, parametric statistical tests can include one-way analysis of variance, two-way analysis of variance, or repeated measures analysis of variance. In other embodiments, statistical significance is determined using a nonparametric statistical test. Examples include, but are not limited to, the Wilcoxon signed-rank test, the Mann-Whitney test, the Kruskal-Wallis test, the Friedman test, Spearman's rank correlation coefficient, Kendall's tau analysis, and nonparametric regression tests. In some embodiments, statistical significance is determined by a p-value of less than about 0.05, 0.01, 0.005, 0.001, 0.0005 or 0.0001.Although the microarray system used in the method described herein may assay thousands of transcripts, data analysis only needs to be performed on the transcript of interest, thereby reducing the problem of multiple comparisons inherent in performing multiple statistical tests.P-values can also be corrected for multiple comparisons, for example, by using Bonferroni correction, its modifications, or other techniques known to those skilled in the art, such as Hochberg correction, Holm-Bonferroni correction, Sidak correction, or Dunnett correction.The degree of differential expression can also be taken into consideration. For example, a gene can be considered differentially expressed if the fold change in expression compared to the control level is at least 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.2, 2.5, 2.7, 3.0, 4, 5, 6, 7, 8, 9, or 10-fold different in the sample compared to the control. Differential expression considers both overexpression and underexpression.If differential expression meets statistical threshold, change fold threshold, or both, gene or gene product can be considered to be up-regulated or down-regulated.For example, the criteria for identifying differential expression can include both a p-value of 0.001 and a change fold of at least 1.5 times (up or down).Those skilled in the art will understand that such statistical measure and threshold measure can be applied to determine differential expression by any molecular profiling technique disclosed herein.
[0272] Various methods described herein utilize many types of microarrays to detect the presence and potentially the quantity of biological entities in a sample. Arrays typically contain addressable moieties that allow the presence of an entity in a sample to be detected, for example, by a binding event. Microarrays include, but are not limited to, DNA microarrays, such as cDNA microarrays, oligonucleotide microarrays and SNP microarrays, microRNA arrays, protein microarrays, antibody microarrays, tissue microarrays, cell microarrays (also called transfection microarrays), chemical compound microarrays, and carbohydrate arrays (glycoarrays). DNA arrays typically contain addressable nucleotide sequences that can bind to sequences present in a sample. MicroRNA arrays, such as the MMChips array from the University of Louisville or commercially available systems from Agilent, can be used to detect microRNAs. Protein microarrays can be used to identify protein-protein interactions, including, but not limited to, identifying substrates of protein kinases, transcription factor protein activation, or targets of biologically active small molecules. Protein arrays may contain arrays of nucleotide sequences that bind to different protein molecules, typically antibodies, or proteins of interest. Antibody microarrays contain antibodies spotted on a protein chip that are used as capture molecules to detect proteins or other biological substances from samples, such as cell or tissue lysates. For example, antibody arrays can be used to detect biomarkers from bodily fluids, such as serum or urine, for diagnostic applications. Tissue microarrays contain separate tissue cores assembled in an array format to enable multiplexed tissue analysis. Cell microarrays, also known as transfection microarrays, contain various capture agents, such as antibodies, proteins, or lipids, that interact with cells to facilitate their capture at addressable locations.Chemical compound microarrays include arrays of chemical compounds that can be used to detect proteins or other biological substances that bind to the compounds. Carbohydrate arrays (glycoarrays) include arrays of carbohydrates that can detect, for example, proteins that bind to sugar moieties. Those skilled in the art will recognize that similar techniques or improvements can be used in accordance with the methods described herein.
[0273] Certain embodiments of the present methods involve multi-well reaction vessels, including, but not limited to, multi-well plates or multi-chamber microfluidic devices, in which multiplex amplification reactions and, in some embodiments, detection, are typically performed in parallel. In certain embodiments, one or more multiplex reactions to generate amplicons are performed in the same reaction vessel, including, but not limited to, multi-well plates such as 96-well, 384-well, or 1536-well plates; or microfluidic devices, such as, but not limited to, TaqMan™ low-density arrays (Applied Biosystems, Foster City, CA). In some embodiments, the massively parallel amplification step involves a plate containing multiple reaction wells, such as, but not limited to, a multi-well reaction vessel, including, but not limited to, a 24-well, 96-well, 384-well, or 1536-well plate; or a multi-chamber microfluidic device, such as, but not limited to, a low-density array, in which each chamber or well contains the appropriate primers, primer sets, and / or reporter probes, as appropriate. Typically, such amplification steps occur in a series of parallel singleplex, 2-plex, 3-plex, 4-plex, 5-plex, or 6-plex reactions, although higher levels of parallel multiplexing are also within the contemplated scope of the present teachings. These methods can include PCR methodologies, such as RT-PCR, in each of the wells or chambers to amplify and / or detect the nucleic acid molecules of interest.
[0274] Low-density arrays can include arrays that detect tens or hundreds of molecules, as opposed to thousands of molecules. These arrays can be more sensitive than high-density arrays. In one embodiment, a low-density array, such as a TaqMan™ low-density array, is used to detect one or more genes or gene products in any of Tables 5-12 of WO2018175501. For example, a low-density array can be used to detect at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, or 100 genes or gene products selected from any of Tables 5-12 of WO2018175501.
[0275] In some embodiments, the disclosed methods involve a microfluidic device, a "lab on a chip," or a micro total analysis system (pTAS). In some embodiments, sample preparation is performed using a microfluidic device. In some embodiments, an amplification reaction is performed using a microfluidic device. In some embodiments, a sequencing or PCR reaction is performed using a microfluidic device. In some embodiments, the nucleotide sequence of at least a portion of the amplification product is obtained using a microfluidic device. In some embodiments, the detecting step involves a microfluidic device, including, but not limited to, a low-density array, such as a TaqMan™ low-density array. Descriptions of exemplary microfluidic devices can be found, inter alia, in published PCT application numbers WO / 0185341 and WO04 / 011666; Kartalov and Quake, Nucl. Acids Res. 32:2873-79, 2004; and Fiorini and Chiu, BioTechniques 38:429-46, 2005.
[0276] Any suitable microfluidic device can be used in the methods described herein. Examples of microfluidic devices that may be used or adapted for use with molecular profiling include U.S. Patent Nos. 7,591,936, 7,581,429, 7,579,136, 7,575,722, 7,568,399, 7,552,741, 7,544,506, 7,541,578, 7,518,726, 7,488,596, 7,485,214, and 7,467,928. No. 7,452,713, No. 7,452,509, No. 7,449,096, No. 7,431,887, No. 7,422,725, No. 7,422,669, No. 7,419,822, No. 7,419,639, 7,413,709, 7,411,184, 7,402,229, 7,390,463, 7,381,471, 7,357,864, 7,351,592, 7,351,380, No. 7,338,637, No. 7,329,391, No. 7,323,140, No. 7,261,824, No. 7,258,837, No. 7,253,003, No. 7,238,324, No. 7,238,255, No. 7 ,233,865, 7,229,538, 7,201,881, 7,195,986, 7,189,581, 7,189,580, 7,189,368, 7,141,978, 7,1 38,062, 7,135,147, 7,125,711, 7,118,910, 7,118,661, 7,640,947, 7,666,361, 7,704,735; U.S. Patent Application Publication No. 20060035243; and International Patent Publication No. WO2010 / 072410, each of which patents or applications is incorporated herein by reference in its entirety.Another example for use with the methods disclosed herein is described in Chen et al., "Microfluidic isolation and transcriptome analysis of serum vesicles," Lab on a Chip, Dec. 8, 2009 DOI: 10.1039 / b916199f.
[0277] Gene expression analysis by massively parallel signature sequencing (MPSS) This method, described by Brenner et al. (2000) Nature Biotechnology 18:630-634, combines non-gel-based signature sequencing with the in vitro cloning of millions of templates on individual microbeads. First, a microbead library of DNA templates is constructed by in vitro cloning. This is followed by the assembly of high-density planar arrays of template-containing microbeads in a flow cell. The free ends of the cloned templates on each microbead are simultaneously analyzed using a fluorescence-based signature sequencing method that does not require DNA fragment separation. This method has been shown to simultaneously and accurately generate hundreds of thousands of gene signature sequences from a cDNA library in a single run.
[0278] MPSS data have many applications. The expression level of nearly every transcript can be quantitatively determined; the abundance of a signature represents the expression level of the gene in the analyzed tissue. Quantitative methods for analyzing tag frequencies and detecting differences between libraries have been published and incorporated into public databases for SAGE™ data, making them applicable to MPSS data. The availability of complete genome sequences allows for direct comparison of signatures to genome sequences, further broadening the utility of MPSS data. Because targets for MPSS analysis are not preselected (as with microarrays), MPSS data can characterize the full complexity of the transcriptome. This is analogous to sequencing millions of ESTs at once, allowing genome sequence data to be used so that the source of MPSS signatures can be easily identified by computational means.
[0279] Serial Analysis of Gene Expression (SAGE) Serial analysis of gene expression (SAGE) is a method that allows for the simultaneous quantitative analysis of multiple gene transcripts without the need to provide individual hybridization probes for each transcript. First, short sequence tags (e.g., approximately 10-14 bp) containing sufficient information to uniquely identify a transcript are generated, provided that the tag is derived from a unique location within each transcript. Multiple transcripts are then linked together to form long, continuous molecules that can be sequenced, revealing the identities of multiple tags simultaneously. The expression pattern of any population of transcripts can be quantitatively assessed by determining the abundance of individual tags and identifying the gene corresponding to each tag. See, e.g., Velculescu et al. (1995) Science 270:484-487; and Velculescu et al. (1997) Cell 88:243-51.
[0280] DNA copy number profiling As long as the resolution is sufficient to identify the copy number variation in biomarkers as described herein, any method that can determine the DNA copy number profile of a particular sample can be used for molecular profiling according to the method described herein.Those skilled in the art will recognize and can use several different platforms to examine whole genome copy number changes with sufficient resolution to identify the copy number of one or more biomarkers of the method described herein.Some of the platforms and techniques are described in the following embodiments.In some embodiments as described herein, next-generation sequencing or ISH technique as described herein or known in the art is used to determine copy number / gene amplification.
[0281] In some embodiments, copy number profile analysis involves amplification of whole genomic DNA using whole genome amplification techniques, which can use strand-displacing polymerases and random primers.
[0282] In some aspects of these embodiments, copy number profile analysis involves hybridization of whole genome amplified DNA with a high-density array. In more specific aspects, the high-density array has 5,000 or more different probes. In another specific aspect, the high-density array has 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, or 1,000,000 or more different probes. In another specific aspect, each of the different probes on the array is an oligonucleotide having a length of about 15 to 200 bases. In another specific aspect, each of the different probes on the array is an oligonucleotide having a length of about 15 to 200, 15 to 150, 15 to 100, 15 to 75, 15 to 60, or 20 to 55 bases.
[0283] In some embodiments, microarrays are employed to help determine copy number profiles for cells from a sample, e.g., a tumor. Microarrays typically contain multiple oligomers (e.g., DNA or RNA polynucleotides or oligonucleotides, or other polymers) synthesized or deposited in an array pattern on a substrate (e.g., a glass support). The support-bound oligomers are "probes" that function to hybridize or bind to sample material (e.g., nucleic acids prepared or obtained from a tumor sample) in a hybridization experiment. The reverse situation can also be applied, in which the sample can be bound to the microarray substrate and the oligomer probes are in solution for hybridization. In use, the array surface is contacted with one or more targets under conditions that promote specific, high-affinity binding of the target to one or more of the probes. In some configurations, the sample nucleic acid is labeled with a detectable label, such as a fluorescent tag, so that the hybridized sample and probes can be detected using a scanning instrument. DNA array technology offers the potential to analyze DNA copy number profiles using a large number (e.g., hundreds of thousands) of different oligonucleotides. In some embodiments, the substrate used for the array is a surface-derivatized glass or silica, or a polymer membrane surface (see, e.g., Z. Guo, et al., Nucleic Acids Res, 22, 5456-65 (1994); U. Maskos, E.M. Southern, Nucleic Acids Res, 20, 1679-84 (1992), and E.M. Southern, et al., Nucleic Acids Res, 22, 1368-73 (1994), each of which is incorporated herein by reference). Modification of the array substrate surface can be achieved by a number of techniques.For example, silicic acid-containing surfaces or metal oxide surfaces can be derivatized with bifunctional silanes, i.e., silanes having a first functional group that allows covalent bonding with the surface (e.g., Si-halogen or Si-alkoxy groups, as found in --SiCl3 or --Si(OCH3)3, respectively), and a second functional group that can provide the desired chemical and / or physical modification to the surface, so as to covalently or non-covalently bind ligands and / or polymers or monomers for biological probe arrays.Silylation derivatization and other surface derivatizations (see, for example, U.S. Patent No. 5,624,711 to Sundberg, U.S. Patent No. 5,266,222 to Willis, and U.S. Patent No. 5,137,765 to Farnsworth, each of which is incorporated herein by reference) are known in the art.Another process for preparing arrays is described in U.S. Patent No. 6,649,348 to Bass et al., assigned to Agilent Corp., which discloses DNA arrays produced by in situ synthesis.
[0284] Polymer array synthesis has also been widely described in the literature, including WO 00 / 58516, U.S. Patent Nos. 5,143,854, 5,242,974, 5,252,743, 5,324,633, 5,384,261, 5,405,783, 5,424,186, 5,451,683, 5,482,867, 5,491, ,074, 5,527,681, 5,550,215, 5,571,639, 5,578,832, 5,593,839, 5,599,695, 5, 624,711, 5,631,734, 5,795,716, 5,831,070, 5,837,832, 5,856,101, 5,858,659, Nos. 5,936,324, 5,968,740, 5,974,164, 5,981,185, 5,981,956, 6,025,601, 6,033,860, 6,040,193, 6,090,555, 6,136,269, 6,269,846 and 6,428,752, 5,412,087, 6,14 Nos. 7,205, 6,262,216, 6,310,189, 5,889,165, and 5,959,098, PCT Application Nos. PCT / US99 / 00730 (International Publication No. WO99 / 36760) and PCT / US01 / 04285 (International Publication No. WO01 / 58593), all of which are incorporated herein by reference in their entirety for all purposes.
[0285] Nucleic acid arrays useful in the present disclosure include, but are not limited to, those commercially available from Affymetrix (Santa Clara, Calif.) under the trade name GeneChip™. Exemplary arrays are shown on the affymetrix.com website. Another microarray supplier is Illumina, Inc. of San Diego, Calif., and exemplary arrays are shown on the illumina.com website.
[0286] In some embodiments, the method of the present invention provides sample preparation.Depending on microarray and the experiment to be carried out, sample nucleic acid can be prepared in several ways by methods known to those skilled in the art.In some aspects as described herein, before or at the same time as genotyping (analysis of copy number profile), sample can be amplified by several mechanisms.The most common amplification procedure used involves PCR. See, e.g., PCR Technology: Principles and Applications for DNA Amplification (Ed. H.A. Erlich, Freeman Press, NY, NY, 1992); PCR Protocols: A Guide to Methods and Applications (Eds. Innis, et al., Academic Press, San Diego, Calif., 1990); Mattila et al., Nucleic Acids Res. 19, 4967 (1991); Eckert et al., PCR Methods and Applications 1, 17 (1991); PCR (Eds. McPherson et al., IRL Press, Oxford); and U.S. Pat. Nos. 4,683,202, 4,683,195, 4,800,159, 4,965,188, and 5,333,675, each of which is incorporated by reference in its entirety for all purposes. In some embodiments, the sample may be amplified on an array (eg, US Pat. No. 6,300,070, incorporated herein by reference).
[0287] Other suitable amplification methods include ligase chain reaction (LCR) (e.g., Wu and Wallace, Genomics 4, 560 (1989), Landegren et al., Science 241, 1077 (1988), and Barringer et al. Gene 89:117 (1990)), transcription amplification (Kwoh et al., Proc. Natl. Acad. Sci. USA 86, 1173 (1989) and WO88 / 10315), self-sustained sequence replication (Guatelli et al., Proc. Nat. Acad. Sci. USA, 87, 1874 (1990) and WO 90 / 06995), selective amplification of target polynucleotide sequences (U.S. Pat. No. 6,410,276), consensus sequence primed polymerase chain reaction (CP-PCR) (U.S. Pat. No. 4,437,975), arbitrarily primed polymerase chain reaction (AP-PCR) (U.S. Pat. Nos. 5,413,909, 5,861,245), and nucleic acid based sequence amplification (NABSA) (see U.S. Pat. Nos. 5,409,818, 5,554,517, and 6,063,603, each of which is incorporated herein by reference). Other amplification methods that may be used are described in US Pat. Nos. 5,242,794, 5,494,810, 4,988,617 and US patent application Ser. No. 09 / 854,317, each of which is incorporated herein by reference.
[0288] Additional methods of sample preparation and techniques for reducing the complexity of nucleic acid samples are described in Dong et al., Genome Research 11, 1418 (2001), U.S. Patent Nos. 6,361,947, 6,391,592, and U.S. Patent Application Nos. 09 / 916,135, 09 / 920,491 (U.S. Patent Application Publication No. 20030096235), 09 / 910,292 (U.S. Patent Application Publication No. 20030082543), and 10 / 013,598.
[0289] The method for carrying out polynucleotide hybridization assay has been fully developed in the art.The procedure and conditions of hybridization assay used in the method described herein vary according to application, and can be selected according to known general binding methods, including the method mentioned in Maniatis et al.Molecular Cloning: A Laboratory Manual (2nd Ed.Cold Spring Harbor, NY, 1989); Berger and Kimmel Methods in Enzymology, Vol. 152, Guide to Molecular Cloning Techniques (Academic Press, Inc., San Diego, Calif., 1987); Young and Davism, PNAS, 80: 1194 (1983). Methods and apparatus for performing repeated and controlled hybridization reactions are described in U.S. Patent Nos. 5,871,928, 5,874,219, 6,045,996, and 6,386,749 and 6,391,623, each of which is incorporated herein by reference.
[0290] Methods as described herein may also involve signal detection of hybridization between the ligands after (and / or during) hybridization. See U.S. Patent Nos. 5,143,854, 5,578,832; 5,631,734; 5,834,758; 5,936,324; 5,981,956; 6,025,601; 6,141,096; 6,185,030; 6,201,639; 6,218,803; and 6,225,625, U.S. Patent Application No. 10 / 389,194, and PCT Application No. PCT / US99 / 06097 (published as WO99 / 47964), each of which is also incorporated by reference herein in its entirety for all purposes.
[0291] Methods and apparatus for signal detection and intensity data processing are described, for example, in U.S. Patent Nos. 5,143,854; 5,547,839; 5,578,832; 5,631,734; 5,800,992; 5,834,758; 5,856,092; 5,902,723; 5,936,324; 5,981,956; 6,025,601; 6,090,555. Nos. 6,141,096, 6,185,030, 6,201,639; 6,218,803; and 6,225,625, U.S. patent application Ser. Nos. 10 / 389,194, 60 / 493,495, and PCT application PCT / US99 / 06097 (published as WO99 / 47964), each of which is also incorporated herein by reference in its entirety for all purposes.
[0292] Immuno-based assays Protein-based detection molecular profiling techniques include immunoaffinity assays based on antibodies selectively immunoreactive with the protein encoded by the mutant gene according to the present method. These techniques include, but are not limited to, immunoprecipitation, Western blot analysis, molecular binding assays, enzyme-linked immunosorbent assays (ELISAs), enzyme-linked immunofiltration assays (ELIFA), fluorescence-activated cell sorting (FACS), etc. For example, any method for detecting the expression of a biomarker in a sample includes contacting the sample with an antibody against the biomarker, or an immunoreactive fragment of the antibody, or a recombinant protein containing the antigen-binding region of the antibody against the biomarker; and then detecting the binding of the biomarker in the sample. Methods for producing such antibodies are known in the art. Antibodies can be used to immunoprecipitate specific proteins from a solution sample or, for example, to immunoblot proteins separated by polyacrylamide gel. Immunocytochemistry can also be used to detect specific protein polymorphisms in tissues or cells. Other well-known antibody-based techniques can also be used, including, for example, ELISA, radioimmunoassays (RIA), immunoradiometric assays (IRMA), and immunoenzymatic assays (IEMA), including sandwich assays using monoclonal or polyclonal antibodies. See, e.g., U.S. Patent Nos. 4,376,110 and 4,486,530, both of which are incorporated herein by reference.
[0293] In an alternative method, a sample may be contacted with an antibody specific for the biomarker under conditions sufficient for the formation of an antibody-biomarker complex, which may then be detected. The presence of a biomarker may be detected in several ways, for example, by Western blotting and ELISA procedures for assaying a wide variety of tissues and samples, including plasma or serum. A wide range of immunoassay techniques using such assay formats are available. See, for example, U.S. Patent Nos. 4,016,043, 4,424,279, and 4,018,653. These include both traditional competitive binding assays as well as non-competitive single-site and two-site or "sandwich" assays. These assays also include direct binding of labeled antibodies to target biomarkers.
[0294] Several variations of the sandwich assay technique exist, and all are intended to be encompassed by the present method. Briefly, in a typical forward assay, an unlabeled antibody is immobilized on a solid substrate, and the test sample is contacted with the bound molecule. After a suitable incubation period sufficient to allow the formation of an antibody-antigen complex, a second antibody specific for the antigen, labeled with a reporter molecule capable of producing a detectable signal, is then added and incubated, allowing sufficient time for the formation of another antibody-antigen-labeled antibody complex. Any unreacted material is washed away, and the presence of the antigen is determined by observation of the signal produced by the reporter molecule. Results can be either qualitative, by simple observation of the visible signal, or quantitated by comparing with a control sample containing known amounts of biomarker.
[0295] Variations on forward assays include simultaneous assays, in which both the sample and labeled antibody are added simultaneously to the bound antibody. These techniques, including any minor variations that will be readily apparent, are well known to those skilled in the art. In a typical forward sandwich assay, a first antibody specific for a biomarker is bound to a solid surface, either covalently or passively. The solid surface is typically glass or a polymer; the most commonly used polymers are cellulose, polyacrylamide, nylon, polystyrene, polyvinyl chloride, or polypropylene. The solid support can be in the form of a tube, beads, a microplate disk, or any other surface suitable for performing immunoassays. The binding process is well known in the art and generally consists of a crosslinking, covalent binding, or physical adsorption step, followed by washing of the polymer-antibody complex in preparation for the test sample. An aliquot of the test sample is then added to the solid phase complex and incubated under appropriate conditions (e.g., room temperature to 40°C, e.g., between 25°C and 32°C, inclusive) for a period of time sufficient to bind any subunits present in the antibody (e.g., 2-40 minutes or, more conveniently, overnight). Following the incubation period, the antibody subunit solid phase is washed, dried, and incubated with a second antibody specific for a portion of the biomarker. The second antibody is linked to a reporter molecule that is used to indicate binding of the second antibody to the molecular marker.
[0296] An alternative method involves immobilizing the target biomarker in a sample and then exposing the immobilized target to a specific antibody, which may or may not be labeled with a reporter molecule. Depending on the amount of target and the signal strength of the reporter molecule, the bound target may be detectable by direct labeling with the antibody. Alternatively, a second, labeled antibody specific to the first antibody is exposed to the target-first antibody complex to form a target-first antibody-second antibody ternary complex. This complex is detected by the signal emitted by the reporter molecule. As used herein, "reporter molecule" refers to a molecule whose chemical nature provides an analytically identifiable signal that allows the antibody bound to the antigen to be detected. The most commonly used reporter molecules in this type of assay are either enzymes, fluorophores, or radionuclide-containing molecules (i.e., radioisotopes) and chemiluminescent molecules.
[0297] In enzyme immunoassays, the enzyme is typically conjugated to the second antibody using glutaraldehyde or periodate. However, as will be readily appreciated, there are a wide variety of different conjugation techniques readily available to those skilled in the art. Commonly used enzymes include horseradish peroxidase, glucose oxidase, β-galactosidase, and alkaline phosphatase, among others. The substrate to be used with a specific enzyme is generally chosen for the production of a detectable color change upon hydrolysis by the corresponding enzyme. Examples of suitable enzymes include alkaline phosphatase and peroxidase. It is also possible to employ fluorogenic substrates that yield fluorescent products rather than the chromogenic substrates mentioned above. In any case, the enzyme-labeled antibody is added to the first antibody-molecular marker complex and allowed to bind, and then excess reagent is washed away. A solution containing the appropriate substrate is then added to the antibody-antigen-antibody complex. The substrate reacts with an enzyme linked to a second antibody to give a qualitative visible signal, which may then be quantified, typically spectrophotometrically, to provide an indication of the amount of biomarker present in the sample. Alternatively, fluorescent compounds such as fluorescein and rhodamine may be chemically coupled to antibodies without altering the antibody's binding capacity. When activated by illumination with light of a specific wavelength, the fluorochrome-labeled antibody absorbs the light energy, inducing an excited state in the molecule, followed by emission of light of a characteristic color that is visually detectable with a light microscope. As in EIA, a fluorescently labeled antibody is allowed to bind to the first antibody-molecular marker complex. After washing away unbound reagents, the remaining ternary complex is then exposed to light of the appropriate wavelength, and the observed fluorescence indicates the presence of the molecular marker of interest. Both immunofluorescence and EIA techniques are very well established in the art. However, other reporter molecules, such as radioisotopes, chemiluminescent, or bioluminescent molecules, may also be employed.
[0298] Immunohistochemistry (IHC) IHC is a process of identifying the location of an antigen (e.g., a protein) in the cells of a tissue using an antibody that specifically binds to the antigen in the tissue. The antigen-binding antibody can be conjugated or fused to a tag that allows its detection, for example, by visualization. In some embodiments, the tag is an enzyme, such as alkaline phosphatase or horseradish peroxidase, that can catalyze a color reaction. The enzyme can be fused to the antibody or non-covalently bound, for example, using the biotin-avidin system. Alternatively, the antibody can be tagged with a fluorophore, such as fluorescein, rhodamine, DyLight Fluor, or Alexa Fluor. The antigen-binding antibody can be directly tagged, or a detection antibody carrying a tag can recognize the antigen-binding antibody itself. IHC can be used to detect one or more proteins. The expression of a gene product can be related to its staining intensity compared to a control level. In some embodiments, a gene product is considered to be differentially expressed if its staining varies by at least 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.2, 2.5, 2.7, 3.0, 4, 5, 6, 7, 8, 9, or 10 fold in a sample compared to a control.
[0299] IHC involves the application of antigen-antibody interactions to histochemistry. In an illustrative example, tissue sections are mounted on slides and incubated with antigen-specific antibodies (polyclonal or monoclonal) (primary reaction). The antigen-antibody signal is then amplified using a second antibody conjugated to peroxidase antiperoxidase (PAP), avidin-biotin-peroxidase (ABC), or avidin-biotin alkaline phosphatase complexes. In the presence of a substrate and chromogen, the enzyme forms a colored deposit at the antibody-antigen binding site. Immunofluorescence is an alternative method for visualizing antigens. In this technique, the primary antigen-antibody signal is amplified using a second antibody conjugated to a fluorescent dye. When UV light is absorbed, the fluorescent dye itself emits light of a longer wavelength (fluorescence), thus enabling the location of the antibody-antigen complex to be identified.
[0300] Epigenetic conditions The molecular profiling method according to the present disclosure also includes measuring epigenetic changes, i.e., gene modifications caused by epigenetic mechanisms, such as changes in methylation status or histone acetylation. Frequently, epigenetic changes result in changes in gene expression levels, which can be detected (at the RNA or protein level, as appropriate) as an indicator of epigenetic changes. Often, epigenetic changes result in the silencing or downregulation of genes, referred to as "epigenetic silencing." The epigenetic changes most frequently investigated with the methods described herein involve determining the DNA methylation status of genes, where increased methylation levels are typically associated with related cancers (as this can cause downregulation of gene expression). Abnormal methylation, sometimes referred to as hypermethylation, of one or more genes can be detected. Typically, the methylation status is determined within appropriate CpG islands, which are often found in the promoter regions of genes. The terms "methylation," "methylation status," or "methylation state" can refer to the presence or absence of 5-methylcytosine at one or more CpG dinucleotides within a DNA sequence. CpG dinucleotides are typically concentrated in the promoter regions and exons of human genes.
[0301] When the reduced gene expression is determined by the methylation status of gene, it can be examined by DNA methylation status or by expression level.One method for detecting epigenetic silencing is to determine that the gene that is expressed in normal cells is less expressed or not expressed in tumor cells.Therefore, the present disclosure provides a molecular profiling method, comprising detecting epigenetic silencing.
[0302] Various assay procedures for directly detecting methylation are known in the art and can be used with the present method. These assays rely on two distinct approaches: bisulfite conversion-based and non-bisulfite-based. Non-bisulfite-based DNA methylation analysis methods rely on the inability of methylation-sensitive enzymes to cleave methylated cytosines at their restriction sites. Bisulfite conversion relies on the treatment of DNA samples with sodium bisulfite, which converts unmethylated cytosines to uracil while maintaining methylated cytosines (Furuichi Y, Wataya Y, Hayatsu H, Ukita T. Biochem Biophys Res Commun. 1970 Dec 9;41(5):1185-91). This conversion results in a change in the original DNA sequence.Methods for detecting such alterations include MS AP-PCR (methylation-sensitive arbitrarily primed polymerase chain reaction), a technique that uses CpG-rich primers to allow a global scan of the genome to focus on regions most likely to contain CpG dinucleotides, as described by Gonzalgo et al., Cancer Research 57:594-599, 1997; Eads et al., Cancer Res. 59:2302-2306, MethyLight™ refers to the art-recognized fluorescence-based real-time PCR technique described by Herman et al., 1999; HeavyMethyl™ assay, an assay in which a methylation-specific blocking probe (also referred to herein as a blocker) that spans the CpG positions between or is spanned by the amplification primers, in the embodiment performed herein, allows for methylation-specific selective amplification of a nucleic acid sample; HeavyMethyl™ MethyLight™, a modification of the MethyLight™ assay in which the MethyLight™ assay is combined with a methylation-specific blocking probe that spans the CpG positions between the amplification primers; Ms-SNuPE (methylation-sensitive single-nucleotide primer extension), an assay described by Herman et al., Proc. Natl. Acad. Sci. USA 93:9821-9826, 1996 and U.S. Pat. No. 5,786,146; COBRA (Multiplex Bisulfite Restriction Analysis), a methylation assay described by Xiong & Laird, Nucleic Acids Res. 25:2532-2534, 1997; and MCA (Methylated CpG Island Amplification), a methylation assay described by Toyota et al., Cancer Res. 59:2307-12, 1999 and WO00 / 26401A1.
[0303] Other techniques for DNA methylation analysis include sequencing, methylation-specific PCR (MS-PCR), melting curve methylation-specific PCR (McMS-PCR), MLPA with or without bisulfite treatment, QAMA, MSRE-PCR, MethyLight, ConLight-MSP, bisulfite conversion-specific methylation-specific PCR (BS-MSP), COBRA (which relies on the use of restriction enzymes to reveal methylation-dependent sequence differences in PCR products of sodium bisulfite-treated DNA), and methylation-sensitive monosodium bisulfite (MSM). These include MS-primer extension conformation analysis (MS-SNuPE), methylation-sensitive single-strand conformation analysis (MS-SSCA), melting curve combined bisulfite restriction analysis (McCOBRA), PyroMethA, HeavyMethyl, MALDI-TOF, MassARRAY, quantification of methylated alleles (QAMA), enzymatic region methylation assay (ERMA), QBSUPT, MethylQuant, quantitative PCR sequencing and oligonucleotide-based microarray systems, pyrosequencing, and Meth-DOP-PCR. Reviews of some useful techniques are provided in Nucleic Acids Research, 1998, Vol. 26, No. 10, 2255-2264; Nature Reviews, 2003, Vol. 3, 253-266; Oral Oncology, 2006, Vol. 42, 5-13, which are incorporated herein in their entireties. Any of these techniques may be used in accordance with the present methods, as appropriate. Other techniques are described in U.S. Patent Application Publication Nos. 20100144836; and 20100184027, which are incorporated by reference in their entireties.
[0304] The DNA binding function of histone proteins is tightly regulated through the activity of various acetylases and deacetylases. Furthermore, histone acetylation and histone deacetylation are associated with malignant progression. See Nature, 429: 457-63, 2004. Methods for analyzing histone acetylation are described in U.S. Patent Application Publication Nos. 20100144543 and 20100151468, which are incorporated herein by reference in their entirety.
[0305] Sequence analysis Molecular profiling according to the present disclosure includes methods for genotyping one or more biomarkers by determining whether an individual has one or more nucleotide variants (or amino acid variants) in one or more genes or gene products. Genotyping one or more genes according to the methods described herein can, in some embodiments, provide more evidence for selecting a treatment.
[0306] Biomarkers as described herein can be analyzed by any method useful for determining alterations in the nucleic acids or proteins they encode. According to one embodiment, one skilled in the art can analyze one or more genes for mutations, including deletion mutants, insertion mutants, frameshift mutants, nonsense mutants, missense mutants, and splice mutants.
[0307] The nucleic acid used for analyzing one or more genes can be isolated from cells in a sample according to standard methodology (Sambrook et al., 1989). For example, the nucleic acid can be genomic DNA, fractionated or whole cell RNA, or miRNA obtained from exosomes or cell surface. When RNA is used, it may be desirable to convert the RNA into complementary DNA. In one embodiment, the RNA is whole cell RNA; in another embodiment, it is poly-A RNA; in another embodiment, it is exosomal RNA. Usually, nucleic acid is amplified. Depending on the assay format for analyzing one or more genes, the specific nucleic acid of interest is identified from the sample directly using amplification, or after amplification, using a second known nucleic acid. The identified product is then detected. In certain applications, detection may be performed by visual means (e.g., ethidium bromide staining of gel). Alternatively, detection may involve indirect identification of the product through chemiluminescence, radioactive or fluorescent labeling, radioscintigraphy, or even through systems that use electrical or thermal impulse signals (Affymax Technology; Bellus, 1994).
[0308] Various types of defects are known to occur in biomarkers such as those described herein. These modifications include, but are not limited to, deletions, insertions, point mutations, and duplications. Point mutations can be silent or result in stop codons, frameshift mutations, or amino acid substitutions. Mutations can occur within and outside the coding regions of one or more genes and can be analyzed according to the methods described herein. Target sites in nucleic acids of interest can contain regions of sequence variation. Examples include, but are not limited to, polymorphisms that exist in different forms, such as single-nucleotide mutations, nucleotide repeats, multi-base deletions (where more than one nucleotide is deleted from a consensus sequence), multi-base insertions (where more than one nucleotide is inserted from a consensus sequence), microsatellite repeats (a small number of nucleotide repeats, typically with 5-1000 repeat units), di-nucleotide repeats, tri-nucleotide repeats, sequence rearrangements (including translocations and duplications), and chimeric sequences (where two sequences from different genetic sources are fused together). Among sequence polymorphisms, the most frequent polymorphisms in the human genome are single-nucleotide variations, also called single nucleotide polymorphisms (SNPs). SNPs are abundant, stable, and widely distributed across the genome.
[0309] Molecular profiling includes methods for haplotyping one or more genes. A haplotype is a set of genetic determinants located on a single chromosome, and typically contains a specific combination of alleles (all alternative sequences of a gene) in a chromosomal region. In other words, a haplotype is the phased sequence information on an individual chromosome. In most cases, the phased SNPs on a chromosome define a haplotype. The combination of haplotypes on a chromosome can determine the genetic profile of a cell. It is the haplotype that determines the association between a specific genetic marker and a disease mutation. Haplotyping can be performed by any method known in the art. Typical methods for scoring SNPs include hybridization microarray or direct gel sequencing, as reviewed in Landgren et al., Genome Research, 8:769-776, 1998. For example, one copy of one or more genes can be isolated from an individual, and the nucleotide at each variant position is determined. Alternatively, allele-specific PCR or similar methods can be used to amplify only one copy of one or more genes in individuals, and the SNP at the variant position of the present disclosure is determined.The Clark method known in the art can also be adopted for haplotyping.High-throughput molecular haplotyping method is also disclosed in Tost et al., Nucleic Acids Res., 30(19):e96 (2002), which is incorporated herein by reference.
[0310] Thus, as will be apparent to those skilled in the art of genetics and haplotyping, additional variants in linkage disequilibrium with the variants and / or haplotypes of the present disclosure can be identified by haplotyping methods known in the art. Additional variants in linkage disequilibrium with the variants or haplotypes of the present disclosure can also be useful for a variety of applications, such as those described below.
[0311] For genotyping and haplotyping purposes, both genomic DNA and mRNA / cDNA can be used, both collectively referred to herein as "genes."
[0312] Numerous techniques for detecting nucleotide variants are known in the art, and all can be used for the methods of the present disclosure. These techniques can be protein-based or nucleic acid-based. In either case, the technique used must be sensitive enough to accurately detect small nucleotide or amino acid variations. Probes labeled with detectable markers are often used. Unless otherwise specified, the specific techniques below can use any suitable marker known in the art, including, but not limited to, radioisotopes, fluorescent compounds, biotin detectable using streptavidin, enzymes (e.g., alkaline phosphatase), enzyme substrates, ligands, and antibodies. See Jablonski et al., Nucleic Acids Res., 14:6115-6128 (1986); Nguyen et al., Biotechniques, 13:116-123 (1992); Rigby et al., J. Mol. Biol., 113:237-251 (1977).
[0313] In nucleic acid-based detection methods, a target DNA sample, i.e., a sample containing genomic DNA, cDNA, mRNA, and / or miRNA corresponding to one or more genes, must be obtained from the individual being tested. Any tissue or cell sample containing genomic DNA, miRNA, mRNA, and / or cDNA (or portions thereof) corresponding to one or more genes can be used. For this purpose, a tissue sample containing cell nuclei and therefore genomic DNA can be obtained from the individual. Blood samples can also be useful, except that only white blood cells and other lymphocytes have nuclei, whereas red blood cells do not have nuclei and contain only mRNA or miRNA. Nevertheless, miRNA and mRNA are also useful because they can be analyzed for the presence of nucleotide variants in their sequences or serve as templates for cDNA synthesis. Tissue or cell samples can be analyzed directly with little or no processing. Alternatively, nucleic acids containing target sequences can be extracted, purified, and / or amplified before being subjected to the various detection procedures described below. In addition to tissue or cell samples, cDNA or genomic DNA from cDNA or genomic DNA libraries constructed using test tissue or cell samples obtained from an individual is also useful.
[0314] To determine the presence or absence of specific nucleotide variants, sequencing of target genomic DNA or cDNA, particularly the region encompassing the nucleotide variant locus to be detected.Various sequencing techniques, including Sanger sequencing and Gilbert chemistry, are generally known and widely used in the art.Pyrosequencing uses a luminometric detection system to monitor DNA synthesis in real time.Pyrosequencing has been shown to be effective for analyzing genetic polymorphisms such as single nucleotide polymorphisms, and can also be used in this method.See Nordstrom et al., Biotechnol.Appl.Biochem.,31(2):107-112 (2000); Ahmadian et al., Anal.Biochem.,280:103-110 (2000).
[0315] Nucleic acid variants can be detected by an appropriate detection process. Non-limiting examples of methods for detection, quantification, sequencing, etc. include mass detection of mass-modified amplicons (e.g., matrix-assisted laser desorption / ionization (MALDI) mass spectrometry and electrospray (ES) mass spectrometry), primer extension methods (e.g., iPLEX™; Sequenom, Inc.), microsequencing (e.g., modifications of primer extension methodology), ligase sequencing (e.g., U.S. Pat. Nos. 5,679,524 and 5,952,174, and WO 01 / 27326), mismatch sequencing (e.g., U.S. Pat. Nos. 5,851,770; 5,958,692; 6,110,684; and 6,183,958), direct DNA sequencing, fragment analysis (FA), restriction fragment length polymorphism (RFLP analysis), allele-specific oligonucleotide (ASO) analysis, methylation-specific PCR (MSPCR), pyrosequencing analysis, acycloprime analysis, reverse dot blot, GeneChip microarray, dynamic allele-specific hybridization (DASH), peptide nucleic acid (PNA) and locked nucleic acid (LNA) probes, TaqMan, molecular beacons, intercalating dyes dye), FRET primers, AlphaScreen, SNPstream, genetic bit analysis (GBA), multiplex minisequencing, SNaPshot, GOOD assay, microarray miniseq, arrayed primer extension (APEX), microarray primer extension (e.g., microarray sequencing), Tag array, coded microspheres, template-dependent incorporation (TDI), fluorescence polarization, colorimetric oligonucleotide ligation assay (OLA), sequence-coded OLA, microarray ligation, ligase chain reaction, Padlock probe, Invader assay, hybridization methods (e.g., hybridization using at least one probe, hybridization using at least one fluorescently labeled probe, etc.), conventional dot blot analysis, single-strand conformation polymorphism analysis (SSCP, e.g., U.S. Pat. Nos. 5,891,625 and 6,013,499; Orita et al., Proc. Natl. Acad. Sci. USA 86: 27776-2770 (1989)), denaturing gr...
Claims
1. A process for obtaining data representative of a test entity, wherein the obtained data comprises: MYC, EP300, U2AF1, ASXL1, MAML2, CNTRL, WRN, CDX2, BCL9, PBX1, PRRX1, INHBA, YWHAE, GNAS, LHFPL6, CASP8, FH, CRKL, MLF1, TRRAP, AKT3, ACKR3, MSI2, PCM1, MNX1, AURKA, GAS7, MN1, SOX10, TCL1A, LMO1, BRD3, SMARCA4, PER1, P comprising data for one or more biomarkers selected from AX7, SBDS, SEPT5, PDGFB, AKT2, TERT, KEAP1, ETV6, TOP1, TLX3, COX6C, NFIB, ARFRP1, ARID1A, MAP2K4, NFKBIA, WWTR1, ZNF217, IL2, NSD3, BRIP1, SDC4, EWSR1, FLT3, FLT1, FAS, CCNE1, RUNX1T1, EZR, FCRL4, BIRC3, and HOXA11; For each of multiple machine learning models that are each trained on the same set of training data, providing the obtained data as input to the machine learning model, the machine learning model being trained to determine a particular class of one or more training entities from a plurality of distinct entity classes based on processing of input data representing each of the one or more training entities, the plurality of distinct entity classes being (1) a response class of test entities that respond to a treatment comprising 5-fluorouracil and leucovorin in combination with oxaliplatin (FOLFOX), and (2) a non-response class of test entities that do not respond to a treatment comprising 5-fluorouracil and leucovorin in combination with oxaliplatin (FOLFOX); processing the provided data through the machine learning model to generate output data; obtaining output data generated by the machine learning model based on processing the provided data by the machine learning model, the obtained output data indicating a particular class of the plurality of different entity classes as an initial classification for the test entity; obtaining output data obtained for each of the plurality of machine learning models, the provided output data including data representative of the initial classification determination for the test entity by each of the plurality of machine learning models; and determining a most likely entity class for the test entity based on the provided output data, wherein the most likely entity class is the reactive class or the non-responsive class.
10. A method for classification of test entities for treatment, comprising:
2. determining a most likely entity class for the test entity based on the provided output data, determining the number of occurrences of each initial classification of the test entity into the particular class of the plurality of different entity classes; and selecting one class of the plurality of different entity classes having the greatest number of occurrences in an initial classification as a most likely entity class for the test entity; 2. The method of claim 1, comprising:
3. accessing a confidence score for each of the plurality of machine learning models; and adjusting the output data generated by each machine learning model based on a confidence score corresponding to each machine learning model. The method of claim 1 further comprising:
4. The method of claim 3 , wherein the confidence score for each of the plurality of machine learning models indicates a historical accuracy of each of the plurality of machine learning models.
5. adjusting the output data generated by each machine learning model based on a confidence score corresponding to the respective machine learning model, increasing a weighting value of output data generated by a first machine learning model of the plurality of machine learning models based on a confidence score corresponding to the first machine learning model; 4. The method of claim 3, comprising:
6. adjusting the output data generated by each machine learning model based on a confidence score corresponding to the respective machine learning model, decreasing a weighting value of output data generated by a first machine learning model of the plurality of machine learning models based on a confidence score corresponding to the first machine learning model.
4. The method of claim 3, comprising:
7. 10. The method of claim 1, wherein at least one machine learning model of the plurality of machine learning models comprises a random forest classification algorithm, a support vector machine, a logistic regression, a k-nearest neighbor model, an artificial neural network, a naive Bayes model, a quadratic discriminant analysis, or a Gaussian process model.
8. 10. The method of claim 1, wherein each machine learning model of the plurality of machine learning models comprises a random forest classification algorithm, a support vector machine, a logistic regression, a k-nearest neighbor model, an artificial neural network, a naive Bayes model, a quadratic discriminant analysis, or a Gaussian process model.
9. The method of claim 1 , wherein at least two of the plurality of machine learning models comprise machine learning models of the same type.
10. 10. The method of claim 9, wherein the same type of machine learning model comprises a random forest classification algorithm, a support vector machine, a logistic regression, a k-nearest neighbor model, an artificial neural network, a naive Bayes model, a quadratic discriminant analysis, or a Gaussian process model.
11. 2. The method of claim 1, wherein at least one of the one or more biomarkers is selected from MYC, EP300, U2AF1, ASXL1, MAML2, CNTRL, WRN, and CDX2.
12. 2. The method of claim 1, wherein at least one of the one or more biomarkers is selected from BCL9, PBX1, PRRX1, INHBA, and YWHAE.
13. 2. The method of claim 1, wherein at least one of the one or more biomarkers is selected from BCL9, PBX1, GNAS, LHFPL6, CASP8, ASXL1, FH, CRKL, MLF1, TRRAP, AKT3, ACKR3, MSI2, PCM1, and MNX1.
14. 2. The method of claim 1, wherein at least one of the one or more biomarkers is selected from GNAS, AURKA, CASP8, ASXL1, CRKL, MLF1, GAS7, MN1, SOX10, TCL1A, LMO1, BRD3, SMARCA4, PER1, PAX7, SBDS, SEPT5, PDGFB, AKT2, TERT, KEAP1, ETV6, TOP1, TLX3, COX6C, NFIB, ARFRP1, ARID1A, MAP2K4, NFKBIA, WWTR1, ZNF217, IL2, NSD3, BRIP1, SDC4, EWSR1, FLT3, FLT1, FAS, CCNE1, RUNX1T1, and EZR.
15. 2. The method of claim 1, wherein at least one of the one or more biomarkers is selected from BCL9, PBX1, PRRX1, INHBA, YWHAE, GNAS, LHFPL6, FCRL4, BIRC3, AURKA, and HOXA11.
16. The method of claim 1 , wherein at least two of the plurality of machine learning models comprise different types of machine learning models.
17. One or more computer readable storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform the method of any one of claims 1 to 16.
18. A system comprising one or more processors configured to perform the method of any one of claims 1 to 16.
19. 1. A data processing apparatus for generating an input data structure for use in training a machine learning model for predicting a subject's therapeutic responsiveness to a particular treatment, the data processing apparatus comprising: one or more processors; and one or more storage devices that store instructions that, when executed by the one or more processors, cause the one or more processors to perform operations, the operations comprising: obtaining, by the data processing device, from a first distributed data source, a first data structure that structures data representing a set of one or more biomarkers associated with the subject, wherein the set of one or more biomarkers is selected from the group consisting of MYC, EP300, U2AF1, ASXL1, MAML2, CNTRL, WRN, CDX2, BCL9, PBX1, PRRX1, INHBA, YWHAE, GNAS, LHFPL6, CASP8, FH, CRKL, MLF1, TRRAP, AKT3, ACKR3, MSI2, PCM1, MNX1, AURKA, GAS7, MN1, S OX10, TCL1A, LMO1, BRD3, SMARCA4, PER1, PAX7, SBDS, SEPT5, PDGFB, AKT2, TERT, KEAP1, ETV6, TOP1, TLX3, COX6C, NFIB, ARFRP1, ARID1A, MAP2K4, NFKBIA, WWTR1, ZNF217, IL2, NSD3, BRIP1, SDC4, EWSR1, FLT3, FLT1, FAS, CCNE1, RUNX1T1, EZR, FCRL4, BIRC3, and HOXA11, and the first data structure includes a key-value pair identifying the subject; storing, by the data processing apparatus, the first data structure in one or more memory devices; obtaining, by the data processing device, from a second distributed data source, a second data structure that structures data representing outcome data for subjects having the set of one or more biomarkers, the outcome data including data identifying a disease or disorder, a treatment, and an indicator of efficacy of the treatment, the second data structure also including key-values that identify the subjects, and the treatment including 5-fluorouracil and leucovorin in combination with oxaliplatin (FOLFOX); storing, by the data processing apparatus, the second data structure in the one or more memory devices; generating, by the data processing device, a labeled training data structure using the first data structure and the second data structure stored in the memory device, the labeled training data structure comprising (i) data representative of the set of one or more biomarkers, the disease or disorder, and a treatment, and (ii) a label providing an indication of the effectiveness of a treatment for the disease or disorder, wherein generating, by the data processing device, using the first data structure and the second data structure includes correlating, by the data processing device, the first data structure structuring data representative of the set of one or more biomarkers associated with the subject based on the key-value identifying the subject, and the second data structure representing outcome data for the subject having the set of one or more biomarkers; and training, by the data processing apparatus, the machine learning model using the labeled training data structure, wherein training the machine learning model using the labeled training data structure comprises providing, by the data processing apparatus, the labeled training data structure to the machine learning model as an input to the machine learning model. The data processing device comprising:
20. The operation is obtaining, by the data processing apparatus, from the machine learning model, an output generated by the machine learning model based on processing by the machine learning model of the labeled training data structure; and determining, by the data processing device, the difference between the output generated by the machine learning model and a label that provides an indication of the effectiveness of a treatment for the disease or disorder.
20. The data processing apparatus of claim 19, further comprising:
21. The operation is adjusting, by the data processing device, one or more parameters of the machine learning model based on the determined difference between the output generated by the machine learning model and a label that provides an indication of the effectiveness of a treatment for the disease or disorder.
21. The data processing apparatus of claim 20, further comprising:
Citation Information
Patent Citations
Medical analysis system
JP2011520206A
Anticancer agent sensitivity-determining marker
JP2016166881A
Microbiome composition as a marker of responsiveness to chemotherapy, and use of microbial modulators (pre-, pro- or symbiotics) to improve the effectiveness of cancer treatment
JP2016539119A
Anticancer agent sensitivity-determining marker
JP2018054621A
Methods for predicting drug responsiveness in cancer patients
US20180202004A1