Next-generation molecular profiling
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- CARIS MPI INC
- Filing Date
- 2025-12-25
- Publication Date
- 2026-04-20
AI Technical Summary
Existing cancer treatments are often one-size-fits-all, leading to inefficiencies and side effects due to lack of personalized approaches based on molecular profiling, particularly in colorectal cancer where treatments like FOLFOX have limited effectiveness and significant side effects.
A machine learning model utilizing comprehensive molecular profiling data and a voting methodology to predict treatment efficacy by combining multiple classifier models, identifying biomarker signatures that indicate response to treatments such as FOLFOX or alternatives like FOLFIRI.
Enhances treatment prediction accuracy by identifying personalized treatment options for colorectal cancer, reducing side effects and improving patient outcomes through targeted therapy selection.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] Claim of priority This application claims the benefits of U.S. Provisional Patent Application No. 62 / 744,082, filed November 30, 2018; U.S. Provisional Patent Application No. 62 / 788,689, filed January 4, 2019; and U.S. Provisional Patent Application No. 62 / 789,495, filed January 7, 2019. The entire contents of the above applications are incorporated herein by reference.
[0002] Technical field This disclosure relates to data structures, data processing and machine learning, and their use in precision medicine, for example, the use of molecular profiling to guide personalized treatment recommendations for victims of various diseases and disorders, including cancer. [Background technology]
[0003] background Drug therapy for cancer patients has long been a challenge. Traditionally, when a patient was diagnosed with cancer, the treating physician would typically select from a predetermined list of treatment options that conventionally corresponded to the patient's observable clinical factors, such as the type and stage of cancer. As a result, cancer patients generally received the same treatment as other patients with the same type and stage of cancer. Because patients with the same type and stage of cancer often respond differently to the same treatment, the efficacy of such treatments becomes determined through trial and error. Moreover, if a patient does not immediately respond to any such "one-size-fits-all" treatment, or if previously successful treatments cease to work, the physician's treatment choices will often be based, at best, on anecdotal evidence.
[0004] Until the late 2000s, limited molecular testing was available to assist physicians in making more informed choices from a list of conventional treatments corresponding to a patient's cancer type, also known as "cancer lineage." For example, a physician of a breast cancer patient, presented with a list of conventional treatment options including Herceptin®, could have tested the patient's tumor for overexpression of the HER2 / neu gene. HER2 / neu was known at the time to be associated with breast cancer and responsiveness to Herceptin®. Approximately one-third of breast cancer patients whose tumors were known to overexpress the HER2 / neu gene showed an initial response to Herceptin® treatment, but the majority of these began to progress within a year. See, for example, Bartsch, R. et al., Trastuzumab in the management of early and advanced stage breast cancer, Biologies. 2007 Mar; 1(1): 19-31 (Non-Patent Literature 1). While this type of molecular testing has helped explain why known treatments for certain types of cancer are more effective than others in treating some patients with that type of cancer, it has neither identified nor ruled out any further treatment options for patients.
[0005] Dissatisfied with the one-size-fits-all approach to treating cancer patients, and facing the reality that many patients' tumors progress and ultimately exhaust all conventional therapies, oncologist Daniel Von Hoff sought to identify further unconventional treatment options for patients. Recognizing the limitations of making treatment decisions based on clinical observation and the limitations of systemic molecular testing, and believing that these limitations were causing effective treatment options to be overlooked, Von Hoff and his colleagues developed a system and method for determining personalized treatment regimens for cancer based on a comprehensive assessment of the molecular characteristics of tumors. Their approach to “molecular profiling” involves collecting molecular information from a patient’s tumor using various testing techniques to create a unique molecular profile regardless of the type of cancer. Physicians can then use the results of this molecular profile to help select candidate therapies for patients, regardless of the stage, anatomical location, or anatomical origin of the cancer cells. See Von Hoff DD, et al., Pilot study using molecular profiling of patients' tumors to find potential targets and select treatments for their refractory cancers. J Clin Oncol. 2010 Nov 20;28(33):4877-83 (Non-Patent Literature 2). Such molecular profiling techniques can suggest promising benefits of treatments that might otherwise be overlooked by treating physicians, and similarly, suggest unpromising benefits of certain treatments, thereby avoiding the time, cost, disease progression, and side effects associated with ineffective treatments. Molecular profiling can be particularly useful in “salvage therapy” settings when a patient has not responded to or developed resistance to multiple treatment regimens. In addition, such techniques can also be used to guide decision-making for first-line and other standard treatment regimens.
[0006] Colorectal cancer (CRC) is the second most common cancer in women and the third most common cancer in men. In 2015, there were 835,000 deaths worldwide due to CRC (see Global Burden of Disease Cancer Collaboration, JAMA Oncol. 2017;3(4):524 (Non-Patent Literature 3)). Surgery is the first-line treatment, but systemic therapy, including 5-fluorouracil / leucovorin in combination with oxaliplatin (FOLFOX) or irinotecan (FOLFIRI), has been shown to be effective in some patients, especially those with metastatic colorectal cancer (Mohelnikova-Duchonova et al., World J Gastroenterol. 2014 Aug 14; 20(30): 10316-10330 (Non-Patent Literature 4)).
[0007] FOLFOX is the standard treatment for metastatic and adjuvant-treated CRCs, but only about half of patients respond to treatment. In addition, 20-100% of FOLFOX-treated patients experience at least one of the following: hair loss, pain or peeling of the palms and soles of the feet, rash, diarrhea, nausea, vomiting, constipation, loss of appetite, dysphagia, mouth pain, heartburn, infection with low white blood cell count, anemia, bruising or bleeding, headache, fatigue, numbness, tingling or pain in the limbs, shortness of breath, cough, and fever; 4-20% experience chest pain, abnormal heartbeat, syncope, injection site reaction, urticaria, weight gain, weight loss, abdominal pain, and internal bleeding (black stool, vomit, or urine). Patients experience at least one of the following side effects (including blood, hemoptysis, vaginal or testicular bleeding, or cerebral bleeding): altered taste, blood clots, liver damage, yellowing of the eyes and skin, allergic reactions, voice changes, confusion, dizziness, weakness, visual impairment, photophobia, tics or spasms, difficulty with motor skills (walking, use of hands, mouth opening, speaking, balance, hearing, smell, eating, sleeping, urination), and hearing loss; less than 3% experience serious side effects, including at least one of cardiac damage and the development of another cancer induced by the treatment.
[0008] A machine learning model can be configured to analyze labeled training data and derive inferences from that training data. Once a machine learning model is trained, a set of unlabeled data may be provided to the machine learning model as input. The machine learning model may process the input data, such as molecular profiling data, and make predictions about the input based on the inferences learned during training. This disclosure provides a “voting” methodology for combining multiple classifier models to achieve more accurate classification than can be achieved by using a single model.
[0009] Comprehensive molecular profiling provides rich data on the molecular state of patient samples. The inventors performed such profiling on well over 100,000 tumor patients from virtually all cancer strains, and tracked patient outcomes and responses to treatment in thousands of these patients. For example, by comparing the inventors' molecular profiling data with patient benefit or lack thereof to treatment and processing it using machine learning algorithms, such as a “voting” methodology, it is possible to identify further biomarker signatures that predict the effectiveness of various treatments. Here, this “next-generation profiling” (NGP) methodology is applied to identify biomarker signatures that predict the benefit of the FOLFOX treatment regimen in patients with colorectal cancer. [Prior art documents] [Non-patent literature]
[0010] [Non-Patent Document 1] Bartsch, R. et al., Trastuzumab in the management of early and advanced stage breast cancer, Biologies. 2007 Mar; 1(1): 19-31 [Non-Patent Document 2] Von Hoff DD, et al., Pilot study using molecular profiling of patients' tumors to find potential targets and select treatments for their refractory cancers. J Clin Oncol. 2010 Nov 20;28(33):4877-83 [Non-Patent Document 3] Global Burden of Disease Cancer Collaboration, JAMA Oncol. 2017;3(4):524 [Non-Patent Document 4] Mohelnikova-Duchonova et al., World J Gastroenterol. 2014 Aug 14; 20(30): 10316-10330 [Overview of the project]
[0011] overview Comprehensive molecular profiling provides rich data on the molecular state of patient samples. By comparing such data with patient responses to treatment, biomarker signatures that predict response or non-response to such treatment can be identified. This technique has been applied to identify biomarker signatures that correlate with the benefit or lack thereof of FOLFOX treatment regimens in patients with colorectal cancer.
[0012] Described herein are methods for training machine learning models to predict the effectiveness of treatments for a disease or disorder of interest that has a specific set of biomarkers.
[0013] Provided herein is a data processing device for generating input data structures for use in training a machine learning model for predicting the effectiveness of treatment for a disease or disorder of interest, wherein the data processing device comprises one or more processors and one or more storage devices that store instructions causing one or more processors to perform an operation when executed by the one or more processors, the operation comprising: the data processing device obtaining one or more biomarker data structures and one or more outcome data structures; the data processing device extracting first data representing one or more biomarkers associated with a subject from one or more biomarker data structures; second data representing a disease or disorder and treatment from one or more outcome data structures; and third data representing the outcome of treatment for the disease or disorder. A data processing device comprising the steps of: extracting data; generating a data structure for input to a machine learning model based on first data representing one or more biomarkers and second data representing a disease or disorder and treatment; providing the generated data structure as input to the machine learning model; obtaining an output generated by the machine learning model based on the machine learning model's processing of the generated data structure; determining the difference between third data representing the treatment outcome for the disease or disorder and the output generated by the machine learning model; and adjusting one or more parameters of the machine learning model based on the difference between the third data representing the treatment outcome for the disease or disorder and the output generated by the machine learning model.
[0014] In some embodiments, a set of one or more biomarkers includes one or more biomarkers listed in any one of Tables 2 to 8. In some embodiments, a set of one or more biomarkers includes each of the biomarkers in Tables 2 to 8. In some embodiments, a set of one or more biomarkers includes at least one of the biomarkers in Tables 2 to 8, and optionally, a set of one or more biomarkers includes biomarkers in Tables 5, 6, 7, 8, or any combination thereof.
[0015] Also provided herein is a data processing device for generating input data structures for use in training a machine learning model for predicting the therapeutic response of a subject to a particular treatment, comprising one or more processors and one or more storage devices for storing instructions causing one or more processors to perform an operation when executed by the one or more processors, the operation comprising: the data processing device obtaining a first data structure from a first distributed data source that structures data representing a set of one or more biomarkers associated with a subject (the first data structure includes key values that identify the subject); the data processing device storing the first data structure in one or more memory devices; and the data processing device obtaining a second data structure from a second distributed data source that structures data representing outcome data of a subject having one or more biomarkers (the outcome data includes disease or disability, The first data structure includes data identifying a treatment and indicators of the treatment's effectiveness, and the second data structure also includes key-values for identifying the subject; the second data structure is stored in one or more memory devices by the data processing device; the third data processing device uses the first and second data structures stored in the memory devices to generate a labeled training data structure which includes (i) data representing one or more sets of biomarkers, a disease or disorder, and a treatment, and (ii) labels providing indicators of the effectiveness of the treatment for the disease or disorder (the third data processing device uses the first and second data structures to generate a labeled training data structure which includes the first data structure which structures data representing one or more sets of biomarkers associated with a subject based on key-values for identifying the subject, and the second data structure which represents outcome data for subjects having one or more biomarkers);and a step of training a machine learning model using the generated labeled training data structure by a data processing device (the step of training a machine learning model using the generated labeled training data structure includes providing, by the data processing device, the generated labeled training data structure as an input to the machine learning model to the machine learning model);
[0016] In some embodiments, the operation further includes obtaining, by the data processing device, an output generated by the machine learning model based on the processing of the machine learning model of the generated labeled training data structure from the machine learning model; and determining, by the data processing device, a difference between the output generated by the machine learning model and a label providing an indicator of the effectiveness of treatment for a disease or disorder.
[0017] In some embodiments, the operation further includes adjusting, by the data processing device, one or more parameters of the machine learning model based on the determined difference between the output generated by the machine learning model and a label providing an indicator of the effectiveness of treatment for a disease or disorder.
[0018] In some embodiments, the set of one or more biomarkers includes one or more biomarkers described in any one of Tables 2 to 8. In some embodiments, the set of one or more biomarkers includes each of the biomarkers in Tables 2 to 8. In some embodiments, the set of one or more biomarkers includes at least one of the biomarkers in Tables 2 to 8, and optionally, the set of one or more biomarkers includes the biomarkers in Tables 5, 6, 7, 8 or any combination thereof.
[0019] In related aspects, provided herein is a method including steps corresponding to each of the operations of the above data processing apparatus. Still further, provided herein is a system including one or more computers and one or more storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform each of the operations described with reference to the above data processing apparatus. Still further, provided herein is a non-transitory computer-readable medium storing software that is executable by one or more computers and that includes instructions that, when executed in such a manner, cause the one or more computers to perform the operations described with reference to the above data processing apparatus.
[0020] In another aspect, provided herein is a method for classifying entities, the method including, for each particular machine learning model of a plurality of machine learning models: i) providing input data representing the type of entity to be classified to the particular machine learning model trained to determine a prediction or classification; ii) obtaining output data representing an entity classification of the entity to an initial entity class of a plurality of candidate entity classes generated by the particular machine learning model based on processing of the input data by the particular machine learning model; providing the output data obtained for each of the plurality of machine learning models to a voting unit (the provided output data includes data representing the initial entity classes determined by each of the plurality of machine learning models); and determining, by the voting unit, an actual entity class for the entity based on the provided output data.
[0021] In some embodiments, the actual entity class for the entity is determined by applying a majority voting principle to the provided output data.
[0022] In some embodiments, the process by which a voting unit determines an actual entity class for an entity based on provided output data includes: determining the occurrence count of each initial entity class among a plurality of candidate entity classes; and selecting the initial entity class having the highest occurrence count among the plurality of candidate entity classes.
[0023] In some embodiments, each machine learning model in a plurality of machine learning models includes a random forest classification algorithm, a support vector machine, a logistic regression, a k-nearest neighbor model, an artificial neural network, a simple Bayesian model, a quadratic discriminant analysis, or a Gaussian process model.
[0024] In some aspects, each machine learning model in a group of machine learning models includes a random forest classification algorithm.
[0025] In some aspects, multiple machine learning models include multiple representations of the same type of classification algorithm.
[0026] In some embodiments, the input data represents (i) entity attributes and (ii) the type of treatment for a disease or disorder.
[0027] In some embodiments, multiple candidate entity classes include reactive or non-reactive classes.
[0028] In some embodiments, the entity attribute includes one or more biomarkers for the entity.
[0029] In some embodiments, one or more biomarkers comprise a panel of fewer genes than all known genes of the entity.
[0030] In some embodiments, one or more biomarkers include a panel of genes containing all known genes for an entity.
[0031] In some embodiments, the input data further includes data representing the type of disease or disorder.
[0032] In connection therewith, provided herein is a system comprising one or more computers and one or more storage media that, when executed by one or more computers, store instructions causing one or more computers to perform each of the operations described with reference to the methods for classifying the entities described above. Furthermore, provided herein is a non-temporary computer-readable medium that is executable by one or more computers and, when executed therewith, stores software that includes instructions causing one or more computers to perform the operations described with reference to the methods for classifying the entities described above.
[0033] In yet another aspect, the foregoing provides a method comprising the steps of: obtaining a biological sample containing cancer-derived cells in a subject; and performing an assay to evaluate at least one biomarker in the biological sample, wherein the biomarker is (a) Group 1, which includes all 1, 2, 3, 4, 5 or 6 of MYC, EP300, U2AF1, ASXL1, MAML2, and CNTRL; (b) Group 2, which includes all 1, 2, 3, 4, 5, 6, 7 or 8 of MYC, EP300, U2AF1, ASXL1, MAML2, CNTRL, WRN, and CDX2; (c) Group 3, including all 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or 14 of BCL9, PBX1, PRRX1, INHBA, YWHAE, GNAS, LHFPL6, FCRL4, HOXA11, AURKA, BIRC3, IKZF1, CASP8 and EP300; (d) Group 4, including all 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or 13 of PBX1, BCL9, INHBA, PRRX1, YWHAE, GNAS, LHFPL6, FCRL4, AURKA, IKZF1, CASP8, PTEN, and EP300; (e) Group 5, which includes all 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12 of BCL9, PBX1, PRRX1, INHBA, GNAS, YWHAE, LHFPL6, FCRL4, PTEN, HOXA11, AURKA, and BIRC3; (f) Group 6, which includes all 1, 2, 3, 4 or 5 of BCL9, PBX1, PRRX1, INHBA, and YWHAE; (g) Group 7, which includes all 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 of BCL9, PBX1, GNAS, LHFPL6, CASP8, ASXL1, FH, CRKL, MLF1, TRRAP, AKT3, ACKR3, MSI2, PCM1 and MNX1; (h)BX1, GNAS, AURKA, CASP8, ASXL1, CRKL, MLF1, GAS7, MN1, SOX10, TCL1A, LMO1, BRD3, SMARCA4, PER1, PAX7, SBDS, SEPT5, PDGFB, AKT2, TERT, KEAP1, ETV6, TOP1, TLX3, COX6C, NFIB, ARFRP1, ARID1A, MAP2K4, NFKBIA, WWTR1, ZNF217, IL2, NSD3, CR Group 8 includes all of EB1, BRIP1, SDC4, EWSR1, FLT3, FLT1, FAS, CCNE1, RUNX1T1 and EZR 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44 or 45; and (i) Group 9, which includes all 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or 11 of BCL9, PBX1, PRRX1, INHBA, YWHAE, GNAS, LHFPL6, FCRL4, BIRC3, AURKA, and HOXA11. A method that includes at least one of the following.
[0034] In some embodiments, biological samples include formalin-fixed paraffin-embedded (FFPE) tissue, fixed tissue, core needle biopsy, aspiration fluid, unstained slides, fresh frozen (FF) tissue, formalin samples, tissue contained in a solution preserving nucleic acids or protein molecules, fresh samples, malignant fluid, body fluids, tumor samples, tissue samples, or any combination thereof.
[0035] In some embodiments, the biological sample includes cells from a solid tumor.
[0036] In some embodiments, the biological sample contains bodily fluids.
[0037] In some embodiments, the body fluids include malignant fluids, pleural fluid, peritoneal fluid, or any combination thereof.
[0038] In some embodiments, body fluids include peripheral blood, serum, plasma, ascites, urine, cerebrospinal fluid (CSF), sputum, saliva, bone marrow, synovial fluid, aqueous humor, amniotic fluid, earwax, breast milk, bronchoalveolar lavage fluid, semen, prostatic fluid, Cowper's gland fluid, bulbourethral gland fluid, female ejaculate, sweat, feces, tears, cystic fluid, pleural fluid, peritoneal fluid, pericardial fluid, lymph, erosion, chyle, bile, interstitial fluid, menstrual secretions, pus, sebum, vomit, vaginal secretions, mucosal secretions, watery stool, pancreatic juice, nasal lavage fluid, bronchopulmonary aspirate, blastocoel fluid, or umbilical cord blood.
[0039] In some embodiments, the assessment involves determining the presence, level, or state of a protein or nucleic acid for each biomarker, optionally including deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or a combination thereof. In some embodiments, (a) the presence, level, or state of a protein is determined using immunohistochemistry (IHC), flow cytometry, immunoassay, antibody or functional fragment thereof, aptamer, or any combination thereof; and / or (b) the presence, level, or state of a nucleic acid is determined using polymerase chain reaction (PCR), in-situ hybridization, amplification, hybridization, microarray, nucleic acid sequencing, di-terminator sequencing, pyrosequencing, next-generation sequencing (NGS; high-throughput sequencing), or any combination thereof.
[0040] In some embodiments, the state of a nucleic acid includes sequence, mutation, polymorphism, deletion, insertion, substitution, translocation, fusion, cleavage, duplication, amplification, repeat, copy number, copy number polymorphism (CNV; copy number variation; CNA), or any combination thereof.
[0041] In some embodiments, the state of a nucleic acid includes the copy number.
[0042] In some embodiments, the method includes the step of performing an assay to determine the copy number of all members of group 1 (i.e., MYC, EP300, U2AF1, ASXL1, MAML2, and CNTRL) or genomic regions adjacent to them.
[0043] In some embodiments, the method includes the step of performing an assay to determine the copy number of all members of group 2 (i.e., MYC, EP300, U2AF1, ASXL1, MAML2, CNTRL, WRN, and CDX2) or genomic regions adjacent to them.
[0044] In some embodiments, the method includes the step of performing an assay to determine the copy number of all members of group 3 (i.e., BCL9, PBX1, PRRX1, INHBA, YWHAE, GNAS, LHFPL6, FCRL4, HOXA11, AURKA, BIRC3, IKZF1, CASP8, and EP300) or genomic regions adjacent to them.
[0045] In some embodiments, the method includes the step of performing an assay to determine the copy number of all members of group 4 (i.e., PBX1, BCL9, INHBA, PRRX1, YWHAE, GNAS, LHFPL6, FCRL4, AURKA, IKZF1, CASP8, PTEN, and EP300) or genomic regions adjacent to them.
[0046] In some embodiments, the method includes the step of performing an assay to determine the copy number of all members of group 5 (i.e., BCL9, PBX1, PRRX1, INHBA, GNAS, YWHAE, LHFPL6, FCRL4, PTEN, HOXA11, AURKA, and BIRC3) or genomic regions adjacent to them.
[0047] In some embodiments, the method includes the step of performing an assay to determine the copy number of all members of group 6 (i.e., BCL9, PBX1, PRRX1, INHBA, and YWHAE) or genomic regions adjacent to them.
[0048] In some embodiments, the method includes the step of performing an assay to determine the copy number of all members of group 7 (i.e., BCL9, PBX1, GNAS, LHFPL6, CASP8, ASXL1, FH, CRKL, MLF1, TRRAP, AKT3, ACKR3, MSI2, PCM1, and MNX1) or genomic regions adjacent to them.
[0049] In some embodiments, the method includes the step of performing an assay to determine the copy number of all members of group 8 (i.e., BX1, GNAS, AURKA, CASP8, ASXL1, CRKL, MLF1, GAS7, MN1, SOX10, TCL1A, LMO1, BRD3, SMARCA4, PER1, PAX7, SBDS, SEPT5, PDGFB, AKT2, TERT, KEAP1, ETV6, TOP1, TLX3, COX6C, NFIB, ARFRP1, ARID1A, MAP2K4, NFKBIA, WWTR1, ZNF217, IL2, NSD3, CREB1, BRIP1, SDC4, EWSR1, FLT3, FLT1, FAS, CCNE1, RUNX1T1, and EZR) or genomic regions adjacent to them.
[0050] In some embodiments, the method includes the step of performing an assay to determine the copy number of all members of group 9 (i.e., BCL9, PBX1, PRRX1, INHBA, YWHAE, GNAS, LHFPL6, FCRL4, BIRC3, AURKA, and HOXA11) or genomic regions adjacent to them.
[0051] In some embodiments, the method includes the step of performing an assay to determine the copy number of (a) at least one or all members of group 1 and group 2 or genomic regions adjacent thereto; (b) at least one or all members of group 3 or genomic regions adjacent thereto; or (c) at least one or all members of group 2, group 6, group 7, group 8 and group 9 or genomic regions adjacent thereto.
[0052] In some embodiments, the method further includes the step of comparing the copy number of a biomarker with a reference copy number (e.g., diploid) to identify a biomarker having copy number polymorphism (CNV).
[0053] In some embodiments, the method further includes the step of generating a molecular profile that identifies a gene or region adjacent thereto that has a CNV.
[0054] In some embodiments, the presence or level of PTEN protein is determined, and optionally, the presence or level of PTEN protein is determined using immunohistochemistry (IHC).
[0055] In some embodiments, the method further includes a step of determining the levels of proteins comprising TOPO1 and one or more mismatch repair proteins (e.g., MLH1, MSH2, MSH6, and PMS2), and optionally, the presence or level of PTEN proteins is determined using immunohistochemistry (IHC).
[0056] In some embodiments, the method further includes the step of comparing the level of one or more proteins to a reference level of those proteins.
[0057] In some embodiments, the method further includes the step of generating molecular profiles that identify proteins having levels different from a reference level, for example, levels significantly different from the reference level.
[0058] In some embodiments, the method further includes a step of selecting a promising benefit treatment based on evaluated biomarkers, optionally the treatment including 5-fluorouracil / leucovorin in combination with oxaliplatin (FOLFOX) or an alternative treatment, and optionally the alternative treatment including 5-fluorouracil / leucovorin in combination with irinotecan (FOLFIRI).
[0059] In some embodiments, the step of selecting a treatment with a promising benefit is based on (a) the copy number determined with respect to the above group; and / or (b) the molecular profile determined as described above.
[0060] In some embodiments, the step of selecting a promising benefit treatment based on the copy number determined for the above group includes the use of a voting module.
[0061] In some embodiments, the voting module is a voting module provided herein.
[0062] In some embodiments, the voting module includes the use of at least one random forest model.
[0063] In some embodiments, the use of the voting module involves applying a machine learning classification model to the number of copies obtained for each of Group 2, Group 6, Group 7, Group 8, and Group 9 (see above), optionally each machine learning classification model being a random forest model, and optionally the random forest models being the random forest models listed in Table 10 below.
[0064] In some embodiments, the subjects have not been previously treated with a treatment of promising benefit.
[0065] In some embodiments, cancer includes metastatic cancer, recurrent cancer, or a combination thereof.
[0066] In some embodiments, the subjects have never received cancer treatment before.
[0067] In some embodiments, the method further includes the step of administering a substance to a target of a promising benefit.
[0068] In some embodiments, progression-free survival (PFS), disease-free survival (DFS), or lifespan is extended by the administration of the treatment.
[0069] In some aspects, cancer includes acute lymphoblastic leukemia; acute myeloid leukemia; adrenocortical carcinoma; AIDS-related cancer; AIDS-related lymphoma; anal cancer; appendiceal cancer; astrocytoma; atypical teratomatoid / rhabdoid tumor; basal cell carcinoma; bladder cancer; brainstem glioma; brain tumors, brainstem gliomas, atypical teratomatoid / rhabdoid tumors of the central nervous system, germ blastomas of the central nervous system, astrocytomas, craniopharyngiomas, ependymocytes, ependymomas, medulloblastomas, medullary epitheliomas, intermediate pineal parenchymal tumors, supratentorial primitive neuroectodermal tumors and pineoblastomas; breast cancer; bronchomas Cervus; Burkitt lymphoma; Carcinoid tumor; Carcinoma of unknown primary origin (CUP); Carcinoid tumor; Carcinoma of unknown primary origin; Atypical teratoid / rhabdoid tumor of the central nervous system; Germ blastoma of the central nervous system; Cervical cancer; Childhood cancer; Chordoma; Chronic lymphocytic leukemia; Chronic myeloproliferative disorder; Colon cancer; Colorectal cancer; Craniopharyngioma; Cutaneous T-cell lymphoma; Endocrine islet cell tumor; Endometrial cancer; Ependymoblastoma; Ependymoma; Esophageal cancer; Nasal neuroblastoma; Ewing's sarcoma; Extracranial germ cell tumor; Extragonadal germ cell tumor; Extrahepatic bile duct cancer; Gallbladder cancer; Gastric cancer (Stomach) cancer; gastrointestinal carcinoid tumor; gastrointestinal stromal cell tumor; gastrointestinal stromal tumor (GIST); gestational trophoblastic neoplasm; glioma; pilocytic cell leukemia; head and neck cancer; heart cancer; Hodgkin lymphoma; hypopharyngeal cancer; intraocular melanoma; pancreatic islet tumor; Kaposi's sarcoma; kidney cancer; Langerhans cell histiocytosis; laryngeal cancer; lip cancer; liver cancer; malignant fibrous histiocytoma bone cancer; medulloblastoma; medullary epithelioma; melanoma; Merkel cell carcinoma; Merkel cell carcinoma; mesothelioma; metastatic squamous cell carcinoma of unknown primary origin; oral cancer; multiple endocrine neoplasia syndrome; multiple myeloma; multiple myeloma / plasmacytic neoplasm; mycosis fungoides; myelodysplastic syndrome; myeloproliferative neoplasm; nasal cavity cancer; nasopharyngeal cancer; neuroblastoma; non-Hodgkin lymphoma; non-melanoma skin cancer; non-small cell lung cancer; oral cancer cancer; oral cavity cancer; oropharyngeal cancer; osteosarcoma; other brain and spinal cord tumors; ovarian cancer; ovarian epithelial carcinoma; ovarian germ cell tumor; low-grade ovarian tumor; pancreatic cancer; papillomatosis; paranasal sinus cancer; parathyroid cancer; pelvic cancer; penile cancer; pharyngeal cancer; intermediate pineal parenchymal tumor; pineoblastoma; pituitary tumor; plasma cell tumor / multiple myeloma; pleuropulmonary blastoma; primary central nervous system (CNS) lymphoma;Primary hepatocellular carcinoma; prostate cancer; rectal cancer; kidney cancer; renal cell carcinoma; renal cell carcinoma; airway cancer; retinoblastoma; rhabdomyosarcoma; salivary gland cancer; Sézary syndrome; small cell lung cancer; small intestine cancer; soft tissue sarcoma; squamous cell carcinoma; cervical squamous cell carcinoma; stomach (gastric) cancer; supratentorial primitive neuroectodermal tumor; T-cell lymphoma; testicular cancer; throat cancer; thymic cancer; thymoma; thyroid cancer; transitional cell carcinoma; transitional cell carcinoma of the renal pelvis and ureter; trophoblastic neoplasm; ureteral cancer; urethral cancer; uterine cancer; uterine sarcoma; vaginal cancer; vulvar cancer; Waldenström type macroglobulinemia; or Wilms' tumor.
[0070] In some aspects, cancer includes acute myeloid leukemia (AML), breast cancer, cholangiocarcinoma, colorectal adenocarcinoma, extrahepatic cholangiocarcinoma, female genital malignancies, gastric adenocarcinoma, gastroesophageal adenocarcinoma, gastrointestinal stromal tumors (GIST), glioblastoma, head and neck squamous cell carcinoma, leukemia, hepatocellular carcinoma, low-grade glioma, bronchoalveolar carcinoma (BAC), non-small cell lung cancer (NSCLC), small cell lung cancer (SCLC), lymphoma, and male reproductive cancer. This includes malignant tumors of the organs, malignant solitary fibrous tumors of the pleura (MSFT), melanoma, multiple myeloma, neuroendocrine tumors, diffuse large B-cell lymphoma, non-epithelial ovarian cancer (non-EOC), ovarian surface epithelial carcinoma, pancreatic adenocarcinoma, pituitary cancer, oligodendroglioma, prostate adenocarcinoma, retroperitoneal or peritoneal cancer, retroperitoneal or peritoneal sarcoma, malignant tumors of the small intestine, soft tissue tumors, thymic carcinoma, thyroid cancer, or uveal melanoma.
[0071] In some aspects, cancer includes colorectal cancer.
[0072] Further provided herein is a method for selecting a treatment for a subject having colorectal cancer, comprising the steps of: obtaining a biological sample containing cells derived from colorectal cancer; performing next-generation sequencing on genomic DNA from the biological sample, thereby obtaining (a) all 1, 2, 3, 4, 5, 6, 7 or 8 of MYC, EP300, U2AF1, ASXL1, MAML2, CNTRL, WRN, and CDX2, Group 2; and (b) all 1, 2, 3, 4, or 5 of BCL9, PBX1, PRRX1, INHBA, and YWHAE. Group 6, (c) BCL9, PBX1, GNAS, LHFPL6, CASP8, ASXL1, FH, CRKL, MLF1, TRRAP, AKT3, ACKR3, MSI2, PCM1 and MNX1, including all 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 of the above; Group 7, (d) BX1, GNAS, AURKA, CASP8, ASXL1, CRKL, MLF1, GAS7, MN1, SOX10, TCL1A, LMO1, BRD3, SMARCA4, PER1, PA X7, SBDS, SEPT5, PDGFB, AKT2, TERT, KEAP1, ETV6, TOP1, TLX3, COX6C, NFIB, ARFRP1, ARID1A, MAP2K4, NFKBIA, WWTR1, ZNF217, IL2, NSD3, CREB1, BRIP1, SDC4, EWSR1, FLT3, FLT1, FAS, CCNE1, RUNX1T1 and EZR 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 A step of determining the copy number for each of the genes or genomic regions adjacent thereto in group 8, which includes all of (e)BCL9, PBX1, PRRX1, INHBA, YWHAE, GNAS, LHFPL6, FCRL4, BIRC3, AURKA, and HOXA11, which includes all of (e)BCL9, PBX1, PRRX1, INHBA, YWHAE, GNAS, LHFPL6, FCRL4, BIRC3, AURKA, and HOXA11;The method comprises the steps of: applying a machine learning classification model to the copy numbers obtained for each of groups 2, 6, 7, 8, and 9 (optionally, each machine learning classification model is a random forest model, and optionally, the random forest models are the random forest models listed in Table 10); obtaining an index from each machine learning classification model of how likely a subject is to benefit from 5-fluorouracil / leucovorin combined with oxaliplatin (FOLFOX); selecting FOLFOX if the majority of the machine learning classification models indicate that the subject is likely to benefit from the treatment, and selecting an alternative treatment to FOLFOX if the majority of the machine learning classification models indicate that the subject is unlikely to benefit from FOLFOX (optionally, the alternative treatment is 5-fluorouracil / leucovorin combined with irinotecan (FOLFIRI)). In some embodiments, the method further comprises the step of administering the selected treatment to the subject.
[0073] Further provided herein are methods for generating molecular profiling reports, comprising the step of preparing a report summarizing the results of performing the above method. In some embodiments, the report includes (a) treatments of promising benefits determined as disclosed above; or (b) selected treatments determined as disclosed above. In some embodiments, the report is computer-generated; a printed report or a computer file; or accessible via a web portal.
[0074] In connection therewith, provided herein is a system for identifying treatments for cancer in a subject, comprising: (a) at least one host server; (b) at least one user interface for accessing data and accessing at least one host server for inputting data; (c) at least one processor for processing the input data; (d) at least one memory coupled to the processor for storing the processed data and instructions for (1) accessing the results of analyzing a biological sample as described above; and (2) determining a treatment of a promising benefit as described above or a selected treatment as described above; and (e) at least one display for displaying a cancer treatment (the treatment being FOLFOX or an alternative thereto, e.g., FOLFIRI).
[0075] In some embodiments, at least one display includes a report containing the results of an analysis of a biological sample and a treatment that has a promising benefit for treating cancer or has been selected for treating cancer.
[0076] In addition, provided herein is a method for providing recommendations for cancer treatment to provide longer progression-free survival, longer disease-free survival, longer overall survival, or life extension, comprising the steps of: obtaining a biological sample containing nucleic acids and / or proteins from an individual diagnosed with cancer; performing molecular testing on the biological sample to determine one or more molecular characteristics selected from the group consisting of nucleic acid sequences of a set of target genes or parts thereof; the presence of copy number variations of a set of target genes; the presence of gene fusions or other genomic alterations; one or more levels of a set of proteins and / or transcripts; and / or epigenetic status of a set of target genes as described herein, thereby generating a molecular profile of cancer; comparing the molecular profile of cancer with a reference molecular profile of cancer of that type; generating a list of molecular characteristics showing differences, e.g., significant differences, when compared with the reference molecular profile; and generating a list of one or more treatment recommendations for the individual based on the list of molecular characteristics showing differences when compared with a reference sequence profile of target genes.
[0077] In some aspects, molecular testing is at least one of next-generation sequencing, Sanger sequencing, ISH, fragment analysis, PCR, IHC, and immunoassays.
[0078] In some embodiments, the biological sample includes cells, tissue samples, blood samples, or combinations thereof.
[0079] In some aspects, molecular testing detects at least one of mutations, polymorphisms, deletions, insertions, substitutions, translocations, fusions, cleavages, duplications, amplifications, or repeats.
[0080] In some embodiments, the nucleic acid sequence includes a deoxyribonucleic acid sequence.
[0081] In some embodiments, the nucleic acid sequence includes a ribonucleic acid sequence.
[0082] Unless otherwise specified, all scientific and technical terms used herein have the same meaning as those commonly understood by those skilled in the art to which this invention pertains. While methods and materials for use in this invention are described herein, other suitable methods and materials known in the art may also be used. Materials, methods, and examples are illustrative and not intended to be limiting. All publications, patent applications, patents, sequences, database entries, and other references cited herein are incorporated herein by reference as a whole. In the event of any conflict, this specification, including definitions, shall prevail.
[0083] [Invention 1001] A data processing device for generating input data structures for use in training a machine learning model to predict the effectiveness of treatment for a target disease or disorder, The data processing device includes one or more processors and one or more storage devices that store instructions causing the one or more processors to perform an operation when executed by the one or more processors, The operation is, A step of obtaining one or more biomarker data structures and one or more outcome data structures using the data processing device; The data processing device performs the following steps: extracts first data representing one or more biomarkers associated with a subject from one or more biomarker data structures; extracts second data representing a disease or disorder and treatment from one or more outcome data structures; and extracts third data representing the outcome of treatment for the disease or disorder; A data processing device generates a data structure for input to a machine learning model based on first data representing one or more biomarkers and second data representing the disease or disorder and treatment; A step of providing the generated data structure as input to the machine learning model using the data processing device; A step of obtaining output generated by a machine learning model based on the processing of the generated data structure by the data processing device; The data processing device performs the step of determining the difference between a third data representing the outcome of treatment for the disease or disorder and the output generated by the machine learning model; and The process involves adjusting one or more parameters of a machine learning model based on the difference between a third set of data representing the outcome of treatment for the disease or disorder and the output generated by the machine learning model, using the data processing device. The data processing device, including the data processing device. [Invention 1002] A data processing apparatus according to the present invention 1001, wherein one or more biomarkers comprise one or more biomarkers listed in any one of Tables 2 to 8. [Invention 1003] A data processing device according to the present invention 1001, wherein one or more sets of biomarkers each include the biomarkers of the present invention 1002. [Invention 1004] A data processing device according to the present invention 1001, wherein one or more sets of biomarkers include at least one of the biomarkers of the present invention 1002, and optionally, one or more sets of biomarkers include the markers in Tables 5, 6, 7, and 8 or any combination thereof. [Invention 1005] A data processing device for generating input data structures for use in training a machine learning model to predict the treatment response of a subject to a specific treatment, The data processing device includes one or more processors and one or more storage devices that store instructions causing the one or more processors to perform an operation when executed by the one or more processors, The operation is, A data processing device comprising the steps of obtaining a first data structure from a first distributed data source, which structures data representing one or more sets of biomarkers associated with a subject, wherein the first data structure includes key values that identify the subject; The data processing device stores the first data structure in one or more memory devices; A step of using a data processing device to obtain a second data structure from a second distributed data source, which structures data representing outcome data for a subject having one or more biomarkers, wherein the outcome data includes data identifying a disease or disorder, a treatment, and an indicator of the effectiveness of the treatment, and the second data structure also includes key values that identify the subject; The data processing device stores the second data structure in one or more memory devices; A data processing device generates a labeled training data structure using the first data structure and the second data structure stored in the memory device, the data processing device generating a labeled training data structure comprising (i) data representing one or more sets of biomarkers, the disease or disorder, and treatment, and (ii) labels providing an indicator of the effectiveness of treatment for the disease or disorder, wherein the data processing device generates using the first data structure and the second data structure, the data processing device correlates a first data structure that structures data representing one or more sets of biomarkers associated with the subject based on key values that identify the subject, with a second data structure that represents outcome data for the subject having the one or more biomarkers; and A step of training a machine learning model using the generated labeled training data structure, wherein the step of training a machine learning model using the generated labeled training data structure includes providing the generated labeled training data structure to the machine learning model as input to the machine learning model. The data processing device, including the data processing device. [Invention 1006] The operation is, A data processing device obtains, from a machine learning model, the output generated by the machine learning model based on the machine learning model's processing of the generated labeled training data structure; and The data processing device determines the difference between the output generated by the machine learning model and a label that provides an indicator of the effectiveness of treatment for a disease or disorder. The data processing apparatus of the present invention 1005 further includes the present invention 1005. [Invention 1007] The operation is, A data processing device adjusts one or more parameters of a machine learning model based on the determined difference between the output generated by the machine learning model and a label that provides an indicator of the effectiveness of treatment for a disease or disorder. A data processing apparatus according to the present invention 1006, further comprising: [Invention 1008] A data processing apparatus according to the present invention 1005, wherein one or more biomarkers include one or more biomarkers listed in any one of Tables 2 to 8, and optionally, one or more biomarkers include markers in Tables 5, 6, 7, and 8 or any combination thereof. [Invention 1009] A data processing apparatus according to the present invention 1005, wherein one or more sets of biomarkers each include the biomarkers of the present invention 1008. [Invention 1010] A data processing device according to the present invention 1005, wherein one or more biomarkers include one of the biomarkers of the present invention 1008. [Invention 1011] A method comprising a step corresponding to each of the operations described in 1001 to 1010 of the present invention. [Invention 1012] A system comprising one or more computers and one or more data storage media that store instructions causing the one or more computers to perform any of the operations 1001 to 1010 of the present invention when executed by the one or more computers. [Invention 1013] Instructions that are executable by one or more computers and, when executed in such manner, cause one or more computers to perform any of the operations described in items 1001 to 1010 of the present invention. A non-temporary computer-readable medium that stores software containing such software. [Invention 1014] A method for classifying entities, Regarding each specific machine learning model among multiple machine learning models, Provide a specific machine learning model, trained to make predictions or classifications, with input data that represents the type of entity to be classified. A step of obtaining output data representing the entity classification of multiple candidate entity classes into initial entity classes, based on the processing of input data by the specific machine learning model; A step of providing output data obtained for each of the plurality of machine learning models to a voting unit, wherein the provided output data includes data representing the initial entity class determined by each of the plurality of machine learning models; and The voting unit then determines the actual entity class for the entity based on the provided output data. The method, including the method described above. [Invention 1015] The method of the present invention 1014, wherein the actual entity class for an entity is determined by applying a majority rule principle to the provided output data. [Invention 1016] The voting unit performs the process of determining the actual entity class for an entity based on the provided output data. The voting unit determines the occurrence count of each initial entity class for multiple candidate entity classes; and The voting unit selects the initial entity class with the highest number of occurrences among the multiple candidate entity classes. A method of the present invention 1014 or 1015, including the method of the present invention. [Invention 1017] The method according to any of items 1014 to 1016 of the present invention, wherein each of the multiple machine learning models includes a random forest classification algorithm, a support vector machine, a logistic regression, a k-nearest neighbor model, an artificial neural network, a simple Bayesian model, a quadratic discriminant analysis, or a Gaussian process model. [Invention 1018] A method according to any of the present invention 1014 to 1016, wherein each of the multiple machine learning models includes a random forest classification algorithm. [Invention 1019] A method according to any of the invention 1014 to 1018, wherein multiple machine learning models include multiple representations of the same type of classification algorithm. [Invention 1020] Any method of the present invention 1014 to 1018 wherein the input data represents (i) entity attributes and (ii) the type of treatment for a disease or disorder. [Invention 1021] The method of the present invention 1020, wherein multiple candidate entity classes include a reactive class or a non-reactive class. [Invention 1022] The method of the present invention 1020 or 1021, wherein the entity attribute includes one or more biomarkers for the entity. [Invention 1023] The method of the present invention 1022, wherein one or more biomarkers comprise a panel of fewer genes than all known genes of the entity. [Invention 1024] The method of the present invention 1022, wherein one or more biomarkers comprise a panel of genes containing all known genes for an entity. [Invention 1025] Any method of the present invention 1020 to 1024, wherein the input data further includes data representing the type of disease or disorder. [Invention 1026] A system comprising one or more computers and one or more data storage media that store instructions causing the one or more computers to perform any of the operations 1014 to 1025 of the present invention when executed by the one or more computers. [Invention 1027] An instruction that is executable by one or more computers, and when executed in such a manner, causes one or more computers to perform any of the operations described in items 1014 to 1025 of the present invention. A non-temporary computer-readable medium that stores software containing such software. [Invention 1028] A process for obtaining a biological sample containing cancer-derived cells in the target; and A step of performing an assay to evaluate at least one biomarker in the biological sample. A method including, The biomarker is (a) Group 1, which includes all 1, 2, 3, 4, 5 or 6 of MYC, EP300, U2AF1, ASXL1, MAML2, and CNTRL; (b) Group 2, which includes all 1, 2, 3, 4, 5, 6, 7 or 8 of MYC, EP300, U2AF1, ASXL1, MAML2, CNTRL, WRN, and CDX2; (c) Group 3, including all 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 or 14 of BCL9, PBX1, PRRX1, INHBA, YWHAE, GNAS, LHFPL6, FCRL4, HOXA11, AURKA, BIRC3, IKZF1, CASP8 and EP300; (d) Group 4, including all 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or 13 of PBX1, BCL9, INHBA, PRRX1, YWHAE, GNAS, LHFPL6, FCRL4, AURKA, IKZF1, CASP8, PTEN, and EP300; (e) Group 5, which includes all 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12 of BCL9, PBX1, PRRX1, INHBA, GNAS, YWHAE, LHFPL6, FCRL4, PTEN, HOXA11, AURKA, and BIRC3; (f) Group 6, which includes all 1, 2, 3, 4 or 5 of BCL9, PBX1, PRRX1, INHBA, and YWHAE; (g) Group 7, which includes all 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 of BCL9, PBX1, GNAS, LHFPL6, CASP8, ASXL1, FH, CRKL, MLF1, TRRAP, AKT3, ACKR3, MSI2, PCM1 and MNX1; (h)BX1, GNAS, AURKA, CASP8, ASXL1, CRKL, MLF1, GAS7, MN1, SOX10, TCL1A, LMO1, BRD3, SMARCA4, PER1, PAX7, SBDS, SEPT5, PDGFB, AKT2, TERT, KEAP1, ETV6, TOP1, TLX3, COX6C, NFIB, ARFRP1, ARID1A, MAP2K4, NFKBIA, WWTR1, ZNF217, IL2, NSD3, CR Group 8 includes all of EB1, BRIP1, SDC4, EWSR1, FLT3, FLT1, FAS, CCNE1, RUNX1T1 and EZR 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44 or 45; and (i) Group 9, which includes all 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or 11 of BCL9, PBX1, PRRX1, INHBA, YWHAE, GNAS, LHFPL6, FCRL4, BIRC3, AURKA, and HOXA11. The method comprising at least one of the following. [Invention 1029] The method of the present invention 1028, wherein the biological sample includes formalin-fixed paraffin-embedded (FFPE) tissue, fixed tissue, core needle biopsy, aspiration fluid, unstained slide, fresh frozen (FF) tissue, formalin sample, tissue contained in a solution for preserving nucleic acids or protein molecules, fresh sample, malignant fluid, body fluid, tumor sample, tissue sample, or any combination thereof. [Invention 1030] The method of the present invention 1028 or 1029, wherein the biological sample contains cells from a solid tumor. [Invention 1031] The method according to invention 1028 or 1029, wherein the biological sample contains bodily fluids. [Invention 1032] A method according to any of items 1028 to 1031 of the present invention, wherein the body fluid includes malignant fluid, pleural fluid, peritoneal fluid, or any combination thereof. [Invention 1033] A method according to any of items 1028 to 1032 of the present invention, wherein the body fluids include peripheral blood, serum, plasma, ascites, urine, cerebrospinal fluid (CSF), sputum, saliva, bone marrow, synovial fluid, aqueous humor, amniotic fluid, earwax, breast milk, bronchoalveolar lavage fluid, semen, prostatic fluid, Cowper's gland fluid, bulbourethral gland fluid, female ejaculate, sweat, feces, tears, cystic fluid, pleural fluid, peritoneal fluid, pericardial fluid, lymph, erosion, chyle, bile, interstitial fluid, menstrual secretions, pus, sebum, vomit, vaginal secretions, mucosal secretions, watery stool, pancreatic juice, nasal lavage fluid, bronchopulmonary aspirate, blastocoel fluid, or umbilical cord blood. [Invention 1034] The evaluation comprises determining the presence, level, or state of a protein or nucleic acid for each biomarker, and optionally, the nucleic acid comprises deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or a combination thereof, any method of the present invention 1028-1033. [Invention 1035] (a) The presence, level, or state of the protein is determined using immunohistochemistry (IHC), flow cytometry, immunoassay, antibody or functional fragment thereof, aptamer, or any combination thereof; and / or (b) The method of the present invention 1034, wherein the presence, level, or state of nucleic acid is determined using polymerase chain reaction (PCR), in situ hybridization, amplification, hybridization, microarray, nucleic acid sequencing, diterminator sequencing, pyrosequencing, next-generation sequencing (NGS; high-throughput sequencing), or any combination thereof. [Invention 1036] The method of the present invention 1035, wherein the nucleic acid state includes sequence, mutation, polymorphism, deletion, insertion, substitution, translocation, fusion, cleavage, duplication, amplification, repeat, copy number, copy number polymorphism (CNV; copy number variation; CNA), or any combination thereof. [Invention 1037] The method of the present invention 1036, wherein the state of the nucleic acid includes the copy number. [Invention 1038] The method of the present invention 1037, comprising the step of performing an assay to determine the copy number of all members of group 1 (i.e., MYC, EP300, U2AF1, ASXL1, MAML2, and CNTRL) or genomic regions adjacent to them. [Invention 1039] The method of the present invention 1037, comprising the step of performing an assay to determine the copy number of all members of group 2 (i.e., MYC, EP300, U2AF1, ASXL1, MAML2, CNTRL, WRN, and CDX2) or genomic regions adjacent thereto. [Invention 1040] The method of the present invention 1037, comprising the step of performing an assay to determine the copy number of all members of group 3 (i.e., BCL9, PBX1, PRRX1, INHBA, YWHAE, GNAS, LHFPL6, FCRL4, HOXA11, AURKA, BIRC3, IKZF1, CASP8, and EP300) or genomic regions adjacent thereto. [Invention 1041] The method of the present invention 1037, comprising the step of performing an assay to determine the copy number of all members of group 4 (i.e., PBX1, BCL9, INHBA, PRRX1, YWHAE, GNAS, LHFPL6, FCRL4, AURKA, IKZF1, CASP8, PTEN, and EP300) or genomic regions adjacent to them. [Invention 1042] The method of the present invention 1037, comprising the step of performing an assay to determine the copy number of all members of group 5 (i.e., BCL9, PBX1, PRRX1, INHBA, GNAS, YWHAE, LHFPL6, FCRL4, PTEN, HOXA11, AURKA, and BIRC3) or genomic regions adjacent thereto. [Invention 1043] The method of the present invention 1037, comprising the step of performing an assay to determine the copy number of all members of group 6 (i.e., BCL9, PBX1, PRRX1, INHBA, and YWHAE) or genomic regions adjacent to them. [Invention 1044] The method of the present invention 1037, comprising the step of performing an assay to determine the copy number of all members of group 7 (i.e., BCL9, PBX1, GNAS, LHFPL6, CASP8, ASXL1, FH, CRKL, MLF1, TRRAP, AKT3, ACKR3, MSI2, PCM1, and MNX1) or genomic regions adjacent thereto. [Invention 1045] The method of the present invention 1037, comprising the step of performing an assay to determine the copy number of all members of group 8 (i.e., BX1, GNAS, AURKA, CASP8, ASXL1, CRKL, MLF1, GAS7, MN1, SOX10, TCL1A, LMO1, BRD3, SMARCA4, PER1, PAX7, SBDS, SEPT5, PDGFB, AKT2, TERT, KEAP1, ETV6, TOP1, TLX3, COX6C, NFIB, ARFRP1, ARID1A, MAP2K4, NFKBIA, WWTR1, ZNF217, IL2, NSD3, CREB1, BRIP1, SDC4, EWSR1, FLT3, FLT1, FAS, CCNE1, RUNX1T1, and EZR) or genomic regions adjacent thereto. [Invention 1046] The method of the present invention 1037, comprising the step of performing an assay to determine the copy number of all members of group 9 (i.e., BCL9, PBX1, PRRX1, INHBA, YWHAE, GNAS, LHFPL6, FCRL4, BIRC3, AURKA, and HOXA11) or genomic regions adjacent thereto. [Invention 1047] (a) at least one or all members of group 1 and group 2, or genomic regions adjacent to them; (b) at least one or all members of group 3 or genomic regions adjacent to them; or (c) At least one or all members of Group 2, Group 6, Group 7, Group 8, and Group 9, or genomic regions adjacent to them. The method of the present invention 1037, comprising the step of performing an assay to determine the copy number of . [Invention 1048] Any method of the present invention 1037 to 1047, further comprising the step of comparing the copy number of a biomarker with a reference copy number (e.g., diploid) to identify a biomarker having copy number polymorphism (CNV). [Invention 1049] The method of the present invention 1048, further comprising the step of generating a molecular profile that identifies a gene or region adjacent thereto having a CNV. [Invention 1050] Any method of the present invention 1028-1049, wherein the presence or level of PTEN protein is determined, and optionally, the presence or level of said PTEN protein is determined using immunohistochemistry (IHC). [Invention 1051] A method of the present invention, any of items 1028-1050, further comprising the step of determining the level of a protein comprising TOPO1 and one or more mismatch repair proteins (e.g., MLH1, MSH2, MSH6, and PMS2), wherein optionally the presence or level of the PTEN protein is determined using immunohistochemistry (IHC). [Invention 1052] The method of the present invention 1050 or 1051, further comprising the step of comparing the level of one or more proteins with the respective reference levels of the one or more proteins. [Invention 1053] The method of the present invention 1052 further comprises the step of generating a molecular profile that identifies proteins having levels different from a reference level, for example, levels significantly different from the reference level. [Invention 1054] The method of any of the invention 1028 to 1053, further comprising the step of selecting a treatment of promising benefit based on evaluated biomarkers, wherein the treatment optionally includes a treatment of 5-fluorouracil / leucovorin in combination with oxaliplatin (FOLFOX) or an alternative treatment, and optionally the alternative treatment includes a treatment of 5-fluorouracil / leucovorin in combination with irinotecan (FOLFIRI). [Invention 1055] The process of selecting a treatment with promising benefits is (a) A determined number of copies of any of invention 1037 to 1047; and / or (b) Molecular profile of Invention 1049 or 1053 The method of the present invention 1054, based on this invention. [Invention 1056] The method of Invention 1055, comprising the step of selecting a treatment with promising benefits based on a determined copy number in any of Invention 1037-1047, wherein the step includes the use of a voting module. [Invention 1057] The method of the present invention 1056, wherein the voting module is one of the inventions 1014 to 1025. [Invention 1058] The method of the present invention 1056 or 1057, wherein the voting module includes the use of at least one random forest model. [Invention 1059] Any method of the Invention 1056-1058, wherein the use of the voting module comprises applying a machine learning classification model to the number of copies obtained for each of Group 2, Group 6, Group 7, Group 8, and Group 9, wherein optionally each machine learning classification model is a random forest model, and optionally the random forest models are listed in Table 10. [Invention 1060] A method according to any of items 1054 to 1059 of the present invention, wherein the subject has not been previously treated with a treatment of promising benefit. [Invention 1061] A method according to any one of items 1028 to 1060 of the present invention, wherein the cancer includes metastatic cancer, recurrent cancer, or a combination thereof. [Invention 1062] A method according to any of the present invention 1028 to 1061, wherein the subject has never previously received cancer treatment. [Invention 1063] Any method of the present invention 1054 to 1062, further comprising the step of administering to a target of a promising benefit. [Invention 1064] The method of the present invention 1063, wherein progression-free survival (PFS), disease-free survival (DFS), or lifespan is extended by the administration of the treatment. [Invention 1065] Cancers include acute lymphoblastic leukemia; acute myeloid leukemia; adrenocortical carcinoma; AIDS-related cancer; AIDS-related lymphoma; anal cancer; appendiceal cancer; astrocytoma; atypical teratoid / rhabdoid tumor; basal cell carcinoma; bladder cancer; brainstem glioma; brain tumors, brainstem gliomas, atypical teratoid / rhabdoid tumors of the central nervous system, germ blastomas of the central nervous system, astrocytomas, craniopharyngiomas, ependymoblastomas, ependymolds, medulloblastomas, medullary epitheliomas, intermediate pineal parenchymal tumors, supratentorial primitive neuroectodermal tumors and pineoblastomas; breast cancer; bronchial tumors; Burkitt Trimphoma; Cancer of Unknown Primary Classification (CUP); Carcinoid Tumor; Carcinoma of Unknown Primary Classification; Atypical Teratoma-like / Rhabdoid Tumor of the Central Nervous System; Central Nervous System Germ Bomboma; Cervical Cancer; Childhood Cancer; Chordoma; Chronic Lymphocytic Leukemia; Chronic Myeloproliferative Disorder; Colon Cancer; Colorectal Cancer; Craniopharyngioma; Cutaneous T-Cell Lymphoma; Endocrine Islet Cell Tumor; Endometrial Cancer; Eependymoma; Esophageal Cancer; Nasal Neuroblastoma; Ewing's Sarcoma; Extracranial Germ Cell Tumor; Extragonadal Germ Cell Tumor; Extrahepatic Bile Duct Cancer; Gallbladder Cancer; Gastric Cancer (Stomach) cancer; gastrointestinal carcinoid tumor; gastrointestinal stromal cell tumor; gastrointestinal stromal tumor (GIST); gestational trophoblastic neoplasm; glioma; pilocytic cell leukemia; head and neck cancer; heart cancer; Hodgkin lymphoma; hypopharyngeal cancer; intraocular melanoma; pancreatic islet tumor; Kaposi's sarcoma; kidney cancer; Langerhans cell histiocytosis; laryngeal cancer; lip cancer; liver cancer; malignant fibrous histiocytoma bone cancer; medulloblastoma; medullary epithelioma; melanoma; Merkel cell carcinoma; Merkel cell carcinoma; mesothelioma; metastatic squamous cell carcinoma of unknown primary origin; oral cancer; multiple endocrine neoplasia syndrome; multiple myeloma; multiple myeloma / plasmacytic neoplasm; mycosis fungoides; myelodysplastic syndrome; myeloproliferative neoplasm; nasal cavity cancer; nasopharyngeal cancer; neuroblastoma; non-Hodgkin lymphoma; non-melanoma skin cancer; non-small cell lung cancer; oral cancer cancer; oral cavity cancer; oropharyngeal cancer; osteosarcoma; other brain and spinal cord tumors; ovarian cancer; ovarian epithelial carcinoma; ovarian germ cell tumor; low-grade ovarian tumor; pancreatic cancer; papillomatosis; paranasal sinus cancer; parathyroid cancer; pelvic cancer; penile cancer; pharyngeal cancer; intermediate pineal parenchymal tumor; pineoblastoma; pituitary tumor; plasma cell tumor / multiple myeloma; pleuropulmonary blastoma; primary central nervous system (CNS) lymphoma; primary hepatocellular carcinoma;Any method of the present invention 1028-1064, comprising prostate cancer; rectal cancer; kidney cancer; renal cell carcinoma; renal cell carcinoma; airway cancer; retinoblastoma; rhabdomyosarcoma; salivary gland cancer; Sézary syndrome; small cell lung cancer; small intestine cancer; soft tissue sarcoma; squamous cell carcinoma; cervical squamous cell carcinoma; stomach (gastric) cancer; supratentorial primitive neuroectodermal tumor; T-cell lymphoma; testicular cancer; pharyngeal cancer; thymic cancer; thymoma; thyroid cancer; transitional cell carcinoma; transitional cell carcinoma of the renal pelvis and ureter; trophoblastic neoplasm; ureteral cancer; urethral cancer; uterine cancer; uterine sarcoma; vaginal cancer; vulvar cancer; Waldenström type macroglobulinemia; or Wilms' tumor. [Invention 1066] Cancers include acute myeloid leukemia (AML), breast cancer, bile duct cancer, colorectal adenocarcinoma, extrahepatic bile duct adenocarcinoma, female genital malignancies, gastric adenocarcinoma, gastroesophageal adenocarcinoma, gastrointestinal stromal tumors (GIST), glioblastoma, head and neck squamous cell carcinoma, leukemia, hepatocellular carcinoma, low-grade glioma, bronchoalveolar carcinoma (BAC), non-small cell lung cancer (NSCLC), small cell lung cancer (SCLC), lymphoma, male reproductive organ malignancies, and malignant solitary fibrous pleura. A method according to any of items 1028 to 1064 of the present invention, comprising tumors (MSFT), melanoma, multiple myeloma, neuroendocrine tumors, diffuse large B-cell lymphoma, non-epithelial ovarian cancer (non-EOC), ovarian surface epithelial carcinoma, pancreatic adenocarcinoma, pituitary cancer, oligodendroglioma, prostatic adenocarcinoma, retroperitoneal or peritoneal cancer, retroperitoneal or peritoneal sarcoma, malignant tumors of the small intestine, soft tissue tumors, thymic carcinoma, thyroid cancer, or uveal melanoma. [Invention 1067] A method according to any one of the present invention 1028 to 1064, wherein the cancer includes colorectal cancer. [Invention 1068] A method for selecting treatment for a patient with colorectal cancer, A process for obtaining a biological sample containing cells derived from colorectal cancer; Next-generation sequencing is performed on the genomic DNA from the biological sample. (a) Group 2, which includes all 1, 2, 3, 4, 5, 6, 7 or 8 of MYC, EP300, U2AF1, ASXL1, MAML2, CNTRL, WRN, and CDX2; (b) Group 6, which includes all 1, 2, 3, 4 or 5 of BCL9, PBX1, PRRX1, INHBA, and YWHAE; (c) Group 7, including all 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 or 15 of BCL9, PBX1, GNAS, LHFPL6, CASP8, ASXL1, FH, CRKL, MLF1, TRRAP, AKT3, ACKR3, MSI2, PCM1 and MNX1; (d)BX1, GNAS, AURKA, CASP8, ASXL1, CRKL, MLF1, GAS7, MN1, SOX10, TCL1A, LMO1, BRD3, SMARCA4, PER1, PAX7, SBDS, SEPT5, PDGFB, AKT2, TERT, KEAP1, ETV6, TOP1, TLX3, COX6C, NFIB, ARFRP1, ARID1A, MAP2K4, NFKBIA, WWTR1, ZNF217, IL2, NSD3, CR Group 8 includes all of EB1, BRIP1, SDC4, EWSR1, FLT3, FLT1, FAS, CCNE1, RUNX1T1 and EZR 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44 or 45; and (e) Group 9, including all 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or 11 of BCL9, PBX1, PRRX1, INHBA, YWHAE, GNAS, LHFPL6, FCRL4, BIRC3, AURKA, and HOXA11. The process of determining the copy number for each of the genes or genomic regions adjacent to them; A step of applying a machine learning classification model to the copy numbers obtained for each of groups 2, 6, 7, 8, and 9, wherein optionally each machine learning classification model is a random forest model, and optionally the random forest models are listed in Table 10; A process of obtaining an indicator from each machine learning classification model of whether the subject is likely to benefit from treatment with 5-fluorouracil / leucovorin combined with oxaliplatin (FOLFOX); and A step of selecting FOLFOX if the majority of machine learning classification models indicate that the subject is likely to benefit from the treatment, and selecting an alternative treatment to FOLFOX if the majority of machine learning classification models indicate that the subject is unlikely to benefit from FOLFOX, wherein the alternative treatment is optionally 5-fluorouracil / leucovorin combined with irinotecan (FOLFIRI). The method, including the method described above. [Invention 1069] The method of the present invention 1068, further comprising the step of administering a selected treatment. [Invention 1070] A method for generating a molecular profiling report, comprising the step of creating a report summarizing the results of performing any of the methods described in 1028 to 1069 of the present invention. [Invention 1071] The report, (a) Therapy for any of the promising benefits of invention 1054-1059; or (b) Selected treatment of the present invention 1068 or 1069 The method of the present invention 1070, including the method of the present invention. [Invention 1072] The method of the present invention 1070 or 1071, wherein the report is computer-generated; is a printed report or computer file; or is accessible via a web portal. [Invention 1073] A system for identifying treatments for cancer in a given subject, (a) at least one host server; (b) at least one user interface for accessing the at least one host server in order to access and input data; (c) at least one processor for processing input data; (d) Processed data and, (1) Access the results of analyzing any biological sample according to any of the inventions 1028 to 1069, and (2) Determine the treatment of any of the promising benefits of Invention 1054 to 1059 or the selected treatment of Invention 1068 or 1069. Commands for and At least one memory connected to the processor for storing; and (e) At least one display for showing cancer treatments such as FOLFOX or its substitute, e.g., FOLFIRI. The system including the above. [Invention 1074] A system of the present invention 1073, wherein at least one display includes a report containing the results of an analysis of a biological sample and a treatment that has a promising benefit for treating cancer or has been selected for treating cancer. Other features and advantages of the present invention will become apparent from the following detailed description and drawings and the appended claims. [Brief explanation of the drawing]
[0084] [Figure 1A] This is a block diagram of an example of a conventional system for training machine learning models. [Figure 1B] This is a block diagram of a system that generates training data structures for training machine learning models to predict the effectiveness of treatments for a target disease or disorder that has a specific set of biomarkers. [Figure 1C] This is a block diagram of a system for using machine learning models trained to predict the effectiveness of treatments for a target disease or disorder that has a specific set of biomarkers. [Figure 1D] This is a flowchart illustrating the process of generating training data for training a machine learning model to predict the effectiveness of treatments for a target disease or disorder that has a specific set of biomarkers. [Figure 1E]This is a flowchart of the process using a machine learning model trained to predict the effectiveness of treatments for a target disease or disorder that has a specific set of biomarkers. [Figure 1F] This is a block diagram of a system for predicting the effectiveness of treatments for a target disease or disorder having a specific set of biomarkers by interpreting the outputs generated by multiple machine learning models using a voting unit. [Figure 1G] These are block diagrams of system components that can be used to implement the systems shown in Figures 2-5. [Figure 1H] A block diagram of an exemplary embodiment of a system for determining personalized medical interventions for cancer using molecular profiling of patient biological specimens is shown. [Figure 2A] This method utilizes molecular profiling of patient biological specimens to determine personalized medical interventions for cancer. [Figure 2B] This is a method for identifying signatures or molecular profiles that can be used to predict the benefits from a treatment. [Figure 2C] This is a flowchart illustrating an exemplary embodiment of an alternative version of (B). [Figure 3A] These are a pair of hazard ratio graphs showing model performance using CNV profiling of eight markers in the case of treatment with FOLFOX. CNA = copy number change. The eight markers were MYC, EP300, U2AF1, ASXL1, MAML2, CNTRL, WRN, and CDX2. [Figure 3B] These are a pair of hazard ratio graphs showing model performance using CNV profiling of eight markers in the case of treatment with FOLFIRI. CNA = copy number change. The eight markers were MYC, EP300, U2AF1, ASXL1, MAML2, CNTRL, WRN, and CDX2. [Figure 3C]These are a pair of hazard ratio graphs showing model performance using CNV profiling of six markers in the case of treatment with FOLFOX. The six markers were MYC, EP300, U2AF1, ASXL1, MAML2, and CNTRL. [Figure 3D] These are a pair of hazard ratio graphs showing model performance using CNV profiling of six markers in the case of treatment with FOLFIRI. The six markers were MYC, EP300, U2AF1, ASXL1, MAML2, and CNTRL. [Figure 3E] Figures 3A and 3B show exemplary random forest decision trees for the eight marker signatures shown. [Figure 4A] This paper describes the development of biosignatures to predict the benefits of the FOLFOX regimen in patients with metastatic colorectal cancer. [Figure 4B] This paper describes the development of biosignatures to predict the benefits of the FOLFOX regimen in patients with metastatic colorectal cancer. [Figure 4C] This paper describes the development of biosignatures to predict the benefits of the FOLFOX regimen in patients with metastatic colorectal cancer. [Figure 4D] This paper describes the development of biosignatures to predict the benefits of the FOLFOX regimen in patients with metastatic colorectal cancer. [Figure 4E] This paper describes the development of biosignatures to predict the benefits of the FOLFOX regimen in patients with metastatic colorectal cancer. [Figure 4F] This paper describes the development of biosignatures to predict the benefits of the FOLFOX regimen in patients with metastatic colorectal cancer. [Figure 4G] This paper describes the development of biosignatures to predict the benefits of the FOLFOX regimen in patients with metastatic colorectal cancer. [Figure 4H] This paper describes the development of biosignatures to predict the benefits of the FOLFOX regimen in patients with metastatic colorectal cancer. [Figure 4I]This paper describes the development of biosignatures to predict the benefits of the FOLFOX regimen in patients with metastatic colorectal cancer. [Figure 4J] This paper describes the development of biosignatures to predict the benefits of the FOLFOX regimen in patients with metastatic colorectal cancer. [Figure 4K] This paper describes the development of biosignatures to predict the benefits of the FOLFOX regimen in patients with metastatic colorectal cancer. [Figure 4L] This paper describes the development of biosignatures to predict the benefits of the FOLFOX regimen in patients with metastatic colorectal cancer. [Figure 4M] This paper describes the development of biosignatures to predict the benefits of the FOLFOX regimen in patients with metastatic colorectal cancer. [Figure 4N] This paper describes the development of biosignatures to predict the benefits of the FOLFOX regimen in patients with metastatic colorectal cancer. [Figure 4O] This paper describes the development of biosignatures to predict the benefits of the FOLFOX regimen in patients with metastatic colorectal cancer. [Figure 5A] This paper describes the development of biosignatures for predicting the benefits of the FOLFOX regimen in patients with colorectal cancer. [Figure 5B] This paper describes the development of biosignatures for predicting the benefits of the FOLFOX regimen in patients with colorectal cancer. [Figure 5C] This paper describes the development of biosignatures for predicting the benefits of the FOLFOX regimen in patients with colorectal cancer. [Modes for carrying out the invention]
[0085] Detailed explanation This specification includes systems, methods, apparatus, and computer programs for training machine learning models and then using the trained machine learning models to predict the effectiveness of treatments for a disease or disorder of interest, as well as methods and systems for identifying therapeutic substances for use in personalized-based therapy by using molecular profiling. In some embodiments, the system may include one or more computer programs on one or more computers located in one or more locations, configured for use in, for example, the methods described herein.
[0086] Aspects of this disclosure relate to a system that generates a set of one or more training data structures that can be used to train a machine learning model to provide various classifications, such as characterizing the phenotype of a biological sample. Characterizing the phenotype may include providing diagnostic, prognostic, theranosis, or other related classifications. For example, the classification may be a classification that predicts the effectiveness of treatment for a disease state or disorder of a subject having a particular set of biomarkers. When trained, the trained machine learning model can be used to process the input data provided by the system and make predictions based on the processed input data. The input data may include a set of features related to the subject, such that the data represents one or more subject biomarkers, and the data represents a disease or disorder. In some embodiments, the input data may further include features representing a proposed treatment type and make predictions that describe a promising response of the subject to the treatment. The predictions may include data output by the machine learning model based on the machine learning model's processing of a particular set of features provided to the machine learning model as input. The data may optionally include data representing one or more subject biomarkers, data representing a disease or disorder, and data representing a proposed treatment type.
[0087] An innovative aspect of this disclosure involves the extraction of specific data from an incoming data stream for use in generating a training data structure. Crucially, this involves the selection of a specific set of one or more biomarkers to include in the training data structure, because the presence, absence, or state of a particular biomarker may indicate a desired classification. For example, a specific biomarker may be selected to determine whether a treatment for a particular disease or disorder is effective or ineffective. In practice, in this disclosure, the applicants present a specific set of biomarkers that, when used in training a machine learning model, produces a trained model that can predict treatment efficacy more accurately than using different sets of biomarkers. See Examples 2-4.
[0088] The system is configured to obtain output data generated by a trained machine learning model based on the processing of data by a machine learning model. In various embodiments, the data includes biological data representing one or more biomarkers, data representing a disease or disorder, and data representing a treatment type. The system can then predict the effectiveness of a treatment for a subject having a particular set of biomarkers. In some embodiments, the disease or disorder may include certain types of cancer, and the treatment for the subject may include one or more therapeutic substances, such as small molecule drugs, biologics, and various combinations thereof. In this setting, the output of the trained machine learning model, generated based on the processing of the input data including the set of biomarkers, disease or disorder, and treatment type, includes data representing the level of responsiveness the subject shows to treatment for the disease or disorder.
[0089] In some embodiments, the output data generated by the trained machine learning model may include probabilities of a desired classification. For example, such probabilities may be the probability that a subject will respond favorably to treatment for a disease or disorder. In other embodiments, the output data may include any output data generated by the trained machine learning model based on the trained machine learning model's processing of the input data. In some embodiments, the input data may include a set of biomarkers, data representing a disease or disorder, and data representing a treatment type.
[0090] In some embodiments, the training data structures generated by this disclosure may include a plurality of training data structures, each containing a field representing a feature vector corresponding to a particular training sample. The feature vector contains a set of features that originate from and represent the training sample. The training sample may include, for example, one or more biomarkers of interest, a disease or disorder of interest, and a proposed treatment for the disease or disorder. The training data structures are flexible because each training data structure may be assigned weights representing each feature of the feature vector. Thus, each training data structure of the plurality of training data structures can be specifically configured so that a particular inference is made by the machine learning model during training.
[0091] Consider a non-limiting example in which a model is trained to predict the promising benefits of a particular treatment for a disease or disorder. As a result, the novel training data structures generated in accordance with this specification are designed to improve the performance of machine learning models. This is because they can be used to train machine learning models to predict the effectiveness of treatments for a disease or disorder of interest having a specific set of biomarkers. As an example, a machine learning model that was unable to make predictions about the effectiveness of treatments for a disease or disorder of interest having a specific set of biomarkers before being trained using the training data structures, systems, and operations described herein can learn to make predictions about the effectiveness of treatments for a disease or disorder of interest by being trained using the training data structures, systems, and operations described herein. Thus, this process employs a otherwise general-purpose machine learning model and transforms that general-purpose machine learning model into a specialized computer for performing the specific task of predicting the effectiveness of treatments for a disease or disorder of interest having a specific set of biomarkers.
[0092] Figure 1A is a block diagram of an example of a conventional technology system 100 for training a machine learning model 110. In some embodiments, the machine learning model may be, for example, a support vector machine. Alternatively, the machine learning model may include neural network models, linear regression models, random forest models, logistic regression models, naive Bayes models, quadratic discriminant analysis models, k-nearest neighbor models, support vector machines, etc. The machine learning model training system 100 may be implemented as a computer program on one or more computers in one or more locations, where the systems, components, and techniques described below can be realized. The machine learning model training system 100 trains the machine learning model 110 using training data items from a database (or dataset) 120 of training data items. A training data item may contain multiple feature vectors. Each training vector may contain multiple values, each corresponding to a specific feature of the training sample that the training vector represents. Training features are sometimes called independent variables. In addition, the system 100 maintains a weight for each feature contained in the feature vector.
[0093] The machine learning model 110 is configured to receive input training data items 122 and process them to produce an output 118. The input training data items may contain multiple features (or independent variables "X") and training labels (or dependent variables "Y"). The machine learning model can be trained using the training items and, if trained, can predict X=f(Y).
[0094] To enable the machine learning model 110 to produce accurate outputs for the received data items, the machine learning model training system 100 can adjust the values of the machine learning model 110's parameters, for example, by training the machine learning model 110 to determine trained values of the parameters from initial values. These parameters derived from the training process may include weights that can be used during the prediction phase using the fully trained machine learning model 110.
[0095] When training the machine learning model 110, the machine learning model training system 100 uses training data items stored in a database (dataset) 120 of labeled training data items. The database 120 stores sets of multiple training data items, and each training data item in a set of multiple training items is associated with its respective label. Generally, the label for a training data item identifies the correct classification (or prediction) for the training data item, i.e., the classification that should be identified as the classification of the training data item by the output value generated by the machine learning model 110. Referring to Figure 1A, a training data item 122 may be associated with a training label 122a.
[0096] The machine learning model training system 100 trains the machine learning model 110 to optimize an objective function. Optimizing the objective function may include, for example, minimizing a loss function 130. Generally, the loss function 130 is a function dependent on (i) the output 118 produced by the machine learning model 110 by processing a given training data item 122, and (ii) the label 122a for the training data item 122, i.e., the target output that the machine learning model 110 should have produced by processing the training data item 122.
[0097] A conventional machine learning model training system 100 can train a machine learning model 110 to minimize a (cumulative) loss function 130 by repeatedly adjusting the parameter values of the machine learning model 110 by performing multiple iterations of conventional machine learning model training techniques, such as hinge loss, stochastic gradient descent, and stochastic gradient descent with backpropagation, on training data items from a database 120. The fully trained machine learning model 110 can then be deployed as a predictive model that can be used to make predictions based on unlabeled input data.
[0098] Figure 1B is a block diagram of a system 200 that generates a training data structure for training a machine learning model to predict the effectiveness of treatments for a target disease or disorder having a specific set of biomarkers.
[0099] System 200 includes two or more distributed computers 210, 310, a network 230, and an application server 240. The application server 240 includes an extraction unit 242, a memory unit 244, a vector generation unit 250, and a machine learning model 270. The machine learning model 270 may include one or more of the following: a vector support machine, a neural network model, a linear regression model, a random forest model, a logistic regression model, a simple Bayes model, a quadratic discriminant analysis model, a k-nearest neighbor model, a support vector machine, etc. Each distributed computer 210, 310 may include a smartphone, tablet computer, laptop computer, or desktop computer, etc. Alternatively, each distributed computer 210, 310 may include a server computer that receives data input by one or more terminals 205, 305, etc. Terminal computers 205, 305 may include any user device, including smartphones, tablet computers, laptop computers, desktop computers, etc. Network 230 may include one or more networks 230, such as a LAN, WAN, wired Ethernet network, wireless network, cellular network, internet, or any combination thereof.
[0100] The application server 240 is configured to obtain, or otherwise receive, data records 220, 222, 224, and 320 provided by one or more distributed computers, such as the first distributed computer 210 and the second distributed computer 310, using the network 230. In some embodiments, each of the distributed computers 210, 310 may provide different types of data records 220, 222, 224, and 320. For example, the first distributed computer 210 may provide biomarker data records 220, 222, and 224 representing the biomarkers of interest, and the second distributed computer 310 may provide outcome data 320 representing the outcome data of interest obtained from the outcome database 312.
[0101] Biomarker data records 220, 222, and 224 may contain any type of biomarker data describing the biometric attributes of the subject. For example, Figure 1B shows a biomarker data record containing data records representing a DNA biomarker 220, a protein biomarker 222, and an RNA data biomarker 224. Each of these biomarker data records may contain a data structure having fields that structure information 220a, 222a, and 224a describing the biomarker of the subject, e.g., the DNA biomarker 220a, the protein biomarker 222a, or the RNA biomarker 224a. However, the disclosure is not limited to this. For example, biomarker data records 220, 222, and 224 may contain next-generation sequencing data, such as DNA alterations. Such next-generation sequencing data may include single variants, insertions and deletions, substitutions, translocations, fusions, cleavages, duplications, amplifications, loss, copy numbers, repeats, total gene mutations, microsatellite instability, and the like. Alternatively or additionally, biomarker data records 220, 222, and 224 may also include in-situ hybridization data, such as DNA copies. Such in-situ hybridization data may include gene copies, gene translocations, etc. Alternatively or additionally, biomarker data records 220, 222, and 224 may also include RNA data, such as gene expression or gene fusion, including whole transcriptome sequencing. Alternatively or additionally, biomarker data records 220, 222, and 224 may include protein expression data, such as that obtained using immunohistochemistry (IHC). Alternatively or additionally, biomarker data records 220, 222, and 224 may include ADAPT data, such as complex numbers.
[0102] In some embodiments, a set of one or more biomarkers includes one or more biomarkers listed in any one of Tables 2-8. However, the disclosure is not limited in this way, and other types of biomarkers may be used instead. For example, biomarker data may be obtained by whole exome sequencing, whole transcriptome sequencing, or a combination thereof.
[0103] An outcome data record 320 may describe the outcome of treatment for a subject. For example, an outcome data record 320 obtained from an outcome database 312 may include one or more data structures having fields that structure the subject's data attributes, such as disease or disorder 320a, treatment 320a received by the subject for the disease or disorder, treatment outcome 320a, or a combination of both. In addition, the outcome data record 320 may also include fields that structure data attributes describing the details of the treatment and the subject's response to the treatment. Examples of diseases or disorders may include, for example, certain types of cancer. Types of treatment may include, for example, the type of drug, biologic or other treatment received by the subject for the disease or disorder included in the outcome data record 320. Treatment outcomes may include data representing the subject's outcome of the treatment regimen, such as benefit, moderate benefit, or no benefit. In some embodiments, treatment outcomes may include the type of cancerous tumor at the end of treatment, such as the amount the tumor shrank, the overall size of the tumor after treatment, etc. Alternatively or additionally, treatment outcomes may include the number or ratio of white blood cells, red blood cells, etc. Treatment details may include dosage, e.g., the amount of medication taken, drug regimen, number of missed doses, etc. Thus, while the example in Figure 1B shows that outcome data may include disease or disorder, treatment, and treatment outcomes, outcome data may also include other types of information as described herein. Furthermore, outcome data does not need to be limited to human “patients.” Instead, outcome data records 220, 222, 224 and biometric data record 320 may be associated with any desired subject, including any non-human organism.
[0104] In some embodiments, each of the data records 220, 222, 224, and 320 may include keyed data that enables the application server 240 to correlate the data records from each distributed computer. The keyed data may include, for example, data representing a subject identifier. The subject identifier may include any form of data that identifies the subject, which allows the subject's biomarkers to be associated with the subject's outcome data.
[0105] The first distributed computer 210 may provide biomarker data records 220, 222, and 224 to the application server 240 (208). The second distributed computer 310 may provide outcome data record 320 to the application server 240 (210). The application server 240 may provide biomarker data record 220 and outcome data records 220, 222, and 224 to the extraction unit 242.
[0106] The extraction unit 242 can process the received biomarker data 220, 222, 224 and outcome data record 320 to extract data 220a-1, 222a-1, 224a-1, 320a-1, 320a-2, and 320a-3 that can be used to train a machine learning model. For example, the extraction unit 242 can obtain data structured by fields in the data structure of biometric data records 220, 222, 224, data structured by fields in the data structure of outcome data record 320, or a combination thereof. The extraction unit 242 can perform one or more information extraction algorithms, such as keyed data extraction, pattern matching, and natural language processing, to identify and obtain data 220a-1, 222a-1, 224a-1, 320a-1, 320a-2, and 320a-3 from biometric data records 220, 222, 224 and outcome data record 320, respectively. The extraction unit 242 may provide the extracted data to the memory unit 244. The extracted data unit may be stored in the memory unit 244, such as flash memory (unlike a hard disk), to improve data access time and reduce latency in accessing the extracted data, thereby improving system performance. In some embodiments, the extracted data may be stored in the memory unit 244 as an in-memory data grid.
[0107] More specifically, the extraction unit 242 may be configured to filter portions of the biomarker data records 220, 222, 224 and outcome data records 320, which are used to generate an input data structure 260 for processing by the machine learning model 270, from portions of the outcome data records 320 that are used as labels for the generated input data structure 260. Such filtering involves the extraction unit 242 separating the biomarker data from a first portion of outcome data, which includes disease or disorder, treatment, treatment details, or a combination thereof, from the treatment outcome. The application server 240 can then generate the input data structure 260 using the biomarker data 220a-1, 222a-1, 224a-1, 320a-1, 320a-2 and the first portion of outcome data, which includes disease or disorder 320a-1, treatment 320a-2, treatment details (not shown in Figure 1B), or a combination thereof. In addition, the application server 240 can also use the second part of the outcome data describing the treatment results 320a-3 as labels for the generated data structure.
[0108] The application server 240 processes the extracted data stored in the memory unit 244 and can correlate the biomarker data 220a-1, 222a-1, and 224a-1 extracted from biomarker data records 220, 222, and 224 with the first portions of the outcome data 320a-1 and 320a-2. The purpose of this correlation is to cluster the biomarker data with the outcome data so that the target outcome data is clustered with the target biomarker data. In some embodiments, the correlation between the biomarker data and the first portions of the outcome data may be based on keyed data associated with each of the biomarker data records 220, 222, and 224 and the outcome data record 320. For example, the keyed data may include a target identifier.
[0109] The application server 240 provides the extracted biomarker data 220a-1, 222a-1, 224a-1 and the extracted first portion of the outcome data 320a-1, 320a-2 as input to the vector generation unit 250. The vector generation unit 250 is used to generate a data structure based on the extracted biomarker data 220a-1, 222a-1, 224a-1 and the extracted first portion of the outcome data 320a-1, 320a-2. The generated data structure is a feature vector 260 containing multiple values that numerically represent the extracted first portion of the extracted biomarker data 220a-1, 222a-1, 224a-1 and the outcome data 320a-1, 320a-2. The feature vector 260 may contain fields for each type of biomarker and each type of outcome data. For example, feature vector 260 may include one or more fields corresponding to (i) one or more types of next-generation sequencing data, e.g., single variants, insertions and deletions, substitutions, translocations, fusions, cleavages, duplications, amplifications, losses, copy numbers, repeats, total gene mutations, microsatellite instability, (ii) one or more types of in-situ hybridization data, e.g., DNA copies, gene copies, gene translocations, (iii) one or more types of RNA data, e.g., gene expression or gene fusion, (iv) one or more types of protein data, such as obtained using immunohistochemistry, (v) one or more types of ADAPT data, e.g., complex numbers, and (vi) one or more types of outcome data, e.g., disease or disorder, treatment type, details of each treatment type, etc.
[0110] The vector generation unit 250 is configured to assign a weight to each field of the feature vector 260 that indicates the extent to which the extracted first portion of the extracted biomarker data 220a-1, 222a-1, 224a-1 and outcome data 320a-1, 320a-2 contains the data represented by each field. In one embodiment, for example, the vector generation unit 250 may assign "1" to each field of the feature vector corresponding to features found in the extracted first portion of the extracted biomarker data 220a-1, 222a-1, 224a-1 and outcome data 320a-1, 320a-2. In such an embodiment, the vector generation unit 250 may also assign "0" to each field of the feature vector corresponding to features not found in the extracted first portion of the extracted biomarker data 220a-1, 222a-1, 224a-1 and outcome data 320a-1, 320a-2. The output of the vector generation unit 250 may include data structures such as feature vectors 260, which can be used to train the machine learning model 270.
[0111] The application server 240 can label the training feature vectors 260. Specifically, the application server can use the extracted second portion of patient outcome data 320a-3 to label the generated feature vectors 260 with treatment outcomes 320a-3. The labels of the training feature vectors 260 generated based on treatment outcomes 320a-3 can provide an indicator of the effectiveness of treatment 320a-2 for a target disease or disorder 320a-1 determined by a specific set of biomarkers 220a-1, 222a-1, and 224a-1 (each described in the training data structure 260).
[0112] The application server 240 can train the machine learning model 270 by providing the feature vector 260 as input to the machine learning model 270. The machine learning model 270 can process the generated feature vector 260 and produce an output 272. The application server 240 can use the loss function 280 to determine the amount of error between the output 272 of the machine learning model 280 and the value specified by the training label (generated based on the second part of the extracted patient outcome data describing the treatment outcome 320a-3). The output 282 of the loss function 280 can be used to tune the parameters of the machine learning model 282.
[0113] In some embodiments, tuning the parameters of the machine learning model 270 may include manual tuning of the machine learning model parameters. Alternatively, in some embodiments, the parameters of the machine learning model 270 may be automatically tuned by one or more algorithms executed by the application server 242.
[0114] The application server 240 may perform multiple iterations of the process described above, with reference to Figure 1B, for each outcome data record 320 stored in the outcome database corresponding to the set of biomarker data of interest. This may include hundreds, thousands, tens of thousands, hundreds of thousands, millions, or more iterations, until all outcome data records 320 with corresponding sets of biomarker data of interest stored in the outcome database 312 are exhausted, until the machine learning model 270 is trained to a specific error range, or a combination thereof. The machine learning model 270 is trained to a specific error range, for example, when it can predict the effectiveness of a treatment for a subject with biomarker data based on a set of unlabeled biomarker data, disease or disability data, and treatment data. Effectiveness may include, for example, probability, a general indicator of whether the treatment is successful or unsuccessful.
[0115] Figure 1C is a block diagram of a system for using a machine learning model trained to predict the effectiveness of treatments for a target disease or disorder that has a specific set of biomarkers.
[0116] The machine learning model 370 includes a machine learning model trained using the process described with reference to the system in Figure 1B above. The trained machine learning model 370 can predict the level of effectiveness of a treatment when treating a target disease or disorder having a biomarker, based on input feature vectors representing one or more biomarkers, a disease or disorder, and a treatment. In some embodiments, “treatment” may include a drug, treatment details (e.g., dosage, regimen, missed doses, etc.) or any combination thereof.
[0117] The application server 240 hosting the machine learning model 370 is configured to receive unlabeled biomarker data records 320, 322, and 324. The biomarker data records 320, 322, and 324 include one or more data structures having fields that structure data representing one or more specific biomarkers, such as a DNA biomarker 320a, a protein biomarker 322a, an RNA biomarker 324a, or any combination thereof. As described above, the received biomarker data record may include biomarkers of types not shown in Figure 1C, such as (i) one or more types of next-generation sequencing data, e.g., single variants, insertions and deletions, substitutions, translocations, fusions, cleavages, duplications, amplifications, loss, copy number, repeats, total gene mutations, microsatellite instability; (ii) one or more types of in-situ hybridization data, e.g., DNA copies, gene copies, gene translocations; (iii) one or more types of RNA data, e.g., gene expression or gene fusion; (iv) one or more types of protein data, such as those obtained using immunohistochemistry; or (v) one or more types of ADAPT data, e.g., complex numbers.
[0118] The application server 240 hosting the machine learning model 370 is also configured to receive data representing proposed treatment data 422a for a disease or disorder described by disease or disorder data 420a for a subject having the biomarker represented by the received biomarker data records 320, 322, and 324. The proposed treatment data 422a for the disease or disorder 422a is also unlabeled and is merely a suggestion for treating a subject having the biomarker represented by the biomarker data records 320, 322, and 324.
[0119] In some embodiments, disease or disorder data 420a and proposed treatment 422a are provided by terminal 405 via network 230 (305), and biomarker data is obtained from a second distributed computer 310. Biomarker data may be derived from experimental equipment used to perform various assays. In other embodiments, disease or disorder data 420a, proposed treatment 422a, and biomarker data 320, 322, 324 may be received from terminal 405, respectively. For example, terminal 405 may be a user device of a physician, an employee working for a physician or a physician's agent, or another person who inputs data representing disease or disorder, data representing proposed treatment, and data representing one or more biomarkers of a person having the disease or disorder. In some embodiments, treatment data 422 may include a data structure that structures fields of data representing proposed treatment described by drug name. In other embodiments, treatment data 422 may include a data structure that structures fields of data representing more complex treatment data, such as dosage, medication regimen, and acceptable number of missed doses.
[0120] The application server 240 receives biomarker data records 320, 322, 324, disease or disorder data 420, and treatment data 422. The application server 240 provides the biomarker data records 320, 322, 324, disease or disorder data 420, and treatment data 422 to an extraction unit 242, which is configured to extract (i) specific biomarker data, e.g., DNA biomarker data 320a-1, protein expression data 322a-1, 324a-1, (ii) disease or disorder data 420a-1, and (iii) proposed treatment data 420a-1 from the fields of the biomarker data records 320, 322, 324 and outcome data records 420, 422. In some embodiments, the extracted data is stored in a memory unit 244 as a buffer, cache, etc., and then provided to a vector generation unit 250 as input when the vector generation unit 250 has bandwidth to receive input for processing. In other embodiments, the extracted data is provided directly to the vector generation unit 250 for processing. For example, in some embodiments, multiple vector generation units 250 may be used to enable parallel processing of inputs in order to reduce waiting times.
[0121] The vector generation unit 250 can generate data structures such as feature vectors 360, which include multiple fields, one or more fields for each type of biomarker data and one or more fields for each type of outcome data. For example, each field of the feature vector 360 may correspond to (i) extracted biomarker data of each type that can be extracted from biomarker data records 320, 322, and 324, such as next-generation sequencing data of each type, in-situ hybridization data of each type, RNA data of each type, immunohistochemistry data of each type, and ADAPT data of each type, and (ii) outcome data of each type that can be extracted from outcome data records 420 and 422, such as each type of disease or disorder, each type of treatment, and each type of treatment details.
[0122] The vector generation unit 250 is configured to assign a weight to each field of the feature vector 360 that indicates the extent to which the extracted biomarker data 320a-1, 322a-1, 324a-1, extracted disease or disorder 420a-1, and extracted treatment 422a-1 contain the data represented by each field. In one embodiment, for example, the vector generation unit 250 may assign "1" to each field of the feature vector 360 corresponding to features found in the extracted biomarker data 320a-1, 322a-1, 324a-1, extracted disease or disorder 420a-1, and extracted treatment 422a-1. In such an embodiment, the vector generation unit 250 may also assign "0" to each field of the feature vector corresponding to features not found in the extracted biomarker data 320a-1, 322a-1, 324a-1, extracted disease or disorder 420a-1, and extracted treatment 422a-1. The output of the vector generation unit 250 may include data structures such as feature vectors 360, which can be provided as input to the trained machine learning model 370.
[0123] The trained machine learning model 370 processes the generated feature vectors 360 based on the adjusted parameters determined during the training phase and described with reference to Figure 1B. The output 272 of the trained machine learning model provides an indicator of the effectiveness of treatment 422a-1 for a target disease or disorder 420a-1 having the biomarkers 320a-1, 322a-1, and 324a-1. In some embodiments, the output 272 may include probabilities indicating the effectiveness of treatment 422a-1 for a target disease or disorder 420a-1 having the biomarkers 320a-1, 322a-1, and 324a-1. In such embodiments, the output 272 may be provided to a terminal 405 using the network 230 (311). The terminal 405 may then generate an output on the user interface 420 indicating the predicted level of effectiveness of treatment for a person's disease or disorder having the biomarkers represented by the feature vectors 360.
[0124] In other embodiments, output 272 may be provided to a prediction unit 380 configured to decode the meaning of output 272. For example, the prediction unit 380 may be configured to map output 272 to one or more categories of effectiveness. The output of the prediction unit 328 may then be used as part of a message 390 provided to terminal 305 via the network 230 for review by subjects, their guardians, nurses, doctors, etc.
[0125] Figure 1D is a flowchart of process 400 for generating training data to train a machine learning model to predict the effectiveness of treatment for a target disease or disorder having a specific set of biomarkers. In one phase, process 400 may include the steps of: obtaining a first data structure from a first distributed data source (410) which includes fields for structuring data representing one or more biomarkers associated with a target; storing the first data structure in one or more memory devices (420); obtaining a second data structure from a second distributed data source (430) which includes fields for structuring data representing outcome data for a target having one or more biomarkers; storing the second data structure in one or more memory devices (440); generating a labeled training data structure based on the first and second data structures (450) which includes (i) data representing one or more biomarkers, (ii) disease or disorder, (iii) treatment, and (iv) the effectiveness of treatment for the disease or disorder; and training a machine learning model using the generated labeled training data (460).
[0126] Figure 1E is a flowchart of process 500, which uses a machine learning model trained to predict the effectiveness of treatment for a target disease or disorder having a specific set of biomarkers. In one phase, process 500 may include steps of: obtaining a data structure representing one or more biomarkers associated with a target (510); obtaining data representing the target disease or disorder type (520); obtaining data representing the treatment type for the target (530); generating a data structure for input to a machine learning model, representing (i) one or more biomarkers, (ii) the disease or disorder, and (iii) the treatment type (540); providing the generated data structure as input to a machine learning model trained using the obtained biomarkers, one or more treatment types, and one or more disease or disorder (550); obtaining an output generated by the machine learning model based on the machine learning model's processing of the provided data structure (560); and determining a predicted outcome for treatment of the target disease or disorder having one or more biomarkers based on the obtained output generated by the machine learning model (570).
[0127] Provided herein is a method for improving classification performance using multiple machine learning models. Traditionally, a single model is selected to perform a desired prediction / classification. For example, during the training phase, various model parameters or model types, such as random forests, support vector machines, logistic regression, k-nearest neighbors, artificial neural networks, naive Bayes, quadratic discriminant analysis, or Gaussian process models, may be compared to identify the model with the best desired performance. The applicants have realized that selecting a single model does not necessarily provide optimal performance in all settings. Instead, multiple models can be trained to perform prediction / classification, and classification can be performed using joint prediction. In this scenario, each model is allowed to "vote," and the classification that receives the majority of votes is considered the winner.
[0128] The voting strategy disclosed herein can be applied to any machine learning classification, including both model building (e.g., using training data) and applications for classifying naive samples. Such settings include, but are not limited to, data in the fields of biology, finance, communications, media, and entertainment. In some preferred embodiments, the data is high-dimensional “big data.” In some embodiments, the data includes biological data, including biological data obtained by molecular profiling as described herein. See, for example, Example 1. Molecular profiling data may include, but are not limited to, high-dimensional next-generation sequencing data for a specific biomarker panel (see, for example, Example 1) or whole exome and / or whole transcriptome data. The classification can be any classification useful, for example, to characterize a phenotype. For example, the classification may provide diagnosis (e.g., diseased or healthy), prognosis (e.g., predicting a good or bad outcome) or theranosis (e.g., predicting or monitoring therapeutic efficacy or lack thereof). Examples of applications of the voting strategy are provided herein in Examples 2-4.
[0129] Figure 1F is a block diagram of system 600, which interprets the outputs generated by multiple machine learning models using a voting unit. System 600 is similar to system 300 in Figure 1C. However, instead of a single machine learning model 370, system 600 includes multiple machine learning models 370-0, 370-1...370-x (where x is any non-zero integer greater than 1). In addition, system 600 also includes a voting unit 480.
[0130] As a non-limiting example, System 600 can be used to predict the effectiveness of treatment for a target disease or disorder having a specific set of biomarkers. See Examples 2-4.
[0131] Each machine learning model 370-0, 370-1, 370-x can include a machine learning model trained to classify a specific type of input data 320-0, 320-1...320-x (where x is any non-zero integer greater than 1, equal to the number of machine learning models x). In some embodiments, each of the machine learning models 370-0, 370-1, 370-x can be of the same type. For example, each of the machine learning models 370-0, 370-1, 370-x can be, for example, a random forest classification algorithm trained with various parameters. In other embodiments, the machine learning models 370-0, 370-1, 370-x can be of different types. For example, they can be one or more random forest model classifiers, one or more neural networks, one or more k-nearest neighbor classifiers, other types of machine learning models, or any combination thereof.
[0132] Input data such as input data 0 (320-0), input data 1 (320-1), and input data x (320-x) can be obtained by the application server 240. In some embodiments, input data 320-0, 320-1, and 320-x are obtained from one or more distributed computers 310, 405 via the network 230. As an example, one or more of the input data items 320-0, 320-1, and 320-x can be generated by correlating data from multiple different data sources 210, 405. In such embodiments, (i) first data describing the biomarker of interest can be obtained from the first distributed computer 310, and (ii) second data describing the disease or disorder and associated treatments can be obtained from the second computer 405. The application server 240 can correlate the first data and the second data to generate input data structures such as input data structure 320-0. This process is illustrated in more detail in Figure 1C. The input data items 320-0, 320-1, and 320-x can be provided sequentially, one at a time, as input to, for example, a vector generation unit. The vector generation unit can generate input vectors 360-0, 360-1, and 36-x corresponding to each input data 320-0, 320-1, and 320-x. While some embodiments may generate vectors 360-0, 360-1, and 360-x sequentially, this disclosure is not limited to such a method.
[0133] Alternatively, in some embodiments, the vector generation unit 250 can be configured to operate multiple parallel vector generation units that can parallelize the vector generation process. In such embodiments, the vector generation unit 250 can receive input data 320-0, 320-1, and 320-x concurrently, process the input data 320-0, 320-1, and 320-x concurrently, and generate vectors 360-0, 360-1, and 360-x, each corresponding to one of the input data 320-0, 320-1, and 320-x, concurrently.
[0134] In some embodiments, vectors 360-0, 360-1, and 360-x can each be generated based on corresponding input data such as input data 320-0, 320-1, and 320-x. That is, vector 360-0 is generated based on input data 320-0 and represents input data 320-0. Similarly, vector 360-1 is generated based on input data 320-1 and represents input data 320-1. Similarly, vector 360-x is generated based on input data 320-x and represents input data 320-x.
[0135] In some embodiments, each input data structure 320-0, 320-1, 320-x may include data representing biomarkers of the subject, data describing diseases or disorders associated with the subject, data describing proposed treatments for the subject, or any combination thereof. Data representing biomarkers of the subject may include data describing a specific subset or panel of genes from the subject. Alternatively, in some embodiments, data representing biomarkers of the subject may include data representing a complete set of known genes for the subject. A complete set of known genes for the subject may include all of the subject's genes. In some embodiments, each of the machine learning models 370-0, 370-1, 370-x is a machine learning model of the same type, for example, a neural network trained to classify input data vectors as corresponding to subjects that are likely to respond to or not respond to treatments identified as associated by the vectors processed by the machine learning model. In such an embodiment, each of the machine learning models 370-0, 370-1, and 370-x is of the same type, but each of the machine learning models 370-0, 370-1, and 370-x may be trained in a different way. The machine learning models 370-1, 370-1, and 370-x can generate output data 272-0, 272-1, and 272-x, respectively, representing whether the subject associated with the input vectors 360-0, 360-1, and 360-x is likely to respond to or not respond to the treatment associated with the input vectors 360-0, 360-1, and 360-x. In this example, the input datasets and their corresponding input vectors are the same. For example, each set of input data has the same biomarker, the same disease or disorder, the same treatment, or any combination thereof.Nevertheless, considering the various training methods used to train each machine learning model 370-0, 370-1, and 370-x, as shown in Figure 1F, each machine learning model 370-0, 370-1, and 370-x that processes the input vectors 360-0, 361-1, and 361-x can produce different outputs 272-0, 272-1, and 272-x, respectively.
[0136] Alternatively, each of the machine learning models 370-0, 370-1, and 370-x could be a different type of machine learning model, trained or otherwise configured to classify input data as representing subjects likely to respond to or not respond to treatment for a disease or disorder. For example, the first machine learning model 370-1 could include a neural network, machine learning model 370-1 could include a random forest classification algorithm, and machine learning model 370-x could include a k-nearest neighbors algorithm. In this example, each of these different types of machine learning models 370-0, 370-1, and 370-x could be trained or otherwise configured to receive and process an input vector and determine whether the input vector is associated with subjects likely to respond to or not respond to the treatment associated with the same input vector. In this example, the input datasets and their corresponding input vectors could be the same. For example, each set of input data might have the same biomarker, the same disease or disorder, the same treatment, or any combination. Therefore, machine learning model 370-0 can be a neural network trained to process input vector 360-0 and generate output data 272-0 indicating whether the subject associated with input vector 360-0 is likely to respond to or not respond to the treatment also associated with input vector 360-0. In addition, machine learning model 370-1 can be a random forest classification algorithm trained to process input vector 360-1, which in this example is the same as input vector 360-0, and generate output data 272-1 indicating whether the subject associated with input vector 360-1 is likely to respond to or not respond to the treatment also associated with input vector 360-1. This input vector analysis method can be continued for each of x inputs, x input vectors, and x machine learning models.Continuing this example with reference to Figure 1F, the machine learning model 370-x can be a k-nearest neighbors algorithm trained to process input vector 360-x, which in this example is the same as input vectors 360-0 and 360-1, and to generate output data 272-x indicating whether the object associated with input vector 360-x is likely to respond to or not to the treatment also associated with input vector 360-x.
[0137] Alternatively, each of the machine learning models 370-0, 370-1, and 370-x can be the same type of machine learning model, or they can be different types of machine learning models configured to receive different inputs. For example, the input to the first machine learning model 370-0 may include a vector 360-0 containing data representing a first subset or first panel of the target gene, and then, based on the processing of vector 360-0 by machine learning model 370-0, it can predict whether the target is likely to respond to a treatment or not. In addition, in this example, the input to the second machine learning model 370-1 may include a vector 360-1 containing data representing a second subset or second panel of the target gene, which is different from the first subset or first panel of the gene. The second machine learning model can then generate second output data 272-1 indicating whether the target associated with input vector 360-1 is likely to respond to a treatment associated with input vector 360-2 or not. This input vector analysis method can be continued for each of the x inputs, x input vectors, and x machine learning models. The input to the x-th machine learning model 370-x may include a vector 360-x containing data representing the x-th subset or x-th panel of the target gene, which is different from (i) at least one, (iii) two or more, or (iii) each of the other x-1 input data vectors 370-0 to 370-x-1. In some embodiments, at least one of the x input data vectors may include data representing the complete set of genes from the target. The x-th machine learning model 370-x can then generate a second output data 272-x, which indicates whether the target associated with input vector 360-x is likely to respond to or not respond to the treatment associated with input vector 360-x.
[0138] The multiple embodiments of the System 400 described above are not intended to be limiting, but rather are merely examples of multiple machine learning models 370-0, 370-1, 370-x and the configuration of their respective inputs that may be used when using this disclosure. When referring to these examples, the subject can be any human, non-human animal, plant, or other subject. As described above, the input feature vectors are generated based on the input data and can represent the input data. Thus, each input vector can represent data including one or more biomarkers, diseases or disorders and treatments, and the level of effectiveness of a treatment when treating a disease or disorder in a subject having a biomarker. "Treatment" may include data describing any therapeutic substance, e.g., small molecule drugs or biologics, treatment details (e.g., dosage, regimen, missed doses, etc.) or any combination thereof.
[0139] In the embodiment shown in Figure 1F, the output data 272-0, 272-1, and 272-x can be analyzed using the voting unit 480. For example, the output data 272-0, 272-1, and 272-x can be input to the voting unit 480. In some embodiments, the output data 272-0, 272-1, and 272-x can be data indicating whether a subject associated with an input vector processed by a machine learning model is likely to respond to or not respond to a treatment associated with the vector processed by the machine learning model. The data indicating a subject associated with an input vector and generated by each machine learning model can include "0" or "1". A "0" generated by machine learning model 370-0 based on machine learning model 370-0's processing of input vector 360-0 may indicate that a subject associated with input vector 360-0 is likely not to respond to a treatment associated with input vector 360-0. Similarly, a “1” generated by machine learning model 360-0 based on the processing of input vector 360-0 by machine learning model 370-0 may indicate that the subject associated with input vector 360-0 is likely to respond to the treatment associated with input vector 360-0. This example uses “0” to represent “not responding” and “1” to represent “responding,” but the disclosure is not limited in this way. Alternatively, any values can be generated as output data to represent “responding” and “non-responding” classes. For example, in some embodiments, one could use “1” to represent the “non-responding” class and “0” to represent the “responding” class. In yet other embodiments, the output data 272-0, 272-1, 272-x may include probabilities indicating the likelihood that the subject associated with the input vector processed by the machine learning model is associated with either the “responding” or “non-responding” class. In such embodiments, for example, the generated probabilities could be applied to a threshold, and if the threshold is met, it could be determined that the subject associated with the input vector processed by the machine learning model is in the “responding” class.
[0140] The voting unit 480 can evaluate the received output data 270-0, 272-1, and 272-x and determine whether the subject associated with the processed input vectors 360-0, 360-1, and 360-x is likely to respond to or not respond to the treatment associated with the processed input vectors 360-0, 360-1, and 360-x. Then, based on the received set of output data 270-0, 272-1, and 272-x, the voting unit 480 can determine whether the subject associated with the input vectors 360-0, 360-1, and 360-x is likely to respond to the treatment associated with the input vectors 360-0, 360-2, and 360-x. In some embodiments, the voting unit 480 can apply the "majority rule principle". Applying the principle of majority rule, the voting unit 480 can aggregate outputs 272-0, 272-1, and 272-x indicating that the subject responds, and outputs 272-0, 272-1, and 272-x indicating that the subject does not respond. Then, the class with the majority of predictions or votes (e.g., respond or non-response classes) is selected as the appropriate classification for the subject associated with the input vectors 360-0, 360-1, and 360-x. This selected class can be called the real entity class, and each of the predictions or votes output by the machine learning models 370-0, 370-1, and 370-x can be called the initial entity class.
[0141] Therefore, in some embodiments, the majority of predictions or votes can be determined by the voting unit 480 aggregating the number of predictions or votes for each initial entity class. For example, the system 600 can determine how many times each initial entity class will be predicted or voted for by machine learning models 370-0, 370-1, 370-x, and then select the entity class associated with the highest number of predictions or votes.
[0142] In some embodiments, the voting unit 480 can perform more nuanced analysis. For example, in some embodiments, the voting unit 480 can store confidence scores for each machine learning model 370-0, 370-1, and 370-x. These confidence scores for each machine learning model 370-0, 370-1, and 370-x can initially be set to default values such as 0 and 1. Then, for each round of processing of the input vector, the voting unit 480 or other modules of the application server 240 can adjust the confidence scores of the machine learning models 370-0, 370-1, and 370-x based on whether the machine learning model accurately predicted the target classification selected by the voting unit 480 during the previous iteration. Thus, the stored confidence scores for each machine learning model can provide an indicator of the historical accuracy for each machine learning model.
[0143] In a more nuanced approach, the voting units 480 can adjust the output data 272-0, 272-0, and 272-x generated by each machine learning model 370-0, 370-1, and 370-x, respectively, based on confidence scores calculated for the machine learning models. Thus, the values of the output data generated by the machine learning models can be boosted using confidence scores indicating that the machine learning models are historically accurate. Similarly, the values of the output data generated by the machine learning models can be decreased using confidence scores indicating that the machine learning models are historically inaccurate. Such boosting or decreasing the values of the output data generated by the machine learning models can be achieved, for example, by using the confidence scores as a multiplier that is less than 1 in the case of a decrease and greater than 1 in the case of a boost. Alternatively, the values of the output data can also be adjusted using other operations, for example, by subtracting the confidence score from the values of the output data to decrease the values of the output data, or by adding the confidence score to the values of the output data to boost the values of the output data. The use of confidence scores to boost or reduce the values of output data generated by a machine learning model is particularly useful when the machine learning model is configured to output probabilities that apply to one or more thresholds for determining whether a subject will respond to treatment or not. This is because confidence scores can be used to adjust the output of a machine learning model, moving the generated output values above or below class thresholds, thereby changing the machine learning model's predictions based on its historical accuracy.
[0144] The use of voting units 480 to evaluate the outputs of multiple machine learning models can lead to higher accuracy in predicting the efficacy of treatments for a particular set of target biomarkers, because it allows us to evaluate the consensus among multiple machine learning models instead of just the output of a single model.
[0145] Figure 1G is a block diagram of system components that can be used to implement the systems shown in Figures 2 and 3.
[0146] Computing device 600 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Computing device 650 is intended to represent various forms of mobile devices, such as personal digital assistants, mobile phones, smartphones, and other similar computing devices. In addition, computing device 600 or 650 may include a USB (Universal Serial Bus) flash drive. The USB flash drive may store an operating system and other applications. The USB flash drive may include input / output components such as a wireless transmitter or USB connector that can be inserted into a USB port of another computing device. The components shown herein, their connections and relationships, and their functions are illustrative only and are not intended to limit the embodiments of the invention described herein and / or claimed herein.
[0147] The computing device 600 includes a processor 602, memory 604, storage device 608, a high-speed interface 608 connected to memory 604 and a high-speed expansion port 610, and a low-speed bus 614 and a low-speed interface 612 connected to storage device 608. Each of the components 602, 604, 608, 608, 610, and 612 can be interconnected using various buses and implemented on a common motherboard or mounted in other appropriate ways. The processor 602 can process instructions to be executed within the computing device 600, including instructions stored in memory 604 or storage device 608 for displaying graphical information for a GUI on an external input / output device, such as a display 616 coupled to the high-speed interface 608. In other embodiments, multiple processors and / or multiple buses may be used, along with multiple memories and memory types as appropriate. Also, multiple computing devices 600, each providing a portion of the required operation, may be connected as, for example, a server bank, a group of blade servers, or a multiprocessor system.
[0148] Memory 604 stores information within the computing device 600. In one embodiment, memory 604 is one or more volatile memory units. In another embodiment, memory 604 is one or more non-volatile memory units. Memory 604 may also be another form of computer-readable medium, such as a magnetic or optical disk.
[0149] The storage device 608 can provide large-capacity storage for the computing device 600. In one embodiment, the storage device 608 may be, or include, a computer-readable medium, such as a floppy disk device, a hard disk device, an optical disk device or tape device, a flash memory or other similar solid-state memory device or an array of devices including such devices in a storage area network or other configuration. A computer program product may be tangibly embodied in the information carrier. The computer program product may also include instructions that, when executed, perform one or more of the methods described above. The information carrier is a computer-readable or machine-readable medium such as memory 604, the storage device 608 or on-processor memory 602.
[0150] The high-speed control unit 608 manages bandwidth-intensive operations for the computing device 600, and the low-speed control unit 612 manages low-bandwidth-intensive operations. Such function assignments are merely illustrative. In one embodiment, the high-speed control unit 608 is coupled to memory 604, a display 616 (e.g., via a graphics processor or accelerator), and a high-speed expansion port 610 that can accept various expansion cards (not shown). In an embodiment, the low-speed control unit 612 is coupled to the storage device 608 and the low-speed expansion port 614. The low-speed expansion port, which can include various communication ports, such as USB, Bluetooth, Ethernet, and wireless Ethernet, can be coupled, for example, via a network adapter, to one or more input / output devices, such as a keyboard, pointing device, microphone / speaker pair, scanner, or networking device, such as a switch or router. The computing device 600 can be implemented in several different forms, as illustrated. For example, it can be implemented as a standard server 620, or multiplexed as a group of such servers. It can also be implemented as part of a rack server system 624. In addition, it can be implemented as a personal computer, such as a laptop computer 622. Alternatively, components from computing device 600 can be combined with other components in a mobile device (not shown), such as device 650. Each of such devices may contain one or more computing devices 600, 650, or the entire system may consist of multiple computing devices 600, 650 communicating with each other.
[0151] The computing device 600 can be implemented in several different forms, as shown in the drawings. For example, it can be implemented as a standard server 620, or multiplexed as a group of such servers. It can also be implemented as part of a rack server system 624. In addition, it can be implemented as a personal computer, such as a laptop computer 622. Alternatively, components from the computing device 600 can be combined with other components in a mobile device (not shown), such as device 650. Each of such devices may contain one or more computing devices 600, 650, or the entire system may consist of multiple computing devices 600, 650 communicating with each other.
[0152] The computing device 650 includes, among other things, a processor 652, memory 664, and input / output devices, such as a display 654, a communication interface 666, and a transceiver 668. Device 650 may also include storage devices, such as a microdrive or other devices, to provide further storage. Each of the components 650, 652, 664, 654, 666, and 668 are interconnected using various buses, and some of the components can be mounted on a common motherboard or in other suitable ways.
[0153] The processor 652 can execute instructions within the computing device 650, including instructions stored in memory 664. The processor can be implemented as a chipset of chips containing separate and multiple analog and digital processors. In addition, the processor can be implemented using one of several architectures. For example, the processor 610 can be a CISC (Complex Instruction Set Computers) processor, a RISC (Reduced Instruction Set Computer) processor, or a MISC (Minimal Instruction Set Computer) processor. The processor can, for example, coordinate other components of the device 650, such as controlling the user interface, providing applications executed by the device 650, and providing wireless communication by the device 650.
[0154] The processor 652 can communicate with the user via a control interface 658 and a display interface 656 coupled to the display 654. The display 654 may be, for example, a TFT (Thin-Film-Transistor Liquid Crystal Display) display, an OLED (Organic Light Emitting Diode) display, or other suitable display technology. The display interface 656 may include appropriate circuitry for driving the display 654 to present graphical and other information to the user. The control interface 658 can receive commands from the user and translate them for submission to the processor 652. In addition, an external interface 662 communicating with the processor 652 may be provided to enable near-field communication between device 650 and other devices. The external interface 662 may, for example, provide wired communication in some embodiments, provide wireless communication in other embodiments, or multiple interfaces may be used.
[0155] Memory 664 stores information within the computing device 650. Memory 664 can be implemented as one or more computer-readable media, volatile memory units, or non-volatile memory units. Additionally, an expansion memory 674 may be provided and connected to the device 650 via an expansion interface 672, which may include, for example, a SIMM (Single In Line Memory Module) card interface. Such an expansion memory 674 may provide extra storage space for the device 650, or it may store applications or other information for the device 650. Specifically, the expansion memory 674 may contain instructions for executing or supplementing the above processes, or it may contain security information. Therefore, for example, the expansion memory 674 may be provided as a security module for the device 650 and programmed with instructions that allow secure use of the device 650. Furthermore, a secure application may be provided via a SIMM card, along with additional information, by placing, for example, identification information on the SIMM card in a hack-proof manner.
[0156] The memory may include, for example, flash memory and / or NVRAM memory, as detailed below. In one embodiment, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more of the methods described above. The information carrier is a computer-readable or machine-readable medium, such as memory 664, extended memory 674, or on-processor memory 652, which can be received, for example, via transceiver 668 or external interface 662.
[0157] Device 650 can communicate wirelessly via a communication interface 666, which may include digital signal processing circuits if necessary. The communication interface 666 can provide communication under various modes or protocols, such as GSM voice call, SMS, EMS or MMS messaging, CDMA, TDMA, PDC, WCDMA®, CDMA2000, or GPRS. Such communication can be carried out, for example, via a radio frequency transceiver 668. In addition, short-range communication can be carried out using, for example, Bluetooth, Wi-Fi, or other such transceivers (not shown). Furthermore, a GPS (Global Positioning System) receiver module 670 can provide device 650 with additional navigation-related and location-related radio data, which can be appropriately used by applications operating on device 650.
[0158] Device 650 can also communicate audibly using an audio codec 660 that can receive voice information from the user and convert it into usable digital information. The audio codec 660 can similarly generate audible sounds for the user, such as through the speaker in the device 650's handset. Such sounds may include sounds from telephone calls, recordings such as voice messages, music files, or sounds generated by applications running on device 650.
[0159] The computing device 650 can be implemented in several different forms, as shown in the illustration. For example, it can be implemented as a mobile phone 680. It can also be implemented as part of a smartphone 682, a personal digital assistant, or other similar mobile device.
[0160] Various embodiments of the systems and methods described herein can be realized as digital electronic circuits, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations of such embodiments. These various embodiments may include embodiments as one or more computer programs executable and / or interpretable on a programmable system which includes a storage system, at least one input device and at least one output device, and coupled to send data and instructions to the storage system, at least one input device and at least one output device, and which may be dedicated or general-purpose programmable processors.
[0161] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages and / or assembly / machine language. As used herein, the terms “machine-readable medium” and “computer-readable medium” include any computer program product, apparatus, and / or device used to provide machine instructions and / or data to a programmable processor, such as magnetic disks, optical disks, memory, and programmable logic devices (PLDs), including machine-readable medium that receives machine instructions as machine-readable signals. The term “machine-readable signal” includes any signal used to provide machine instructions and / or data to a programmable processor.
[0162] To provide user interaction, the systems and technologies described herein can be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, and a keyboard and pointing device, such as a mouse or trackball, on which the user can provide input to the computer. Other types of devices may also be used to provide user interaction. For example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; input from the user may be received in any form, including acoustic, voice, or tactile input.
[0163] The systems and technologies described herein may be implemented as computing systems including backend components, such as data servers, or middleware components, such as application servers, or frontend components, such as client computers having a graphical user interface or web browser on which a user can interact with embodiments of the systems and technologies described herein, or any combination of such backend, middleware, or frontend components. The components of the system may be interconnected by digital data communications in any form or medium, such as communication networks. Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), and the Internet.
[0164] A computing system can include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship between a client and a server arises thanks to computer programs that run on each computer and have a client-server relationship with each other.
[0165] Computer system The implementation of this method may also utilize computer-related software and systems. Computer software products such as those described herein typically include a computer-readable medium having computer-executable instructions for performing the logical steps of the method described herein. Suitable computer-readable media include floppy disks, CD-ROMs / DVDs / DVD-ROMs, hard disk drives, flash memory, ROM / RAM, magnetic tape, and the like. Computer-executable instructions may be written in a suitable computer language or a combination of several languages. Basic computational biology methods are described, for example, in Setubal and Meidanis at al., Introduction to Computational Biology Methods (PWS Publishing Company, Boston, 1997); Salzberg, Searles, Kasif, (Ed.), Computational Methods in Molecular Biology (Elsevier, Amsterdam, 1998); Rashidi and Buehler, Bioinformatics Basics: Application in Biological Science and Medicine (CRC Press, London, 2000); and Ouelette and Bzevanis, Bioinformatics: A Practical Guide for Analysis of Gene and Proteins (Wiley & Sons, Inc., 2nd ed., 2001). See U.S. Patent No. 6,420,108.
[0166] This method can also utilize a variety of computer program products and software for diverse purposes, such as probe design, data management, analysis, and instrument operation. See U.S. Patents 5,593,839, 5,795,716, 5,733,729, 5,974,164, 6,066,454, 6,090,555, 6,185,561, 6,188,783, 6,223,127, 6,229,911, and 6,308,170.
[0167] In addition, the present invention relates to embodiments of methods for providing genetic information via a network such as the Internet, as described in U.S. Patent Applications No. 10 / 197,621, No. 10 / 063,559 (Publication No. 20020183936), No. 10 / 065,856, No. 10 / 065,868, No. 10 / 328,818, No. 10 / 328,872, No. 10 / 423,403 and No. 60 / 482,389. For example, one or more molecular profiling techniques can be performed in one location, e.g., a city, state, country, or continent, and the results can be transmitted to a different city, state, country, or continent. Treatment selection can then be made in whole or in part at a second location. Methods as described herein include the processing of information transfer between different locations.
[0168] The conventional data networking, application development, and other functional aspects of the system (and the components of the individual operating components of the system) may not be described in detail herein, but are part of those described herein. Furthermore, the connection lines shown in the various drawings contained herein are intended to represent exemplary functional relationships and / or physical connections between various elements. It should be noted that many alternative or additional functional relationships or physical connections may exist in actual systems.
[0169] The various system components detailed herein may include: a host server or other computing system including a processor for processing digital data; processor-coupled memory for storing digital data; processor-coupled input digitizers for inputting digital data; application programs stored in memory and accessible by the processor for instructing the processor on processing digital data; a display device coupled to the processor and memory for displaying information derived from the digital data processed by the processor; and one or more databases. The various databases used herein may include patient data, e.g., family history, age group and environmental data; biological sample data, previous treatment and protocol data; patient clinical data; molecular profiling data of biological samples; data on therapeutic drugs and / or investigational drugs; gene libraries; disease libraries; drug libraries; patient follow-up data; file management data; financial management data; billing data and / or similar data useful for the operation of the system. As those skilled in the art will understand, a user computer may include an operating system (e.g., Windows NT, 95 / 98 / 2000, OS2, UNIX®, Linux®, Solaris, MacOS, etc.) as well as various conventional support software and drivers that typically accompany the computer. Examples of computers include any suitable personal computer, network computer, workstation, minicomputer, mainframe, etc. A user computer may be located in a home or medical / business environment with network access. In exemplary embodiments, access may be network access or internet access via a commercially available web browser software package.
[0170] As used herein, the term “network” includes any electronic means of communication that incorporates both hardware and software components. Communication between parties may be achieved through any suitable communication channel, e.g., telephone networks, extranets, intranets, internets, points of interaction devices, personal digital assists (e.g., Palm Pilot®, Blackberry®), mobile phones, kiosks, etc.), online communications, satellite communications, offline communications, wireless communications, transponder communications, local area networks (LANs), wide area networks (WANs), network-connected or linked devices, keyboards, mice, and / or any suitable communication or data input modalities. Furthermore, while systems are often described herein as being implemented using the TCP / IP communication protocol, systems may be implemented using IPX, Appletalk, IP-6, NetBIOS, OSI, or any existing or future protocol. If a network has the nature of a public network such as the internet, it may be advantageous to assume that the network is insecure and susceptible to interception. Specific information relating to protocols, standards, and application software used in connection with the Internet is generally known to those skilled in the art and therefore does not need to be detailed herein. For example, see Dilip Naik, Internet Standards and Protocols (1998); Java 2 Complete, various authors, (Sybex 1999); Deborah Ray and Eric Ray, Mastering HTML 4.0 (1997); and Loshin, TCP / IP Clearly Explained (1997) and David Gourley and Brian Totty, HTTP, The Definitive Guide (2002), the contents of which are incorporated herein by reference.
[0171] Various system components may be appropriately combined, individually, or collectively, independently with networks via data links, including, for example, connections to an Internet Service Provider (ISP) via local loops typically used in conjunction with standard modem communications, cable modems, dish networks, ISDN, DSL (Digital Subscriber Line), or various wireless communication methods. See, for example, Gilbert Held, Understanding Data Communications (1996), incorporated herein by reference. Note that the network may be implemented as other types of networks, such as interactive television (ITV) networks. Furthermore, the system considers the use, sale, or distribution of any goods, services, or information on any network having similar functionality to those described herein.
[0172] As used herein, “transmission” may include sending electronic data from one system component to another via a network connection. In addition, as used herein, “data” may include information such as commands, queries, files, and data for storage, whether digital or in any other form.
[0173] The system considers uses related to web services, utility computing, pervasive and personalized computing, security and identity solutions, autonomic computing, commodity computing, mobility and wireless solutions, open source, biometrics, grid computing and / or mesh computing.
[0174] Any database detailed herein may include relational, hierarchical, graphical, or object-oriented structures and / or any other database configuration. Common database products that may be used to implement a database include IBM's (White Plains, NY) DB2, various commercially available database products from Oracle Corporation (Redwood Shores, CA), Microsoft Access or Microsoft SQL Server from Microsoft Corporation (Redmond, Washington), or any other suitable database product. Furthermore, a database may be organized in any suitable way, such as data tables or lookup tables. Each record may be a single file, a set of files, a linked set of data fields, or any other data structure. The association of specific data may be achieved by any desired data association technique, such as those known or implemented in the art. For example, associations may be achieved manually or automatically. Examples of automatic association techniques include database search, database merge, GREP, AGREP, SQL, the use of key fields in tables to speed up searches, sequential searches of all tables and files, and sorting of records in files according to a known order to simplify searches. The association process can be achieved, for example, by a database merge function using a pre-selected "key field" of a database or data sector.
[0175] More specifically, a "key field" divides the database according to the superclass of objects defined by the key field. For example, a particular type of data may be designated as a key field in multiple related data tables, in which case the data tables may be linked based on the type of data in the key field. The data corresponding to the key field in each of the linked data tables is preferably the same or of the same type. However, data tables that have similar but not identical data in their key fields can also be linked, for example, using AGREP. In one aspect, data may be stored without a standard format using any suitable data storage technique. Datasets may be stored using any appropriate technology, including, for example, storing individual files using the ISO / IEC 7816-4 file structure; realizing a domain in which dedicated files are selected to expose one or more underlying files containing one or more datasets; using datasets stored in individual files using a hierarchical filing system; using datasets stored as records in a single file (compressed, SQL accessible, hashed one or more keys, numeric, alphabet by first tuple, etc.); BLOB (Binary Large Object); being stored as ungrouped data elements coded using ISO / IEC 7816-6 data elements; being stored as ungrouped data elements coded using ISO / IEC Abstract Syntax Notation (ASN.1) as in ISO / IEC 8824 and 8825; and / or using other proprietary technologies, including fractal compression schemes, image compression methods, etc.
[0176] In one exemplary embodiment, the ability to store diverse information in different formats is facilitated by storing information as blobs. Thus, arbitrary binary information can be stored in the storage space associated with the dataset. The blob method can store datasets as ungrouped data elements formatted as blocks of binary data via fixed memory offsets, using either fixed memory allocation, circular queue techniques, or best practices for memory management (e.g., paged memory, least recently used). The ability to store diverse datasets with various formats by using the blob method facilitates data storage by multiple unrelated owners of the dataset. For example, a first dataset to be stored may be provided by a first party, a second dataset to be stored may be provided by an unrelated second party, and a third dataset to be stored may be provided by a third party unrelated to the first and second parties. Each of these three exemplary datasets may contain different information stored using different data storage formats and / or techniques. Furthermore, each dataset may contain a subset of data that is also different from other subsets.
[0177] As described above, in various embodiments, data can be stored regardless of a common format. However, in one exemplary embodiment, a dataset (e.g., a BLOB) may be annotated in a standard manner when provided for manipulation. The annotation may include a short header, trailer, or other appropriate indicator associated with each dataset, configured to carry information useful when managing various datasets. For example, the annotation may be referred to herein as a “conditional header,” “header,” “trailer,” or “status,” and may include an indication of the status of the dataset, or an identifier correlated with a particular issuer or owner of the data. Subsequent bytes of the data may be used to indicate, for example, the ID of the issuer or owner of the data, a user, a transaction / membership account identifier, etc. Each of these conditional annotations will be further described herein.
[0178] Dataset annotations may also be used for other types of status information and various other purposes. For example, dataset annotations may include security information that establishes access levels. Access levels may be configured such that only certain individuals, employee levels, companies, or other entities are allowed to access the dataset, or they may be configured to grant access to specific datasets based on transactions, data issuers or owners, users, etc. Furthermore, security information may restrict / permit only specific actions such as access to the dataset, modification of its, and / or deletion of it. In one example, a dataset annotation may indicate that only the dataset owner or user is allowed to delete the dataset, various identified users may be allowed to access the dataset for reading, and all other users are excluded from accessing the dataset. However, other access restriction parameters may be used to allow various entities to access the dataset as appropriate at various permission levels. Data including headers or trailers may be received by standalone interactive devices configured to add, delete, modify, or augment data according to the headers or trailers.
[0179] Those skilled in the art will also understand, for security reasons, that any database, system, device, server, or other component of a system may consist of any combination thereof in one or more locations, and that each database or system may include any of various appropriate security mechanisms such as firewalls, access codes, encryption, decryption, compression, and decompression.
[0180] The computing unit of the web client may further comprise an internet browser connected to the internet or intranet using standard dial-up, cable, DSL, or any other internet protocol known in the art. Transactions occurring in the web client may pass through firewalls to prevent unauthorized access from users on other networks. Furthermore, additional firewalls may be deployed between the various components of the CMS to further enhance security.
[0181] A firewall may include any hardware and / or software appropriately configured to protect CMS components and / or enterprise computing resources from users on other networks. Furthermore, the firewall may be configured to limit or restrict access to various systems and components behind it, in the case of web clients connecting via the web server. Firewalls can exist in a variety of configurations, including, among others, stateful inspection, proxy-based firewalls, and packet filtering. The firewall may be integrated with the web server or any other CMS component, or it may exist as a separate entity.
[0182] The computers detailed herein may provide a suitable website or other internet-based graphical user interface accessible to the user. In one embodiment, Microsoft Internet Information Server (IIS), Microsoft Transaction Server (MTS), and Microsoft SQL Server are used together with the Microsoft operating system, Microsoft NT web server software, the Microsoft SQL Server database system, and Microsoft Commerce Server. In addition, components such as Access or Microsoft SQL Server, Oracle, Sybase, Informix MySQL®, and Interbase may be used to provide an Active Data Object (ADO) compliant database management system.
[0183] Any of the communications, inputs, storages, databases, or displays detailed herein can be facilitated through websites that have web pages. The term “web page” as used herein is not intended to limit the types of documents and applications that may be used to interact with users. For example, a typical website may include, in addition to standard HTML documents, various forms, Java applets, JavaScript, Active Server Pages (ASP), Common Gateway Interface (CGI) scripts, Extensible Markup Language (XML), Dynamic HTML, Cascading Style Sheets (CSS), helper applications, and plugins. A server may include a web service that receives requests from a web server, including URLs (http: / / yahoo.com / stockquotes / ge) and IP addresses (123.56.789.234). The web server searches for the appropriate web page and sends the data or application for the web page to the IP address. A web service is an application that can interact with other applications via means of communication such as the Internet. Web services are typically based on standards or protocols such as XML, XSLT, SOAP, WSDL, and UDDI. The web services method is well-known in this technical field and is covered in many standard textbooks. For example, see Alex Nghiem, IT Web Services: A Roadmap for the Enterprise (2003), which is incorporated herein by reference.
[0184] The web-based clinical database for the system and method of this method preferably has the ability to upload and store clinical data files in native format and is searchable by any clinical parameter. The database is also extensible and can input clinical annotations from any study using an EAV data model (metadata) for easy integration with other studies. In addition, the web-based clinical database can be flexible and can be enabled with XML and XSLT so that users can dynamically add customized questions. Furthermore, the database includes an export function to CDISC ODM.
[0185] Implementers will also understand that there are many ways to display data within browser-based documents. The data may be displayed as standard text or within fixed lists, scrollable lists, dropdown lists, editable text fields, fixed text fields, popup windows, etc. Similarly, there are many ways available to change data within a web page, such as free text input using the keyboard, selection of menu items, checkboxes, option boxes, etc.
[0186] Systems and methods may be described herein with respect to functional block components, screenshots, arbitrary selections, and various processing steps. It should be understood that such functional blocks can be implemented by any number of hardware and / or software components configured to perform a specified function. For example, the system may use various integrated circuit components, such as memory elements, processing elements, logic elements, lookup tables, etc., that can perform a variety of functions under the control of one or more microprocessors or other control devices. Similarly, the software elements of the system may be implemented in any programming or scripting language, such as C, C++, Macromedia Cold Fusion, Microsoft Active Server Pages, Java, COBOL, assembler, Perl, Visual Basic, SQL Stored Procedures, or Extensible Markup Language (XML), and various algorithms may be implemented in any combination of data structures, objects, processes, routines, or other programming elements. Furthermore, it should be noted that the system may use several prior arts for data transmission, signaling, data processing, network control, etc. Moreover, this system can also be used to detect or prevent security issues in client-side scripting languages such as JavaScript and VBScript.For an introduction to the fundamentals of encryption and network security, please refer to one of the following references, all of which are incorporated herein by reference: (1) "Applied Cryptography: Protocols, Algorithms, And Source Code In C," by Bruce Schneier, published by John Wiley & Sons (second edition, 1995); (2) "Java Cryptography" by Jonathan Knudson, published by O'Reilly & Associates (1998); (3) "Cryptography & Network Security: Principles & Practice" by William Stallings, published by Prentice Hall.
[0187] The terms "end user", "consumer", "customer", "client", "healthcare provider", "hospital", or "business" as used herein may be used interchangeably with one another and shall each mean any person, entity, machine, hardware, software, or business. Each participant is equipped with a computing device to interact with the system and facilitate online data access and data entry. The customer has a computing unit in the form of a personal computer, although other types of computing units may be used, including laptops, notebooks, handheld computers, set-top boxes, mobile phones, touch-tone phones, etc. The owner / operator of the system and method of the present method has a computing unit implemented in the form of a computer server, although other embodiments are contemplated by systems including computing centers shown as mainframe computers, minicomputers, PC servers, networks of computers at the same or different geographical locations, etc. Moreover, the system contemplates the use, sale, or distribution of any goods, services, or information on any network having similar functions as described herein.
[0188] In one exemplary embodiment, each client customer may be issued an “account” or “account number.” Examples of accounts or account numbers as used herein include any device, code, number, letter, symbol, digital certificate, smart chip, digital signal, analog signal, biometric authentication or other identifier / sign (e.g., one or more such as authentication / access codes, personal identification numbers (PINs), internet codes, or other identification codes) appropriately configured to allow a consumer to access, interact with, or communicate with the system. An account number may optionally reside on or be associated with a charge card, credit card, debit card, prepaid card, embossed card, smart card, magnetic stripe card, barcode card, transponder, radio frequency card, or associated account. The system may include, but is not limited to, any of the aforementioned cards or devices, or a fob having a transponder and an RFID reader that communicates RF with the fob. The system may include, but is not limited to, a fob configuration. In fact, the system may include any device having a transponder configured to communicate with an RFID reader via RF communication. Typical devices may include, for example, keyrings, tags, cards, mobile phones, watches, or any such form that can be presented for inquiry. Furthermore, the systems, computing units, or devices detailed herein may include “pervasive computing devices,” which may include conventional non-computerized devices into which computing units are embedded. Account numbers may be distributed and stored in any form of plastic, electronic, magnetic, radio frequency, wireless, audio, and / or optical device that can transmit or download data from itself to a second device.
[0189] As will be understood by those skilled in the art, a system can be embodied as a customized version of an existing system, an add-on product, upgraded software, a standalone system, a distributed system, a method, a data processing system, a device for data processing, and / or a computer program product. Thus, a system can take the form of an all-software configuration, an all-hardware configuration, or a combination of both software and hardware aspects. Furthermore, a system may also take the form of a computer program product on a computer-readable storage medium having computer-readable program code means embodied in the storage medium. Any suitable computer-readable storage medium may be used, including hard disks, CD-ROMs, optical storage devices, magnetic storage devices, and the like.
[0190] Systems and methods are described herein in various aspects with reference to screenshots, block diagrams, and flowcharts of methods, apparatus (e.g., systems), and computer program products. It will be understood that each functional block in the block diagrams and flowcharts, as well as combinations of functional blocks in the block diagrams and flowcharts, can be implemented by computer program instructions.
[0191] These computer program instructions may be loaded into a general-purpose computer, a dedicated computer, or other programmable data processing device for manufacturing a machine, thereby creating means for instructions executed on the computer or other programmable data processing device to realize a function specified in one or more flowchart blocks. These computer program instructions may also be stored in computer-readable memory that can be instructed to function in a particular way, thereby manufacturing a product that includes instruction means for instructions stored in computer-readable memory to realize a function specified in one or more flowchart blocks. The computer program instructions may also be loaded into a computer or other programmable data processing device to manufacture a computer realization process such that a set of operating steps executed on the computer or other programmable device provide steps for instructions executed on the computer or other programmable device to perform a function specified in one or more flowchart blocks.
[0192] Therefore, functional blocks in block diagrams and flowcharts support combinations of means for performing a specified function, combinations of steps for performing a specified function, and program instruction means for performing a specified function. It will also be understood that each functional block in block diagrams and flowcharts, as well as combinations of functional blocks in block diagrams and flowcharts, can be implemented by a dedicated hardware-based computer system that performs the specified function or step, or by an appropriate combination of dedicated hardware and computer instructions. Furthermore, the examples and descriptions of process flows may refer to user windows, web pages, websites, web forms, prompts, etc. The implementer will understand that the illustrated steps described herein may include several configurations, including the use of windows, web pages, web forms, pop-up windows, prompts, etc. Furthermore, it should be understood that multiple steps illustrated and described may be combined into a single web page and / or window, but are expanded for clarity. In other cases, a step illustrated and described as a single process step may be divided into multiple web pages and / or windows, but are combined for clarity.
[0193] Molecular profiling Molecular profiling techniques provide methods for selecting individual candidate therapies that can improve the clinical course of an individual with a disease or disorder such as cancer. Molecular profiling techniques provide individual clinical benefits, such as identifying treatment regimens that offer longer progression-free survival (PFS), longer disease-free survival (DFS), longer overall survival (OS), or longer lifespan. Methods and systems described herein relate to individual-based molecular profiling of cancer that can identify optimal treatment regimens. Molecular profiling provides an individualized approach for selecting candidate therapies that are likely to provide benefits to cancer. Using the molecular profiling methods described herein, treatment can be guided in any desired setting, including the setting of first-line / standard treatment, or in cases of patients with a poor prognosis, such as patients with metastatic disease, or patients whose cancer has progressed with standard first-line treatment, or patients whose cancer has progressed with previous chemotherapy or hormone therapy.
[0194] Using the systems and methods of the invention, patients may be classified as likely or unlikely to benefit from or respond to a variety of treatments. Unless otherwise specified, the terms “responder” or “non-responder” as used herein mean any appropriate indicator that a treatment provides a patient with benefit (“responder” or “benefitter”) or lacks benefit to the patient (“non-responder” or “non-benefitter”). Such indicators may be determined using accepted clinical response criteria, such as the standard RECIST (Response Evaluation Criteria in Solid Tumors) criteria, or other useful patient response criteria, such as progression-free survival (PFS), progression-free time (TTP), disease-free survival (DFS), time to initiation of the next treatment (TNT, TTNT), tumor reduction or disappearance, etc. RECIST is a set of rules published by an international consortium that defines whether a tumor in a cancer patient will improve ("respond"), remain unchanged ("stabilize"), or worsen ("progress") during treatment. Where used herein, unless otherwise specified, patient "benefit" from treatment may refer to any appropriate measure of improvement, including a RECIST response or a longer PFS / TTP / DFS / TNT / TTNT, and "lack of benefit" from treatment may refer to any appropriate measure of disease progression during treatment. Generally, disease stabilization is considered a benefit, but in certain circumstances, stabilization may be considered a lack of benefit, where it is specifically stated herein. Where there is no acceptable level of prediction of benefit or lack of benefit, a predicted or indicated benefit may be marked as "uncertain." In some cases, a benefit may be considered uncertain if it is not possible to calculate the benefit, for example, due to the lack of necessary data.
[0195] Personalized medicine based on pharmacogenetic insights, such as those provided by molecular profiling as described herein, is increasingly taken for granted by some practitioners and general journals, and forms the basis for hopes of improving cancer treatment. However, the molecular profiling taught herein represents a fundamental departure from conventional approaches to tumor treatment, where patients are largely grouped and treated using methods based on findings from optical microscopy and disease staging. Traditionally, differential responses to specific treatment strategies have been determined only after treatment has been administered, i.e., retrospectively. "Standard" approaches to disease treatment rely on being generally true with respect to a given cancer diagnosis, and treatment responses are scrutinized by randomized Phase III clinical trials to form "standard treatments" in medical practice. The results of these trials are compiled into consensus statements by guideline organizations such as the National Comprehensive Cancer Information Network and the American Society of Clinical Oncology. The NCCN Compendium® contains authoritative, scientifically derived information designed to support decision-making regarding the appropriate use of drugs and biologics in cancer patients. NCCN Compendium® is recognized by the Centers for Medicare & Medicaid Services (CMS) and UnitedHealthcare as an authoritative reference for oncology insurance coverage. On Compendium treatments are recommended by such guides. The biostatistical methods used to validate the results of clinical trials rely on minimizing patient-to-patient variability and are based on declaring the possibility of error that one method may be superior to another for a patient group determined solely by light microscopy and disease staging (and not by individual variability in the tumor). The molecular profiling methods described herein take advantage of such individual variability. This method can provide candidate therapies that can then be selected by a physician to treat a patient.
[0196] Molecular profiling can be used to provide a comprehensive view of the biological state of a sample. In some embodiments, molecular profiling is used for whole-tumor profiling. Thus, several molecular techniques are used to assess the state of a tumor. Whole-tumor profiling can be used to select candidate therapies for a tumor. Molecular profiling can be used to select candidate therapeutic substances for any sample at any stage of disease. In some embodiments, methods such as those described herein are used to profile newly diagnosed cancers. Candidate therapies indicated by molecular profiling can be used to select therapies for treating newly diagnosed cancers. In other embodiments, methods such as those described herein are used, for example, to profile cancers that have already been treated with one or more standard therapies. In some embodiments, the cancer is resistant to previous treatments. For example, the cancer may be resistant to standard therapies for cancer. The cancer may be metastatic or other recurrent cancer. The therapy may be on-compendium or off-compendium therapy.
[0197] Molecular profiling can be performed by any known means for detecting molecules in a biological sample. Molecular profiling includes methods such as nucleic acid sequencing, e.g., DNA sequencing or RNA sequencing; immunohistochemistry (IHC); in-situ hybridization (ISH); fluorescence in-situ hybridization (FISH); chromogenic in-situ hybridization (CISH); PCR amplification (e.g., qPCR or RT-PCR); various types of microarrays (mRNA expression arrays, low-density arrays, protein arrays, etc.); various types of sequencing (Sanger, pyrosequencing, etc.); comparative genomic hybridization (CGH); high-throughput or next-generation sequencing (NGS); Northern blotting; Southern blotting; immunoassays; and any other suitable techniques for assaying the presence or quantity of biomolecules of interest. In various embodiments, one or more of these methods can be used concurrently or sequentially with respect to each other to evaluate target genes disclosed herein.
[0198] Molecular profiling of individual samples is used to select one or more candidate therapies for the disorder in a subject, for example, by identifying targets for drugs that may be effective against a given cancer. For example, candidate therapies may be therapies, experimental drugs, government or regulatory-approved drugs, or any combination of such drugs, that are known to affect cells that differentially express genes identified by molecular profiling techniques (the biological sample may have been taken and studied and approved for the same or different specific indications as the subject being molecularly profiled).
[0199] When multiple biomarker targets are identified by evaluating target genes through molecular profiling, one or more decision rules can be applied to prioritize the selection of specific therapeutic agents for the treatment of an individual on a personalized basis. Rules such as those described herein support the prioritization of treatment, e.g., the direct results of molecular profiling, the expected efficacy of the therapeutic agent, prior treatment history (same or other), expected side effects, availability of the therapeutic agent, cost of the therapeutic agent, drug interactions, and other factors considered by the treating physician. Based on the recommended and prioritized therapeutic agent targets, the physician can determine the course of treatment for a particular individual. Thus, molecular profiling methods and systems such as those described herein allow for the selection of candidate therapies based on the individual characteristics of diseased cells, e.g., tumor cells, and other individualizing factors in the subject requiring treatment, in contrast to relying on conventional, universally applicable methods used to treat individuals suffering from diseases, particularly cancer. In some cases, the recommended therapy is one that is not typically used to treat the disease or disorder afflicting the subject. In some cases, the recommended therapy is used after standard treatments no longer provide sufficient efficacy.
[0200] The treating physician can use the results of molecular profiling to optimize the treatment regimen for the patient. Candidate therapies identified by methods such as those described herein can be used to treat the patient, although such therapies are not required by the method. In fact, the analysis of molecular profiling results and the identification of candidate therapies based on such results can be automated and do not require the involvement of a physician.
[0201] Biological entities Nucleic acids include deoxyribonucleotides or ribonucleotides and polymers thereof in either single-stranded or double-stranded forms, or their complements. Nucleic acids may contain known nucleotide analogs or modified skeletal residues or bonds, which are synthetic, natural, and unnatural, and which have similar binding properties to the reference nucleic acid and are metabolized in a similar manner to the reference nucleotide. Examples of such analogs include, but are not limited to, phosphorothioates, phosphoramides, methylphosphonates, chiral-methylphosphonates, 2-O-methylribonucleotides, and peptide-nucleic acids (PNAs). Nucleic acid sequences may include explicitly stated sequences in addition to their conservatively modified variants (e.g., degenerate codon substitutions) and complementary sequences. Specifically, degenerate codon substitution can be achieved by generating sequences in which the third position of one or more selected (or all) codons is replaced with a mixed base and / or a deoxyinosine residue (Batzer et al., Nucleic Acid Res. 19:5081 (1991); Ohtsuka et al., J. Biol. Chem. 260:2605-2608 (1985); Rossolini et al., Mol. Cell Probes 8:91-98 (1994)). The term nucleic acid can be used interchangeably with gene, cDNA, mRNA, oligonucleotide, and polynucleotide.
[0202] A specific nucleic acid sequence may implicitly include nucleic acid sequences encoding specific sequences as well as “splice variants” and cleavage forms. Similarly, a specific protein encoded by a nucleic acid may include any protein encoded by a splice variant or cleavage form of that nucleic acid. A “splice variant,” as its name suggests, is the product of alternative splicing of a gene. After transcription, the original nucleic acid transcript may be spliced so that different (alternative) nucleic acid splice products encode different polypeptides. The mechanisms of splice variant production vary, but include alternative splicing of exons. Alternative polypeptides obtained by read-through transcription from the same nucleic acid are also included by this definition. Any product of a splicing reaction, including recombinant splice products, is included in this definition. Nucleic acids can be cleaved at their 5' or 3' end. Polypeptides can be cleaved at their N-terminus or C-terminus. Cleavage versions of nucleic acid or polypeptide sequences can be native or produced using recombinant techniques.
[0203] The terms “gene variant” and “nucleotide variant” are used interchangeably herein to refer to changes or alterations to a reference human gene or cDNA sequence at a particular locus, including but not limited to deletions, insertions, inversions, and substitutions of nucleotide bases in coding and non-coding regions. A deletion may be a single nucleotide base, a portion or region of a gene's nucleotide sequence, or an entire gene sequence. An insertion may be one or more nucleotide bases. Gene variants or nucleotide variants may occur in transcriptional regulatory regions, untranslated regions of mRNA, exons, introns, exon / intron junctions, etc. Potentially, gene variants or nucleotide variants may result in stop codons, frameshifts, amino acid deletions, altered gene transcript splice forms, or altered amino acid sequences.
[0204] An allele or genetic allele generally includes a natural gene having a reference sequence or a gene containing a specific nucleotide variant.
[0205] A haplotype refers to a combination of gene (nucleotide) variants in a region of mRNA or genomic DNA on a chromosome found in an individual. Thus, a haplotype typically includes several genetically linked polymorphic variants that are inherited together as a unit.
[0206] As used herein, the term "amino acid variant" is used to refer to an amino acid change to a reference human protein sequence resulting from a gene variant or nucleotide variant relative to a reference human gene encoding the reference protein. The term "amino acid variant" is intended to include not only a single amino acid substitution in the amino acid sequence of the reference protein, but also amino acid deletions, insertions, and other significant changes.
[0207] As used herein, the term "genotype" means the nature of the nucleotides at a particular nucleotide variant marker (or locus) in one or both alleles of a gene (or a particular chromosomal region). For a particular nucleotide position of a gene of interest, the nucleotides at that locus or its equivalent in one or both alleles form the genotype of the gene at that locus. A genotype can be homozygous or heterozygous. Thus, "genotyping" means determining the genotype, i.e., the nucleotides at a particular gene locus. Genotyping can also be performed by determining amino acid variants at specific positions of a protein that can be used to infer the corresponding nucleotide variants.
[0208] The term "locus" refers to a specific position or site in a gene sequence or protein. Therefore, a particular gene locus may contain one or more consecutive nucleotides, or a particular locus in a polypeptide may contain one or more amino acids. Furthermore, a locus may refer to a specific location in a gene where one or more nucleotides are deleted, inserted, or inverted.
[0209] Unless otherwise specified or understood by those skilled in the art, the terms “polypeptide,” “protein,” and “peptide” are used interchangeably herein to refer to amino acid chains in which amino acid residues are linked by covalent peptide bonds. Amino acid chains can be of any length, including full-length proteins, consisting of at least two amino acids. Unless otherwise specified, polypeptides, proteins, and peptides also encompass their various modified forms, including, but not limited to, glycosylated forms, phosphorylated forms, and so on. Polypeptides, proteins, or peptides may also be referred to as gene products.
[0210] A list of genes and gene products that can be assayed by molecular profiling techniques is presented herein. Lists of genes may be presented in relation to molecular profiling techniques for detecting gene products (e.g., mRNA or protein). Those skilled in the art will understand that this implies the detection of gene products of the listed genes. Similarly, lists of gene products may be presented in relation to molecular profiling techniques for detecting gene sequences or copy numbers. Those skilled in the art will understand that this implies the detection of the gene corresponding to the gene product, including, for example, the DNA encoding the gene product. As recognized by those skilled in the art, “biomarker” or “marker” includes genes and / or gene products, depending on the context.
[0211] The terms “label” and “detectable label” can refer to any composition detectable by spectroscopic, photochemical, biochemical, immunochemical, electrical, optical, chemical or similar methods. Such labels include biotin for staining with labeled streptavidin conjugates, magnetic beads (e.g., DYNABEADS®), fluorescent dyes (e.g., fluorescein, Texas Red, rhodamine, green fluorescent protein, etc.), radiolabels (e.g., 3 H, 125 I, 35 S, 14 C, or 32These include colorimetric labels such as enzymes (e.g., horseradish peroxidase, alkaline phosphatase, and others commonly used in ELISA), and colloidal gold or colored glass or plastic (e.g., polystyrene, polypropylene, latex, etc.) beads. Patents teaching the use of such labels include U.S. Patents 3,817,837; 3,850,752; 3,939,350; 3,996,345; 4,277,437; 4,275,149; and 4,366,241. Means for detecting such labels are well known to those skilled in the art. For example, radioactive labels may be detected using photographic film or a scintillation counter, and fluorescent markers may be detected using a photodetector to detect emitted light. Enzymatic labeling is typically detected by providing a substrate to an enzyme and detecting the reaction product produced by the enzyme's action on the substrate, while colorimetric labeling is detected simply by visualizing the colored label. Labels can include, for example, ligands, fluorophores, chemiluminescent agents, enzymes, and antibodies that can serve as members of a binding pair specific to the labeled antibody. An overview of labeling, labeling procedures, and label detection can be found in Polak and Van Noorden, Introduction to Immunocytochemistry, 2nd ed., Springer Verlag, NY (1997); and the combined handbook and catalog, Haugland Handbook of Fluorescent Probes and Research Chemicals (1996), published by Molecular Probes, Inc.
[0212] Detectable labels include, but are not limited to, nucleotides (labeled or unlabeled), compomers, sugars, peptides, proteins, antibodies, chemical compounds, conductive polymers, binding sites (e.g., biotin), mass tags, colorimetric agents, luminescent agents, chemiluminescent agents, light scattering agents, fluorescent tags, radioactive tags, charge tags (charge or magnetic charge), volatile tags, and hydrophobic tags, as well as biomolecules (e.g., bound antibody / antigen, antibody / antibody, antibody / antibody fragment, antibody / antibody receptor, antibody / protein A or protein G, hapten / antihapten, biotin / avidin, biotin / streptavidin, folic acid / folate-binding protein, vitamin B12 / intrinsic factor, chemical reactive groups / complementary chemical reactive groups (e.g., sulfhydryl / maleimide, sulfhydryl / haloacetyl derivatives, amine / isotriocyanate, amine / succinimidyl esters, and amine / sulfonyl halide members).
[0213] The terms “primer,” “probe,” and “oligonucleotide” are used interchangeably herein to refer to relatively short nucleic acid fragments or sequences. They may include DNA, RNA, or hybrids thereof, or chemically modified analogs or derivatives thereof. Typically, they are single-stranded. However, they may also be double-stranded, having two complementary strands that can be separated by denaturation. Primers, probes, and oligonucleotides typically have a length of about 8 to about 200 nucleotides, preferably about 12 to about 100 nucleotides, and more preferably about 18 to about 50 nucleotides. They may be labeled with detectable markers or modified using conventional methods for various molecular biological applications.
[0214] When the term “isolated” is used in relation to nucleic acids (e.g., genomic DNA, cDNA, mRNA, or fragments thereof), it is intended to mean that a nucleic acid molecule exists in a form substantially separated from other naturally occurring nucleic acids that are normally associated with that molecule. Since naturally occurring chromosomes (or their viral equivalents) contain long nucleic acid sequences, an isolated nucleic acid can be a nucleic acid molecule that has only a portion of the nucleic acid sequence in a chromosome, but does not have one or more other portions present in the same chromosome. More specifically, an isolated nucleic acid can contain a naturally occurring nucleic acid sequence adjacent to the nucleic acid in a naturally occurring chromosome (or its viral equivalent). An isolated nucleic acid can be substantially separated from other naturally occurring nucleic acids on different chromosomes of the same organism. An isolated nucleic acid can also be a composition in which a particular nucleic acid molecule is significantly concentrated to constitute at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or at least 99% of the total nucleic acid in the composition.
[0215] Isolated nucleic acids can be hybrid nucleic acids having a specific nucleic acid molecule covalently linked to one or more nucleic acid molecules, where one or more nucleic acid molecules are not naturally occurring nucleic acids adjacent to the specific nucleic acid. For example, isolated nucleic acids can be present in a vector. In addition, the specific nucleic acid may have a nucleotide sequence identical to that of a naturally occurring nucleic acid, or a modified form of that nucleic acid with one or more mutations, such as nucleotide substitutions, deletions / insertions, or inversions, or to a mutain.
[0216] Isolated nucleic acids can be prepared from recombinant host cells (in which nucleic acids are recombinantly amplified and / or expressed), or they can be chemically synthesized nucleic acids having a native nucleotide sequence or an artificially modified form thereof.
[0217] When used in relation to nucleic acid hybridization, the term "high stringency hybridization conditions" includes hybridization carried out overnight at 42°C in a solution containing 50% formamide, 5×SSC (750mM NaCl, 75mM sodium citrate), 50mM sodium phosphate, pH 7.6, 5×Denhart solution, 10% dextran sulfate, and 20 micrograms / ml of denatured fragmented salmon sperm DNA, wherein the hybridization filter is washed in 0.1×SSC at approximately 65°C. When used in relation to nucleic acid hybridization, the term "moderately stringent hybridization conditions" includes hybridization performed overnight at 37°C in a solution containing 50% formamide, 5×SSC (750mM NaCl, 75mM sodium citrate), 50mM sodium phosphate, pH 7.6, 5×Denhart solution, 10% dextran sulfate, and 20 micrograms / ml of denatured fragmented salmon sperm DNA, with the hybridization filter being washed in 1×SSC at approximately 50°C. It should be noted, as will be apparent to those skilled in the art, that similar stringent hybridization conditions can be achieved using many other hybridization methods, solutions, and temperatures.
[0218] For the purpose of comparing two different nucleic acid or polypeptide sequences, one sequence (test sequence) may be described as being identical to another sequence (comparison sequence) by a specific percentage. This percentage of identicality can be determined by the algorithm described by Karlin and Altschul, Proc. Natl. Acad. Sci. USA, 90:5873-5877 (1993), which is incorporated into various BLAST programs. The percentage of identicality can also be determined using the "BLAST 2 Sequences" tool available on the National Center for Biotechnology Information (NCBI) website. See also Tatusova and Madden, FEMS Microbiol. Lett., 174(2):247-250 (1999). For DNA-DNA pair comparisons, the BLASTN program is used with default parameters (e.g., match: 1; mismatch: -2; open gap: 5 penalty; extended gap: 2 penalty; gap x_dropoff: 50; expected value: 10; and word size: 11, with filter). For protein-protein sequence pair comparisons, the BLASTP program can be employed with default parameters (e.g., matrix: BLOSUM62; open gap: 11; extended gap: 1; x_dropoff: 15; expected value: 10.0; and word size: 3, with filter). The percentage of identical sequences is calculated by using BLAST to align the test sequence with the comparison sequence, determining the number of amino acids or nucleotides in the aligned test sequence that are identical to the amino acids or nucleotides at the same position in the comparison sequence, and dividing the number of identical amino acids or nucleotides by the number of amino acids or nucleotides in the comparison sequence. When BLAST is used to compare two sequences, BLAST aligns the sequences and yields a percentage of identical sequences across a given aligned region. If two sequences are aligned along their entire length, the identical percentages obtained by BLAST are identical percentages of these two sequences.If BLAST does not align two sequences over their entire length, the number of identical amino acids or nucleotides in the unaligned regions of the test and comparison sequences is considered zero. The percentage of identical amino acids or nucleotides is calculated by adding up the number of identical amino acids or nucleotides in the aligned regions and dividing that number by the length of the comparison sequence. Various versions of the BLAST program, such as BLAST 2.1.2 or BLAST+ 2.2.22, can be used to compare sequences.
[0219] The subject or individual may be any animal that may benefit from the methods described herein, including, for example, humans and non-human mammals, such as primates, rodents, horses, dogs and cats. Subjects include, but are not limited to, eukaryotes, most preferably mammals, such as primates, such as chimpanzees or humans, cattle; dogs; cats; rodents, such as guinea pigs, rats, mice; rabbits; or birds; reptiles; or fish. Subjects specifically intended for treatment using the methods described herein include humans. Subjects may also be referred to herein as individuals or patients. In these methods, the subject is someone having colorectal cancer, for example, someone diagnosed with colorectal cancer. Methods for identifying subjects having colorectal cancer, such as using biopsy, are known in the art. For example, see Fleming et al., J Gastrointest Oncol. 2012 Sep; 3(3): 153-173; Chang et al., Dis Colon Rectum. 2012; 55(8):831-43.
[0220] Treatment of a disease or individual by the methods described herein is a method for obtaining beneficial or desired medical outcomes, including clinical outcomes, but not necessarily a method for obtaining a cure. Beneficial or desired clinical outcomes for the methods described herein include, but are not limited to, reduction or recovery of one or more symptoms, whether detectable or undetectable; reduction of disease severity; stabilization of the disease (i.e., no worsening); prevention of disease progression; delay or slowing of disease progression; recovery or mitigation of the disease; and remission (whether partial or total remission). Treatment also includes extending survival compared to predicted survival without treatment or with different treatments. Treatment may include administration of either the FOLFOX or FOLFIRI regimen. A biomarker generally refers to a molecule that, when detected in a tissue or cell, has characteristics that can provide predictive, diagnostic, prognostic, and / or theranostic information about sensitivity or resistance to a candidate treatment, including but not limited to genes or their products, nucleic acids (e.g., DNA, RNA), proteins / peptides / polypeptides, glycan structures, lipids, and glycolipids.
[0221] Biological samples As used herein, samples include any relevant biological samples that can be used for molecular profiling, such as biopsies or tissues, body fluids, autopsy specimens taken during surgical or other procedures, and tissue sections such as frozen sections taken for histological purposes. Such samples include blood and blood fractions or products (e.g., serum, buffy coat, plasma, platelets, red blood cells, etc.), sputum, malignant exudate, buccal cell tissue, cultured cells (e.g., primary cultures, explants, and transformed cells), feces, urine, other biological fluids or body fluids (e.g., prostatic fluid, gastric fluid, intestinal fluid, renal fluid, pulmonary fluid, cerebrospinal fluid, etc.), and others. Samples may include biological materials that are fresh-frozen and formalin-fixed paraffin-embedded (FFPE) blocks, formalin-fixed paraffin-embedded, or in RNA preservative + formalin fixative. More than one sample of more than one type may be used for each patient. In a preferred embodiment, samples include fixed tumor samples.
[0222] The specimens used in the system and method of the present invention may be formalin-fixed paraffin-embedded (FFPE) specimens. FFPE specimens may be one or more of fixed tissue, unstained slides, bone marrow cores or clots, core needle biopsies, malignant fluids, and fine-needle aspiration (FNA). In one embodiment, fixed tissue includes a tumor-containing formalin-fixed paraffin-embedded (FFPE) block from surgery or biopsy. In another embodiment, unstained slides include unstained, charged, unbaked slides from a paraffin block. In yet another embodiment, bone marrow cores or clots include a decalcified core. Formalin-fixed cores and / or clots may be paraffin-embedded. In yet another embodiment, core needle biopsies include one, two, three, four, five, six, seven, eight, nine, ten or more, for example, three to four, paraffin-embedded biopsy specimens. 18-gauge needle biopsies may be used. The malignant fluid may contain a sufficient volume of fresh pleural / peritoneal fluid to produce a 5 × 5 × 2 mm cell pellet. The fluid can be formalin-fixed in paraffin block form. In one embodiment, core needle biopsies may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more, for example, 4 to 6, paraffin-embedded aspirates.
[0223] Samples may be processed according to techniques understood by those skilled in the art. Samples may be, but are not limited to, fresh, frozen, or fixed cells or tissues. In some embodiments, samples may include formalin-fixed paraffin-embedded (FFPE) tissue, fresh tissue, or fresh-frozen (FF) tissue. Samples may include cultured cells, including primary or immortalized cell lines derived from the subject sample. Samples may also refer to extracts from samples derived from the subject. For example, a sample may include DNA, RNA, or proteins extracted from tissue or body fluids. Many techniques and commercially available kits are available for such purposes. Fresh samples from individuals may be treated with an active substance to preserve RNA before further processing, such as cell lysis and extraction. Samples may include frozen samples collected for other purposes. Samples may be associated with relevant information such as age, sex, and clinical symptoms present in the subject; origin of the sample; and how the sample was collected and stored. Samples are typically obtained from the subject.
[0224] A biopsy is the process of obtaining a tissue sample for diagnosis or prognosis, and includes the tissue specimen itself. Any biopsy technique known in the art can be applied to the molecular profiling methods of this disclosure. The biopsy technique applied may depend on several factors, including the type of tissue to be evaluated (e.g., colon, prostate, kidney, bladder, lymph node, liver, bone marrow, blood cells, lung, breast, etc.), the size and type of tumor (e.g., solid or floating, blood or ascites). Representative biopsy techniques include, but are not limited to, excisional biopsy, incisional biopsy, needle biopsy, surgical biopsy, and bone marrow biopsy. An “excisional biopsy” refers to the removal of the entire tumor mass, along with a small margin of surrounding normal tissue. An “incisional biopsy” refers to the removal of a wedge-shaped tissue containing the cross-sectional diameter of the tumor. Molecular profiling may use a “core needle biopsy” of the tumor mass, or a “fine-needle aspiration biopsy” which generally obtains a suspension of cells from within the tumor mass. Biopsy techniques are discussed, for example, in Chapter 70 and throughout Part V of Harrison's Principles of Internal Medicine, Kasper, et al., eds., 16th ed., 2005.
[0225] Unless otherwise specified, the “sample” as referred herein for the molecular profiling of a patient may include more than one physical specimen. In an unrestricted example, a “sample” may include multiple sections from a tumor, e.g., multiple sections of an FFPE block or multiple core needle biopsy sections. In another unrestricted example, a “sample” may include multiple biopsy specimens, e.g., one or more surgical biopsy specimens, one or more core needle biopsy specimens, one or more fine-needle aspiration biopsy specimens, or any useful combination thereof. In yet another unrestricted example, a molecular profile may be generated for a subject using a “sample” that includes solid tumor specimens and fluid specimens. In some embodiments, a sample is a unit sample, i.e., a single physical specimen.
[0226] Standard molecular biological techniques known in the art and not specifically described are generally referred to in Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York (1989), Ausubel et al., Current Protocols in Molecular Biology, John Wiley and Sons, Baltimore, Md. (1989), Perbal, A Practical Guide to Molecular Cloning, John Wiley & Sons, New York (1988), Watson et al., Recombinant DNA, Scientific American Books, New York, and Birren et al (eds) Genome Analysis: A Laboratory Manual Series, Vols. 1-4 Cold Spring Harbor Laboratory Press, New York. (1998), and the methodologies set forth in U.S. Patents No. 4,666,828; No. 4,683,202; No. 4,801,531; No. 5,192,659 and No. 5,272,057, which are incorporated herein by reference. Polymerase chain reaction (PCR) can generally be carried out as described in PCR Protocols: A Guide to Methods and Applications, Academic Press, San Diego, Calif. (1990).
[0227] Small vesicles The sample may contain vesicles. Methods as described herein may include examining one or more vesicles, including examining a population of vesicles. When used herein, vesicles are membrane vesicles that have been shed from a cell. Vesicles or membrane vesicles include, but are not limited to, circulating microvesicles (cMVs), microvesicles, exosomes, nanovesicles, dexosomes, blebs, blebbies, prostasomes, microparticles, intraluminal vesicles, membrane fragments, intraluminal endosomal vesicles, endosomal-like vesicles, exocytosis vehicles, endosome vesicles, endosomal vesicles, apoptotic bodies, multivesicles, secretory vesicles, phospholipid vesicles, liposomal vesicles, argosomes, texasomes, secretomes, tolerosomes, melanosomes, oncosomes, or exocytosis vehicles. Furthermore, although vesicles may be produced by different cellular processes, the methods described herein are not limited to or dependent on any single mechanism, as long as such vesicles are present in a biological sample and can be characterized by the methods disclosed herein. Unless otherwise specified, methods utilizing one type of vesicle can be applied to other types of vesicles. Vesicles consist of a spherical structure having a cell membrane-like lipid bilayer surrounding an internal compartment that may contain soluble components sometimes referred to as a payload. In some embodiments, the methods described herein utilize exosomes, which are small secretory vesicles with a diameter of approximately 40–100 nm. For a review of membrane vesicles, including types and characterization, see Thery et al., Nat Rev Immunol. 2009 Aug;9(8):581–93. Some characteristics of different types of vesicles include those listed in Table 1:
[0228] (Table 1) Characteristics of vesicles TIFF2026048991000002.tif127169 Abbreviation: Phosphatidylserine (PPS); Electron microscopy (EM)
[0229] Vesicles contain expelled membrane-bound particles or "microparticles" derived from either the plasma membrane or the intomer membrane. Vesicles can be released from cells into the extracellular environment. Cells releasing vesicles include, but are not limited to, cells derived from or originating from the ectoderm, endoderm, or mesoderm. Cells may have undergone genetic, environmental, and / or any other variation or modification. For example, a cell may be a tumor cell. Vesicles can reflect any changes in the source cell and thereby reflect changes in the originating cell, for example, a cell with various genetic mutations. In one mechanism, vesicles are generated within a cell when a segment of the cell membrane spontaneously invaginates and eventually undergoes exocytosis (see, e.g., Keller et al., Immunol. Lett. 107 (2): 102-8 (2006)). Vesicles also include cell-derived structures bound by a lipid bilayer resulting from both the separation of extruded blebings and the sealing of plasma membrane portions, or from the extrusion of any intracellular membrane-bound vesicular structure containing various membrane-related proteins of tumor origin, which include surface-bound molecules derived from host circulation that selectively bind to tumor-derived proteins together with molecules contained in the vesicular lumen, including but not limited to tumor-derived microRNAs or intracellular proteins. Blebs and blebings are further described in Charras et al., Nature Reviews Molecular and Cell Biology, Vol. 9, No. 11, p. 730-736 (2008). Vesicles extruded from tumor cells into circulation or into body fluids are sometimes referred to as “circulating tumor-derived vesicles.” If such vesicles are exosomes, they may be referred to as circulating tumor-derived exosomes (CTEs). In some cases, vesicles may originate from specific cellular origins. CTEs, like cell-origin-specific vesicles, typically have one or more unique biomarkers that allow for the isolation of CTEs or cell-origin-specific vesicles, for example, from body fluids, in a sometimes specific manner.For example, cell or tissue-specific markers are used to identify cell origin. Examples of such cell or tissue-specific markers are disclosed herein and can be further accessed from the Tissue-specific Gene Expression and Regulation (TiGER) database, available at bioinfo.wilmer.jhu.edu / tiger / ; Liu et al. (2008) TiGER: a database for tissue-specific gene expression and regulation. BMC Bioinformatics. 9:271; and TissueDistributionDBs, available at genome.dkfz-heidelberg.de / menu / tissue_db / index.html.
[0230] Vesicles can have diameters greater than approximately 10 nm, 20 nm, or 30 nm. Vesicles can have diameters greater than 40 nm, 50 nm, 100 nm, 200 nm, 500 nm, 1000 nm, or greater than 10,000 nm. Vesicles can have diameters of approximately 30–1000 nm, approximately 30–800 nm, approximately 30–200 nm, or approximately 30–100 nm. In some embodiments, vesicles have diameters of 10,000 nm, 1000 nm, 800 nm, 500 nm, 200 nm, 100 nm, 50 nm, 40 nm, 30 nm, less than 20 nm, or less than 10 nm. As used herein, the term “approximately” in relation to a number means that a variation of more or less than 10% above or below the number is within the range attributable to a particular value. Typical sizes for various types of vesicles are shown in Table 1. Vesicles can be examined to measure the diameter of a single vesicle or any number of vesicles. For example, the range of diameters of a vesicle population or the average diameter of a vesicle population can be determined. The diameter of a vesicle can be examined using imaging techniques known in the art, such as electron microscopy. In some embodiments, the diameter of one or more vesicles is determined using optical particle detection. See, for example, U.S. Patent No. 7,751,053, issued July 6, 2010, entitled "Optical Detection and Analysis of Particles"; and U.S. Patent No. 7,399,600, issued July 15, 2010, entitled "Optical Detection and Analysis of Particles".
[0231] In some embodiments, vesicles are assayed directly from a biological sample without prior isolation, purification, or concentration from the biological sample. For example, the amount of vesicles in a sample can itself provide a biosignature that provides diagnostic, prognostic, or theranostic determination. Alternatively, vesicles in a sample may be isolated, captured, purified, or concentrated from the sample before analysis. As described above, isolation, capture, or purification, as used herein, includes partial isolation, partial capture, or partial purification separated from other components in the sample. Isolation of vesicles can be performed using a variety of techniques, such as those described herein or known in the art, including, but not limited to, size exclusion chromatography, density gradient centrifugation, fractional centrifugation, nanomembrane ultrafiltration, immunoadsorption capture, affinity purification, affinity capture, immunoassay, immunoprecipitation, microfluidic separation, flow cytometry, or combinations thereof.
[0232] By examining vesicles and comparing their characteristics to a standard, phenotypic characterization can be provided. In some embodiments, surface antigens on vesicles are examined. Vesicles or vesicle populations possessing a specific marker can be referred to as positive (biomarker+) vesicles or vesicle populations. For example, a DLL4+ population refers to a population of vesicles bound to DLL4. Conversely, a DLL4- population is one that is not bound to DLL4. Surface antigens can provide anatomical and / or cellular origin information of vesicles as well as other phenotypic information, such as indicators of tumor status. For example, vesicles found in a patient's sample can be examined for surface antigens indicating colorectal origin and the presence of cancer, thereby identifying vesicles associated with colorectal cancer cells. Surface antigens may include any biological entities that provide detectable information on the vesicle membrane surface, including, but not limited to, surface proteins, lipids, carbohydrates, and other membrane components. For example, positive detection of vesicles obtained from the colon expressing a tumor antigen can indicate that the patient has colorectal cancer. Thus, by using methods such as those described herein, for example, by examining disease-specific and cell-specific biomarkers in one or more vesicles obtained from a subject, any disease or condition related to anatomical or cellular origin can be characterized.
[0233] In various embodiments, one or more vesicle payloads are examined to provide phenotypic characterization. The payloads containing vesicles include, but are not limited to, any informative biological entities that can be detected as encapsulated within the vesicle, including, but not limited to, proteins and nucleic acids, such as genomes or cDNA, mRNA, or functional fragments thereof, as well as microRNAs (miRs). In addition, methods such as those described herein are directed to provide phenotypic characterization by detecting vesicle surface antigens (in addition to or exclusively with respect to the vesicle payload). For example, vesicles can be characterized using a binder specific to the vesicle surface antigen (e.g., an antibody or aptamer), and the bound vesicles can be further examined to identify one or more payload components disclosed herein. As described herein, phenotypes can be characterized by comparing the level of vesicles containing the surface antigen of interest or the payload of interest to a reference. For example, overexpression of cancer-related surface antigens or vesicle payloads, such as tumor-related mRNA or microRNA, in a sample compared to a reference can indicate the presence of cancer in the sample. The biomarkers being examined may be present or absent, increased or decreased, based on the selection of the desired target sample and the comparison of the target sample with the desired reference sample. Non-limiting examples of target samples include disease; treated / untreated; and different time points in a longitudinal study, for example; non-limiting examples of reference samples include non-disease; normal; different time points; and those sensitive or resistant to the candidate treatment.
[0234] In some embodiments, molecular profiling as described herein includes the analysis of microvesicles, such as circulating microvesicles.
[0235] microRNA Various biomarker molecules can be examined in biological samples or vesicles obtained from such biological samples. MicroRNAs constitute one class of biomarkers that can be examined by methods such as those described herein. MicroRNAs, also referred to herein as miRNAs or miRs, are short RNA chains approximately 21–23 nucleotides long. miRNAs include non-coding RNAs, which are encoded by genes that are transcribed from DNA but not translated into proteins. miRs are processed from primary transcripts known as pri-miRNAs into short stem-loop structures called pre-miRNAs, and finally into the resulting single-stranded miRNAs. Pre-miRNAs typically form structures that fold over themselves in a self-complementary region. These structures are then processed by the nuclease Dicer in animals or DCL1 in plants. Mature miRNA molecules are partially complementary to one or more messenger RNA (mRNA) molecules and can function to regulate protein translation. Identified miRNA sequences can be accessed from publicly available databases such as www.microRNA.org, www.mirbase.org, or www.mirz.unibas.ch / cgi / miRNA.cgi.
[0236] miRNAs are generally assigned numbers according to the naming convention "mir-[number]". The numbers of miRNAs are assigned according to the order of their discovery relative to previously identified miRNA species. For example, if the last published miRNA was mir-121, the next discovered miRNA would be named mir-122, and so on. If a miRNA is found to be homologous to a known miRNA from a different organism, it can be given an optional bioidentifier in the form of [biological identifier]-mir-[number]. Identifiers include hsa for Homo sapiens and mmu for mouse (Mus Musculus). For example, a human homolog of mir-121 might be called hsa-mir-121, while a mouse homolog might be called mmu-mir-121.
[0237] Mature microRNAs are typically named with the prefix "miR," while genes or precursor miRNAs are named with the prefix "mir." For example, mir-121 is a precursor for miR-121. When different miRNA genes or precursors process into the same mature miRNA, the genes / precursors can be described with numbered suffixes. For example, mir-121-1 and mir-121-2 can refer to distinct genes or precursors that process into miR-121. Lettered suffixes are used to indicate closely related mature sequences. For example, mir-121a and mir-121b can process into closely related miRNAs, miR-121a and miR-121b, respectively. In connection with this disclosure, any microRNA named herein with the prefix mir-* or miR-* (miRNA or miR) is understood to include both precursors and / or mature species unless otherwise specified.
[0238] Occasionally, two mature miRNA sequences are observed to originate from the same precursor. If one sequence is more abundant than the other, the less common variant can be named using the suffix "*". For example, miR-121 is the primary product, while miR-121* is the less common variant found on the opposite arm of the precursor. If the primary variant cannot be identified, miRs can be identified by the suffix "5p" for variants from the 5' arm of the precursor and "3p" for variants from the 3' arm. For example, miR-121-5p originates from the 5' arm of the precursor, while miR-121-3p originates from the 3' arm. Less commonly, the 5p and 3p variants are referred to as sense ("s") and antisense ("as") forms, respectively. For example, miR-121-5p may be referred to as miR-121-s, while miR-121-3p may be referred to as miR-121-as.
[0239] The above nomenclature rules have developed over time and are more of a general guideline than an absolute rule. For example, the let and lin families of miRNAs continue to be referred to by their nicknames. The mir / miR rules for precursor / mature forms are also guidelines, and the context should be taken into consideration when deciding which form to refer to. Further details on miR nomenclature can be found at www.mirbase.org or in Ambros et al., A uniform system for microRNA annotation, RNA 9:277-279 (2003).
[0240] Plant miRNAs follow different nomenclature rules, as described in Meyers et al., Plant Cell. 2008 20(12):3186-3190.
[0241] Several miRNAs are involved in gene regulation, and miRNAs are part of the expanding class of non-coding RNAs currently recognized as a major hierarchy of gene regulation. In some cases, miRNAs can disrupt translation by binding to regulatory sites embedded in the 3'-UTR of target mRNA, resulting in translational repression. Target recognition involves complementary base pairing between the target site and the miRNA's seed region (positions 2-8 at the 5' end of the miRNA). However, the exact degree of seed complementarity is not precisely determined and can be modified by 3' pairing. In other cases, miRNAs function like small interfering RNAs (siRNAs), binding to perfectly complementary mRNA sequences to disrupt target transcripts.
[0242] Characterization of several miRNAs has shown that they influence a variety of processes, including early development, cell proliferation and death, apoptosis, and adipogenesis. For example, some miRNAs, such as lin-4, let-7, mir-14, mir-23, and bantam, have been shown to play important roles in cell differentiation and histogenesis. Others are also thought to play similarly important roles due to their differential spatial and temporal expression patterns.
[0243] The miRNA database available at miRBase (www.mirbase.org) includes a searchable database of published miRNA sequences and annotations. Further information on miRBase can be found in the following references, each incorporated herein by reference in its entirety: Griffiths-Jones et al., miRBase: tools for microRNA genomics. NAR 2008 36(Database Issue):D154-D158; Griffiths-Jones et al., miRBase: microRNA sequences, targets and gene nomenclature. NAR 2006 34(Database Issue):D140-D144; and Griffiths-Jones, S. The microRNA Registry. NAR 2004 32(Database Issue):D109-D111. Representative miRNAs included in miRBase Release 16 became available in September 2010.
[0244] As described herein, microRNAs are known to be involved in cancer and other diseases and can be examined to characterize the phenotype in a sample. See, for example, Ferracin et al., Micromarkers: miRNAs in cancer diagnosis and prognosis, Exp Rev Mol Diag, Apr 2010, Vol. 10, No. 3, Pages 297-308; Fabbri, miRNAs as molecular biomarkers of cancer, Exp Rev Mol Diag, May 2010, Vol. 10, No. 4, Pages 435-444.
[0245] In some embodiments, molecular profiling as described herein includes the analysis of microRNAs.
[0246] Techniques for isolating and characterizing vesicles and miRs are known to those skilled in the art. In addition to the methodologies presented herein, additional methods include U.S. Patent No. 7,888,035, issued on February 15, 2011, entitled "METHODS FOR ASSESSING RNA PATTERNS"; U.S. Patent No. 7,897,356, issued on March 1, 2011, entitled "METHODS AND SYSTEMS OF USING EXOSOMES FOR DETERMINING PHENOTYPES"; International Patent Publication WO / 2011 / 066589, issued on November 30, 2010, entitled "METHODS AND SYSTEMS FOR ISOLATING, STORING, AND ANALYZING VESICLES"; WO / 2011 / 088226, issued on January 13, 2011, entitled "DETECTION OF GASTROINTESTINAL DISORDERS"; and "BIOMARKERS FOR These can be found in WO / 2011 / 109440, issued on March 1, 2011, under the title “THERANOSTICS” and in WO / 2011 / 127219, issued on April 6, 2011, under the title “CIRCULATING BIOMARKERS FOR DISEASE,” each of these applications is incorporated herein by reference in its entirety.
[0247] Circulating biomarkers Circulating biomarkers include detectable biomarkers in body fluids, such as blood, plasma, and serum. Examples of circulating cancer biomarkers include cardiac troponin T (cTnT), prostate-specific antigen (PSA) for prostate cancer, and CA125 for ovarian cancer. Circulating biomarkers according to this disclosure include any suitable detectable biomarkers in body fluids, non-limitingly including proteins, nucleic acids, such as DNA, mRNA, and microRNA, lipids, carbohydrates, and metabolites. Circulating biomarkers may include biomarkers that are not associated with cells, such as membrane-bound biomarkers, biomarkers embedded in membrane fragments, biomarkers that are part of a biological complex, or biomarkers that are free in solution. In one embodiment, the circulating biomarker is a biomarker associated with one or more vesicles present in the biological fluid of interest.
[0248] Circulating biomarkers have been identified for use in characterizing various phenotypes, such as in the detection of cancer. For example, Ahmed N, et al., Proteomic-based identification of haptoglobin-1 precursor as a novel circulating biomarker of ovarian cancer. Br. J. Cancer 2004; Mathelin _et al., Circulating proteinic biomarkers and breast cancer, Gynecol Obstet Fertil. 2006 Jul-Aug;34(7-8):638-46. Epub 2006 Jul 28; Ye et al., Recent technical strategies to identify diagnostic biomarkers for ovarian cancer. Expert Rev Proteomics. 2007 Feb;4(1):121-31; Carney, Circulating oncoproteins HER2 / neu, EGFR and CAIX (MN) as novel cancer biomarkers. Expert Rev Mol Diagn. 2007 May;7(3):309-19; Gagnon, Discovery and application of protein biomarkers for ovarian cancer, Curr Opin Obstet Gynecol. 2008 Feb;20(1):9-13; Pasterkamp et al., Immune regulatory cells: circulating biomarker factories in cardiovascular disease. Clin Sci (Lond). 2008 Aug;115(4):129-31; Fabbri, miRNAs as biomarker moleculars of cancer, Exp Rev Mol Diag, May 2010, Vol. 10, No.4, Pages 435-444; PCT Patent Publication WO / 2007 / 088537; U.S. Patent Nos. 7,745,150 and 7,655,479; U.S. Patent Application Publications 20110008808, 20100330683, 20100248290, 20100222230, 20100203566, 20100173788, 20090291932, and 2 See publications 0090239246, 20090226937, 20090111121, 20090004687, 20080261258, 20080213907, 20060003465, 20050124071, and 20040096915, each of which is incorporated herein by reference in its entirety. In some embodiments, molecular profiling as described herein includes the analysis of circulating biomarkers.
[0249] Gene expression profiling Methods and systems described herein include expression profiling, which involves examining the differential expression of one or more target genes disclosed herein. Differential expression may include overexpression and / or underexpression of biological products, e.g., genes, mRNA, or proteins, compared to a control (or baseline). The control may include cells similar to the sample but without the disease (e.g., expression profiles obtained from a sample from a healthy individual). The control may be a previously determined level indicating the efficacy of a drug target associated with a particular disease and a specific drug target. The control may be derived from the same patient, e.g., from a normal adjacent part of the same organ as the affected cell, or the control may be obtained from healthy tissue from another patient, or it may be a previously determined threshold indicating that the disease responds to or does not respond to a particular drug target. The control may also be a control found in the same sample, e.g., a housekeeping gene or its product (e.g., mRNA or protein). For example, the control nucleic acid may be one that is known not to differ depending on whether the cell is cancerous or non-cancerous. The expression levels of control nucleic acids can be used to normalize signal levels in the test population and the reference population. Exemplary control genes include, but are not limited to, β-actin, glyceraldehyde 3-phosphate dehydrogenase, and ribosomal protein P1. Multiple controls or types of controls may be used. The causes of differential expression can vary. For example, the gene copy number may increase in a cell, thereby resulting in increased gene expression. Alternatively, gene transcription may be modified by, for example, chromatin remodeling, differential methylation, differential expression, or activity of transcription factors. Translation may also be modified by, for example, differential expression of factors that degrade, translate, or silence mRNA, such as microRNA or siRNA. In some embodiments, differential expression includes differential activity. For example, a protein may harbor mutations that increase its activity, such as constitutive activation, which contribute to a disease condition.Molecular profiling, which reveals changes in activity, can be used to guide treatment selection.
[0250] Gene expression profiling methods include those based on polynucleotide hybridization analysis and those based on polynucleotide sequencing. Commonly known methods in the art for quantifying mRNA expression in a sample include Northern blotting and in-situ hybridization (Parker & Barnes (1999) Methods in Molecular Biology 106:247-283); RNase protection assays (Hod (1992) Biotechniques 13:852-854); and reverse transcription polymerase chain reaction (RT-PCR) (Weis et al. (1992) Trends in Genetics 8:263-264). Alternatively, antibodies capable of recognizing specific double helixes, including DNA double helixes, RNA double helixes, and DNA-RNA hybrid double helixes or DNA-protein double helixes, may be employed. Representative methods for sequencing-based gene expression analysis include Serial Analysis of Gene Expression (SAGE), massively parallel signature sequencing (MPSS), and / or next-generation sequencing.
[0251] RT-PCR Reverse transcription polymerase chain reaction (RT-PCR) is a variation of polymerase chain reaction (PCR). This technique involves reverse transcription of an RNA strand into its DNA complement (i.e., complementary DNA, or cDNA) using an enzyme called reverse transcriptase, and the resulting cDNA is then amplified using PCR. Real-time polymerase chain reaction is another PCR variation, also known as quantitative PCR, Q-PCR, qRT-PCR, or sometimes RT-PCR. Either reverse transcription PCR or real-time PCR may be used for molecular profiling in accordance with this disclosure, and RT-PCR may be expressed as is understood by those skilled in the art, unless otherwise specified.
[0252] RT-PCR can be used to determine RNA levels of biomarkers described herein, such as mRNA or miRNA levels. RT-PCR can be used to compare such RNA levels of biomarkers described herein in different sample populations, in normal and tumor tissues, with or without drug treatment, to characterize gene expression patterns, to identify closely related RNAs, and to analyze RNA structure.
[0253] The first step is the isolation of RNA, e.g., mRNA, from the sample. The starting material can be total RNA isolated from a human tumor or tumor cell line, and from the corresponding normal tissue or cell line, respectively. Thus, RNA can be isolated from a sample, e.g., tumor cells or tumor cell line, and compared to pooled DNA from a healthy donor. If the origin of mRNA is a primary tumor, mRNA can be extracted, for example, from a frozen tissue sample or a preserved tissue sample that has been paraffin-embedded and fixed (e.g., formalin-fixed).
[0254] General methods for mRNA extraction are well known in the art and are disclosed in standard molecular biology textbooks, including Ausubel et al. (1997) Current Protocols of Molecular Biology, John Wiley and Sons. Methods for extracting RNA from paraffin-embedded tissues are disclosed, for example, in Rupp & Locker (1987) Lab Invest. 56:A67 and De Andres et al., BioTechniques 18:42044 (1995). In particular, RNA isolation can be performed using purification kits, buffer sets, and proteases from commercial manufacturers such as Qiagen, according to the manufacturer's instructions (QIAGEN Inc., Valencia, CA). For example, total RNA can be isolated from cultured cells using Qiagen RNeasy mini-columns. Numerous RNA isolation kits are commercially available and can be used in methods such as those described herein.
[0255] Alternatively, the first step is the isolation of miRNA from the target sample. The starting material is typically total RNA isolated from human tumors or tumor cell lines and corresponding normal tissues or cell lines, respectively. Thus, RNA can be isolated from a variety of primary tumors or tumor cell lines, along with pooled DNA from healthy donors. If the origin of the miRNA is a primary tumor, the miRNA can be extracted, for example, from frozen tissue samples or preserved tissue samples embedded and fixed in paraffin (e.g., formalin).
[0256] General methods for miRNA extraction are well known in the art and are disclosed in standard molecular biology textbooks, including Ausubel et al. (1997) Current Protocols of Molecular Biology, John Wiley and Sons. Methods for extracting RNA from paraffin-embedded tissues are disclosed, for example, in Rupp & Locker (1987) Lab Invest. 56:A67 and De Andres et al., BioTechniques 18:42044 (1995). In particular, RNA isolation can be performed using purification kits, buffer sets, and proteases from commercial manufacturers such as Qiagen, according to the manufacturer's instructions. For example, total RNA can be isolated from cultured cells using Qiagen RNeasy mini-columns. Numerous miRNA isolation kits are commercially available and can be used in methods such as those described herein.
[0257] Gene expression profiling by RT-PCR, whether the RNA includes mRNA, miRNA, or other types of RNA, can involve reverse transcription of the RNA template into cDNA, followed by amplification in a PCR reaction. Commonly used reverse transcriptases include, but are not limited to, avian myeloblastosis virus reverse transcriptase (AMV-RT) and Moloney mouse leukemia virus reverse transcriptase (MMLV-RT). The reverse transcription step is typically primed using specific primers, random hexamers, or oligo-dT primers, depending on the context and goals of the expression profiling. For example, extracted RNA can be reverse transcribed using the GeneAmp RNA PCR kit (Perkin Elmer, Calif., USA) according to the manufacturer's instructions. The resulting cDNA can then be used as a template in a subsequent PCR reaction.
[0258] While the PCR process can utilize a variety of heat-stable DNA-dependent DNA polymerases, it typically employs Taq DNA polymerase, which possesses 5'-3' nuclease activity but lacks 3'-5' proofreading end nuclease activity. TaqMan PCR typically uses the 5'-nuclease activity of Taq or Tth polymerase to hydrolyze the hybridization probe bound to the target amplicon, but any enzyme with equivalent 5' nuclease activity can be used. Two oligonucleotide primers are used to generate the amplicon typical of the PCR reaction. A third oligonucleotide, or probe, is designed to detect the nucleotide sequence located between the two PCR primers. The probe cannot be extended by the Taq DNA polymerase enzyme and is labeled with a reporter fluorescent dye and a quencher fluorescent dye. If the two dyes are located very close to each other on the probe, any laser-induced emission from the reporter dye is quenched by the quencher dye. During the amplification reaction, the Taq DNA polymerase enzyme cleaves the probe in a template-dependent manner. The resulting probe fragments dissociate in solution, and the signal from the released reporter dye is freed from the quenching effect of the second fluorophore. Since one molecule of reporter dye is released for each new molecule synthesized, detecting the unquenched reporter dye provides a basis for quantitatively interpreting the data.
[0259] TaqMan® RT-PCR can be performed using commercially available instruments, such as the ABI PRISM 7700® Sequence Detection System® (Perkin-Elmer-Applied Biosystems, Foster City, Calif., USA) or the LightCycler (Roche Molecular Biochemicals, Mannheim, Germany). In a particular embodiment, the 5' nuclease procedure is performed using a real-time quantitative PCR device such as the ABI PRISM 7700 Sequence Detection System. This system consists of a thermocycler, laser, charge-coupled device (CCD), camera, and computer. The system amplifies the sample in a 96-well format using the thermocycler. During amplification, laser-induced fluorescence signals are collected in real time for all 96 wells via fiber optic cables and detected by the CCD. The system includes software for operating the instrument and analyzing the data.
[0260] TaqMan data is initially represented as a Ct, or threshold cycle. As mentioned above, fluorescence values are recorded between each cycle, representing the amount of product amplified up to that point in the amplification reaction. The threshold cycle (Ct) is the point where the fluorescence signal is first recorded as statistically significant.
[0261] To minimize errors and inter-sample variability, RT-PCR is typically performed using an internal standard. An ideal internal standard is expressed at a consistent level across different tissues and is unaffected by experimental processing. The RNAs most frequently used to normalize gene expression patterns are mRNAs for housekeeping genes, glyceraldehyde-3-phosphate dehydrogenase (GAPDH), and β-actin.
[0262] Real-time quantitative PCR (also known as quantitative real-time polymerase chain reaction, QRT-PCR, or Q-PCR) is a more recent variation of the RT-PCR technique. Q-PCR can measure the accumulation of PCR products through a double-labeled fluorescence-generating probe (i.e., a TaqMan probe). Real-time PCR is compatible with both quantitative competitive PCR, where internal competitors for each target sequence are used for normalization, and quantitative comparative PCR, which uses normalization genes contained in the sample or housekeeping genes for RT-PCR. See, for example, Held et al. (1996) Genome Research 6:986-994.
[0263] Protein-based detection techniques are also useful for molecular profiling, especially when nucleotide variants cause amino acid substitutions, deletions, insertions, or frameshifts that affect the primary, secondary, or tertiary structure of a protein. Protein sequencing techniques may be used to detect amino acid variations. For example, a protein or fragment corresponding to a gene can be synthesized by recombinant expression using a DNA fragment isolated from the test individual. Preferably, a cDNA fragment of 100-150 base pairs or less containing the polymorphic locus to be determined is used. The amino acid sequence of the peptide can then be determined by conventional protein sequencing methods. Alternatively, HPLC-microscopy tandem mass spectrometry can be used to determine amino acid sequence variations. In this technique, proteolytic digestion is performed on the protein, and the resulting peptide mixture is separated by reverse-phase chromatography. Tandem mass spectrometry is then performed, and the collected data is analyzed. See Gatlin et al., Anal. Chem., 72:757-763 (2000).
[0264] microarray Biomarkers such as those described herein can also be identified, confirmed, and / or measured using microarray techniques. Thus, expression profile biomarkers can be measured in cancer samples using microarray techniques. In this method, the polynucleotide sequence of interest is plated or arrayed on a microchip substrate. The arrayed sequence is then hybridized with a specific DNA probe from the cell or tissue of interest. The mRNA source can be total RNA isolated from a sample, e.g., a human tumor or tumor cell line and a corresponding normal tissue or cell line. Thus, RNA can be isolated from a variety of primary tumors or tumor cell lines. When the mRNA source is a primary tumor, mRNA can be extracted, for example, from frozen tissue samples or paraffin-embedded and fixed (e.g., formalin-fixed) preserved tissue samples, which are routinely prepared and stored in daily clinical practice.
[0265] Biomarker expression profiles can be measured using microarray techniques in either fresh or paraffin-embedded tumor tissue or body fluids. In this method, the polynucleotide sequence of interest is plated or arrayed on a microchip substrate. The arrayed sequence is then hybridized with a specific DNA probe from the cells or tissue of interest. Similar to RT-PCR, the miRNA source is typically total RNA isolated from human tumor or tumor cell lines and corresponding normal tissue or cell lines, including body fluids such as serum, urine, tears, and exosomes. Thus, RNA can be isolated from a variety of sources. When the miRNA source is a primary tumor, miRNA can be extracted, for example, from frozen tissue samples, which are routinely prepared and stored in daily clinical practice.
[0266] cDNA microarray techniques, also known as biochips, DNA chips, or gene arrays, enable the identification of gene expression levels in biological samples. Each cDNA or oligonucleotide representing a given gene is immobilized and tagged on a substrate, such as a small chip, bead, or nylon membrane, and serves as a probe indicating whether they are expressed in the biological sample of interest. The simultaneous expression of thousands of genes can be monitored.
[0267] In certain embodiments of microarray techniques, PCR-amplified inserts of cDNA clones are applied to a substrate in the form of a high-density array. In one plane, at least 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 1,500, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 15,000, 20,000, 25,000, 30,000, 35,000, 40,000, 45,000, or at least 50,000 nucleotide sequences are applied to the substrate. Each sequence can correspond to a different gene, or multiple sequences can be arrayed per gene. Microarrayed genes immobilized on a microchip are suitable for hybridization under stringent conditions. Fluorescently labeled cDNA probes can be generated through the incorporation of fluorescent nucleotides via reverse transcription of RNA extracted from the tissue of interest. The labeled cDNA probes, applied to a chip, specifically hybridize to each DNA spot on the array. After stringent washing to remove non-specifically bound probes, the chip is scanned by confocal laser microscopy or another detection method such as a CCD camera. Quantifying the hybridization of each arrayed element allows for the examination of the corresponding mRNA abundance. cDNA probes generated from two RNA sources and separately labeled with dichromatic fluorescence are hybridized in pairs into the array. Therefore, the relative abundance of transcripts from the two sources corresponding to each specific gene is determined simultaneously. Miniaturized-scale hybridization provides a conveniently rapid assessment of expression patterns for a large number of genes. Such methods have been shown to have the sensitivity necessary to detect rare transcripts expressed at a few copies per cell and to reproducibly detect differences in expression levels of at least approximately twofold (Schena et al. (1996) Proc. Natl. Acad. Sci. USA 93(2):106-149).Microarray analysis can be performed using commercially available instruments in accordance with the manufacturer's protocol, including but not limited to the Affymetrix GeneChip technique (Affymetrix, Santa Clara, CA), Agilent (Agilent Technologies, Inc., Santa Clara, CA), or Illumina (Illumina, Inc., San Diego, CA) microarray techniques.
[0268] The development of microarray methods for large-scale gene expression analysis will enable the systematic search for molecular markers for cancer classification and outcome prediction in diverse tumor types.
[0269] In some embodiments, the Agilent Whole Human Genome Microarray Kit (Agilent Technologies, Inc., Santa Clara, CA) is used. This system can analyze more than 41,000 unique human genes and transcripts, all of which are public domain annotated. The system is used according to the manufacturer's instructions.
[0270] In some embodiments, the Illumina Whole Genome DASL assay (Illumina Inc., San Diego, CA) is used. This system provides a method for simultaneously profiling over 24,000 transcripts in a high-throughput manner from minimal RNA inputs from both fresh-frozen (FF) and formalin-fixed paraffin-embedded (FFPE) tissue sources.
[0271] Microarray expression analysis involves identifying whether a gene or gene product is upregulated or downregulated compared to a baseline. This identification can be performed using statistical tests to determine the statistical significance of any observed differential expression. In some embodiments, statistical significance is determined using parametric statistical tests. Parametric statistical tests may include, for example, partially factorial designs, analysis of variance (ANOVA), t-tests, least squares, Pearson correlation, linear simple regression, nonlinear regression, multiple linear regression, or multiple nonlinear regression. Alternatively, parametric statistical tests may include one-way analysis of variance, two-way analysis of variance, or repeated measures analysis of variance. In other embodiments, statistical significance is determined using nonparametric statistical tests. Examples include, but are not limited to, the Wilcoxon signed-rank test, Mann-Whitney test, Kruskal-Wallis test, Friedman test, Spearman's rank correlation coefficient, Kendall's tau analysis, and nonparametric regression tests. In some embodiments, statistical significance is determined by p-values of approximately 0.05, 0.01, 0.005, 0.001, 0.0005, or less than 0.0001. Although the microarray systems used in methods such as those described herein may assay thousands of transcripts, data analysis needs to be performed only on the transcript of interest, thereby reducing the problem of multiple comparisons inherent in performing multiple statistical tests. The p-values can also be corrected for multiple comparisons using, for example, the Bonferroni correction, its variations, or other techniques known to those skilled in the art, such as the Hochberg correction, Holm-Bonferroni correction, Sidak correction, or Dunnett correction. The degree of differential expression can also be taken into consideration. For example, a gene can be considered differentially expressed if the ratio of change in expression compared to the control level differs by at least 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.2, 2.5, 2.7, 3.0, 4, 5, 6, 7, 8, 9, or 10 times in the sample compared to the control. Differential expression includes both overexpression and underexpression.A gene or gene product can be considered upregulated or downregulated if differential expression satisfies a statistical threshold, a change factor threshold, or both. For example, criteria for identifying differential expression may include both a p-value of 0.001 and a change factor of at least 1.5 times (up or down). Those skilled in the art will understand that differential expression can be determined by any molecular profiling technique disclosed herein, by adapting such statistical and threshold measures.
[0272] Various methods described herein utilize many types of microarrays to detect the presence and potential amount of biological entities in a sample. Arrays typically contain addressable portions in which the presence of entities in the sample can be detected, for example, by binding events. Microarrays include, but are not limited to, DNA microarrays, e.g., cDNA microarrays, oligonucleotide microarrays and SNP microarrays, microRNA arrays, protein microarrays, antibody microarrays, tissue microarrays, cell microarrays (also called transfection microarrays), chemical compound microarrays, and carbohydrate arrays (glycoarrays). DNA arrays typically contain addressable nucleotide sequences that can bind to sequences present in the sample. MicroRNA arrays, for example, can detect microRNAs using MMChips arrays from the University of Louisville or commercially available systems from Agilent. Protein microarrays can be used to identify protein-protein interactions, including, but not limited to, identifying substrates for protein kinases, transcription factor protein activations, or targets for biologically active small molecules. Protein arrays may contain arrays of nucleotide sequences that bind to different protein molecules, commonly antibodies, or proteins of interest. Antibody microarrays contain antibodies spotted on a protein chip, which are used as capture molecules to detect proteins or other biomolecules from a sample, such as a cell or tissue lysate. For example, antibody arrays can be used to detect biomarkers from body fluids, such as serum or urine, for diagnostic applications. Tissue microarrays contain separate tissue cores assembled in an array-like manner to enable multiplex tissue analysis. Cell microarrays, also called transfection microarrays, contain various capture agents, such as antibodies, proteins, or lipids, that interact with cells and facilitate capture at a location that can be referenced by address.Chemical compound microarrays comprise an array of chemical compounds that can be used to detect proteins or other biomolecules that bind to these compounds. Glycoarrays comprise an array of carbohydrates that can, for example, detect proteins that bind to sugar moieties. Those skilled in the art will recognize that similar techniques or improvements can be used in accordance with methods such as those described herein.
[0273] Certain embodiments of this method include, but are not limited to, multiwell reaction vessels, including multiwell plates or multichamber microfluidic devices in which multiple amplification reactions and, in some embodiments, detection are typically carried out in parallel. In certain embodiments, one or more multiplex reactions for generating amplicons are carried out in the same reaction vessel, including, but not limited to, multiwell plates such as 96-well, 384-well, or 1536-well plates; or microfluidic devices, such as, but not limited to, TaqMan® low-density arrays (Applied Biosystems, Foster City, CA). In some embodiments, the ultra-parallel amplification steps include, but are not limited to, multiwell reaction vessels containing plates with multiple reaction wells, such as, but not limited to, 24-well plates, 96-well plates, 384-well plates; or multichamber microfluidic devices, such as, but not limited to, low-density arrays, where each chamber or well contains, as appropriate, suitable primers, primer sets, and / or reporter probes. Typically, such amplification steps occur in a series of parallel singleplex, 2-plex, 3-plex, 4-plex, 5-plex, or 6-plex reactions, but higher levels of parallel multiplexing are also within the scope of this instruction. These methods may include PCR methodologies in each well or chamber for amplifying and / or detecting nucleic acid molecules of interest, such as RT-PCR.
[0274] Low-density arrays can include arrays that detect tens or hundreds of molecules, as opposed to thousands of molecules. These arrays can be more sensitive than high-density arrays. In one embodiment, a low-density array, such as a TaqMan® low-density array, is used to detect one or more genes or gene products in any of Tables 5-12 of WO2018175501. For example, a low-density array can be used to detect at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, or 100 genes or gene products selected from any of Tables 5-12 of WO2018175501.
[0275] In some embodiments, the disclosed method includes a microfluidic device, a "lab on a chip," or a micrototal analysis system (pTAS). In some embodiments, sample preparation is performed using a microfluidic device. In some embodiments, amplification reactions are performed using a microfluidic device. In some embodiments, sequencing or PCR reactions are performed using a microfluidic device. In some embodiments, at least a portion of the nucleotide sequence of the amplification product is obtained using a microfluidic device. In some embodiments, the detection step includes, but is not limited to, a microfluidic device including a low-density array such as a TaqMan® low-density array. Descriptions of exemplary microfluidic devices can be found, among other things, in published PCT applications WO / 0185341 and WO04 / 011666; Kartalov and Quake, Nucl. Acids Res. 32:2873-79, 2004; and Fiorini and Chiu, Bio Techniques 38:429-46, 2005.
[0276] Any suitable microfluidic device can be used in the manner described herein. Examples of microfluidic devices used in or adapted for use with molecular profiling include U.S. Patents No. 7,591,936, 7,581,429, 7,579,136, 7,575,722, 7,568,399, 7,552,741, 7,544,506, 7,541,578, 7,518,726, 7,488,596, 7,485,214, and 7,467,928. No. 7,452,713, No. 7,452,509, No. 7,449,096, No. 7,431,887, No. 7,422,725, No. 7,422,669, No. 7,419,822, No. 7,419,639, 7,413,709, 7,411,184, 7,402,229, 7,390,463, 7,381,471, 7,357,864, 7,351,592, 7,351,380, No. 7,338,637, No. 7,329,391, No. 7,323,140, No. 7,261,824, No. 7,258,837, No. 7,253,003, No. 7,238,324, No. 7,238,255, No. 7 ,233,865, 7,229,538, 7,201,881, 7,195,986, 7,189,581, 7,189,580, 7,189,368, 7,141,978, 7,1 This includes, but is not limited to, patents or applications described in U.S. Patent Publication No. 38,062, No. 7,135,147, No. 7,125,711, No. 7,118,910, No. 7,118,661, No. 7,640,947, No. 7,666,361, and No. 7,704,735; U.S. Patent Application Publication No. 20060035243; and International Patent Publication WO2010 / 072410, each of these patents or applications is incorporated herein by reference in whole.Another example of use with the methods disclosed herein is described in Chen et al., "Microfluidic isolation and transcriptome analysis of serum vesicles," Lab on a Chip, Dec. 8, 2009 DOI: 10.1039 / b916199f.
[0277] Gene expression analysis using large-scale parallel signature sequencing (MPSS) The method described by Brenner et al. (2000) Nature Biotechnology 18:630-634 is a sequencing technique that combines non-gel-based signature sequencing with in vitro cloning of millions of templates on separate microbeads. First, a microbead library of DNA templates is constructed by in vitro cloning. Subsequently, a high-density planar array of template-containing microbeads is assembled in a flow cell. The free ends of the cloned templates on each microbead are simultaneously analyzed using a fluorescence-based signature sequencing method that does not require separation of DNA fragments. This method has been shown to provide hundreds of thousands of gene signature sequences from a cDNA library simultaneously and accurately in a single operation.
[0278] MPSS data has many applications. The expression levels of almost all transcripts can be quantified; the abundance of signatures represents the expression level of genes in the analyzed tissue. Quantitative methods for analyzing tag frequencies and detecting differences between libraries are published and incorporated into the public database for SAGE™ data, making them applicable to MPSS data. The availability of complete genome sequences allows for direct comparison between signatures and genome sequences, further expanding the usefulness of MPSS data. Because targets for MPSS analysis are not pre-selected (unlike microarrays), MPSS data can characterize the complete complexity of the transcriptome. This is analogous to sequencing millions of ESTs at once, and genome sequence data can be used so that MPSS signature sources can be easily identified by computational means.
[0279] Sequential gene expression analysis (SAGE) Sequential gene expression analysis (SAGE) is a method that enables the simultaneous quantitative analysis of multiple gene transcripts without the need to provide individual hybridization probes for each transcript. First, short sequence tags (e.g., approximately 10-14 bp) containing enough information to uniquely identify a transcript are generated, provided that the tags are obtained from a unique location within each transcript. Then, multiple transcripts are ligated together to form long sequential molecules, which can be sequenced to simultaneously reveal the identity of multiple tags. The expression patterns of any population of transcripts can be quantitatively evaluated by determining the abundance of individual tags and identifying the genes corresponding to each tag. See, for example, Velculescu et al. (1995) Science 270:484-487; and Velculescu et al. (1997) Cell 88:243-51.
[0280] DNA copy number profiling Any method capable of determining the DNA copy number profile of a particular sample can be used for molecular profiling according to the methods described herein, provided that the resolution is sufficient to identify copy number variations in biomarkers as described herein. Those skilled in the art will recognize that several different platforms can be used to examine whole-genome copy number variations with sufficient resolution to identify the copy number of one or more biomarkers as described herein, and such platforms can be used. Some of the platforms and techniques are described in the embodiments below. In some embodiments as described herein, next-generation sequencing or ISH techniques, as described herein or known in the art, are used to determine copy number / gene amplification.
[0281] In some embodiments, copy number profiling analysis involves amplification of whole-genome DNA using whole-genome amplification methods. Whole-genome amplification methods can utilize strand substitution polymerases and random primers.
[0282] In some aspects of these embodiments, copy number profiling analysis involves hybridization of whole-genome amplified DNA using a high-density array. In more specific aspects, the high-density array has 5,000 or more different probes. In another specific aspect, the high-density array has 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, or 1,000,000 or more different probes. In yet another specific aspect, each of the different probes on the array is an oligonucleotide with a length of approximately 15–200 base pairs. In another specific context, each of the different probes on the array is an oligonucleotide having a length of approximately 15–200, 15–150, 15–100, 15–75, 15–60, or 20–55 base pairs.
[0283] In some embodiments, microarrays are employed to help determine copy number profiles for a sample, e.g., cells from a tumor. A microarray typically contains multiple oligomers (e.g., DNA or RNA polynucleotides or oligonucleotides, or other polymers) synthesized or deposited in an array pattern on a substrate (e.g., a glass support). The oligomers bound to the support are “probes” that function to hybridize or bind to the sample material (e.g., nucleic acids prepared or obtained from a tumor sample) in a hybridization experiment. The reverse situation can also be applied, where the sample is bound to the microarray substrate and the oligomer probes are in a solution for hybridization. In use, the array surface is contacted with one or more targets under conditions that promote specific high-affinity binding of the targets to one or more probes. In some configurations, the hybridized sample and probes are detectable using a scanning instrument by labeling the sample nucleic acid with a detectable label such as a fluorescent tag. DNA array techniques offer the potential to analyze DNA copy number profiles using a large number (e.g., hundreds of thousands) of different oligonucleotides. In some embodiments, the substrate used for the array is a surface-derivatized glass or silica, or a polymer membrane surface (see, for example, Z. Guo, et al., Nucleic Acids Res, 22, 5456-65 (1994); U. Maskos, EM Southern, Nucleic Acids Res, 20, 1679-84 (1992), and EM Southern, et al., Nucleic Acids Res, 22, 1368-73 (1994), each of which is incorporated herein by reference). Modification of the array substrate surface can be achieved by many techniques.For example, a silicic acid-containing surface or a metal oxide surface can be derivatized with a bifunctional silane, i.e., a silane having a first functional group that enables covalent bonding with the surface (e.g., a Si-halogen or Si-alkoxy group, such as those found in --SiCl3 or --Si(OCH3)3, respectively) and a second functional group that can provide the surface with desired chemical and / or physical modifications, thereby covalently or acovalently linking ligands and / or polymers or monomers for biological probe arrays. Silylation derivatization and other surface derivatizations (see, for example, U.S. Patent No. 5,624,711 to Sundberg, U.S. Patent No. 5,266,222 to Willis, and U.S. Patent No. 5,137,765 to Farnsworth, each of which is incorporated herein by reference) are known in the art. Other processes for preparing arrays are described in U.S. Patent No. 6,649,348 to Bass et al., assigned to Agilent Corp., which discloses DNA arrays produced by in-situ synthesis.
[0284] Polymer array synthesis is also widely described in the literature, including: International Publication No. 00 / 58516, U.S. Publication Nos. 5,143,854, 5,242,974, 5,252,743, 5,324,633, 5,384,261, 5,405,783, 5,424,186, 5,451,683, 5,482,867, and 5,491. ,074, 5,527,681, 5,550,215, 5,571,639, 5,578,832, 5,593,839, 5,599,695, 5, 624,711, 5,631,734, 5,795,716, 5,831,070, 5,837,832, 5,856,101, 5,858,659, Nos. 5,936,324, 5,968,740, 5,974,164, 5,981,185, 5,981,956, 6,025,601, 6,033,860, 6,040,193, 6,090,555, 6,136,269, 6,269,846 and 6,428,752, 5,412,087, 6,14 Applications Nos. 7,205, 6,262,216, 6,310,189, 5,889,165, and 5,959,098, PCT application numbers PCT / US99 / 00730 (International Publication No. WO99 / 36760) and PCT / US01 / 04285 (International Publication No. WO01 / 58593), all of which are incorporated herein by reference in their entirety for all purposes.
[0285] Nucleic acid arrays useful in this disclosure include, but are not limited to, those commercially available from Affymetrix (Santa Clara, Calif.) under the trade name GeneChip®. Examples of arrays are shown on the affymetrix.com website. Another microarray supplier is Illumina, Inc., San Diego, Calif., and examples of arrays are shown on the illumina.com website.
[0286] In some embodiments, the methods of the present invention provide sample preparation. Depending on the microarray and the experiment to be performed, the sample nucleic acid can be prepared in several ways by methods known to those skilled in the art. In some aspects, such as those described herein, the sample may be amplified by several mechanisms before or simultaneously with genotyping (analysis of copy number profiles). The most common amplification procedure used involves PCR. For example, see PCR Technology: Principles and Applications for DNA Amplification (Ed. HA Erlich, Freeman Press, NY, NY, 1992); PCR Protocols: A Guide to Methods and Applications (Eds. Innis, et al., Academic Press, San Diego, Calif., 1990); Mattila et al., Nucleic Acids Res. 19, 4967 (1991); Eckert et al., PCR Methods and Applications 1, 17 (1991); PCR (Eds. McPherson et al., IRL Press, Oxford); and U.S. Patents Nos. 4,683,202, 4,683,195, 4,800,159, 4,965,188, and 5,333,675, each of which is incorporated herein by reference in whole for all purposes. In some embodiments, the sample may be amplified on an array (e.g., U.S. Patent No. 6,300,070, incorporated herein by reference).
[0287] Other suitable amplification methods include ligase chain reaction (LCR) (e.g., Wu and Wallace, Genomics 4, 560 (1989), Landegren et al., Science 241, 1077 (1988), and Barringer et al. Gene 89:117 (1990)), transcriptional amplification (Kwoh et al., Proc. Natl. Acad. Sci. USA 86, 1173 (1989), and WO88 / 10315), and autologous persistent sequence replication (Guatelli et al., Proc. Nat. Acad. Sci. USA, 87, 1874). These include (1990) and WO90 / 06995), selective amplification of targeted polynucleotide sequences (U.S. Patent No. 6,410,276), consensus sequence-primed polymerase chain reaction (CP-PCR) (U.S. Patent No. 4,437,975), arbitrary primed polymerase chain reaction (AP-PCR) (U.S. Patents No. 5,413,909, No. 5,861,245), and nucleic acid-based sequence amplification methods (NABSA) (see U.S. Patents No. 5,409,818, No. 5,554,517, and No. 6,063,603, each of which is incorporated herein by reference). Other amplification methods that may be used are described in U.S. Patent Nos. 5,242,794, 5,494,810, 4,988,617 and U.S. Patent Application No. 09 / 854,317, each of which is incorporated herein by reference.
[0288] Additional methods for sample preparation and techniques for reducing the complexity of nucleic acid samples are described in Dong et al., Genome Research 11, 1418 (2001), U.S. Patent Nos. 6,361,947 and 6,391,592, and U.S. Patent Applications Nos. 09 / 916,135, 09 / 920,491 (Publication 20030096235), 09 / 910,292 (Publication 20030082543), and 10 / 013,598.
[0289] Methods for performing polynucleotide hybridization assays are well-developed in the art. The procedures and conditions for hybridization assays used in methods such as those described herein vary depending on the application and are selected according to known general binding methods, including those mentioned in Maniatis et al. Molecular Cloning: A Laboratory Manual (2.sup.nd Ed. Cold Spring Harbor, NY, 1989); Berger and Kimmel Methods in Enzymology, Vol. 152, Guide to Molecular Cloning Techniques (Academic Press, Inc., San Diego, Calif., 1987); and Young and Davism, PNAS, 80: 1194 (1983). Methods and apparatus for carrying out repeated and controlled hybridization reactions are described in U.S. Patents No. 5,871,928, No. 5,874,219, No. 6,045,996, and No. 6,386,749 and No. 6,391,623, each of which is incorporated herein by reference.
[0290] Methods described herein may also involve signal detection of ligand-to-ligand hybridization after (and / or during) hybridization. See U.S. Patent Nos. 5,143,854, 5,578,832; 5,631,734; 5,834,758; 5,936,324; 5,981,956; 6,025,601; 6,141,096; 6,185,030; 6,201,639; 6,218,803; and 6,225,625, U.S. Patent Application No. 10 / 389,194, and PCT Application PCT / US99 / 06097 (published as WO99 / 47964), each of which is also incorporated herein by reference in whole for all purposes.
[0291] Methods and apparatus for signal detection and processing of intensity data are, for example, U.S. Patent Nos. 5,143,854, 5,547,839, 5,578,832, 5,631,734, 5,800,992, 5,834,758; 5,856,092, 5,902,723, 5,936,324, 5,981,956, 6,025,601, and 6,090,555. These are disclosed in U.S. Patent Applications Nos. 6,141,096, 6,185,030, 6,201,639; 6,218,803; and 6,225,625, U.S. Patent Applications Nos. 10 / 389,194, 60 / 493,495 and PCT Application PCT / US99 / 06097 (published as WO99 / 47964), each of which is also incorporated herein by reference in whole for all purposes.
[0292] Immunoassays Protein-based molecular profiling techniques include immunoaffinity assays based on antibodies selectively immunoreactive to proteins encoded by mutant genes according to this method. These techniques include, but are not limited to, immunoprecipitation, Western blotting, molecular binding assays, enzyme-linked immunosorbent assays (ELISA), enzyme-linked immunofiltration assays (ELIFA), and fluorescence-activated cell sorting (FACS). For example, any method for detecting the expression of a biomarker in a sample includes contacting the sample with an antibody against the biomarker, an immunoreactive fragment of that antibody, or a recombinant protein containing the antigen-binding domain of an antibody against the biomarker; and then detecting the binding of the biomarker to the sample. Methods for producing such antibodies are known in the art. Antibodies can be used to immunoprecipitate specific proteins from solution samples, or to immunoblot proteins separated, for example, by a polyacrylamide gel. Immunocytochemistry can also be used to detect specific protein polymorphisms in tissues or cells. Other well-known antibody-based techniques, including ELISA, radioimmunoassay (RIA), immunoradioquantification assay (IRMA), and immunoenzyme assay (IEMA), which include sandwich assays using monoclonal or polyclonal antibodies, can also be used. See, for example, U.S. Patent Nos. 4,376,110 and 4,486,530, both of which are incorporated herein by reference.
[0293] Alternatively, a sample may be contacted with an antibody specific to the biomarker under conditions sufficient for the formation of an antibody-biomarker complex, and then the complex may be detected. The presence of a biomarker may be detected in several ways, for example, by Western blotting and ELISA procedures for assaying a wide variety of tissues and samples, including plasma or serum. A wide range of immunoassay techniques using such assay formats are available. See, for example, U.S. Patents 4,016,043, 4,424,279, and 4,018,653. These include not only conventional competitive binding assays but also both non-competitive single-site and two-site or "sandwich" assays. These assays also include direct binding of labeled antibodies to the target biomarker.
[0294] Several variations of the sandwich assay technique exist, all of which are intended to be encompassed by this method. Briefly, in a typical forward assay, an unlabeled antibody is immobilized on a solid substrate, and the test sample is brought into contact with the bound molecule. After a suitable incubation period sufficient for antibody-antigen complex formation, a second antibody specific to the antigen, labeled with a reporter molecule capable of producing a detectable signal, is then added and incubated, allowing sufficient time for the formation of another antibody-antigen-labeled antibody complex. Any unreacted material is washed away, and the presence of the antigen is determined by observing the signal produced by the reporter molecule. The results may be qualitative, by simple observation of a visible signal, or quantitative, by comparison with a control sample containing a known amount of the biomarker.
[0295] Modifications of forward assays include simultaneous assays in which both the sample and the labeled antibody are added to the conjugated antibody at the same time. These techniques, including any minor modifications that would readily become apparent, are well known to those skilled in the art. In a typical forward sandwich assay, a first antibody specific to a biomarker is conjugated to a solid surface either covalently or passively. The solid surface is typically glass or a polymer, with the most commonly used polymers being cellulose, polyacrylamide, nylon, polystyrene, polyvinyl chloride, or polypropylene. The solid support can be in the form of a tube, beads, a microplate disk, or any other surface suitable for performing an immunoassay. The conjugation process is well known in the art and generally consists of crosslinking, covalent bonding, or physical adsorption steps, and the polymer-antibody complex is washed during the preparation of the test sample. Next, aliquots of the test sample are added to the solid-phase complex and incubated under appropriate conditions (e.g., room temperature to 40°C, e.g., between 25°C and 32°C (including both extreme values)) for a sufficient period (e.g., 2 to 40 minutes or, if more convenient, overnight) to bind any subunits present in the antibody. Following the incubation period, the antibody subunit solid phase is washed, dried, and incubated with a second antibody specific to a portion of the biomarker. The second antibody is then ligated to a reporter molecule used to indicate the binding of the second antibody to a molecular marker.
[0296] The alternative method involves immobilizing a target biomarker in a sample and then exposing the immobilized target to a specific antibody, which may or may not be labeled with a reporter molecule. Depending on the amount of target and the signal intensity of the reporter molecule, the bound target may be detectable by direct labeling with the antibody. Alternatively, a second labeled antibody specific to the first antibody is exposed to the target-first antibody complex to form a ternary complex of target-first antibody-second antibody. This complex is detected by a signal emitted by the reporter molecule. As used herein, “reporter molecule” means a molecule that, by its chemical properties, provides an analytically identifiable signal that makes an antigen-bound antibody detectable. The reporter molecules most commonly used in this type of assay are enzymes, fluorophores, or radionuclide-containing molecules (i.e., radioisotopes), and chemiluminescent molecules.
[0297] In enzyme immunoassays, the enzyme is typically conjugated to a second antibody with glutaraldehyde or periodate. However, as readily recognizable, a wide variety of different conjugation techniques are readily available to those skilled in the art. Commonly used enzymes include, among others, horseradish peroxidase, glucose oxidase, β-galactosidase, and alkaline phosphatase. The substrate to be used with the specific enzyme is generally selected for the production of a detectable color change upon hydrolysis by the corresponding enzyme. Examples of suitable enzymes include alkaline phosphatase and peroxidase. It is also possible to employ a fluorescence-generating substrate that produces a fluorescent product instead of the chromogenic substrates mentioned above. In any case, the enzyme-labeled antibody is added to and conjugated to the first antibody-molecular marker complex, and then any excess reagent is washed away. A solution containing the appropriate substrate is then added to the antibody-antigen-antibody complex. The substrate reacts with an enzyme linked to a second antibody, yielding a qualitative visible signal, which may then be further quantified spectrophotometrically to indicate the amount of biomarker present in the sample. Alternatively, fluorescent compounds such as fluorescein and rhodamine may be chemically coupled to the antibody without altering its binding ability. When activated by illumination with light of a specific wavelength, the fluorescently labeled antibody absorbs light energy, inducing an excited state in the molecule, and subsequently emitting a characteristic color of light that is visiblely detectable under a light microscope. Similar to EIA, the fluorescently labeled antibody is bound to the first antibody-molecular marker complex. After washing away any unbound reagents, the remaining ternary complex is exposed to light of an appropriate wavelength, and the observed fluorescence indicates the presence of the molecular marker of interest. Both immunofluorescence and EIA techniques are very well established in the art. However, other reporter molecules such as radioisotopes, chemiluminescent or bioluminescent molecules may also be employed.
[0298] Immunohistochemistry (IHC) IHC is a process of locating antigens (e.g., proteins) within tissue cells using antibodies that specifically bind to antigens in the tissue. Antigen-binding antibodies can be conjugated or fused to tags that enable their detection, for example, by visualization. In some embodiments, the tag is an enzyme, such as alkaline phosphatase or horseradish peroxidase, that can catalyze a color reaction. The enzyme can fuse to the antibody or be non-covalently bonded, for example, using a biotin-avidin system. Alternatively, antibodies can be tagged with fluorophores such as fluorescein, rhodamine, DyLight Fluor, or Alexa Fluor. Antigen-binding antibodies can be directly tagged, or a detection antibody possessing the tag can recognize the antigen-binding antibody itself. Using IHC, one or more proteins may be detected. The expression of gene products may be related to their staining intensity compared to control levels. In some embodiments, a gene product is considered to be differentially expressed if its staining differs in the sample from that of a control by at least 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.2, 2.5, 2.7, 3.0, 4, 5, 6, 7, 8, 9, or 10 times.
[0299] IHC involves the application of antigen-antibody interactions to histochemical techniques. In an illustrative example, tissue sections are mounted on slides and incubated with antigen-specific antibodies (polyclonal or monoclonal) (primary reaction). The antigen-antibody signal is then amplified using a second antibody conjugated to a peroxidase anti-peroxidase (PAP), avidin-biotin-peroxidase (ABC), or avidin-biotin alkaline phosphatase complex. In the presence of substrates and chromogens, the enzymes form colored deposits at the antibody-antigen binding site. Immunofluorescence is an alternative technique for visualizing antigens. In this technique, the primary antigen-antibody signal is amplified using a second antibody conjugated to a fluorescent dye. When UV light is absorbed, the fluorescent dye itself emits longer wavelength light (fluorescence), thus allowing the location of the antibody-antigen complex to be identified.
[0300] Epigenetic state The molecular profiling methods described herein also include the step of measuring epigenetic changes, i.e., gene modifications caused by epigenetic mechanisms, such as changes in methylation status or histone acetylation. Frequently, epigenetic changes result in changes in gene expression levels, which can be detected (at the RNA or protein level, as appropriate) as an indicator of the epigenetic change. Often, epigenetic changes result in gene silencing or downregulation, which is referred to as “epigenetic silencing.” The epigenetic changes most frequently investigated by methods such as those described herein involve determining the DNA methylation status of a gene, where increased methylation levels are typically associated with related cancers (as this may cause downregulation of gene expression). Abnormal methylation of one or more genes, which may be referred to as hypermethylation, can be detected. Typically, the methylation status is determined in appropriate CpG islands, often found in the promoter region of the gene. The terms “methylation,” “methylation status,” or “methylation state” may refer to the presence or absence of 5-methylcytosine at one or more CpG dinucleotides in a DNA sequence. CpG dinucleotides are typically enriched in the promoter regions and exons of human genes.
[0301] If reduced gene expression is determined by the gene's methylation status, it can be examined by DNA methylation status or by expression level. One method for detecting epigenetic silencing is to determine that genes expressed in normal cells are expressed less or not expressed at all in tumor cells. Accordingly, this disclosure provides a molecular profiling method that includes the step of detecting epigenetic silencing.
[0302] Various assay procedures for directly detecting methylation are known in the art and can be used in conjunction with this method. These assays rely on two distinct methods: bisulfite conversion-based methods and bisulfite-free methods. The bisulfite-free DNA methylation analysis method relies on the fact that methylation-sensitive enzymes cannot cleave methylated cytosines at their restriction sites. Bisulfite conversion relies on treating DNA samples with sodium bisulfite, which converts unmethylated cytosines to uracil while preserving methylated cytosines (Furuichi Y, Wataya Y, Hayatsu H, Ukita T. Biochem Biophys Res Commun. 1970 Dec 9;41(5):1185-91). This conversion results in changes within the original DNA sequence.Methods for detecting such changes include MS AP-PCR (methylation-sensitive optional primed polymerase chain reaction), a technique that uses CG-rich primers to enable a global scan of the genome to focus on the regions most likely to contain CpG dinucleotides, as described by Gonzalgo et al., Cancer Research 57:594-599, 1997; Eads et al., Cancer Res. 59:2302-2306, MethyLight® refers to a fluorescence-based real-time PCR technique approved in the art as described in 1999; the HeavyMethyl® assay, in an embodiment performed herein, is an assay in which a methylation-specific blocking probe (also referred to herein as a blocker) covering the CpG positions between amplification primers or covered by the amplification primers enables methylation-specific selective amplification of nucleic acid samples; HeavyMethyl®MethyLight® is a variation of the MethyLight® assay in which the MethyLight® assay is combined with a methylation-specific blocking probe covering the CpG positions between amplification primers; Ms-SNuPE (methylation-sensitive single-nucleotide primer extension) is an assay described by Gonzalgo & Jones, Nucleic Acids Res. 25:2529-2531, 1997; Herman et al. Proc. Natl. Acad. Sci. USA 93:9821-9826, This includes MSP (methylation-specific PCR), a methylation assay described in 1996 and U.S. Patent No. 5,786,146; COBRA (compound bisulfite restriction analysis), a methylation assay described by Xiong & Laird, Nucleic Acids Res. 25:2532-2534, 1997; and MCA (methylated CpG island amplification), a methylation assay described by Toyota et al., Cancer Res. 59:2307-12, 1999 and WO00 / 26401A1.
[0303] Other techniques for DNA methylation analysis include sequencing, methylation-specific PCR (MS-PCR), melting curve methylation-specific PCR (McMS-PCR), MLPA with or without bisulfite treatment, QAMA, MSRE-PCR, MethyLight, ConLight-MSP, bisulfite conversion-specific methylation-specific PCR (BS-MSP), COBRA (which relies on using restriction enzymes to reveal methylation-dependent sequence differences in PCR products of sodium bisulfite-treated DNA), and methylation-sensitive monomethylation. These include basal primer extension higher-order structure (MS-SNuPE), methylation-sensitive single-chain higher-order structure analysis (MS-SSCA), melting curve composite bisulfite restriction analysis (McCOBRA), PyroMethA, HeavyMethyl, MALDI-TOF, MassARRAY, methylated allele quantitative analysis (QAMA), enzyme region methylation assay (ERMA), QBSUPT, MethylQuant, quantitative PCR sequencing and oligonucleotide-based microarray systems, pyrosequencing, and Meth-DOP-PCR. Reviews of several useful techniques are provided in Nucleic Acids Research, 1998, Vol. 26, No. 10, 2255-2264; Nature Reviews, 2003, Vol. 3, 253-266; and Oral Oncology, 2006, Vol. 42, 5-13, and these references are incorporated herein in their entirety. Any of these techniques may be used in accordance with the present method as appropriate. Other techniques are described in U.S. Patent Publication Nos. 20100144836 and 20100184027, which are incorporated herein by reference in their entirety.
[0304] The DNA-binding function of histone proteins is tightly regulated through the activity of various acetylases and deacetylases. Furthermore, histone acetylation and histone deacetylation are associated with malignant progression. See Nature, 429: 457-63, 2004. Methods for analyzing histone acetylation are described in U.S. Patent Applications Publications 20100144543 and 20100151468, which are incorporated herein by reference in their entirety.
[0305] Sequence analysis Molecular profiling as disclosed herein includes methods for genotyping one or more biomarkers by determining whether an individual has one or more nucleotide variants (or amino acid variants) in one or more genes or gene products. Genotyping one or more genes in accordance with methods such as those described herein can, in some embodiments, provide more evidence for selecting a treatment.
[0306] Biomarkers such as those described herein can be analyzed by any method useful for determining alterations in the nucleic acids or proteins they encode. In one embodiment, a person skilled in the art can analyze one or more genes for mutations including deletion mutants, insertion mutants, frameshift mutants, nonsense mutants, missense mutants, and splice mutants.
[0307] Nucleic acids used for the analysis of one or more genes can be isolated from cells in a sample according to standard methodologies (Sambrook et al., 1989). For example, nucleic acids may be genomic DNA or fractional or whole cellular RNA, or miRNAs acquired from exosomes or the cell surface. When RNA is used, it may be desirable to convert the RNA to complementary DNA. In one embodiment, the RNA is whole cellular RNA; in another embodiment, it is poly-A RNA; in yet another embodiment, it is exosomal RNA. Typically, nucleic acids are amplified. Depending on the assay format for analyzing one or more genes, the specific nucleic acid of interest is identified from the sample directly using amplification, or after amplification using a second known nucleic acid. The identified product is then detected. In certain applications, detection may be performed by visual means (e.g., ethidium bromide staining of a gel). Alternatively, detection may involve indirect identification of the product via chemiluminescence, radioactive or fluorescently labeled radioscintigraphy, or even via systems using electrical or thermal impulse signals (Affymax Technology; Bellus, 1994).
[0308] Various types of deletions are known to occur in biomarkers such as those described herein. These changes include, but are not limited to, deletions, insertions, point mutations, and duplications. Point mutations can be silent or may result in stop codons, frameshift mutations, or amino acid substitutions. Mutations may occur within and outside the coding regions of one or more genes and can be analyzed according to methods such as those described herein. Target sites of nucleic acids of interest may include regions where the sequence is altered. Examples include, but are not limited to, polymorphisms that exist in different forms, such as mononucleotide mutations, nucleotide repeats, polynucleotide deletions (deletion of more than one nucleotide from the consensus sequence), polynucleotide insertions (insertion of more than one nucleotide from the consensus sequence), microsatellite repeats (a small number of nucleotide repeats with a typical 5-1000 repeat units), dinucleotide repeats, trinucleotide repeats, sequence rearrangements (including translocations and duplications), and chimeric sequences (fusion of two sequences from different gene origins). Among sequence polymorphisms, the most frequent polymorphism in the human genome is single nucleotide polymorphism (SNP), also known as a single nucleotide variant. SNPs are abundant, stable, and widely distributed throughout the genome.
[0309] Molecular profiling includes methods for haplotyping one or more genes. A haplotype is a set of genetic determinants located on a single chromosome, typically containing a specific combination of alleles (all selective sequences of a gene) within a region of the chromosome. In other words, a haplotype is phased sequence information on an individual chromosome. Very often, phased SNPs on a chromosome define a haplotype. The combination of haplotypes on a chromosome can determine the genetic profile of a cell. It is haplotypes that determine the association between a particular genetic marker and a disease mutation. Haplotyping can be performed by any method known in the art. Common methods for scoring SNPs include hybridization microarrays or direct gel sequencing, as reviewed in Landgren et al., Genome Research, 8:769-776, 1998. For example, only one copy of one or more genes can be isolated from an individual, and the nucleotides at each variant location can be determined. Alternatively, allele-specific PCR or similar methods can be used to amplify only one copy of one or more genes in an individual, thereby determining the SNP at the variant site of this disclosure. Clarke's method, known in the art, can also be employed for haplotyping. High-throughput molecular haplotyping methods are also disclosed in Tost et al., Nucleic Acids Res., 30(19):e96 (2002), which is incorporated herein by reference.
[0310] Therefore, as will be apparent to those skilled in the art of genetics and haplotyping, the variants of the present disclosure and / or additional variants linked to the haplotype can be identified by haplotyping methods known in the art. The variants of the present disclosure or additional variants linked to the haplotype can also be useful for a variety of applications, such as those listed below.
[0311] Both genomic DNA and mRNA / cDNA can be used for genotyping and haplotyping, and both are collectively referred to as "genes" in this specification.
[0312] Numerous techniques for detecting nucleotide variants are known in the art and all can be used for the methods of this disclosure. These techniques can be protein-based or nucleic acid-based. In either case, the technique used must be sufficiently sensitive to accurately detect small nucleotide or amino acid variations. Probes labeled with detectable markers are frequently used. Unless otherwise specified, any suitable marker known in the art can be used in the specific techniques described below, including but not limited to radioisotopes, fluorescent compounds, biotin detectable using streptavidin, enzymes (e.g., alkaline phosphatase), enzyme substrates, ligands, and antibodies. See Jablonski et al., Nucleic Acids Res., 14:6115-6128 (1986); Nguyen et al., Biotechniques, 13:116-123 (1992); Rigby et al., J. Mol. Biol., 113:237-251 (1977).
[0313] Nucleic acid-based detection methods require obtaining a target DNA sample, i.e., a sample containing genomic DNA, cDNA, mRNA, and / or miRNA corresponding to one or more genes, from the test individual. Any tissue or cell sample containing genomic DNA, miRNA, mRNA, and / or cDNA (or a portion thereof) corresponding to one or more genes can be used. For this purpose, tissue samples containing cell nuclei and therefore genomic DNA can be obtained from the individual. Blood samples can also be useful, except that red blood cells do not have a nucleus and contain only mRNA or miRNA, while only leukocytes and other lymphocytes have cell nuclei. Nevertheless, miRNA and mRNA are also useful because they can be analyzed for the presence of nucleotide variants in their sequence or serve as templates for cDNA synthesis. Tissue or cell samples can be analyzed directly with little or no processing. Alternatively, nucleic acids containing the target sequence can be extracted, purified, and / or amplified before being subjected to the various detection procedures described below. In addition to tissue or cell samples, cDNA or genomic DNA from cDNA or genomic DNA libraries constructed using test tissue or cell samples obtained from the individual are also useful.
[0314] Sequencing of target genomic DNA or cDNA, particularly the region containing the nucleotide variant locus to be detected, to determine the presence or absence of a specific nucleotide variant. Various sequencing techniques, including the Sanger method and the Gilbert chemical method, are generally known and widely used in the art. Pyrosequencing monitors DNA synthesis in real time using a luminometric detection system. Pyrosequencing has been shown to be effective for analyzing genetic polymorphisms such as single nucleotide polymorphisms and can also be used in this method. See Nordstrom et al., Biotechnol. Appl. Biochem., 31(2):107-112 (2000); Ahmadian et al., Anal. Biochem., 280:103-110 (2000).
[0315] Nucleic acid variants can be detected by appropriate detection processes. Non-limiting examples of methods include detection, quantification, and sequencing, such as mass detection of mass-modified amplicons (e.g., matrix-assisted laser desorption / ionization (MALDI) mass spectrometry and electrospray (ES) mass spectrometry), and primer extension methods (e.g., iPLEX®; Sequenom, Inc.).), microsequencing methods (e.g., modifications of primer extension methodologies), ligase sequencing methods (e.g., U.S. Patents No. 5,679,524 and 5,952,174, and International Publication No. 01 / 27326), mismatch sequencing methods (e.g., U.S. Patents No. 5,851,770; 5,958,692; 6,110,684; and 6,183,958), direct DNA sequencing, fragment analysis (FA), restriction fragment length polymorphism (RFLP analysis), allele-specific oligonucleotide (ASO) analysis, methylation-specific PCR (...
Claims
1. A method for selecting a treatment for colorectal cancer in the treatment of a first target, A process of obtaining multiple copies based on output data generated by a next-generation sequencer using one or more computers, The next-generation sequencer generates output data by sequencing a biological sample containing colorectal cancer cells derived from the first target. The aforementioned multiple copy numbers are groups of the following genes or adjacent genomic regions: (a) Group 1, comprising one or more genes selected from the group consisting of MYC, EP300, U2AF1, ASXL1, MAML2, CNTRL, WRN, and CDX2; (b) Group 2, comprising one or more genes selected from the group consisting of BCL9, PBX1, PRRX1, INHBA, and YWHAE; (c) Group 3, comprising one or more genes selected from the group consisting of BCL9, PBX1, GNAS, LHFPL6, CASP8, ASXL1, FH, CRKL, MLF1, TRRAP, AKT3, ACKR3, MSI2, PCM1, and MNX1; (d) Group 4 comprising one or more genes selected from the group consisting of BX1, GNAS, AURKA, CASP8, ASXL1, CRKL, MLF1, GAS7, MN1, SOX10, TCL1A, LMO1, BRD3, SMARCA4, PER1, PAX7, SBDS, SEPT5, PDGFB, AKT2, TERT, KEAP1, ETV6, TOP1, TLX3, COX6C, NFIB, ARFRP1, ARID1A, MAP2K4, NFKBIA, WWTR1, ZNF217, IL2, NSD3, CREB1, BRIP1, SDC4, EWSR1, FLT3, FLT1, FAS, CCNE1, RUNX1T1 and EZR; and (e) Group 5, comprising one or more genes selected from the group consisting of BCL9, PBX1, PRRX1, INHBA, YWHAE, GNAS, LHFPL6, FCRL4, BIRC3, AURKA, and HOXA11. The process includes the number of copies of each; The process of applying a machine learning classification model to the copy numbers obtained for each of the groups: Group 1, Group 2, Group 3, Group 4, and Group 5; A step of obtaining an indicator from each machine learning classification model of whether the subject is likely to benefit from treatment with 5-fluorouracil / leucovorin combined with oxaliplatin (FOLFOX); and The process of selecting FOLFOX if the majority of machine learning classification models indicate that the subject is likely to benefit from the treatment, and selecting an alternative treatment to FOLFOX if the majority of machine learning classification models indicate that the subject is unlikely to benefit from FOLFOX. The method, including the method described above.
2. The next-generation sequencer is Group 6 includes one or more genes selected from the group consisting of MYC, EP300, U2AF1, ASXL1, MAML2, and CNTRL; Group 7 includes one or more genes selected from the group consisting of BCL9, PBX1, PRRX1, INHBA, YWHAE, GNAS, LHFPL6, FCRL4, HOXA11, AURKA, BIRC3, IKZF1, CASP8, and EP300; Group 8 comprises one or more genes selected from the group consisting of PBX1, BCL9, INHBA, PRRX1, YWHAE, GNAS, LHFPL6, FCRL4, AURKA, IKZF1, CASP8, PTEN, and EP300; and Group 9 includes one or more genes selected from the group consisting of BCL9, PBX1, PRRX1, INHBA, GNAS, YWHAE, LHFPL6, FCRL4, PTEN, HOXA11, AURKA, and BIRC3. The method according to claim 1, further comprising generating output data for determining the number of copies.
3. The method according to claim 1, wherein the copy numbers of multiple genes in each group are determined.
4. The method according to claim 1, wherein the copy number of all genes in each group is determined.
5. The method according to claim 1, wherein the treatment further comprises irinotecan or a biological agent.
6. The method according to claim 5, wherein the biological agent comprises bevacizumab or cetuximab.
7. The method according to claim 1, wherein the colorectal cancer is metastatic colorectal cancer.
8. A step of obtaining a second biological sample containing colorectal cancer cells from a second subject using one or more computers; A process of obtaining a plurality of second copy numbers based on second output data generated by a next-generation sequencer using one or more computers, The next-generation sequencer generates second output data by sequencing a second biological sample containing colorectal cancer cells derived from a second target. The aforementioned group of second copy numbers are the following genes or groups of genomic regions adjacent to them: (f) Group 1, comprising one or more genes selected from the group consisting of MYC, EP300, U2AF1, ASXL1, MAML2, CNTRL, WRN, and CDX2; (g) Group 2, comprising one or more genes selected from the group consisting of BCL9, PBX1, PRRX1, INHBA, and YWHAE; (h) Group 3 comprising one or more genes selected from the group consisting of BCL9, PBX1, GNAS, LHFPL6, CASP8, ASXL1, FH, CRKL, MLF1, TRRAP, AKT3, ACKR3, MSI2, PCM1, and MNX1; (i) Group 4 comprising one or more genes selected from the group consisting of BX1, GNAS, AURKA, CASP8, ASXL1, CRKL, MLF1, GAS7, MN1, SOX10, TCL1A, LMO1, BRD3, SMARCA4, PER1, PAX7, SBDS, SEPT5, PDGFB, AKT2, TERT, KEAP1, ETV6, TOP1, TLX3, COX6C, NFIB, ARFRP1, ARID1A, MAP2K4, NFKBIA, WWTR1, ZNF217, IL2, NSD3, CREB1, BRIP1, SDC4, EWSR1, FLT3, FLT1, FAS, CCNE1, RUNX1T1, and EZR; and (j) Group 5, comprising one or more genes selected from the group consisting of BCL9, PBX1, PRRX1, INHBA, YWHAE, GNAS, LHFPL6, FCRL4, BIRC3, AURKA, and HOXA11. The process includes the number of copies of each; The process of applying a machine learning classification model to the second copy number obtained for each of the groups 1, 2, 3, 4, and 5; The process of obtaining an indicator from each machine learning classification model of whether the second subject is likely to benefit from treatment with 5-fluorouracil / leucovorin in combination with oxaliplatin (FOLFOX); and The process of selecting FOLFOX if the majority of machine learning classification models indicate that the second subject is likely to benefit from the aforementioned treatment, and selecting an alternative treatment to FOLFOX if the majority of machine learning classification models indicate that the second subject is unlikely to benefit from FOLFOX. The method according to claim 1, further comprising:
9. The method according to claim 8, wherein the alternative treatment is a treatment comprising a combination of 5-fluorouracil / leucovorin with irinotecan (FOLFIRI).
10. The method according to claim 1, wherein the biological sample comprises formalin-fixed paraffin-embedded (FFPE) tissue, fixed tissue, core needle biopsy, aspiration fluid, unstained slide, fresh frozen (FF) tissue, formalin sample, tissue contained in a solution for preserving nucleic acids or protein molecules, fresh sample, malignant fluid, body fluid, tumor sample, tissue sample, or any combination thereof.
11. The method according to claim 1, wherein the biological sample comprises cells from a solid tumor or body fluid.
12. The method according to claim 11, wherein the biological sample comprises body fluids, and the body fluids include peripheral blood, serum, plasma, ascites, urine, cerebrospinal fluid (CSF), sputum, saliva, bone marrow, synovial fluid, aqueous humor, amniotic fluid, earwax, breast milk, bronchoalveolar lavage fluid, semen, prostatic fluid, Cowper's gland fluid, bulbourethral gland fluid, female ejaculate, sweat, feces, tears, cystic fluid, pleural fluid, peritoneal fluid, pericardial fluid, lymphatic fluid, ovoid, chyle, bile, interstitial fluid, menstrual secretions, pus, sebum, vomit, vaginal secretions, mucosal secretions, watery stool, pancreatic juice, nasal lavage fluid, bronchopulmonary aspirate, blastocoel fluid, or umbilical cord blood.
13. The method according to claim 1, wherein a next-generation sequencer performs whole exome sequencing.
14. The method according to claim 1, wherein each of the machine learning classification models has equal voting.
15. The method according to claim 1, wherein the voting for each machine learning classification model is adjusted based on the confidence score associated with each machine learning classification model.
16. The method according to claim 1, wherein FOLFOX is selected in response to a machine learning classification model that indicates that the subject is likely to benefit from the treatment.
17. A step to obtain data representing test entities, wherein the obtained data is MYC, EP300, U2AF1, ASXL1, MAML2, CNTRL, WRN, CDX2, BCL9, PBX1, PRRX1, INHBA, YWHAE, GNAS, LHFPL6, CASP8, FH, CRKL, MLF1, TRRAP, AKT3, ACKR3, MSI2, PCM1, MNX1, AURKA, GAS7, MN1, SOX10, TCL1A, LMO1, BRD3, SMARCA4, PER1, P A process including data for one or more biomarkers selected from AX7, SBDS, SEPT5, PDGFB, AKT2, TERT, KEAP1, ETV6, TOP1, TLX3, COX6C, NFIB, ARFRP1, ARID1A, MAP2K4, NFKBIA, WWTR1, ZNF217, IL2, NSD3, BRIP1, SDC4, EWSR1, FLT3, FLT1, FAS, CCNE1, RUNX1T1, EZR, FCRL4, BIRC3, and HOXA11, For each machine learning model of one or more machine learning models trained on the same set of training data, The step of providing the obtained data as input to the machine learning model, wherein the machine learning model is trained to determine a specific class of one or more training entities from a plurality of different entity classes based on the processing of input data representing each of one or more training entities, wherein the plurality of different entity classes include (1) a response class of test entities that respond to a treatment comprising 5-fluorouracil and leucovorin in combination with oxaliplatin (FOLFOX), and (2) a non-response class of test entities that do not respond to a treatment comprising 5-fluorouracil and leucovorin in combination with oxaliplatin (FOLFOX); The process of processing the provided data through each layer of the machine learning model in order to generate output data; and A step of obtaining output data generated by the machine learning model based on the processing of the provided data by the machine learning model, wherein the obtained output data indicates a specific class among the plurality of different entity classes as an initial classification for the test entity; A step of obtaining output data obtained for each of the one or more machine learning models, wherein the provided output data includes data representing the initial classification decision for the test entity by each of the one or more machine learning models; and A step of determining the most likely entity class for the test entity based on the provided output data, wherein the most likely entity class is the reactive class or the non-reactive class. A method for classifying experimental entities for therapeutic purposes, including
18. The step of determining the most likely entity class for the test entity based on the provided output data is: Determining the number of occurrences of each initial classification of the test entity in a particular class among the plurality of different entity classes; and Selecting one of the multiple different entity classes that has the highest frequency of occurrence in the initial classification as the most likely entity class for the test entity. The method according to claim 17, including the method described in claim 17.
19. The process of accessing the confidence scores of each of the one or more machine learning models, and The process of adjusting the output data generated by each machine learning model based on the confidence score corresponding to each machine learning model. The method according to claim 17, further comprising:
20. The method according to claim 19, wherein the confidence score of each of the one or more machine learning models indicates the historical accuracy of each of the one or more machine learning models.
21. The process of adjusting the output data generated by each machine learning model based on the confidence score corresponding to each machine learning model is as follows: Increasing the weighting of the output data generated by the first machine learning model among the one or more machine learning models, based on the confidence score corresponding to the first machine learning model. The method according to claim 19, including the method described in claim 19.
22. The process of adjusting the output data generated by each machine learning model based on the confidence score corresponding to each machine learning model is as follows: Decreasing the weighting of the output data generated by the first machine learning model among the one or more machine learning models, based on the confidence score corresponding to the first machine learning model. The method according to claim 19, including the method described in claim 19.
23. The method according to claim 17, wherein at least one of the one or more machine learning models includes a random forest classification algorithm, a support vector machine, logistic regression, a k-nearest neighbor model, an artificial neural network, a simple Bayesian model, quadratic discriminant analysis, or a Gaussian process model.
24. The method according to claim 17, wherein at least one of the one or more biomarkers is selected from MYC, EP300, U2AF1, ASXL1, MAML2, CNTRL, WRN, and CDX2.
25. The method according to claim 17, wherein at least one of the one or more biomarkers is selected from BCL9, PBX1, PRRX1, INHBA, and YWHAE.
26. The method according to claim 17, wherein at least one of the one or more biomarkers is selected from BCL9, PBX1, GNAS, LHFPL6, CASP8, ASXL1, FH, CRKL, MLF1, TRRAP, AKT3, ACKR3, MSI2, PCM1, and MNX1.
27. The method according to claim 17, wherein at least one of the one or more biomarkers is selected from GNAS, AURKA, CASP8, ASXL1, CRKL, MLF1, GAS7, MN1, SOX10, TCL1A, LMO1, BRD3, SMARCA4, PER1, PAX7, SBDS, SEPT5, PDGFB, AKT2, TERT, KEAP1, ETV6, TOP1, TLX3, COX6C, NFIB, ARFRP1, ARID1A, MAP2K4, NFKBIA, WWTR1, ZNF217, IL2, NSD3, BRIP1, SDC4, EWSR1, FLT3, FLT1, FAS, CCNE1, RUNX1T1, and EZR.
28. The method according to claim 17, wherein at least one of the one or more biomarkers is selected from BCL9, PBX1, PRRX1, INHBA, YWHAE, GNAS, LHFPL6, FCRL4, BIRC3, AURKA, and HOXA11.
29. The method according to claim 17, wherein at least two of the one or more machine learning models include different types of machine learning models.
30. One or more computer-readable storage media that, when executed by one or more computers, stores instructions causing one or more computers to perform the method according to any one of claims 1 to 29.
31. A system comprising one or more processors configured to perform the method described in any one of claims 1 to 29.