Methods and systems for predicting success in genomic profiling - Patents.com

JP2025505920A5Pending Publication Date: 2025-12-12FOUNDATION MEDICINE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024537449
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-12-21
Filing Date
2022-12-06
Publication Date
2025-12-12

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method is described for predicting the likelihood of success for performing genomic profiling of a sample from a subject.In some examples, the method may include: receiving data of a plurality of pre-analysis variables associated with the sample; applying the received data to a multivariate model trained to predict the outcome for genomic profiling assay; generating a prediction of the likelihood of success for performing genomic profiling of the sample based on the applied multivariate model; and reporting the prediction of the likelihood of success for performing genomic profiling of the sample.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 292,395, filed December 21, 2021, the contents of which are incorporated by reference herein in their entirety.

[0002] The present disclosure relates generally to genomic profiling methods and systems, and more particularly to methods and systems for predicting the success of performing genomic profiling of an individual sample based on pre-analytical variables of the individual sample. [Background technology]

[0003] Genomic profiling (GP) is a next-generation sequencing (NGS)-based method that allows the analysis of a panel of genes (e.g., tens to hundreds of genes) in a single assay for the purpose of detecting genomic alterations (e.g., variant nucleic acid sequences) that may be diagnostic and / or prognostic for diseases such as cancer. For example, GP can be used to detect the four major classes of genomic alterations known to cause cancer growth (i.e., base substitutions, insertions and deletions, copy number alterations (CNAs), and rearrangements or fusions).

[0004] Despite the potential benefits and success of using GP to identify genomic alterations associated with disease, various challenges remain for the expanded implementation of GP testing.Existing guidelines for physician sample submission do not take into account pre-analytical variables that may affect the success of performing GP.Therefore, a method for more accurately predicting the likelihood of performing GP successfully for a particular sample would be advantageous in terms of both reducing the time and cost of performing GP, and improving the impact of GP test results on health care decisions and health care outcomes. Summary of the Invention

[0005] Disclosed herein are methods and systems for more accurately predicting the likelihood of successfully performing genomic profiling (GP) on a particular sample based on a set of pre-analytical variables associated with the sample. In some examples, the disclosed methods and systems are particularly useful for predicting the likelihood of successfully performing comprehensive genomic profiling (CGP) on a particular sample based on a set of pre-analytical variables associated with the sample. The methods include using a training model to process pre-analytical data for sample input by a physician or other health care provider, and outputting an evidence-based prediction of the likelihood that a GP assay using the sample will produce a reliable result. In some examples, for example, if the predicted likelihood of successfully performing a GP analysis is low, a recommendation may be made to procure a new sample before submitting for GP testing. The disclosed methods and systems are advantageous in terms of reducing both the time and cost of performing GP, and improving the impact of GP testing results on health care decisions and health care outcomes.

[0006] Disclosed herein is a method for predicting a likelihood of success for performing genomic profiling of a sample from a subject, the method including receiving data for a plurality of pre-analytical variables associated with the sample, applying the received data to a multivariate model trained to predict outcomes for genomic profiling assays, generating a prediction of the likelihood of success for performing genomic profiling of the sample based on the applied multivariate model, and reporting, using one or more processors, the prediction of the likelihood of success for performing genomic profiling of the sample.

[0007] In some embodiments, the method further includes providing a plurality of nucleic acid molecules obtained from a sample from the subject based on a prediction of a likelihood of success being equal to or greater than a predetermined threshold; ligating one or more adaptors onto one or more nucleic acid molecules from the plurality of nucleic acid molecules; amplifying one or more ligated nucleic acid molecules from the plurality of nucleic acid molecules; capturing an amplified nucleic acid molecule from the amplified nucleic acid molecules; sequencing the captured nucleic acid molecules with a sequencer to obtain a plurality of sequence reads that represent the captured nucleic acid molecules and overlap with one or more loci within one or more subgenomic intervals in the sample; and generating, by the one or more processors, a genomic profile based on the sequence reads of the sample, the genomic profile comprising sequence read analysis data.

[0008] In some embodiments, the method further comprises training the multivariate model using training data. In some embodiments, the training data comprises data derived from univariate analysis of clinical study data of samples collected from the subject, representing subject age range, subject gender, diagnosis, stage of disease, sample type, sample collection site, sample collection method, sample preparation method, sample storage method, sample age, sample transportation method, or any combination thereof. In some embodiments, the training data further comprises data of ECOG status, subject treatment status, tumor cellularity in the sample, tumor nuclear content in the sample, tissue surface area in the sample, tissue matrix in the sample, previous genomic profiling assay results, or any combination thereof.

[0009] In some embodiments, the plurality of pre-analytical variables comprises subject age, subject sex, diagnosis, stage of disease, sample type, sample collection site, sample collection method, sample preparation method, sample storage method, sample age, sample transportation method, or any combination thereof. In some embodiments, the plurality of pre-analytical variable data is added with data on tumor cellularity in the sample, tumor nuclear content in the sample, tissue surface area in the sample, tissue matrix in the sample, or any combination thereof.

[0010] In some embodiments, the multivariate model comprises a machine learning model. In some embodiments, the machine learning model comprises a supervised learning model. In some embodiments, the machine learning model comprises an unsupervised learning model. In some embodiments, the multivariate model comprises a logistic regression model, a multiple linear regression model, a random forest model, a neural network model, or a deep learning model.

[0011] In some embodiments, the prediction of the likelihood of success comprises a binary value, a percentage, or a score.

[0012] In some embodiments, the method further includes comparing the predicted likelihood of success to a predetermined threshold and, based on a determination that the predicted likelihood of success is equal to or greater than the predetermined threshold, outputting an indication that the sample from the subject is suitable for providing a genomic profile of the subject.

[0013] In some embodiments, the method further includes comparing the predicted likelihood of success to a predetermined threshold and outputting a recommendation to collect a new sample instead of submitting the sample for genomic profiling based on a determination that the predicted likelihood of success is less than the predetermined threshold. In some embodiments, the method further includes outputting a recommendation of a sample type or sample collection site for the new sample. In some embodiments, the method further includes outputting a recommendation to perform an alternative nucleic acid sequencing-based testing method. In some embodiments, the predetermined threshold varies depending on the sample type.

[0014] In some embodiments, data for the plurality of pre-analytical variables is entered by a user via a graphical user interface (GUI) on a display device. In some embodiments, a prediction of the likelihood of success for performing genomic profiling of the sample is reported via a graphical user interface (GUI) on a display device. In some embodiments, the graphical user interface (GUI) is displayed in a web browser.

[0015] In some embodiments, the received data for the plurality of pre-analytical variables includes sample type data. In some embodiments, the remaining pre-analytical variables of the plurality of pre-analytical variables are selected based on the sample type. In some embodiments, the multivariate model is selected based on the sample type.

[0016] In some embodiments, the sample comprises a tissue biopsy sample, a liquid biopsy sample, or a normal control. In some embodiments, the sample is a tissue biopsy sample and comprises bone marrow. In some embodiments, the sample is a liquid biopsy sample and comprises blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva. In some embodiments, the sample is a liquid biopsy sample and comprises circulating tumor cells (CTCs). In some embodiments, the sample is a liquid biopsy sample and comprises cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), or any combination thereof.

[0017] In some embodiments, if the predicted likelihood of success is equal to or exceeds a predetermined threshold, genomic profiling is performed and used to diagnose or confirm a diagnosis of a disease in the subject. In some embodiments, genomic profiling is also used to determine eligibility for therapy based on biomarker status. In some embodiments, the disease is cancer. In some embodiments, the method further comprises selecting an anti-cancer therapy for administration to the subject based on the results of the genomic profiling. In some embodiments, the method further comprises determining an effective amount of the anti-cancer therapy for administration to the subject based on the results of the genomic profiling. In some embodiments, the method further comprises administering an anti-cancer therapy to the subject based on the results of the genomic profiling. In some embodiments, the anti-cancer therapy comprises chemotherapy, radiation therapy, immunotherapy, targeted therapy, or surgery.In some embodiments, the cancer is B-cell cancer (multiple myeloma), melanoma, breast cancer, lung cancer, bronchial cancer, colorectal cancer, prostate cancer, pancreatic cancer, gastric cancer, ovarian cancer, bladder cancer, brain cancer, central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, endometrial cancer, oral cavity cancer, pharyngeal cancer, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, small intestine cancer, appendix cancer, salivary gland cancer, thyroid cancer, adrenal gland cancer, osteosarcoma, chondrosarcoma, blood Tissue cancer, Adenocarcinoma, Inflammatory myofibroblastoma, Gastrointestinal stromal tumor (GIST), Colon cancer, Multiple myeloma (MM), Myelodysplastic syndrome (MDS), Myeloproliferative disorder (MPD), Acute lymphocytic leukemia (ALL), Acute myeloid leukemia (AML), Chronic myeloid leukemia (CML), Chronic lymphocytic leukemia (CLL), Polycythemia Vera, Hodgkin's lymphoma, Non-Hodgkin's lymphoma (NHL), Soft tissue sarcoma, Fibrosarcoma, Myxosarcoma, Lipoma Liposarcoma, osteosarcoma, chordoma, angiosarcoma, endothelial sarcoma, lymphangiosarcoma, lymphangioendothelial sarcoma, synovium, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinoma, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatoma, bile duct carcinoma, choriocarcinoma, seminoma, embryonal carcinoma, Wilms' tumor, bladder cancer, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharynx Cephaloma, ependymoma, pineal cell tumor, glioblastoma, acoustic neuroblastoma, oligodendroglioma, meningioma, neuroblastoma, retinoblastoma, follicular lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, hepatocellular carcinoma, thyroid cancer, gastric cancer, head and neck cancer, small cell carcinoma, essential thrombocythemia, aplastic myeloid metaplasia, hypereosinophilic syndrome, systemic mastocytosis, familial eosinophilia, chronic eosinophilic leukemia, neuroendocrine carcinoma, or carcinoid tumor.

[0018] In some embodiments, genomic profiling of the subject comprises obtaining results from a comprehensive genomic profiling (CGP) test, a gene expression profiling test, a cancer hotspot panel test, a DNA methylation test, a DNA fragmentation test, an RNA fragmentation test, or any combination thereof. In some embodiments, genomic profiling of the subject further comprises obtaining results from a nucleic acid sequencing-based test. In some embodiments, the method further comprises selecting an anti-cancer drug, administering an anti-cancer drug, or administering an anti-cancer treatment to the subject based on the results of the genomic profiling.

[0019] Also disclosed herein is a system comprising: one or more processors; and a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to receive data for a plurality of pre-analytical variables associated with a sample from a subject, apply the received data to a multivariate model trained to predict an outcome for a genomic profiling assay, generate a prediction of the likelihood of success for performing genomic profiling of the sample based on the applied multivariate model, and report the prediction of the likelihood of success for performing genomic profiling of the sample.

[0020] In some embodiments, the instructions further cause the system to compare the predicted likelihood of success to a predetermined threshold and output a recommendation to collect a new sample instead of submitting the sample for genomic profiling based on a determination that the predicted likelihood of success is less than the predetermined threshold. In some embodiments, the data for the plurality of pre-analytical variables is entered by a user via a graphical user interface (GUI) on a display device. In some embodiments, the predicted likelihood of success for performing genomic profiling of the sample is reported via a graphical user interface (GUI) on a display device. In some embodiments, the graphical user interface (GUI) is displayed in a web browser.

[0021] Disclosed herein is a non-transitory computer readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of the system, cause the system to receive data of a plurality of pre-analytical variables associated with a sample from a subject, apply the received data to a multivariate model trained to predict outcomes for a genomic profiling assay, generate a prediction of the likelihood of success for performing genomic profiling of the sample based on the applied multivariate model, and report the prediction of the likelihood of success for performing genomic profiling of the sample. In some embodiments, the instructions further cause the system to compare the predicted likelihood of success with a predetermined threshold, and output a recommendation to collect a new sample instead of submitting the sample for genomic profiling based on a determination that the predicted likelihood of success is less than the predetermined threshold. Incorporation by Reference

[0022] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference in their entirety to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference in its entirety. In the event of a conflict between a term in this specification and a term in an incorporated reference, the term in this specification shall control. [Brief description of the drawings]

[0023] Various aspects of the disclosed methods, devices, and systems are set forth with particularity in the appended claims. A better understanding of the features and advantages of the disclosed methods, devices, and systems will be obtained by reference to the following detailed description of exemplary embodiments and the accompanying drawings.

[0024] [Figure 1] FIG. 1 provides a non-limiting example of a process flow chart for receiving and processing pre-analytical variable data using a multivariate model to predict the likelihood of success for performing a GP on a particular sample. [Diagram 2] A non-limiting example of a process flow chart for displaying and inputting a request for pre-analytical variable data for a particular sample via a graphical user interface (GUI), processing the pre-analytical variable data using a multivariate model to predict the likely likelihood of success for performing a GP on the sample, and displaying the predicted likelihood of success for performing a GP on the sample in the graphical user interface is provided. [Diagram 3] 1 illustrates an exemplary computing device in accordance with some examples of systems described herein. [Figure 4] 1 illustrates an exemplary computer system or computer network in accordance with some examples of the systems described herein. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0025] Disclosed herein are methods and systems for more accurately predicting the likelihood of successfully performing genomic profiling (GP) on a particular sample based on a set of pre-analytical variables associated with the sample. In some examples, the disclosed methods and systems are particularly useful for predicting the likelihood of successfully performing comprehensive genomic profiling (CGP) on a particular sample based on a set of pre-analytical variables associated with the sample. The methods include using a training model to process pre-analytical data of the sample (e.g., data input by a physician or other health care provider) and outputting an evidence-based prediction of the likelihood that a GP assay using the sample will produce a reliable result. In some examples, for example, if the predicted likelihood of successfully performing a GP analysis is low, a recommendation may be made to procure a new sample before submitting for GP testing. The disclosed methods and systems are advantageous in terms of reducing both the time and cost of performing GP and improving the impact of GP testing results on health care decisions and health care outcomes.

[0026] In some examples, methods are described that include, for example, receiving data for a plurality of pre-analytical variables associated with a sample, applying the received data to a multivariate model trained to predict an outcome for a GP assay, generating a prediction of the likelihood of success for performing a GP on the sample based on the applied multivariate model, and reporting the prediction of the likelihood of success for performing a GP on the sample.

[0027] Examples of pre-analytical variables include, but are not limited to, subject age, subject gender, diagnosis, stage of disease, sample type, sample collection site, sample collection method, sample preparation method, sample storage method, sample age, sample transportation method, or any combination thereof.

[0028] In some examples, the method may further include comparing the predicted likelihood of success to a predetermined threshold and, based on a determination that the predicted likelihood of success is less than the predetermined threshold, outputting a recommendation to collect a new sample instead of submitting a sample for GP.

[0029] In some examples, data for the plurality of pre-analytical variables is entered by a user via a graphical user interface (GUI) on a display device. In some examples, a prediction of the likelihood of success for performing a GP of the sample is reported via a graphical user interface (GUI) on a display device. In some examples, the graphical user interface (GUI) is displayed in a web browser.

[0030] definition Unless otherwise defined, all technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.

[0031] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Any reference to "or" herein is intended to include "and / or" unless specifically stated otherwise.

[0032] As used herein, the terms "comprising" (and any form or variation of comprising, such as "comprise" and "comprises"), "having" (and any form or variation of having, such as "have" and "has"), "including" (and any form or variation of including "includes" and "include"), or "containing" (and any form or variation of containing, such as "contains" and "contain") are inclusive or open-ended and do not exclude additional, unlisted additives, components, integers, elements, or method steps.

[0033] As used herein, the term "subgenomic interval" (or "subgenomic sequence interval") refers to a portion of a genome sequence.

[0034] As used herein, the term "interval of interest" refers to a subgenomic interval or an expressed subgenomic interval (eg, a transcribed sequence of a subgenomic interval).

[0035] As used herein, the terms "variant sequence" or "variant" are used interchangeably and refer to a nucleic acid sequence that is modified relative to a corresponding "normal" or "wild-type" sequence. In some examples, the variant sequence may be a "short variant sequence" (or "short variant"), i.e., a variant sequence that is less than about 50 base pairs in length.

[0036] The terms "allele frequency" and "allele fraction" are used interchangeably herein and refer to the fraction of sequence reads that correspond to a particular allele relative to the total number of sequence reads for a genomic locus.

[0037] The terms "variant allele frequency" and "variant allele fraction" are used interchangeably herein and refer to the fraction of sequence reads that correspond to a particular variant allele relative to the total number of sequence reads for a genomic locus.

[0038] As used herein, the term "pre-analytical variables" refers to sample-specific characteristics or parameters related to the patient's history and / or the sample's history. Examples of pre-analytical variables include, but are not limited to, patient's gender, patient's age, patient's diagnosis, stage of disease, sample type, sample collection site, sample preparation method, sample age, sample storage method, sample transport method, or any combination thereof. In some examples, the pre-analytical variable data can be added with additional data, such as Eastern Cooperative Oncology Group (ECOG) status, patient's treatment status, sample imaging-based characteristics, such as tumor cellularity in the sample, tumor nuclear content in the sample, tissue surface area of ​​the sample, tissue matrix in the sample, previous GP assay results, or any combination thereof.

[0039] As used herein, the phrase "likelihood of success for performing genomic profiling" refers to the prediction that a nucleic acid-based assay, including but not limited to a GP assay, a personalized GP assay (e.g., a hotspot panel), or other molecular-based assay, will produce reliable results for a particular individual's sample.

[0040] Any section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0041] Methods for successful genomic profiling: The disclosed method for predicting the likelihood of successfully performing a GP on a particular sample based on a set of pre-analytical variables associated with the sample allows the physician / system to evaluate the predicted sequencing success prediction for available specimens for an individual patient and therefore determine which specimen to send for GP testing, or alternatively whether an alternative specimen should be procured if the currently available specimen is unlikely to yield a successful GP test result.

[0042] There are existing guidelines to follow when submitting samples, but these guidelines do not consider pre-analytical variables that may affect the success of performing GP analysis, and improved methods are needed. The methods described herein take into account multiple pre-analytical variables and utilize a training model (e.g., a multivariate regression model) that uses real-world data to generate an evidence-based success score (e.g., a prediction of the probability or likelihood of performing GP on a sample and obtaining a reliable result). For example, a score above a threshold may indicate that the sample is suitable for performing GP (e.g., that GP analysis of the sample will be successful).

[0043] As indicated above, examples of pre-analytical variables that may affect the likelihood of success for performing a GP include, but are not limited to, patient gender, patient age, patient diagnosis, sample type, sample collection site, sample preparation method, sample age, sample storage method, sample transport method, etc. In some examples, the pre-analytical variable data may be added with additional data, such as features based on sample imaging, such as tumor cellularity in the sample, tumor nuclear content in the sample, tissue surface area of ​​the sample, tissue matrix in the sample, or any combination thereof.

[0044] In some examples, the optimal set of pre-analytical variables considered to predict the likelihood of success for performing a GP and / or the model used may depend on the particular sample type. For example, in the case of some tissue specimens, the set of pre-analytical variables used may include the type of specimen, the patient's age, the patient's sex, the patient's diagnosis, and the specimen collection site, as well as additional sample features such as tumor nucleus cellularity, tissue surface area, tissue matrix, etc. Alternatively, in the case of liquid biopsy samples, the set of pre-analytical variables used may include the patient's age, the patient's sex, the patient's diagnosis, stage of disease, etc. In some examples, the pre-analytical variables may affect the likelihood of GP success, for example, because they are predictors of the amount of DNA available in the sample, the quality of DNA available in the sample, or other biological factors. The amount of circulating cell-free DNA (cfDNA) and circulating tumor DNA (ctDNA) present in a peripheral blood biopsy can vary depending on several preanalytical variables, for example (Huang, et al. (2021) “Circulating Cell-Free DNA Yield and Circulating-Tumor DNA Quantity from Liquid Biopsies of 12,139 Cancer Patients”, Clinical Chem. Oct 9:hvab176. doi:10.1093 / clinchem / hvab176. Epub ahead of print. PMID:34626187). Non-limiting examples of strong predictors of GP success include sample type, size of the biopsy sample, or the amount of tumor nuclei visible within the sample.

[0045] In some examples, a system configured to perform the disclosed methods may be accessed by a physician, for example, by ordering via a graphical user interface displayed on a website. The physician or other health care provider may input data for a number of pre-analytical variables available for a given specimen in the website, and a training model processes the data and generates a probability of successfully performing a GP analysis. In some examples, data for the pre-analytical variables may be automatically retrieved by the system, for example, from one or more databases.

[0046] FIG. 1 provides a non-limiting example of a flow chart of a process 100 for receiving and processing pre-analytical variable data using a multivariate model to predict the likelihood of success for performing a GP on a particular sample. The process 100 may be implemented using, for example, one or more electronic devices (e.g., computers) implementing a software platform. In some examples, the process 100 is implemented using a client-server system, and the blocks of the process 100 are divided in any manner between a server and a client device. In other examples, the blocks of the process 100 are divided between a server and multiple client devices. Thus, although portions of the process 100 are described herein as being implemented by a particular device of a client-server system, it should be understood that the process 100 is not so limited. In other examples, the process 100 is implemented using only a client device or only multiple client devices. In the process 100, some blocks are optionally combined, the order of some blocks is optionally changed, and some blocks are optionally omitted. In some examples, additional steps can be implemented in combination with the process 100. Accordingly, the operations illustrated (and described in more detail below) are exemplary in nature and therefore should not be considered as limiting.

[0047] In step 102 of FIG. 1, data for a number of pre-analytical variables associated with the sample is received (eg, entered by a physician or other health care provider or retrieved from one or more databases).

[0048] In some examples, the plurality of pre-analytical variables includes subject age, subject gender, diagnosis, stage of disease, sample type, sample collection site, sample collection method, sample preparation method, sample storage method, sample age, sample transportation method, or any combination thereof.

[0049] In some examples, data for multiple pre-analytical variables is added along with data for tumor cellularity in the sample, tumor nuclear content in the sample, tissue surface area in the sample, tissue matrix in the sample, or any combination thereof.

[0050] In some examples, the number of pre-analytical variables and / or auxiliary features for which data is received may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more than 20.

[0051] The sample may include any of a variety of sample types. For example, the sample may include a tissue biopsy sample, a liquid biopsy sample, and / or a normal control. In some examples, the sample is a tissue biopsy sample and includes bone marrow. In some examples, the sample is a liquid biopsy sample and includes blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva. In some examples, the sample is a liquid biopsy sample and includes circulating tumor cells (CTCs). In some examples, the sample is a liquid biopsy sample and includes cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), or any combination thereof.

[0052] In step 104 of Figure 1, the received data for the multiple pre-analytical variables and the additional data are processed using a multivariate model (e.g., a multivariate regression model) to generate a prediction (e.g., probability or success score) of the likelihood of performing a successful GP analysis on the sample. In some examples, different multivariate models (e.g., different multivariate regression models) may be used to predict the probability of performing a successful GP analysis for different sample types. In some examples, the prediction of the likelihood of success may include a binary value, a percentage, or a score.

[0053] Any of a variety of statistical analysis and / or machine learning techniques may be used to implement the multivariate model. In some examples, for example, the multivariate model may include a multiple linear regression model or a logistic regression analysis. In some examples, the multivariate model may include a machine learning model, such as a supervised or unsupervised learning model. In some examples, the machine learning model may include a random forest model, a neural network model, or a deep learning model.

[0054] In some examples, the multivariate model may be trained on training data derived from univariate analysis of clinical study data of samples collected from subjects, representing, for example, subject age range, subject gender, diagnosis, stage of disease, sample type, sample collection site, sample collection method, sample preparation method, sample storage method, sample age, sample transportation method, or any combination thereof.

[0055] In some examples, the training data used to train the multivariate model may further include data on ECOG status, treatment status, tumor cellularity in the sample, tumor nuclear content in the sample, tissue surface area in the sample, tissue matrix in the sample, previous GP assay results, or any combination thereof.

[0056] In step 106 of FIG. 1, a predicted likelihood of success for performing a GP on the sample is output (e.g., reported to an ordering physician or other health care provider). In some examples, the process may further include comparing the predicted likelihood of success to a predetermined threshold and outputting a recommendation to collect a new sample instead of submitting the sample for a GP based on a determination that the predicted likelihood of success is less than the predetermined threshold. In some examples, the process may further include outputting a recommendation of a sample type or sample collection site for the new sample. In some embodiments, the process may further include outputting a recommendation to perform an alternative nucleic acid sequencing-based testing method.

[0057] In some examples, the predetermined threshold may vary depending on the sample type. For example, the predetermined threshold may be different for tissue samples and liquid biopsy samples. In some examples, the threshold may be empirically determined by evaluating data for GP result-based health care decisions as a function of candidate thresholds for multiple samples. In some examples, the value of the predetermined threshold (i.e., the "probability of success" threshold) for a given sample type may be 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or 95%, and a recommendation is made to collect a new sample if the predicted likelihood of success for a given sample is below the predetermined threshold. Alternatively, or in addition, the predetermined threshold may be determined and / or modified based on the success rate of one or more previous processes from which the new sample was obtained.

[0058] FIG. 2 provides a non-limiting example flowchart of a process 200 for displaying and inputting a request for pre-analytical variable data for a particular sample via a graphical user interface (GUI), processing the pre-analytical variable data using a multivariate model to predict the likelihood of success for performing a GP on the sample, and displaying the predicted likelihood of success for performing a GP on the sample in the graphical user interface.

[0059] In step 202 of Figure 2, one or more fields of the GUI are displayed to request input of multiple pre-analytical variable data and / or additional data associated with each sample. In some examples, the GUI may include additional input fields, such as, for example, the ordering physician's name, the ordering physician's affiliation and / or address, the patient's name, the patient's billing address, the patient's insurance information, etc. In some examples, the list of requested pre-analytical variable data is updated as the sample type is entered by the user (e.g., a physician or other healthcare provider).

[0060] In step 204 of FIG. 2 , the requested information (e.g., data for the plurality of pre-analytical variables and / or additional data associated with the sample) is entered by a physician or other healthcare provider. In some examples, data for two or more of the plurality of pre-analytical variables (and / or additional data) may be entered within a single GUI field. In some examples, data for each pre-analytical variable (or item of additional data) may be entered in a separate field. In some examples, the requested information (e.g., data for the plurality of pre-analytical variables and / or additional data associated with the sample) may be downloaded from a database (e.g., a cloud-based database) containing the patient's medical records at the time of submission of the lead request. In some examples, the downloaded data may be used to automatically populate data fields in the GUI.

[0061] In step 206 of Figure 2, the entered data and / or additional data for a plurality of pre-analytical variables associated with the sample are processed using a multivariable model (e.g., a multivariable regression model) to generate a prediction of the likelihood of successful GP analysis of the sample. In some examples, the selection of the multivariable model (i.e., selected from a plurality of trained multivariable models) may be based on the sample type entered by a user (e.g., a physician or other health care provider) and / or based on a particular subset of pre-analytical variables for which data was entered by the user.

[0062] As indicated above, any of a variety of statistical analysis and / or machine learning techniques may be used to implement the multivariate model. In some examples, for example, the multivariate model may include a multiple linear regression model or a logistic regression analysis. In some examples, the multivariate model may include a machine learning model, such as a supervised or unsupervised learning model. In some examples, the machine learning model may include a random forest model, a neural network model, or a deep learning model.

[0063] In some examples, the multivariate model may be trained on training data derived from univariate analysis of clinical study data of samples collected from subjects, representing, for example, subject age range, subject gender, diagnosis, stage of disease, sample type, sample collection site, sample collection method, sample preparation method, sample storage method, sample age, sample transportation method, or any combination thereof.

[0064] In some examples, the training data used to train the multivariate model may further include data on ECOG status, treatment status, tumor cellularity in the sample, tumor nuclear content in the sample, tissue surface area in the sample, tissue matrix in the sample, previous GP assay results, or any combination thereof.

[0065] In some instances, the multivariate model may be trained using one or more sets of training data as well as any of a variety of machine learning training methodologies (e.g., steepest descent, Newton's method, conjugate gradient methods, quasi-Newton methods, or Levenberg-Marquardt methods).

[0066] In step 208 of FIG. 2, the predicted likelihood of GP success is displayed in a GUI field. As noted above, in some examples, the process may further include comparing the predicted likelihood of success to a predefined threshold and, based on a determination that the predicted likelihood of success is less than the predefined threshold, displaying a recommendation to collect a new sample instead of submitting a sample for GP. In some examples, the process may further include displaying a recommendation of a sample type or sample collection site for the new sample. In some embodiments, the process may further include displaying a recommendation to perform an alternative nucleic acid sequencing-based testing method. In some examples, the process may include displaying a plurality of alternative recommendations.

[0067] How to use In some examples, the disclosed methods include (i) obtaining a sample from a subject (e.g., a subject suspected of having or determined to have cancer); (ii) extracting nucleic acid molecules (e.g., a mixture of tumor and non-tumor nucleic acid molecules) from the sample; (iii) ligating one or more adapters (e.g., one or more amplification primers, flow cell adapter sequences, substrate adapter sequences, or sample index sequences) to the nucleic acid molecules extracted from the sample; (iv) amplifying the nucleic acid molecules (e.g., using a polymerase chain reaction (PCR) amplification technique, a non-PCR amplification technique, or an isothermal amplification technique); and (v) hybridizing the nucleic acid molecules to one or more bait molecules, each of which comprises one or more nucleic acid molecules that each comprise a region complementary to a region of the captured nucleic acid molecule (e.g., a region complementary to a region of the captured nucleic acid molecule). (vi) capturing the nucleic acid molecules from the amplified nucleic acid molecules (by hybridization); (vi) sequencing the nucleic acid molecules extracted from the sample (or library proxies derived therefrom), e.g., using a next-generation (e.g., massively parallel) sequencer, e.g., using next-generation (e.g., massively parallel) sequencing technology, whole genome sequencing (WGS) technology, whole exome sequencing technology, targeted sequencing technology, direct sequencing technology, or Sanger sequencing technology; and (vii) generating, displaying, transmitting, and / or delivering a report (e.g., electronic report, web-based report, or paper report) to a subject (or patient), a caregiver, a health care provider, a physician, an oncologist, an electronic medical record system, a hospital, a clinic, a medical clinic, a third party payer, an insurance company, or a government agency. In some examples, the report comprises an output from a method described herein. In some examples, all or a portion of the report can be displayed in a graphical user interface of an online or web-based health care portal. In some examples, the report is transmitted over a computer network or a peer-to-peer connection.

[0068] The disclosed methods can be used with any of a variety of samples. For example, in some examples, the sample can include a tissue biopsy sample, a liquid biopsy sample, or a normal control. In some examples, the sample can be a liquid biopsy sample and can include blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva. In some examples, the sample can be a liquid biopsy sample and can include circulating tumor cells (CTCs). In some examples, the sample can be a liquid biopsy sample and can include cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), or any combination thereof.

[0069] In some examples, the nucleic acid molecules extracted from the sample can include a mixture of tumor nucleic acid molecules and non-tumor nucleic acid molecules. In some examples, the tumor nucleic acid molecules can be derived from the tumor part of the heterogeneous tissue biopsy sample, and the non-tumor nucleic acid molecules can be derived from the normal part of the heterogeneous tissue biopsy sample. In some examples, the sample can include a liquid biopsy sample, the tumor nucleic acid molecules can be derived from the circulating tumor DNA (ctDNA) fraction of the liquid biopsy sample, and the non-tumor nucleic acid molecules can be derived from the non-tumor cell-free DNA (cfDNA) fraction of the liquid biopsy sample.

[0070] In some examples, the disclosed methods for predicting GP success ensure that reliable GP results are obtained and can be used to diagnose the presence of a disease or other condition in a subject (e.g., a patient) (e.g., cancer, genetic disorders (such as Down's syndrome and fragile X), neurological disorders, or any other disease type where detection of a variant, e.g., copy number variation, is relevant for diagnosing, treating, or predicting the disease). In some examples, the disclosed methods can be applicable to the diagnosis of any of a variety of cancers, as described elsewhere herein.

[0071] In some examples, the disclosed method for predicting GP success ensures that reliable GP results are obtained, which can then be used to predict genetic disorders in fetal DNA (e.g., for invasive or non-invasive prenatal testing). For example, sequence read data obtained by sequencing fetal DNA extracted from samples obtained using invasive amniocentesis, chorionic villus sampling (cVS), or fetal umbilical cord sampling techniques, or samples obtained using non-invasive sampling of cell-free DNA (cfDNA) samples (including a mixture of maternal and fetal cfDNA), can be processed according to the disclosed method to identify variants, e.g., copy number changes, associated with, for example, Down's syndrome (trisomy 21), trisomy 18, trisomy 13, extra copies or missing copies of X and Y chromosomes.

[0072] In some examples, the disclosed methods for predicting GP success ensure that reliable GP results are obtained, which can then be used to select subjects (e.g., patients) for clinical trials. In some examples, patient selection for clinical trials, for example, based on the identification of one or more variants at one or more loci, can accelerate the development of targeted therapies and improve the medical outcomes of treatment decisions.

[0073] In some examples, the disclosed methods for predicting GP success can be used to ensure that reliable GP results are obtained and then select an appropriate therapy or treatment (e.g., anti-cancer therapy or treatment) for a subject. In some examples, for example, the anti-cancer therapy or treatment can include the use of poly(ADP-ribose) polymerase inhibitors (PARPi), platinum compounds, chemotherapy, radiation therapy, targeted therapy (e.g., immunotherapy), surgery, or any combination thereof.

[0074] In some examples, the disclosed methods for predicting GP success ensure that reliable GP results are obtained, which can then be used in treating a disease (e.g., cancer) in a subject. For example, in response to identifying the presence of a variant sequence, an effective amount of an anti-cancer therapy or treatment can be administered to the subject using a GP assay.

[0075] In some examples, the disclosed methods for predicting GP success ensure that reliable GP results are obtained, which can then be used to monitor disease progression or recurrence (e.g., cancer or tumor progression or recurrence) in a subject. For example, in some examples, a GP assay may be used to detect the presence of one or more variant sequences in a first sample obtained from a subject at a first time point, and a GP assay may be used to detect the presence of one or more variant sequences in a second sample obtained from a subject at a second time point, and a comparison of the first and second determinations allows for monitoring disease progression or recurrence. In some examples, the first time point is selected before the subject is administered a therapy or treatment, and the second time point is selected after the subject is administered a therapy or treatment.

[0076] In some examples, GP analysis may be used to adjust a therapy or treatment (e.g., an anti-cancer therapy or anti-cancer therapy) for a subject, for example, by adjusting a therapeutic dose and / or selecting a different treatment in response to changes in variant allele frequency detected in a sample derived from the subject.

[0077] In some instances, the value of predicting the success of GP is to ensure that a reliable GP result is obtained, which can then be used as a prognostic or diagnostic indicator associated with the sample. For example, in some instances, the prognostic or diagnostic indicator can include an indicator of the presence of a disease (e.g., cancer) in the sample, an indicator of the probability that a disease (e.g., cancer) is present in the sample, an indicator of the probability that the subject from which the sample is derived will develop a disease (e.g., cancer) (i.e., a risk factor), or an indicator of the likelihood that the subject from which the sample is derived will respond to a particular therapy or treatment.

[0078] In some examples, the disclosed methods for predicting GP success can ensure that reliable GP results are obtained. The GP process can include identification of the presence of variant sequences at one or more loci in a sample from a subject as part of the detection, monitoring, risk factor prediction, or treatment selection of a particular disease (e.g., cancer). In some examples, the variant panel selected for GP can include detection of variant sequences at a selected set of loci. In some examples, the variant panel selected for GP can include detection of variant sequences at several loci via a next generation sequencing (NGS) approach used for GP, evaluating hundreds of genes (including relevant cancer biomarkers) in a single assay. The improved reliability of GP results, as facilitated using the disclosed methods for predicting GP success, can improve the efficacy of, for example, disease detection calls and treatment decisions made based on genomic profiles by, for example, independently confirming the presence of variant sequences in a given patient sample.

[0079] In some examples, a genomic profile may include information about the presence of genes (or variant sequences thereof), copy number variations, epigenetic traits, proteins (or modifications thereof), and / or other biomarkers in an individual's genome and / or proteome, as well as the individual's corresponding phenotypic traits, and information about the interactions between genetic or genomic traits, phenotypic traits, and environmental factors.

[0080] In some examples, the subject's genomic profile may include results from a GP test, a nucleic acid sequencing-based test, a gene expression profiling test, a cancer hotspot panel test, a DNA methylation test, a DNA fragmentation test, an RNA fragmentation test, or any combination thereof.

[0081] In some examples, the method may further include administering or applying a treatment or therapy (e.g., an anti-cancer drug, anti-cancer treatment, or anti-cancer therapy) to the subject based on the generated genomic profile. An anti-cancer drug or anti-cancer treatment may refer to a compound that is effective in treating cancer cells. Examples of anti-cancer drugs or anti-cancer therapies include, but are not limited to, alkylating agents, antimetabolites, natural products, hormones, chemotherapy, radiation therapy, immunotherapy, surgery, or therapies configured to target defects in a particular cell signaling pathway, such as defects in the DNA mismatch repair (MMR) pathway.

[0082] sample The disclosed methods and systems may be used with any of a variety of samples (also referred to herein as specimens) containing nucleic acids (e.g., DNA or RNA) collected from a subject (e.g., a patient). Examples include, but are not limited to, a tumor sample, a tissue sample, a biopsy sample, a blood sample (e.g., a peripheral whole blood sample), a plasma sample, a serum sample, a lymph sample, a saliva sample, a sputum sample, a urine sample, a gynecological fluid sample, a circulating tumor cell (CTC) sample, a cerebrospinal fluid (CSF) sample, a pericardial fluid sample, a pleural fluid sample, an ascites (peritoneal fluid) sample, a feces (or stool) sample, or other bodily fluid, secretion, and / or excretion sample (or a cell sample derived therefrom). In certain examples, the sample may be a frozen sample or a formalin-fixed paraffin-embedded (FFPE) sample.

[0083] In some examples, samples may be collected by tissue resection (e.g., surgical resection), needle biopsy, bone marrow biopsy, bone marrow aspirate, skin biopsy, endoscopic biopsy, fine needle aspiration, oral swab, nasal swab, vaginal swab, or cytological smear, scraping, lavage or washing (such as luminal lavage or bronchoalveolar lavage), or the like.

[0084] In some examples, the sample is a liquid biopsy sample and may include, for example, whole blood, plasma, serum, urine, stool, sputum, saliva, or cerebrospinal fluid. In some examples, the sample is a liquid biopsy sample and may include circulating tumor cells (CTCs). In some examples, the sample is a liquid biopsy sample and may include cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), or any combination thereof.

[0085] In some examples, the sample may include one or more pre-malignant or malignant cells. As used herein, a pre-malignant tumor refers to a cell or tissue that is not yet malignant, but is preparing to become malignant. In certain examples, the sample may be obtained from a solid tumor, a soft tissue tumor, or a metastatic lesion. In certain examples, the sample may be obtained from a hematological malignancy or a pre-malignancy. In other examples, the sample may include tissue or cells from a surgical margin. In certain examples, the sample may include tumor-infiltrating lymphocytes. In some examples, the sample may include one or more non-malignant cells. In some examples, the sample may be or be part of a primary tumor or metastasis (e.g., a metastasis biopsy sample). In some examples, the sample may be obtained from a site (e.g., a tumor site) with the highest percentage of tumor (e.g., tumor cells) compared to an adjacent site (e.g., a site adjacent to the tumor). In some examples, a sample may be obtained from a site (e.g., a tumor site) with the largest tumor lesion (e.g., the largest number of tumor cells as viewed under a microscope) compared to an adjacent site (e.g., a site adjacent to the tumor).

[0086] In some examples, the disclosed methods may further include analyzing a primary control (e.g., a normal tissue sample). In some examples, the disclosed methods may further include determining whether a primary control is available and, if so, isolating a control nucleic acid (e.g., DNA) from the primary control. In some examples, the sample may include an optional normal control (e.g., normal adjacent tissue (NAT)) if a primary control is not available. In some examples, the sample may be or include histologically normal tissue. In some examples, the method includes evaluating a sample, e.g., a histologically normal sample (e.g., from a surgical tissue margin), using the methods described herein. In some examples, the disclosed methods may further include obtaining a subsample enriched in non-tumor cells, e.g., by macro-dissecting non-tumor tissue from the NAT in a sample without a primary control. In some examples, the disclosed methods may further include determining that a primary control and NAT are unavailable and marking the sample for analysis without a matching control.

[0087] In some examples, samples obtained from histologically normal tissue (e.g., from otherwise histologically normal tissue margins) may still contain genetic alterations, such as variant sequences described herein. Thus, the method may further include reclassifying the sample based on the presence of the detected genetic alteration. In some examples, multiple samples (e.g., from different subjects) are processed simultaneously.

[0088] The disclosed methods and systems may be applied to the analysis of nucleic acids extracted from various tissue samples (or disease states thereof), such as any of solid tissue samples, soft tissue samples, metastatic lesions, or liquid biopsy samples. Examples of tissues include, but are not limited to, connective tissue, muscle tissue, nervous system tissue, epithelial tissue, and blood. Tissue samples may be collected from any organ within an animal or human body. Examples of human organs include, but are not limited to, the brain, heart, lung, liver, kidney, pancreas, spleen, thyroid, breast, uterus, prostate, large intestine, small intestine, bladder, bone, skin, etc.

[0089] In some examples, the nucleic acid extracted from the sample may include deoxyribonucleic acid (DNA) molecules. Examples of DNA that may be suitable for analysis by the disclosed methods include, but are not limited to, genomic DNA or fragments thereof, mitochondrial DNA or fragments thereof, cell-free DNA (cfDNA), and circulating tumor DNA (ctDNA). Cell-free DNA (cfDNA) is composed of fragments of DNA released from normal and / or cancer cells during apoptosis and necrosis, circulating in the bloodstream, and / or accumulating in other body fluids. Circulating tumor DNA (ctDNA) is composed of fragments of DNA released from cancer cells and tumors, circulating in the bloodstream, and / or accumulating in other body fluids.

[0090] In some examples, DNA is extracted from nucleated cells from the sample. In some examples, the sample has low nucleated cellularity, for example, when the sample is composed mainly of red blood cells, diseased cells containing excess cytoplasm, or tissue with fibrosis. In some examples, a sample with low nucleated cellularity may require more, for example, a larger tissue volume, for DNA extraction.

[0091] In some examples, the nucleic acid extracted from the sample may include ribonucleic acid (RNA) molecules. Examples of RNA that may be suitable for analysis by the disclosed methods include, but are not limited to, total cellular RNA, total cellular RNA after depletion of specific abundance RNA sequences (e.g., ribosomal RNA), cell-free RNA (cfRNA), messenger RNA (mRNA) or fragments thereof, poly(A)-tailed mRNA fraction of total RNA, ribosomal RNA (rRNA) or fragments thereof, transfer RNA (tRNA) or fragments thereof, and mitochondrial RNA or fragments thereof. In some examples, RNA may be extracted from the sample and converted to complementary DNA, for example, using a reverse transcription reaction. In some examples, cDNA is produced by random primed cDNA synthesis. In other examples, cDNA synthesis is initiated at the poly(A) tail of mature mRNA by priming with an oligo(dT)-containing oligonucleotide. Methods for depletion, poly(A) enrichment, and cDNA synthesis are well known to those skilled in the art.

[0092] In some examples, the sample may contain tumor content, including, for example, tumor cells or tumor cell nuclei. In some examples, the sample may contain tumor content with at least 5-50%, 10-40%, 15-25%, or 20-30% tumor cell nuclei. In some examples, the sample may contain tumor content of at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, or at least 50% tumor cell nuclei. In some examples, the percentage of tumor nuclei is determined (e.g., calculated) by dividing the number of tumor cells in the sample by the total number of all cells in the sample that have nuclei. In some examples, for example, when the sample is a liver sample containing hepatocytes, a different tumor content calculation may be required due to the presence of hepatocytes with double or more than double nuclei, the presence of other DNA content, e.g., non-hepatocyte, somatic cell nuclei. In some examples, the sensitivity of detection of genetic alterations, e.g., variant sequences, or the sensitivity of determination of, e.g., microsatellite instability, may depend on the tumor content of the sample. For example, a sample with a lower tumor content may result in a lower sensitivity of detection for a sample of a given size.

[0093] In some instances, as described above, the sample contains nucleic acid (e.g., DNA, RNA (or cDNA derived from RNA), or both), e.g., from a tumor or from normal tissue. In certain instances, the sample may further contain non-nucleic acid components, e.g., cells, proteins, carbohydrates, or lipids, e.g., from the tumor or normal tissue.

[0094] subject In some examples, the sample is obtained (e.g., collected) from a subject (e.g., a patient) having or suspected of having a condition or disease (e.g., a hyperproliferative disease or a non-cancer indication). In some examples, the hyperproliferative disease is cancer. In some examples, the cancer is a solid tumor or a metastatic form thereof. In some examples, the cancer is a blood cancer, e.g., leukemia or lymphoma.

[0095] In some examples, the subject has cancer or is at risk of having cancer. For example, in some examples, the subject has a genetic predisposition to cancer (e.g., having a genetic mutation that increases the baseline risk for developing cancer). In some examples, the subject has been exposed to an environmental perturbation (e.g., radiation or chemicals) that increases the risk of developing cancer. In some examples, the subject needs to be monitored for the development of cancer. In some examples, the subject needs to be monitored for progression or regression of cancer, for example, after being treated with an anti-cancer therapy (or anti-cancer treatment). In some examples, the subject needs to be monitored for recurrence of cancer. In some examples, the subject needs to be monitored for minimal residual disease (MRD). In some examples, the subject has been or is being treated for cancer. In some examples, the subject has not been treated with an anti-cancer therapy (or anti-cancer treatment).

[0096] In some examples, a subject (e.g., a patient) is being treated or has been previously treated with one or more targeted therapies. In some examples, for example, for a patient who has been previously treated with targeted therapy, a post-targeted therapy sample (e.g., specimen) is obtained (e.g., collected). In some examples, a post-targeted therapy sample is a sample obtained after completion of targeted therapy.

[0097] In some cases, the patient has not been previously treated with a targeted therapy. In some cases, e.g., for a patient who has not been previously treated with a targeted therapy, the sample includes a resection, e.g., the original resection, or a resection after a recurrence (e.g., after disease recurrence after therapy).

[0098] cancer In some examples, the sample is obtained from a subject with cancer. Exemplary cancers include, but are not limited to, B-cell cancer (e.g., multiple myeloma), melanoma, breast cancer, lung cancer (such as non-small cell lung cancer or NSCLC), bronchial cancer, colorectal cancer, prostate cancer, pancreatic cancer, gastric cancer, ovarian cancer, bladder cancer, brain or central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, uterine or endometrial cancer, oral or pharyngeal cancer, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, and the like. Cancer of the small intestine or adnexa, cancer of the salivary gland, thyroid cancer, adrenal adenocarcinoma, osteosarcoma, chondrosarcoma, cancer of the blood tissue, adenocarcinoma, inflammatory myofibroblastic tumor, gastrointestinal stromal tumor (GIST), colon cancer, multiple myeloma (MM), myelodysplastic syndrome (MDS), myeloproliferative disorder (MPD), acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic myelogenous leukemia (CML), chronic lymphocytic leukemia (CLL), polycythemia vera, Hodgkin's lymphoma, Non-Hodgkin's lymphoma (NHL), soft tissue sarcoma, fibrosarcoma, myxosarcoma, liposarcoma, osteogenic sarcoma, chordoma, angiosarcoma, endothelial sarcoma, synovium, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinoma, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatocellular carcinoma, cholangiocarcinoma, choriocarcinoma, seminoma, embryonal carcinoma, Wilms' tumor, bladder cancer, epithelial carcinoma, glioma, astrocytic These include medulloblastoma, craniopharyngioma, ependymoma, pinealoma, hemangioblastoma, acoustic neuroma, oligodendroglioma, meningioma, neuroblastoma, retinoblastoma, diffuse large B-cell lymphoma, mantle cell lymphoma, hepatocellular carcinoma, thyroid cancer, gastric cancer, head and neck cancer, small cell carcinoma, essential thrombocythemia, agnogenic myeloid metaplasia, hypereosinophilic syndrome, systemic mastocytosis, familial hypereosinophilia, chronic eosinophilic leukemia, neuroendocrine carcinoma, and carcinoma-like tumors.

[0099] In some examples, the cancer is a hematological malignancy (or pre-malignancy). As used herein, hematological malignancy refers to a tumor of hematopoietic or lymphatic tissue, e.g., a tumor affecting the blood, bone marrow, or lymph nodes. Exemplary hematological malignancies include leukemia (e.g., acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), chronic myelogenous leukemia (CML), hairy cell leukemia, acute monocytic leukemia (AMoL), chronic myelomonocytic leukemia (CMML), juvenile myelomonocytic leukemia (JMML), or large granular lymphocytic leukemia), lymphoma (e.g., AIDS-related lymphoma, cutaneous T-cell lymphoma, Hodgkin's lymphoma (e.g., classical Hodgkin's lymphoma or nodal lymphocytic predominant), and / or pulmonary leukemia (PCL). Hodgkin's lymphoma), mycosis fungoides, non-Hodgkin's lymphoma (e.g., B-cell non-Hodgkin's lymphoma (e.g., Burkitt's lymphoma, small lymphocytic lymphoma (CLL / SLL), diffuse large B-cell lymphoma, follicular lymphoma, immunoblastic large cell lymphoma, precursor B-lymphoblastic lymphoma, or mantle cell lymphoma) or T-cell non-Hodgkin's lymphoma (mycosis fungoides, anaplastic large cell lymphoma, or precursor T-lymphoblastic lymphoma)), primary central nervous system.

[0100] Nucleic Acid Extraction and Processing DNA or RNA can be extracted from tissue samples, biopsy samples, blood samples, or other bodily fluid samples using any of a variety of techniques known to those of skill in the art (see, e.g., Example 1 of International Patent Application Publication No. WO 2012 / 092426; Tan, et al. (2009), "DNA, RNA, and Protein Extraction: The Past and The Present", J. Biomed. Biotech. 2009:574398; the technical literature for the Maxwell® 16 LEV Blood DNA Kit (Promega Corporation, Madison, WI); and the Maxwell 16 Buccal Swab LEV DNA Purification Kit Technical Manual (Promega Literature #TM333, January 1, 2011, Promega Corporation, Madison, WI)). Protocols for RNA isolation are disclosed, for example, in the Maxwell® 16 Total RNA Purification Kit Technical Bulletin (Promega Literature #TB351, August 2009, Promega Corporation, Madison, WI).

[0101] A typical DNA extraction procedure includes, for example, (i) collection of a fluid, cell, or tissue sample from which DNA is to be extracted; (ii) if necessary, disruption of cell membranes to release DNA and other cytoplasmic components (i.e., cell lysis); (iii) treatment of the liquid or lysed sample with a concentrated salt solution to precipitate proteins, lipids, and RNA, followed by centrifugation to separate the precipitated proteins, lipids, and RNA; and (iv) purification of the DNA from the supernatant to remove detergents, proteins, salts, or other reagents used during the cell membrane lysis step.

[0102] Disruption of cell membranes can be performed using a variety of mechanical shearing (e.g., French press or fine needle) or ultrasonic disruption techniques. The cell lysis step often involves the use of detergents and surfactants to dissolve lipid, cell and nuclear membranes. In some instances, the lysis step may further include the use of proteases to break down proteins and / or RNases to digest RNA in the sample.

[0103] Examples of suitable techniques for DNA purification include, but are not limited to, (i) precipitation in ice-cold ethanol or isopropanol followed by centrifugation (which may be enhanced by increasing the ionic strength, e.g., by adding sodium acetate, to precipitate the DNA); (ii) phenol-chloroform extraction followed by centrifugation to separate the aqueous phase containing the nucleic acids from the organic phase containing the denatured proteins; and (iii) solid-phase chromatography, in which the nucleic acids are adsorbed to a solid phase (e.g., silica or other) depending on the pH and salt concentration of the buffer.

[0104] In some instances, cellular and histone proteins bound to DNA can be removed by adding proteases or by precipitating the proteins with sodium acetate or ammonium acetate, or through extraction with a phenol-chloroform mixture prior to the DNA precipitation step.

[0105] In some instances, DNA may be extracted using any of a variety of suitable commercially available DNA extraction and purification kits, including, but not limited to, the QIAamp (for isolating genomic DNA from human samples) and DNAeasy (for isolating genomic DNA from animal or plant samples) kits from Qiagen (Germantown, MD), or the Maxwell® and ReliaPrep™ series from Promega (Madison, WI).

[0106] As noted above, in some examples, the sample may include a formalin-fixed (or formaldehyde-fixed, or paraformaldehyde-fixed), paraffin-embedded (FFPE) tissue preparation. For example, an FFPE sample may be a tissue sample embedded in a matrix, such as an FFPE block. Methods for isolating nucleic acids (e.g., DNA) from formaldehyde-fixed or paraformaldehyde-fixed, paraffin-embedded (FFPE) tissues are described, for example, in Cronin, et al., (2004) Am J Pathol. 164(1):35-42; Masuda, et al., (1999) Nucleic Acids Res. 27(22):4436-4443; Specht, et al., (2001) Am J Pathol. 158(2):419-429; Ambion RecoverAll™ Total Nucleic Acid Isolation Protocol (Ambion, Cat. No. AM1975, September 2008); Maxwell® 16 FFPE Plus LEV DNA Purification Kit Technical Manual (Promega Literature #TM349, February 2011); EZNA® FFPE DNA Kit Handbook (OMEGA bio-tek, Norcross, GA, product numbers D3399-00, D3399-01, and D3399-02, June 2009), as well as the QIAamp® DNA FFPE Tissue Handbook (Qiagen, Cat. No. 37625, October 2007). For example, the RecoverAll™ Total Nucleic Acid Isolation Kit uses xylene at high temperature to solubilize paraffin-embedded samples and captures nucleic acids over glass fiber filters. The Maxwell® 16 FFPE Plus LEV DNA Purification Kit is used with the Maxwell® 16 Instrument to purify genomic DNA from 1-10 μm sections of FFPE tissue.DNA is purified using silica-clad paramagnetic particles (PMPs) and eluted in low elution volumes. The EZNA® FFPE DNA Kit uses spin columns and a buffer system for isolation of genomic DNA. The QIAamp® DNA FFPE Tissue Kit uses QIAamp® DNA Micro technology for purification of genomic and mitochondrial DNA.

[0107] In some examples, the disclosed methods may further include determining or obtaining a yield value of the nucleic acid extracted from the sample and comparing the determined value to a reference value. For example, if the determined or obtained value is less than the reference value, the nucleic acid may be amplified before proceeding with library construction. In some examples, the disclosed methods may further include determining or obtaining a value for the size (or average size) of the nucleic acid fragments in the sample and comparing the determined or obtained value to a reference value, for example, a size (or average size) of at least 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 base pairs (bps). In some examples, one or more parameters described herein may be adjusted or selected in response to this determination.

[0108] After isolation, the nucleic acid is typically dissolved in a slightly alkaline buffer, such as Tris-EDTA (TE) buffer, or in ultrapure water. In some examples, the isolated nucleic acid (e.g., genomic DNA) can be fragmented or sheared by using any of a variety of techniques known to those skilled in the art. For example, genomic DNA can be fragmented by physical shearing, enzymatic cleavage, chemical cleavage, and other methods known to those skilled in the art. Methods for DNA shearing are described, for example, in Example 4 of International Patent Application Publication No. 2012 / 092426. In some examples, alternative methods of DNA shearing can be used to avoid the ligation step during library preparation.

[0109] Library preparation In some examples, the nucleic acids isolated from the sample can be used to construct a library (e.g., a nucleic acid library described herein). In some examples, the nucleic acids are fragmented using any of the methods described above, optionally subjected to strand end damage repair, optionally ligated to synthesize adapters, primers, and / or barcodes (e.g., amplification primers, sequencing adapters, flow cell adapters, substrate adapters, sample barcodes or indexes, and / or unique molecular identifier sequences), size selected (e.g., by preparative gel electrophoresis), and / or amplified (e.g., using PCR, non-PCR amplification techniques, or isothermal amplification techniques). In some examples, the fragmented and adapter-ligated nucleic acid population is used without explicit size selection or amplification prior to hybridization-based selection of target sequences. In some examples, the nucleic acids are amplified by any of a variety of specific or non-specific nucleic acid amplification methods well known to those of skill in the art. In some examples, the nucleic acids are amplified by whole genome amplification methods, such as, for example, random primed strand displacement amplification. Examples of nucleic acid library preparation techniques for next-generation sequencing are described, for example, in van Dijk, et al. (2014), Exp. Cell Research 322:12-20, and in Illumina's genomic DNA sample preparation kit.

[0110] In some examples, the resulting nucleic acid library may contain all or substantially all of the complexity of the genome. The term "substantially all" in this context actually refers to the possibility that there may be some undesired loss of genome complexity during the initial steps of the procedure. The methods described herein are also useful when the nucleic acid library is a portion of a genome, for example, when the complexity of the genome is reduced by design. In some examples, any selected portion of the genome may be used with the methods described herein. For example, in certain embodiments, the entire exome or a subset thereof is isolated. In some examples, the library may contain at least 95%, 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, or 5% genomic DNA. In some examples, the library may consist of cDNA copies of genomic DNA that contain at least 95%, 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, or 5% genomic DNA. In certain examples, the amount of nucleic acid used to generate a nucleic acid library can be less than 5 micrograms, less than 1 microgram, less than 500 ng, less than 200 ng, less than 100 ng, less than 50 ng, less than 10 ng, less than 5 ng, or less than 1 ng.

[0111] In some examples, a library (e.g., a nucleic acid library) comprises a collection of nucleic acid molecules. As described herein, the nucleic acid molecules of the library can include target nucleic acid molecules (e.g., tumor nucleic acid molecules, reference nucleic acid molecules, and / or control nucleic acid molecules, also referred to herein as first, second, and / or third nucleic acid molecules, respectively). The nucleic acid molecules of the library can be derived from a single subject or individual. In some examples, a library can include nucleic acid molecules from two or more subjects (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30 or more subjects). For example, two or more libraries from different subjects can be combined to form a library with nucleic acid molecules from two or more subjects (wherein the nucleic acid molecules from each subject are optionally ligated to a unique sample barcode corresponding to the particular subject). In some examples, the subjects are humans who have or are at risk of having cancer or a tumor.

[0112] In some examples, a library (or a portion thereof) may include one or more subgenomic intervals. In some examples, a subgenomic interval may be a single nucleotide position, e.g., a nucleotide position where a variant at that position is associated (positively or negatively) with a tumor phenotype. In some examples, a subgenomic interval includes two or more nucleotide positions. Such examples include a sequence of at least 2, 5, 10, 50, 100, 150, 250, or more than 250 nucleotide positions in length. A subgenomic interval may include, for example, one or more entire genes (or portions thereof), one or more exons or coding sequences (or portions thereof), one or more introns (or portions thereof), one or more microsatellite regions (or portions thereof), or any combination thereof. A subgenomic interval may include all or a portion of a fragment of a naturally occurring nucleic acid molecule, e.g., a genomic DNA molecule. For example, a subgenomic interval may correspond to a fragment of genomic DNA that is subjected to a sequencing reaction. In some examples, a subgenomic interval is a contiguous sequence from a genomic source. In some instances, the subgenomic interval comprises sequences that are not contiguous in the genome, for example, a subgenomic interval in a cDNA may comprise an exon-exon junction formed as a result of splicing. In some instances, the subgenomic interval comprises a tumor nucleic acid molecule. In some instances, the subgenomic interval comprises a non-tumor nucleic acid molecule.

[0113] Targeting loci for analysis The methods described herein can be used in combination with, or as part of, a method for evaluating a set of intervals of interest (e.g., target sequences) from, for example, a set of genomic loci (e.g., loci or fragments thereof), as described herein.

[0114] In some examples, the set of genomic loci evaluated by the disclosed methods includes multiple, e.g., genes, that, in mutant form, are associated with an effect on cell division, proliferation or survival, or are associated with cancer, e.g., a cancer described herein.

[0115] In some examples, the set of loci evaluated by the disclosed methods includes at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, or more than 100 loci.

[0116] In some examples, the selected locus (also referred to herein as a target locus or target sequence) or a fragment thereof may comprise an interval of interest, including a non-coding sequence, a coding sequence, an intragenic region, or an intergenic region of the genome of interest. For example, the interval of interest may include a non-coding sequence or a fragment thereof (e.g., a promoter sequence, an enhancer sequence, a 5' untranslated region (5'UTR), a 3' untranslated region (3'UTR), or a fragment thereof), a coding sequence of a fragment thereof, an exon sequence or a fragment thereof, an intron sequence or a fragment thereof.

[0117] Target Capture Reagents The methods described herein may include contacting a nucleic acid library with a plurality of target capture reagents to select and capture a plurality of specific target sequences (e.g., gene sequences or fragments thereof) for analysis. In some examples, a target capture reagent (i.e., a molecule that binds to a target molecule, thereby allowing capture of the target molecule) is used to select the interval of interest to be analyzed. For example, the target capture reagent may be a bait molecule, e.g., a nucleic acid molecule (e.g., a DNA molecule or an RNA molecule), that can hybridize to (i.e., be complementary to) a target molecule, thereby allowing capture of the target nucleic acid. In some examples, the target capture reagent, e.g., a bait molecule (or bait sequence), is a capture oligonucleotide (or capture probe). In some examples, the target nucleic acid is a genomic DNA molecule, an RNA molecule, a cDNA molecule derived from an RNA molecule, a microsatellite DNA sequence, or the like. In some examples, the target capture reagent is suitable for solution-phase hybridization to the target. In some examples, the target capture reagent is suitable for solid-phase hybridization to the target. In some instances, the target capture reagent is suitable for both solution phase and solid phase hybridization to the target. The design and construction of target capture reagents is described in more detail, for example, in International Patent Application Publication No. WO 2020 / 236941, the entire contents of which are incorporated herein by reference.

[0118] The methods described herein provide optimized sequencing of multiple genomic loci (e.g., genes or gene products (e.g., mRNA), microsatellite loci, etc.) from samples (e.g., cancer tissue specimens, liquid biopsy samples, etc.) from one or more subjects by appropriate selection of target capture reagents to select the target nucleic acid molecules to be sequenced. In some examples, a target capture reagent may hybridize to a specific target locus, e.g., a specific target locus or fragment thereof. In some examples, a target capture reagent may hybridize to a specific group of target loci, e.g., a specific group of loci or fragments thereof. In some examples, multiple target capture reagents may be used, including a mix of target-specific and / or group-specific target capture reagents.

[0119] In some examples, the number of target capture reagents (e.g., bait molecules) in a plurality of target capture reagents (e.g., bait set) contacted with a nucleic acid library to capture a plurality of target sequences for nucleic acid sequencing is greater than 10, greater than 50, greater than 100, greater than 200, greater than 300, greater than 400, greater than 500, greater than 600, greater than 700, greater than 800, greater than 900, greater than 1,000, greater than 1,250, greater than 1,500, greater than 1,750, greater than 2,000, greater than 3,000, greater than 4,000, greater than 5,000, greater than 10,000, greater than 25,000, or greater than 50,000.

[0120] In some examples, the total length of the target capture reagent sequence can be between about 70 nucleotides and 1000 nucleotides. In one example, the length of the target capture reagent is between about 100 and 300 nucleotides, 110 and 200 nucleotides, or 120 and 170 nucleotides in length. In addition to the above, intermediate oligonucleotide lengths of about 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 300, 400, 500, 600, 700, 800, and 900 nucleotides in length can be used in the methods described herein. In some embodiments, oligonucleotides of about 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220 or 230 bases can be used.

[0121] In some examples, each target capture reagent sequence may include (i) a target-specific capture sequence (e.g., a locus- or microsatellite locus-specific complementary sequence), (ii) an adapter, primer, barcode, and / or unique molecular identifier sequence, and (iii) a universal tail at one or both ends. As used herein, the term "target capture reagent" may refer to a target-specific target capture sequence or an entire target capture reagent oligonucleotide that includes a target-specific target capture sequence.

[0122] In some examples, the target-specific capture sequence in the target capture reagent is about 40 nucleotides to 1000 nucleotides in length. In some examples, the target-specific capture sequence is about 70 nucleotides to 300 nucleotides in length. In some examples, the target-specific sequence is about 100 nucleotides to 200 nucleotides in length. In still other examples, the target-specific sequence is about 120 nucleotides to 170 nucleotides in length, typically 120 nucleotides in length. In addition to the above, target-specific sequences of intermediate lengths, for example, about 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 300, 400, 500, 600, 700, 800, and 900 nucleotides in length, as well as target-specific sequences of lengths between the above lengths, can also be used in the methods described herein.

[0123] In some examples, target capture reagent can be designed to select the target section that contains one or more rearrangements, for example, the intron that contains genomic rearrangement.In such examples, target capture reagent is designed to mask repetitive sequences to improve selection efficiency.In these examples where rearrangement has known linking sequence, complementary target capture reagent can be designed to linking sequence to improve selection efficiency.

[0124] In some examples, the disclosed methods may include the use of target capture reagents designed to capture two or more different target categories, with each category having a different target capture reagent design strategy. In some examples, the hybridization-based capture methods and target capture reagent compositions disclosed herein provide capture and homogenous coverage of a target sequence set, while minimizing coverage of genomic sequences outside the targeted sequence set. In some examples, the target sequences may include entire exomes of genomic DNA or selected subsets thereof. In some examples, the target sequences may include, for example, large chromosomal regions (e.g., entire chromosomal arms). The methods and compositions disclosed herein provide different target capture reagents to achieve different sequencing depths and coverage patterns for a complex target nucleic acid sequence set.

[0125] Typically, DNA molecules are used as target capture reagent sequences, but RNA molecules can also be used. In some instances, DNA molecule target capture reagents can be single-stranded DNA (ssDNA) or double-stranded DNA (dsDNA). In some instances, RNA-DNA duplexes are more stable than DNA-DNA duplexes, and therefore potentially provide better capture of nucleic acids.

[0126] In some examples, the disclosed methods include providing a selected set of nucleic acid molecules (e.g., library catch) captured from one or more nucleic acid libraries. For example, the methods can include providing one or more nucleic acid libraries, each containing a plurality of nucleic acid molecules (e.g., a plurality of target nucleic acid molecules and / or reference nucleic acid molecules) extracted from one or more samples from one or more subjects, contacting the one or more libraries (e.g., in a solution-based hybridization reaction) with one, two, three, four, five, or more than five target capture reagents (e.g., oligonucleotide target capture reagents) to form a hybridization mixture containing a plurality of target capture reagent / nucleic acid molecule hybrids, and separating the plurality of target capture reagent / nucleic acid molecule hybrids from the hybridization mixture, e.g., by contacting the hybridization mixture with a binding entity that allows separation of the plurality of target capture reagent / nucleic acid molecule hybrids from the hybridization mixture, thereby providing a library catch (e.g., a subgroup of selected or enriched nucleic acid molecules from one or more libraries).

[0127] In some examples, the disclosed methods can further include amplifying the library catch (e.g., by performing PCR). In other examples, the library catch is not amplified.

[0128] In some examples, the target capture reagent may be part of a kit, which may include instructions, standards, buffers or enzymes or other reagents as needed.

[0129] Hybridization conditions As noted above, the methods disclosed herein can include contacting a library (e.g., a nucleic acid library) with a plurality of target capture reagents to contact selected library target nucleic acid sequences (i.e., library catch). The contacting step can be performed, for example, by solution-based hybridization. In some examples, the method includes repeating the hybridization step for one or more additional solution-based hybridizations. In some examples, the method further includes subjecting the library catch to one or more additional solution-based hybridizations with the same or a different set of target capture reagents.

[0130] In some examples, the contacting step is performed using a solid support, e.g., an array. Suitable solid supports for hybridization are described, for example, in Albert, TJ et al. (2007) Nat. Methods 4(11):903-5, Hodges, E. et al. (2007) Nat. Genet. 39(12):1522-7, and Okou, DT et al. (2007) Nat. Methods 4(11):907-9, the contents of which are incorporated herein by reference in their entirety.

[0131] Hybridization methods that can be adapted for use in the methods herein are described in the art, for example, as described in International Patent Application Publication No. 2012 / 092426. Methods for hybridizing target capture reagents to multiple target nucleic acids are described in more detail, for example, in International Patent Application Publication No. 2020 / 236941, the entire contents of which are incorporated herein by reference.

[0132] Sequencing methods The methods and systems disclosed herein may be used in conjunction with, or as part of, a method or system for sequencing nucleic acids (e.g., a next-generation sequencing system) to generate multiple sequence reads that overlap one or more loci within a subgenomic interval in a sample, thereby, for example, determining gene allele sequences at multiple loci. "Next-generation sequencing" (or "NGS") as used herein may also be referred to as "massively parallel sequencing," which involves the determination of the nucleotide sequence of individual nucleic acid molecules (e.g., in single-molecule sequencing) or clonally amplified proxies of individual nucleic acid molecules in a high-throughput manner (e.g., in 10 3 , 10 4 , 10 5 , or 10 5 This refers to any sequencing method in which multiple molecules are sequenced simultaneously.

[0133] Next generation sequencing methods are known in the art and are described, for example, in Metzker, M. (2010) Nature Biotechnology Reviews 11:31-46, which is incorporated herein by reference. Other examples of sequencing methods suitable for use in implementing the methods and systems disclosed herein are described, for example, in International Patent Application Publication No. 2012 / 092426. In some examples, sequencing may include, for example, whole genome sequencing (WGS), whole exome sequencing, targeted sequencing, or direct sequencing. In some examples, sequencing may be performed using, for example, Sanger sequencing. In some examples, sequencing may include paired-end sequencing techniques that allow both ends of a fragment to be sequenced and generate high quality alignable sequence data for, for example, detection of genomic rearrangements, repetitive sequence elements, gene fusions, and novel transcripts.

[0134] The disclosed methods and systems may be implemented using a sequencing platform, such as Roche 454, Illumina Solexa, ABI-SOLiD, ION Torrent, Complete Genomics, Pacific Bioscience, Helicos, and / or Polonator platforms. In some examples, the sequencing may include Illumina MiSeq sequencing. In some examples, the sequencing may include Illumina HiSeq sequencing. In some examples, the sequencing may include Illumina NovaSeq sequencing. Optimized methods for sequencing multiple target genomic loci in nucleic acid extracted from a sample are described in more detail, for example, in International Patent Application Publication No. WO 2020 / 236941, the entire contents of which are incorporated herein by reference.

[0135] In certain examples, the disclosed methods include the steps of: (a) obtaining a library from a sample comprising a plurality of normal and / or tumor nucleic acid molecules; (b) simultaneously or sequentially contacting the library with a plurality of 1, 2, 3, 4, 5, or more than 5 target capture reagents under conditions that allow hybridization of the target capture reagents to the target nucleic acid molecules, thereby providing a selected set of captured normal and / or tumor nucleic acid molecules (i.e., library catch); and (c) isolating the selected subset of nucleic acid molecules (e.g., library catch) by, for example, contacting the hybridization mixture with a binding entity that allows separation of the target capture reagent / nucleic acid molecule hybrids from the hybridization mixture. (d) sequencing the library catch to obtain a plurality of reads (e.g., sequence reads) that overlap with one or more intervals of interest (e.g., one or more target sequences) from the library catch that may contain mutations (or alterations), e.g., variant sequences that contain somatic or germline mutations; (e) aligning the sequence reads using an alignment method described elsewhere herein; and / or (f) assigning nucleotide values ​​to nucleotide positions within the intervals of interest from one or more sequence reads of the plurality (e.g., calling mutations using a Bayesian method or other method described herein).

[0136] In some examples, obtaining sequence reads for one or more intervals of interest may include sequencing at least 1, at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1,000, at least 1,250, at least 1,500, at least 1,750, at least 2,000, at least 2,250, at least 2,500, at least 2,750, at least 3,000, at least 3,500, at least 4,000, at least 4,500, or at least 5,000 loci, e.g., genomic loci, loci, microsatellite loci, etc. In some examples, obtaining sequence reads for one or more target intervals can include sequencing the target intervals for any number of loci within the ranges described in this paragraph, for example, at least 2,850 loci.

[0137] In some examples, obtaining sequence reads for one or more intervals of interest includes sequencing the intervals of interest using a sequencing method that provides a sequence read length (or average sequence read length) of at least 20 bases, at least 30 bases, at least 40 bases, at least 50 bases, at least 60 bases, at least 70 bases, at least 80 bases, at least 90 bases, at least 100 bases, at least 120 bases, at least 140 bases, at least 160 bases, at least 180 bases, at least 200 bases, at least 220 bases, at least 240 bases, at least 260 bases, at least 280 bases, at least 300 bases, at least 320 bases, at least 340 bases, at least 360 bases, at least 380 bases, or at least 400 bases. In some examples, obtaining sequence reads for one or more intervals of interest can include sequencing the intervals of interest using a sequencing method that provides a sequence read length (or average sequence read length) of any number of bases within the ranges described in this paragraph, for example, a sequence read length (or average sequence read length) of 56 bases.

[0138] In some examples, obtaining sequence reads for one or more target intervals may include sequencing at an average coverage (or depth) of at least 100x or more. In some examples, obtaining sequence reads for one or more target intervals may include sequencing at an average coverage (or depth) of at least 100x, at least 150x, at least 200x, at least 250x, at least 500x, at least 750x, at least 1,000x, at least 1,500x, at least 2,000x, at least 2,500x, at least 3,000x, at least 3,500x, at least 4,000x, at least 4,500x, at least 5,000x, at least 5,500x, or at least 6,000x or more. In some examples, obtaining sequence reads for one or more target intervals may include sequencing at an average coverage (or depth) having any value within the range of values ​​described in this paragraph, for example, at least 160x.

[0139] In some examples, obtaining sequence reads for one or more target intervals includes sequencing at an average sequencing depth having any value ranging from at least 100x to at least 6,000x for more than about 90%, 92%, 94%, 95%, 96%, 97%, 98%, or 99% of the sequenced loci. For example, in some examples, obtaining reads for a target interval includes sequencing at an average sequencing depth of at least 125x for at least 99% of the sequenced loci. As another example, in some examples, obtaining reads for a target interval includes sequencing at an average sequencing depth of at least 4,100x for at least 95% of the sequenced loci.

[0140] In some examples, the relative abundance of nucleic acid species in a library can be estimated by counting the relative number of occurrences of their cognate sequences in the data generated by the sequencing experiment (e.g., the number of sequence reads for a given cognate sequence).

[0141] In some examples, the disclosed methods and systems provide nucleotide sequences for a set of subject intervals (e.g., loci) as described herein. In certain instances, the sequences are provided without the use of methods that include matched normal controls (e.g., wild-type controls) and / or matched tumor controls (e.g., primary vs. metastatic).

[0142] In some examples, as used herein, the level of sequencing depth (e.g., X-fold level of sequencing depth) refers to the number of reads (e.g., unique reads) obtained after detection and removal of duplicate reads (e.g., PCR duplicate reads). In other examples, duplicate reads are evaluated, for example, to aid in the detection of copy number alterations (CNAs).

[0143] alignment Alignment is the process of matching a read to a location, e.g., a genomic location or locus. In some examples, NGS reads can be aligned to a known reference sequence (e.g., a wild-type sequence). In some examples, NGS reads can be de novo assembled. Methods of sequence alignment for NGS reads are described, for example, in Trapnell, C. and Salzberg, SL Nature Biotech., 2009, 27:455-457. Examples of de novo sequence assembly are described, for example, in Warren R. et al., Bioinformatics, 2007, 23:500-501, Butler J. et al., Genome Res., 2008, 18:810-820, and Zerbino DR and Birney E., Genome Res., 2008, 18:821-829. Optimization of sequence alignment is described in the art, for example, as described in International Patent Application Publication No. 2012 / 092426. Additional description of sequence alignment methods is described in more detail, for example, in International Patent Application Publication No. 2020 / 236941, the entire contents of which are incorporated herein by reference.

[0144] Misalignment (e.g., placement of base pairs from short reads at incorrect locations in the genome), for example, misalignment of reads due to sequence context (e.g., presence of repetitive sequences) around the actual cancer mutation can lead to reduced sensitivity of mutation detection because alternate allele reads can be shifted from the histogram peak of alternate allele reads. Other examples of sequence contexts that can cause misalignment include short tandem repeats, interspersed repetitive sequences, low complexity regions, insertion-deletions (indels), and paralogs. When problematic sequence situations arise in the absence of actual mutations, misalignment can introduce artifactual reads of "mutated" alleles by misplacing reads of the actual reference genome base sequence. Because mutation calling algorithms for multigene analysis must be sensitive to even low abundance mutations, sequence misalignment can increase false positive discovery rates and / or reduce specificity.

[0145] In some examples, the methods and systems disclosed herein may integrate the use of multiple individually tailored alignment methods or algorithms to optimize base calling performance in sequencing methods, particularly those that rely on massively parallel sequencing of multiple diverse genetic events at multiple diverse genomic loci. In some examples, the disclosed methods and systems may include the use of one or more global alignment algorithms. In some examples, the disclosed methods and systems may include the use of one or more local alignment algorithms.Examples of alignment algorithms that may be used include, but are not limited to, the Burrows-Wheeler Alignment (BWA) software bundle (see, e.g., Li, et al. (2009), "Fast and Accurate Short Read Alignment with Burrows-Wheeler Transform", Bioinformatics 25:1754-60; Li, et al. (2010), Fast and Accurate Long-Read Alignment with Burrows-Wheeler Transform", Bioinformatics epub. PMID:20080505), the Smith-Waterman algorithm (see, e.g., Smith, et al. (1981), "Identification of Common Molecular Subsequences", J. Molecular Biology 147(1):195-197), the Stripped Smith-Waterman algorithm (see, e.g., Farrar (2007), "Striped Smith-Waterman Speeds Database Searches Six Times Over Other SIMD Implementations”, Bioinformatics 23(2):156-161), the Needleman-Wunsch algorithm (Needleman, et al. (1970) “A General Method Applicable to the Search for Similarities in the Amino Acid Sequence of Two Proteins”, J. Molecular Biology 48(3):443-53), or any combination thereof.

[0146] In some examples, the methods and systems disclosed herein may also include the use of a sequence assembly algorithm, such as the Arachne sequencing assembly algorithm (see, e.g., Batzoglou, et al. (2002), "ARACHNE: A Whole-Genome Shotgun Assembler", Genome Res. 12:177-189).

[0147] In some examples, the alignment method used to analyze the sequence reads is not individually customized or adjusted for the detection of different variants (e.g., point mutations, insertions, deletions, etc.) at different genomic loci. In some examples, different alignment methods are used to analyze the reads that are individually customized or adjusted for the detection of at least a subset of the different variants detected at different genomic loci. In some examples, different alignment methods are used to analyze the reads that are individually customized or adjusted for the detection of each different variant at different genomic loci. In some examples, the adjustments can be a function of one or more of: (i) the locus to be sequenced (e.g., a gene locus, a microsatellite locus, or other interval of interest), (ii) the tumor type associated with the sample, (iii) the variant to be sequenced, or (iv) the characteristics of the sample or subject. The selection or use of alignment conditions that are individually adjusted to several specific intervals of interest to be sequenced allows for the optimization of speed, sensitivity, and specificity. This method is particularly effective when the alignment of reads to a relatively large number of diverse intervals of interest is optimized. In some instances, the methods include the use of an alignment method optimized for the rearrangement in combination with another alignment method optimized for a target interval not associated with the rearrangement.

[0148] In some examples, the methods disclosed herein further include selecting or using an alignment method for analyzing, e.g., aligning, the sequence reads, where the alignment method is a function of, selected in response to, or optimized for one or more of: (i) the tumor type, e.g., the tumor type in the sample; (ii) the location (e.g., locus) of the interval of interest to be sequenced; (iii) the type of variant (e.g., point mutation, insertion, deletion, substitution, copy number variation (CNV), rearrangement, or fusion) within the interval of interest to be sequenced; (iv) the site (e.g., nucleotide position) being analyzed; (v) the type of sample (e.g., a sample described herein); and / or (vi) the flanking sequences within or near the interval of interest being evaluated (e.g., according to their expected propensity for misalignment of the interval of interest due to the presence of repetitive sequences within or near the interval of interest).

[0149] In some examples, the methods disclosed herein allow for fast and efficient alignment of troublesome reads, e.g., reads with rearrangements.Thus, in some examples where the reads for the target interval include nucleotide positions with rearrangements, e.g., translocations, the method can include using an alignment method that is appropriately adjusted and includes: (i) selecting a rearrangement reference sequence for alignment with the read, where the rearrangement reference sequence aligns with the rearrangement (in some examples, the reference sequence is not identical to the genomic rearrangement), and (ii) comparing, e.g., aligning, the read with the rearrangement reference sequence.

[0150] In some examples, alternative methods can be used to align problematic reads. These methods are particularly effective when alignment of reads to a relatively large number of diverse target intervals is optimized. As an example, a method of analyzing a sample can include (i) performing a comparison (e.g., alignment comparison) of reads using a first set of parameters (e.g., using a first mapping algorithm or by comparison with a first reference sequence) to determine whether the reads meet a first alignment criterion (e.g., the reads can be aligned with the first reference sequence, e.g., with less than a certain number of mismatches); and (ii) performing a second alignment comparison using a second set of parameters if the reads do not meet the first alignment criterion (e.g., performing a second alignment comparison). (e.g., using a mapping algorithm or by comparison to a second reference sequence); and (iii) optionally determining whether the read satisfies a second criterion (e.g., the read can be aligned with the second reference sequence, e.g., with less than a certain number of mismatches), including use of the second reference sequence, where the second parameter set is more likely to result in alignment of the read to a variant (e.g., a rearrangement, insertion, deletion, or translocation), e.g., compared to the first parameter set.

[0151] In some examples, alignment of sequence reads in the disclosed methods may be combined with mutation calling methods described elsewhere herein. As discussed herein, reduced sensitivity for detecting actual mutations may be addressed by assessing (manually or in an automated manner) the quality of the alignment around the expected mutation site of the gene or genomic locus (e.g., locus) being analyzed. In some examples, the site to be assessed may be obtained from a database of human genomes (e.g., HG19 human reference genome) or cancer mutations (e.g., COSMIC). Regions identified as problematic may be repaired, for example, by alignment optimization (or realignment) using a slower but more accurate alignment algorithm such as Smith-Waterman alignment, using an algorithm selected to give better performance in the relevant sequence context. In cases where general alignment algorithms cannot ameliorate the problem, customized alignment approaches can be created by, for example, adjusting different maximum mismatch penalty parameters for genes that are likely to contain substitutions, adjusting specific mismatch penalty parameters based on specific mutation types that are common to particular tumor types (e.g., C→T in melanoma), or adjusting specific mismatch penalty parameters based on specific mutation types that are common to certain sample types (e.g., substitutions that are common to FFPE).

[0152] The loss of specificity (increased false positive rate) of the evaluated target interval due to misalignment can be assessed by manual or automated inspection of all variant calls in the sequencing data. Regions found to be prone to false variant calls due to misalignment can be subjected to alignment improvements as discussed above. If algorithmic improvements are not possible, "variations" from problematic regions can be sorted or screened from a panel of target loci.

[0153] Mutation Call Base calling refers to the raw output of a sequencing device, e.g., the determined sequence of nucleotides in an oligonucleotide molecule. Mutation calling refers to the process of selecting a nucleotide value, e.g., A, G, T, or C, for a given nucleotide position being sequenced. Typically, a sequence read (or base call) for a position will provide more than one value, e.g., some reads will indicate T and some will indicate G. Mutation calling is the process of assigning the correct nucleotide value, e.g., one of those values, to a sequence. Although referred to as a "mutation" call, it can be applied to assign a nucleotide value to any nucleotide position, e.g., a position corresponding to a mutant allele, a wild type allele, an allele not characterized as mutant or wild type, or a position not characterized by variability.

[0154] In some examples, the disclosed methods may include the use of customized or tuned variant calling algorithms or parameters to optimize performance when applied to sequencing data, particularly in methods that rely on massively parallel sequencing of multiple diverse genetic events at multiple diverse genomic loci (e.g., gene loci, microsatellite regions, etc.) in a sample, e.g., from a subject with cancer. Optimization of variant calling has been described in the art, e.g., as described in International Patent Application Publication No. WO 2012 / 092426.

[0155] Methods for variant calling may include one or more of the following: making independent calls based on information at each position in the reference sequence (e.g., examining sequence reads; examining base calls and quality scores; calculating the probability of observed bases and quality scores given potential genotypes; and assigning genotypes (e.g., using Bayes' rule); removing false positives (e.g., using a depth threshold to reject SNPs with read depths much lower or higher than expected; local refinement to remove false positives due to small indels); and performing linkage disequilibrium (LD) / imputation-based analysis to refine the calls.

[0156] The formula used to calculate the genotype likelihood associated with a particular genotype and location is described, for example, in Li H. and Durbin R. Bioinformatics, 2010;26(5):589-95. A priori prediction for a particular mutation in a particular cancer type can be used when evaluating samples from that cancer type. Such likelihoods can be obtained from public databases of cancer mutations, such as Catalogue of Somatic Mutation in Cancer (COSMIC), HGMD (Human Gene Mutation Database), The SNP Consortium, Breast Cancer Mutation Data Base (BIC) and Breast Cancer Gene Database (BCGD).

[0157] Examples of LD / complementation based analyses are described, for example, in Browning, B. L. and Yu, Z. Am. J. Hum. Genet. 2009, 85(6):847-61. Examples of low coverage SNP calling methods are described, for example, in Li, Y., et al., Annu. Rev. Genomics Hum. Genet. 2009, 10:387-406.

[0158] After alignment, detection of substitutions can be performed using a calling method (e.g., a Bayesian mutation calling method), which is applied to each base in each of the target intervals, e.g., exons of the gene or other locus being evaluated, and the presence of alternative alleles is observed. This method compares the probability of observing read data in the presence of a mutation to the probability of observing read data in the presence of base calling errors alone. If this comparison is sufficiently strong in support of the presence of a mutation, the mutation can be called.

[0159] The advantage of a Bayesian mutation detection approach is that the comparison of the probability of the presence of a mutation versus the probability of a base calling error alone can be weighted by a prior expectation of the presence of the mutation at that site. If several reads of alternative alleles are observed at a site that is frequently mutated for a given cancer type, the presence of the mutation can be reliably called even if the amount of evidence for the mutation does not meet the usual threshold. This flexibility can then be used to increase the detection sensitivity of rarer mutations / lower purity samples or to make the test more robust against reduced read coverage. The likelihood of a random base pair in the genome being mutated in cancer is approximately 1e-6. For example, the likelihood of a specific mutation occurring at many sites in a typical polygenic cancer genome panel can be orders of magnitude higher. These likelihoods can be derived from public databases of cancer mutations (e.g., COSMIC).

[0160] Indel calling is the process of finding bases in sequencing data that differ from a reference sequence by insertion or deletion, typically with associated confidence score or statistical evidence indicator. Methods of indel calling can include identifying candidate indels, calculating genotype likelihood by local realignment, and performing LD-based genotype inference and calling. Typically, a Bayesian method is used to obtain potential indel candidates, and then these candidates are tested with reference sequences in a Bayesian framework.

[0161] Algorithms for generating candidate indels are described, for example, in McKenna, A., et al., Genome Res. 2010; 20(9): 1297-303, Ye, K., et al., Bioinformatics, 2009; 25(21): 2865-71, Lunter, G., and Goodson, M., Genome Res. 2011; 21(6): 936-9, and Li, H., et al. (2009), Bioinformatics 25(16): 2078-9.

[0162] Methods for generating indel calling and individual-level genotype likelihood include, for example, the Dindel algorithm (Albers CA et al., Genome Res. 2011; 21(6): 961-73). For example, a Bayesian EM algorithm can be used to analyze reads, make initial indel calls, and generate genotype likelihoods for each candidate indel, followed by imputing genotypes using, for example, QCALL (Le SQ and Durbin R. Genome Res. 2011; 21(6): 952-60). Parameters such as prior expectation of observing an indel can be adjusted (e.g., increased or decreased) based on the size or location of the indel.

[0163] Methods have been developed that address limited deviations from 50% or 100% allele frequency for the analysis of cancer DNA. (See, for example, SNVMix - Bioinformatics. 2010 March 15; 26(6): 730-736.) However, the methods disclosed herein allow for frequencies (or allele fractions) ranging from 1% to 100% (i.e., allele fractions ranging from 0.01 to 1.0), and in particular, consideration of the possible presence of mutant alleles at levels below 50%. This approach is particularly important for the detection of mutations in, for example, low-purity FFPE samples of native (multiclonal) tumor DNA.

[0164] In some examples, the mutation calling method used to analyze the sequence reads is not individually customized or tuned for the detection of different mutations at different genomic loci. In some examples, a different mutation calling method is used that is individually customized or fine-tuned for at least a subset of the different mutations detected at different genomic loci. In some examples, a different mutation calling method is used that is individually customized or fine-tuned for each different mutation detected at each different genomic locus. The customization or adjustment can be based on one or more of the factors described herein, such as the type of cancer in the sample, the gene or locus in which the target interval to be sequenced is located, or the variant to be sequenced. This selection or use of a mutation calling method that is individually customized or fine-tuned for the number of target intervals to be sequenced allows for optimization of the speed, sensitivity, and specificity of mutation calling.

[0165] In some examples, a nucleotide value is assigned to a nucleotide position of each of the X unique target intervals using a unique variant calling method, where X is at least 2, at least 3, at least 4, at least 5, at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 1000, at least 1500, at least 2000, at least 2500, at least 3000, at least 3500, at least 4000, at least 4500, at least 5000 or more. The calling methods can be different and thereby unique, for example, by relying on different Bayesian priors.

[0166] In some instances, assigning the nucleotide value is a function of a value that is or represents an expected value prior to observing (e.g., in the literature) reads that exhibit a variant, e.g., mutation, at that nucleotide position in a tumor of the type.

[0167] In some examples, the methods include assigning nucleotide values ​​(e.g., mutation calls) for at least 10, 20, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1,000 nucleotide positions, where each assignment is a function of a unique (as opposed to other assigned values) value that is or represents an expected value prior (e.g., in the literature) of observing reads that indicate a variant, e.g., mutation, at that nucleotide position in a tumor of the type.

[0168] In some examples, assigning the nucleotide value is a function of a set of values ​​that represent the probability of observing a read that exhibits the variant at that nucleotide position if the variant is present in the sample at a particular frequency (e.g., 1%, 5%, 10%, etc.) and / or if the variant is not present (e.g., observed in the read due to base calling errors only).

[0169] In some examples, a mutation calling method described herein may include: (a) obtaining, for a nucleotide position in each of the X intervals of interest, (i) a first value that is or represents an expectation prior to (e.g., literature) observing a read exhibiting a variant, e.g., a mutation, at that nucleotide position in a tumor of type X; and (ii) a second set of values ​​that represent the likelihood of observing a read exhibiting the variant at that nucleotide position if the variant is present in the sample at a certain frequency (e.g., 1%, 5%, 10%, etc.) and / or if the variant is not present (e.g., observed in a read due to base calling error alone); and (b) in response to the values, assigning to each of the nucleotide positions a nucleotide value from the read (e.g., calling a mutation) by weighting a comparison between the values ​​in the second set using the first value, e.g., by a Bayesian method described herein, thereby analyzing the sample.

[0170] Further description of variant calling methods is described in more detail, for example, in International Patent Application Publication No. WO 2020 / 236941, the entire contents of which are incorporated herein by reference.

[0171] system Also disclosed herein is a system designed to implement any of the disclosed methods for predicting the likelihood of success for performing GP analysis on a sample from a subject.The system may, for example, comprise one or more processors, and a memory communicatively coupled to the one or more processors and configured to store instructions, which, when executed by the one or more processors, cause the system to receive data of a plurality of pre-analysis variables associated with a sample from a subject, apply the received data to a multivariate model trained to predict the outcome for GP assay, generate a prediction of the likelihood of success for performing GP on the sample based on the applied multivariate model, and report the prediction of the likelihood of success for performing GP on the sample.

[0172] In some examples, the disclosed systems may further include a sequencer, such as a next-generation sequencer (also called a massively parallel sequencer). Examples of next-generation (or massively parallel) sequencing platforms include, but are not limited to, Roche 454, Illumina Solexa, ABI-SOLiD, ION Torrent, or Pacific Bioscience sequencing platforms.

[0173] In some examples, the disclosed systems may be used to predict the likelihood of success for performing a GP analysis in any of the various samples described herein (e.g., a tissue sample, a biopsy sample, a blood sample, or a liquid biopsy sample from a subject).

[0174] In some examples, the nucleic acid sequence data is obtained using next generation sequencing technology (also called massively parallel sequencing technology) with a read length of less than 400 bases, less than 300 bases, less than 200 bases, less than 150 bases, less than 100 bases, less than 90 bases, less than 80 bases, less than 70 bases, less than 60 bases, less than 50 bases, less than 40 bases, or less than 30 bases.

[0175] In some examples, a determination of a successful GP assay result may be used to select, initiate, adjust, or terminate treatment for cancer in the subject (e.g., patient) from which the sample was derived, as described elsewhere herein.

[0176] In some cases, the disclosed systems may further include sample processing and library preparation workstations, microplate handling robots, fluid dispensing systems, temperature control modules, environmental control chambers, additional data storage modules, data communication modules (e.g., Bluetooth, WiFi, intranet, or internet communication hardware and associated software), display modules, one or more local and / or cloud-based software packages (e.g., instrument / system control software packages, sequencing data analysis software packages), etc., or any combination thereof. In some cases, the systems may include or be part of a computer system or computer network described elsewhere herein.

[0177] Computer Systems and Networks FIG. 3 illustrates an example of a computing device or system according to an embodiment. The device 300 can be a host computer connected to a network. The device 300 can be a client computer or a server. As shown in FIG. 3, the device 300 can be any suitable type of microprocessor-based device, such as a personal computer, a workstation, a server, or a handheld computing device (portable electronic device, e.g., a phone or tablet). The device can include, for example, one or more processors 310, input devices 320, output devices 330, memory or storage devices 340, communication devices 360, and a nucleic acid sequencer 370. The software 350 resident in the memory or storage devices 340 can include, for example, an operating system, and software for implementing the methods described herein. The input devices 320 and output devices 330 can generally correspond to those described herein and can be connectable to or integrated with the computer.

[0178] The input device 320 may be any suitable device that provides input, such as a touch screen, a keyboard or keypad, a mouse, or a voice recognition device. The output device 330 may be any suitable device that provides output, such as a touch screen, a tactile device, or a speaker.

[0179] Storage 340 may be any suitable device that provides storage (e.g., electrical, magnetic, or optical memory, including RAM (volatile and non-volatile), cache, a hard drive, or a removable storage disk). Communications device 360 ​​may include any suitable device that can send and receive signals over a network, such as a network interface chip or device. The components of the computer may be connected in any suitable manner, for example, via wired media (e.g., a physical system bus 380, an Ethernet connection, or any other wired transmission technology) or wirelessly (e.g., Bluetooth, Wi-Fi, or any other wireless technology).

[0180] The software modules 350 can be stored as executable instructions in the storage 340 and executed by the processor 310 and can include, for example, an operating system and / or processes that embody the functionality of the methods of the present disclosure (e.g., embodied in the devices described above).

[0181] The software module 350 may also be stored and / or transferred in any non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device (e.g., those described herein) that may fetch instructions associated with the software from and execute the instructions. In the context of the present disclosure, a computer-readable storage medium may be any medium, such as the storage 340, that may include or store processes for use by or in connection with an instruction execution system, apparatus, or device. Examples of computer-readable storage media may include memory units, such as hard drives, flash drives, and distribution modules that operate as a single functional unit. Also, the various processes described herein may be embodied as modules configured to operate according to the above embodiments and techniques. Moreover, while the processes may be shown and / or described separately, one skilled in the art will understand that the above processes may be routines or modules within other processes.

[0182] The software module 350 may also be propagated in any transmission medium for use by or in connection with an instruction execution system, apparatus, or device such as those mentioned above, which may fetch instructions associated with the software from and execute the instructions. In the context of this disclosure, a transmission medium may be any medium that may communicate, propagate, or transmit programming for use by or in connection with an instruction execution system, apparatus, or device. Transmission-readable media may include, but are not limited to, electronic, magnetic, optical, electromagnetic, or infrared wired or wireless propagation media.

[0183] The device 300 may be connected to a network (e.g., network 404, shown in FIG. 4 and / or described below), which may be any suitable type of interconnected communications system. The network may implement any suitable communications protocol and may be protected by any suitable security protocol. The network may include any suitable arrangement of network links that may implement transmission and reception of network signals, such as wireless network connections (T1 or T3 lines), cable networks, DSL, or telephone lines.

[0184] Device 300 may be implemented using any operating system, e.g., an operating system suitable for operating on a network. Software modules 350 may be written in any suitable programming language, such as C, C++, Java, or Python. In various embodiments, application software embodying functionality of the present disclosure may be deployed in different configurations (e.g., in a client / server arrangement, or via a web browser as a web-based application or web service). In some embodiments, the operating system is executed by one or more processors, e.g., processor 310.

[0185] The device 300 may further include a sequencer 370, which may be any suitable nucleic acid sequencing instrument.

[0186] 4 illustrates an example of a computing system according to one embodiment. In system 400, device 300 (e.g., as described above and illustrated in FIG. 3) is connected to a network 404, which is also connected to device 406. In some embodiments, device 406 is a sequencer. Exemplary sequencers include, but are not limited to, Roche / 454's Genome Sequencer (GS) FLX System, Illumina / Solexa's Genome Analyzer (GA), Illumina's HiSeq 2500, HiSeq 3000, HiSeq 4000, and NovaSeq 6000 sequencing systems, Life / APG's Support Oligonucleotide Ligation Detection (SOLiD) system, Polonator's G.007 system, HeliScope Gene sequencing system from Helicos BioSciences, or Pacific Biosciences' PacBio RS system.

[0187] The devices 300 and 406 can communicate using a suitable communication interface over a network 404, such as, for example, a local area network (LAN), a virtual private network (VPN), or the Internet. In some embodiments, the network 404 can be, for example, the Internet, an intranet, a virtual private network, a cloud network, a wired network, or a wireless network. The devices 300 and 406 can communicate partially or entirely over wireless or wired communications, such as Ethernet, IEEE 802.11b wireless, etc. Additionally, the devices 300 and 406 can communicate over a second network, such as, for example, a mobile / cellular network, using a suitable communication interface. The communication between the devices 300 and 406 can further include or communicate with various servers, such as a mail server, a mobile server, a media server, a telephone server, etc. In some embodiments, the devices 300 and 406 can communicate directly (instead of or in addition to communication over the network 404), such as, for example, Ethernet, IEEE 802.11b wireless, etc., or wired communications. In some embodiments, devices 300 and 406 communicate via a communication 408, which may be a direct connection or may occur over a network (eg, network 404).

[0188] One or all of devices 300 and 406 generally include logic (e.g., http web server logic) or are programmed to format data accessed from local or remote databases or other sources of data and content to provide and / or receive information over network 404 in accordance with various examples described herein.

[0189] Exemplary Implementations Exemplary implementations of the methods and systems described herein include the following. 1. A method for predicting the likelihood of success for performing genomic profiling of a sample from a subject, comprising: receiving, using one or more processors, data for a plurality of pre-analytical variables associated with the sample; applying, using one or more processors, the received data to a multivariate model trained to predict outcomes for the genomic profiling assay; generating, using the one or more processors, a prediction of the likelihood of success for genomic profiling of the sample based on the applied multivariate model; and reporting, using the one or more processors, a prediction of the likelihood of success for genomic profiling of the sample. 2. Providing a plurality of nucleic acid molecules obtained from a sample from a subject based on a prediction of likelihood of success that is equal to or greater than a predetermined threshold; ligating one or more adaptors onto one or more nucleic acid molecules from the plurality of nucleic acid molecules; amplifying one or more ligated nucleic acid molecules from the plurality of nucleic acid molecules; capturing the amplified nucleic acid molecules from the amplified nucleic acid molecules; sequencing the captured nucleic acid molecules with a sequencer to obtain a plurality of sequence reads that represent the captured nucleic acid molecules and overlap with one or more loci within one or more subgenomic intervals in the sample; 2. The method of claim 1, further comprising generating, by the one or more processors, a genomic profile based on the sequence reads of the sample, the genomic profile comprising sequence read analysis data. 3. The method of claim 1 or 2, further comprising training, by one or more processors, the multivariate model using the training data. 4. The method of claim 3, wherein the training data includes data derived from univariate analysis of clinical research data of samples collected from subjects representing subject age range, subject gender, diagnosis, stage of disease, sample type, sample collection site, sample collection method, sample preparation method, sample storage method, sample age, sample transportation method, or any combination thereof. 5. The method of clause 4, wherein the training data further comprises data on ECOG status, subject treatment status, tumor cellularity in the sample, tumor nuclear content in the sample, tissue surface area in the sample, tissue matrix in the sample, previous genomic profiling assay results, or any combination thereof. 6. The method of any one of clauses 1-5, wherein the plurality of pre-analytical variables comprises subject age, subject sex, diagnosis, stage of disease, sample type, sample collection site, sample collection method, sample preparation method, sample storage method, sample age, sample transportation method, or any combination thereof. 7. The method of any one of clauses 1-6, wherein data for a plurality of pre-analytical variables is added with data for tumor cellularity in the sample, tumor nuclear content in the sample, tissue surface area in the sample, tissue matrix in the sample, or any combination thereof. 8. The method of any one of clauses 1-7, wherein the multivariate model comprises a machine learning model. 9. The method of claim 8, wherein the machine learning model comprises a supervised learning model. 10. The method of claim 8, wherein the machine learning model comprises an unsupervised learning model. 11. The method of any one of clauses 1-10, wherein the multivariate model comprises a logistic regression model, a multiple linear regression model, a random forest model, a neural network model, or a deep learning model. 12. The method of any one of clauses 1-11, wherein the prediction of the likelihood of success comprises a binary value, a percentage, or a score. 13. Comparing the predicted likelihood of success to a predetermined threshold; and 12. The method of any one of clauses 1-11, further comprising outputting an indication that the sample from the subject is suitable for providing a genomic profile of the subject based on a determination that the predicted likelihood of success is equal to or greater than a predetermined threshold. 14. Comparing the predicted likelihood of success to a predetermined threshold; and 12. The method of any one of clauses 1-11, further comprising: outputting a recommendation to collect a new sample instead of submitting the sample for genomic profiling based on a determination that the predicted likelihood of success is less than a predetermined threshold. 15. The method of claim 14, further comprising outputting a sample type or sample collection site recommendation for the new sample. 16. The method of claim 14 or 15, further comprising outputting a recommendation to perform an alternative nucleic acid sequencing-based testing method. 17. The method of any one of clauses 2-16, wherein the pre-determined threshold varies depending on the sample type. 18. The method of any one of clauses 1-17, wherein data for a plurality of pre-analytical variables is entered by a user via a graphical user interface (GUI) on a display device. 19. The method of any one of clauses 1-18, wherein the prediction of the likelihood of success for performing genomic profiling of the sample is reported via a graphical user interface (GUI) on a display device. 20. The method of claim 18 or 19, wherein the graphical user interface (GUI) is displayed in a web browser. 21. The method of any one of clauses 1-20, wherein the data received for a plurality of pre-analytical variables includes sample type data. 22. The method of claim 21, wherein the remaining pre-analytical variables of the plurality of pre-analytical variables are selected based on the sample type. 23. The method of claim 21 or 22, wherein the multivariate model is selected based on sample type. 24. The method of any one of clauses 1-23, wherein the sample comprises a tissue biopsy sample, a liquid biopsy sample, or a normal control. 25. The method of clause 24, wherein the sample is a tissue biopsy sample and comprises bone marrow. 26. The method of clause 24, wherein the sample is a liquid biopsy sample and comprises blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva. 27. The method of clause 24, wherein the sample is a liquid biopsy sample and comprises circulating tumor cells (CTCs). 28. The method of clause 24, wherein the sample is a liquid biopsy sample and contains cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), or any combination thereof. 29. The method of any one of clauses 13 to 28, wherein if the predicted likelihood of success is equal to or greater than a predetermined threshold, then genomic profiling is performed and used to diagnose or confirm a diagnosis of a disease in the subject. 30. The method of claim 29, wherein genomic profiling is also used to determine eligibility for therapy based on biomarker status. 31. The method of clause 29 or clause 30, wherein the disease is cancer. 32. The method of claim 31, further comprising selecting an anti-cancer therapy for administration to the subject based on the results of the genomic profiling. 33. The method of clause 31 or clause 32, further comprising determining an effective amount of an anti-cancer therapy to administer to the subject based on the results of the genomic profiling. 34. The method of claim 33, further comprising administering an anti-cancer therapy to the subject based on the results of the genomic profiling. 35. The method of any one of clauses 32-34, wherein the anti-cancer therapy comprises chemotherapy, radiation therapy, immunotherapy, targeted therapy, or surgery. 36. Cancer: B-cell cancer (multiple myeloma), melanoma, breast cancer, lung cancer, bronchial cancer, colorectal cancer, prostate cancer, pancreatic cancer, stomach cancer, ovarian cancer, bladder cancer, brain cancer, central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, endometrial cancer, oral cancer, pharyngeal cancer, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, small intestine cancer, appendix cancer, salivary gland cancer, thyroid cancer, adrenal gland cancer, osteosarcoma, chondrosarcoma, cancer of the blood tissue, and glandular cancer. Cancer, Inflammatory myofibroblastoma, Gastrointestinal stromal tumor (GIST), Colon cancer, Multiple myeloma (MM), Myelodysplastic syndrome (MDS), Myeloproliferative disorder (MPD), Acute lymphocytic leukemia (ALL), Acute myeloid leukemia (AML), Chronic myeloid leukemia (CML), Chronic lymphocytic leukemia (CLL), Polycythemia vera, Hodgkin's lymphoma, Non-Hodgkin's lymphoma (NHL), Soft tissue sarcoma, Fibrosarcoma, Myxosarcoma, Liposarcoma, Osteosarcoma, Chordoma, Blood Angiosarcoma, endothelial sarcoma, lymphangiosarcoma, lymphangioendothelial sarcoma, synovium, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinoma, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatic carcinoma, bile duct carcinoma, choriocarcinoma, seminoma, embryonal carcinoma, Wilms' tumor, bladder carcinoma, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pinealocytoma, hemangioblastoma 36. The method of any one of clauses 31 to 35, wherein the cancer is idiopathic leukemia, leukemia, leukemia, thyroid cancer, thyroid cancer, gastric cancer, head and neck cancer, small cell carcinoma, essential thrombocythemia, primary myelofibrosis, hypereosinophilic syndrome, systemic mastocytosis, familial eosinophilia, chronic eosinophilic leukemia, neuroendocrine carcinoma, or carcinoid tumor. 37. The method of any one of clauses 29 to 36, wherein the genomic profiling of the subject comprises obtaining results from a comprehensive genomic profiling (CGP) test, a gene expression profiling test, a cancer hotspot panel test, a DNA methylation test, a DNA fragmentation test, an RNA fragmentation test, or any combination thereof. 38. The method of claim 37, wherein the genomic profiling of the subject further comprises obtaining results from a nucleic acid sequencing-based test. 39. The method of clause 37 or clause 38, further comprising selecting an anti-cancer drug, administering an anti-cancer drug, or administering an anti-cancer treatment to the subject based on the results of the genomic profiling. 40. A system comprising: one or more processors; a memory communicatively coupled to the one or more processors and configured to store instructions, the instructions, when executed by the one or more processors, to provide a system with: receiving data for a plurality of pre-analytical variables associated with a sample derived from a subject; applying the received data to a multivariate model trained to predict outcomes for the genomic profiling assay; generating a prediction of the likelihood of success for performing genomic profiling of the sample based on the applied multivariate model; A system comprising: a memory that reports a prediction of the likelihood of success for performing genomic profiling of a sample. 41. The system of clause 40, wherein the instructions further cause the system to compare the predicted likelihood of success to a predetermined threshold and, based on a determination that the predicted likelihood of success is less than the predetermined threshold, output a recommendation to collect a new sample instead of submitting the sample for genomic profiling. 42. The system of claim 40 or 41, wherein data for a plurality of pre-analytical variables is entered by a user via a graphical user interface (GUI) on a display device. 43. The system of any one of clauses 40 to 42, wherein a prediction of the likelihood of success for performing genomic profiling of the sample is reported via a graphical user interface (GUI) on a display device. 44. The system of claim 42 or 43, wherein the graphical user interface (GUI) is displayed in a web browser. 45. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of the system, provide the system with: receiving data for a plurality of pre-analytical variables associated with a sample derived from a subject; applying the received data to a multivariate model trained to predict outcomes for the genomic profiling assay; generating a prediction of the likelihood of success for performing genomic profiling of the sample based on the applied multivariate model; A non-transitory computer readable storage medium that causes a prediction of the likelihood of success for performing genomic profiling of a sample to be reported. 46. ​​The non-transitory computer-readable storage medium of clause 45, wherein the instructions further cause the system to compare the predicted likelihood of success to a predetermined threshold and, based on a determination that the predicted likelihood of success is less than the predetermined threshold, output a recommendation to collect a new sample instead of submitting the sample for genomic profiling.

[0190] From the foregoing, it should be understood that, although specific embodiments of the disclosed method and system have been illustrated and described, various modifications can be made thereto and are contemplated herein. It is also not intended that the present invention be limited by the specific examples provided herein. Although the present invention has been described with reference to the above specification, the description and illustration of the preferred embodiments herein are not meant to be construed in a limiting sense. Furthermore, it should be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative proportions set forth herein, which depend upon various conditions and variables. Various modifications in form and details of the embodiments of the present invention will be apparent to those skilled in the art. Thus, the present invention is also intended to encompass any such modifications, variations, and equivalents.

Claims

1. 1. A method for predicting the likelihood of success for performing genomic profiling of a sample from a subject, comprising: receiving, using one or more processors, data for a plurality of pre-analytical variables associated with the sample; using the one or more processors to apply the received data to a multivariate model trained to predict outcomes for genomic profiling assays; generating, using the one or more processors, a prediction of the likelihood of success for genomic profiling of the sample based on the applied multivariate model; and and reporting, using the one or more processors, the prediction of the likelihood of success for performing genomic profiling of the sample.

2. providing a plurality of nucleic acid molecules obtained from the sample from the subject based on the prediction of the likelihood of success being equal to or greater than a predetermined threshold; ligating one or more adaptors onto one or more nucleic acid molecules from said plurality of nucleic acid molecules; amplifying one or more ligated nucleic acid molecules from said plurality of nucleic acid molecules; capturing the amplified nucleic acid molecule from the amplified nucleic acid molecules; sequencing the captured nucleic acid molecules with a sequencer to obtain a plurality of sequence reads that represent the captured nucleic acid molecules and that overlap with one or more loci within one or more subgenomic intervals in the sample; 10. The method of claim 1, further comprising generating, by one or more processors, a genomic profile based on the sequence reads for the sample, the genomic profile comprising sequence read analysis data.

3. The method of claim 1 or 2, further comprising training, by the one or more processors, the multivariate model using training data.

4. 4. The method of claim 3, wherein the training data comprises data derived from univariate analysis of clinical research data of samples collected from subjects representing subject age range, subject gender, diagnosis, disease stage, sample type, sample collection site, sample collection method, sample preparation method, sample storage method, sample age, sample transportation method, or any combination thereof.

5. 5. The method of claim 4, wherein the training data further comprises data on ECOG status, subject treatment status, tumor cellularity in the sample, tumor nuclear content in the sample, tissue surface area in the sample, tissue matrix in the sample, previous genomic profiling assay results, or any combination thereof.

6. 3. The method of claim 1 or 2, wherein the plurality of pre-analytical variables comprises subject age, subject gender, diagnosis, disease stage, sample type, sample collection site, sample collection method, sample preparation method, sample storage method, sample age, sample transportation method, or any combination thereof.

7. 3. The method of claim 1 or 2, wherein the data for the plurality of pre-analytical variables is supplemented with data for tumor cellularity in the sample, tumor nuclear content in the sample, tissue surface area in the sample, tissue matrix in the sample, or any combination thereof.

8. The method of claim 1 or 2, wherein the multivariate model comprises a machine learning model.

9. The method of claim 8 , wherein the machine learning model comprises a supervised learning model.

10. The method of claim 8 , wherein the machine learning model comprises an unsupervised learning model.

11. 3. The method of claim 1 or 2, wherein the multivariate model comprises a logistic regression model, a multiple linear regression model, a random forest model, a neural network model, or a deep learning model.

12. The method of claim 1 or 2, wherein the prediction of the likelihood of success comprises a binary value, a percentage, or a score.

13. comparing the predicted likelihood of success to a predetermined threshold; 3. The method of claim 1 or 2, further comprising: outputting an indication that the sample from the subject is suitable for providing a genomic profile of the subject based on a determination that the predicted likelihood of success is equal to or greater than a predetermined threshold.

14. comparing the predicted likelihood of success to a predetermined threshold; 3. The method of claim 1 or 2, further comprising: outputting a recommendation to collect a new sample instead of submitting the sample for genomic profiling based on a determination that the predicted likelihood of success is less than the predetermined threshold.

15. 15. The method of claim 14, further comprising outputting a sample type or sample collection site recommendation for the new sample.

16. 15. The method of claim 14, further comprising outputting a recommendation to perform an alternative nucleic acid sequencing-based testing method.

17. The method of claim 2 , wherein the predetermined threshold varies depending on the sample type.

18. 3. The method of claim 1, wherein the data for the plurality of pre-analytical variables is entered by a user via a graphical user interface (GUI) on a display device.

19. 3. The method of claim 1 or 2, wherein the prediction of the likelihood of success for performing genomic profiling of the sample is reported via a graphical user interface (GUI) on a display device.

20. The method of claim 18 , wherein the graphical user interface (GUI) is displayed in a web browser.

21. The method of claim 1 or 2, wherein the data received for the plurality of pre-analytical variables includes sample type data.

22. 22. The method of claim 21, wherein the remaining pre-analytical variables of the plurality of pre-analytical variables are selected based on the sample type.

23. 22. The method of claim 21, wherein the multivariate model is selected based on the sample type.

24. 3. The method of claim 1 or 2, wherein the sample comprises a tissue biopsy sample, a liquid biopsy sample, or a normal control.

25. 25. The method of claim 24, wherein the sample is a tissue biopsy sample and comprises bone marrow.

26. 25. The method of claim 24, wherein the sample is a liquid biopsy sample, including blood, plasma, cerebrospinal fluid, sputum, stool, urine, or saliva.

27. 25. The method of claim 24, wherein the sample is a liquid biopsy sample and comprises circulating tumor cells (CTCs).

28. 25. The method of claim 24, wherein the sample is a liquid biopsy sample and comprises cell-free DNA (cfDNA), circulating tumor DNA (ctDNA), or any combination thereof.

29. 14. The method of claim 13, wherein if the predicted likelihood of success is equal to or greater than the predetermined threshold, the genomic profiling is performed and used to diagnose or confirm a diagnosis of disease in the subject.

30. 30. The method of claim 29, wherein the genomic profiling is also used to determine eligibility for therapy based on biomarker status.

31. 30. The method of claim 29, wherein the disease is cancer.

32. A pharmaceutical comprising an anti-cancer therapy, wherein the anti-cancer therapy is selected based on the results of the genomic profiling used in the method of claim 31.

33. A pharmaceutical comprising an anti-cancer therapy, wherein an effective amount of the anti-cancer therapy is determined based on the results of the genomic profiling used in the method of claim 31.

34. The method of claim 33, wherein the anti-cancer therapy is administered to the subject based on the results of the genomic profiling.

35. The method of claim 32, wherein the anti-cancer therapy comprises chemotherapy, radiation therapy, immunotherapy, targeted therapy, or surgery.

36. The cancer is B-cell cancer (multiple myeloma), melanoma, breast cancer, lung cancer, bronchial cancer, colorectal cancer, prostate cancer, pancreatic cancer, stomach cancer, ovarian cancer, bladder cancer, brain cancer, central nervous system cancer, peripheral nervous system cancer, esophageal cancer, cervical cancer, endometrial cancer, oral cancer, pharyngeal cancer, liver cancer, kidney cancer, testicular cancer, biliary tract cancer, small intestine cancer, appendix cancer, salivary gland cancer, thyroid cancer, adrenal gland cancer, osteosarcoma, chondrosarcoma, cancer of the blood tissue , adenocarcinoma, inflammatory myofibroblastoma, gastrointestinal stromal tumor (GIST), colon cancer, multiple myeloma (MM), myelodysplastic syndrome (MDS), myeloproliferative disorder (MPD), acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic myeloid leukemia (CML), chronic lymphocytic leukemia (CLL), polycythemia vera, Hodgkin's lymphoma, non-Hodgkin's lymphoma (NHL), soft tissue sarcoma, fibrosarcoma, myxosarcoma, liposarcoma, osteosarcoma, spinal Chordoma, angiosarcoma, endothelial sarcoma, lymphangiosarcoma, lymphangioendothelial sarcoma, synovium, mesothelioma, Ewing's tumor, leiomyosarcoma, rhabdomyosarcoma, squamous cell carcinoma, basal cell carcinoma, adenocarcinoma, sweat gland carcinoma, sebaceous gland carcinoma, papillary carcinoma, papillary adenocarcinoma, medullary carcinoma, bronchogenic carcinoma, renal cell carcinoma, hepatocarcinoma, bile duct carcinoma, choriocarcinoma, seminoma, embryonal carcinoma, Wilms' tumor, bladder cancer, epithelial carcinoma, glioma, astrocytoma, medulloblastoma, craniopharyngioma, ependymoma, pinealocytoma , hemangioblastoma, acoustic neuroblastoma, oligodendroglioma, meningioma, neuroblastoma, retinoblastoma, follicular lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, hepatocellular carcinoma, thyroid cancer, gastric cancer, head and neck cancer, small cell carcinoma, essential thrombocythemia, primary myelofibrosis, hypereosinophilic syndrome, systemic mastocytosis, familial eosinophilia, chronic eosinophilic leukemia, neuroendocrine carcinoma, or carcinoid tumor.

37. 30. The method of claim 29, wherein the genomic profiling of the subject comprises obtaining results from a comprehensive genomic profiling (CGP) test, a gene expression profiling test, a cancer hotspot panel test, a DNA methylation test, a DNA fragmentation test, an RNA fragmentation test, or any combination thereof.

38. 38. The method of claim 37, wherein said genomic profiling of said subject further comprises obtaining results from a nucleic acid sequencing-based test.

39. A pharmaceutical comprising an anticancer agent or anticancer therapy, wherein the anticancer agent is selected or administered, or the anticancer therapy is applied to the subject, based on the results of genomic profiling used in the method of claim 37.

40. 1. A system comprising: one or more processors; a memory communicatively coupled to the one or more processors and configured to store instructions that, when executed by the one or more processors, cause the system to: receiving data for a plurality of pre-analytical variables associated with a sample derived from a subject; applying the received data to a multivariate model trained to predict outcomes for the genomic profiling assay; generating a prediction of the likelihood of success for performing genomic profiling of said sample based on the applied multivariate model; and a memory that causes the prediction of the likelihood of success for performing genomic profiling of the sample to be reported.

41. 41. The system of claim 40, wherein the instructions further cause the system to compare the predicted likelihood of success to a predetermined threshold and, based on a determination that the predicted likelihood of success is less than the predetermined threshold, output a recommendation to collect a new sample instead of submitting the sample for genomic profiling.

42. 42. The system of claim 40 or 41, wherein the data for the plurality of pre-analytical variables is entered by a user via a graphical user interface (GUI) on a display device.

43. 42. The system of claim 40 or 41, wherein the prediction of the likelihood of success for performing genomic profiling of the sample is reported via a graphical user interface (GUI) on a display device.

44. 43. The system of claim 42, wherein the graphical user interface (GUI) is displayed in a web browser.

45. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by one or more processors of a system, cause the system to: receiving data for a plurality of pre-analytical variables associated with a sample derived from a subject; applying the received data to a multivariate model trained to predict outcomes for the genomic profiling assay; generating a prediction of the likelihood of success for performing genomic profiling of said sample based on the applied multivariate model; A non-transitory computer-readable storage medium that reports a prediction of the likelihood of success for performing genomic profiling of the sample.

46. 46. ​​The non-transitory computer-readable storage medium of claim 45, wherein the instructions further cause the system to compare the predicted likelihood of success to a predetermined threshold and, based on a determination that the predicted likelihood of success is less than the predetermined threshold, output a recommendation to collect a new sample instead of submitting the sample for genomic profiling.