Somatic variant calling from unpaired biospecimens
The method employs a machine learning model to identify somatic variants in unpaired biological samples by processing nucleic acid sequence data without a normal sample, improving accuracy and reducing costs.
Patent Information
- Application Number
- JP2022526078
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-11-05
- Filing Date
- 2020-11-04
- Publication Date
- 2025-05-21
- Estimated Expiration
- 2040-11-04
AI Technical Summary
Current methods for identifying somatic variants in biological samples rely heavily on contrasting evidence between tumor and normal samples, which is not feasible when normal samples are not available.
A method using a trained machine learning model to process nucleic acid sequence data from unpaired biological samples, aligning it with a reference genome, and generating an attribute table to identify somatic variants without requiring nucleic acid sequencing data from a matched normal sample.
This approach enhances the accuracy of somatic variant identification, improves diagnostic, prognostic, and treatment recommendation reports, and reduces the costs and resources required for tumor analysis.
Smart Images

Figure 0007681016000002 
Figure 0007681016000003 
Figure 0007681016000004
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 62 / 931,100, filed November 5, 2019, which is incorporated by reference in its entirety for all purposes.
[0002] The present disclosure relates generally to systems and methods for identifying somatic variants in a biological sample, and more particularly, but not by way of limitation, the present disclosure relates to identifying somatic variants in a biological sample by filtering false positives from a set of candidate variants detected using a trained machine learning model. [Background technology]
[0003] 2. Background of the Invention Somatic variants of DNA sequences can represent one or more mutations that contribute to the development of cancer. For many analyses of tumor samples, identifying somatic variants can facilitate the improvement of cancer diagnosis, prognosis, treatment decisions, and treatment efficiency. To identify somatic variants in biological samples, germline sequence variants and somatic variants can be distinguished. Traditional somatic variant calling techniques rely heavily on contrasting evidence of differences between tumor samples and corresponding normal samples. However, there are some instances where corresponding normal samples are not available for analysis.
[0004] Thus, there is a need to accurately identify somatic variants in biological samples and distinguish somatic variants from germline variants without relying on normal control samples. Summary of the Invention
[0005] In some embodiments, a method of identifying somatic variants from a biological sample is provided. The method can include obtaining nucleic acid sequence data corresponding to the biological sample of a subject. The method can also include aligning the nucleic acid sequence data with a reference genome (generated based on samples from other subjects). The method can also include identifying a set of candidate variants in the nucleic acid sequence data based on the aligned nucleic acid sequence data. In some examples, the set of candidate variants includes one or more somatic variants and one or more germline variants.
[0006] The method can also include using a trained machine learning model to process the set of candidate variants to identify the somatic variant without using the nucleic acid sequencing data of the subject's corresponding biological sample. The subject's corresponding biological sample shows the absence of tumor. The method can also output a report that identifies the somatic variant.
[0007] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more of the methods disclosed herein.
[0008] In some embodiments, a computer program product is provided that is tangibly embodied in a non-transitory machine-readable storage medium and includes instructions configured to cause one or more data processors to perform some or all of the methods disclosed herein.
[0009] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium and including instructions configured to cause one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein.
[0010] The terms and expressions used are used as terms of description rather than of limitation, and there is no intention to use such terms and expressions to exclude any equivalents of the features shown and described or any portion thereof, but it is understood that various modifications are possible within the scope of the claimed invention. Thus, although the claimed invention has been specifically disclosed by embodiments and optional features, it should be understood that improvements and modifications of the concepts disclosed herein may be made by those skilled in the art, and that such improvements and modifications are considered to be within the scope of the invention as defined by the appended claims. [Brief description of the drawings]
[0011] The features, embodiments, and advantages of the present disclosure will be better understood from the following detailed description of the invention when read in conjunction with the following drawings: The invention or application contains at least one drawing executed in color. Copies of this patent or patent application publication containing color drawing(s) will be provided by the Office upon request and payment of the necessary fee.
[0012] [Figure 1] FIG. 1 illustrates an example interface configured to identify somatic variants in paired tumor / normal sequence data, according to some embodiments. [Diagram 2]FIG. 2 shows a graph highlighting the difference in precision and recall values between a trained gradient boosting decision tree model and a baseline, according to some embodiments. [Diagram 3] FIG. 3 shows two classification models that can be trained to identify somatic variants in unpaired biological samples, according to some embodiments. [Figure 4] FIG. 4 shows precision-recall curves corresponding to a trained filtering model for removing false positives from a set of candidate somatic variants, according to some embodiments. [Diagram 5] FIG. 5 illustrates a Shapley Additive exPlanations (SHAP) graph 500 that identifies which attributes in an attribute table influenced the output of a trained filtering model, according to some embodiments. [Figure 6] FIG. 6 shows precision-recall curves corresponding to a trained rescue model for removing false negatives from a set of candidate somatic variants, according to some embodiments. [Figure 7] FIG. 7 illustrates a SHAP graph identifying which attributes from an attribute table influenced the output of a trained rescue model, according to some embodiments. [Figure 8] FIG. 8 shows a comparison of the performance of a machine learning model with a filtering model and a rescue model before and after training and threshold adjustment, according to some embodiments. [Figure 9] FIG. 9 shows a flowchart for identifying somatic variants in unpaired biological samples according to some embodiments. [Figure 10] FIG. 10 illustrates an example computer system for implementing some of the embodiments disclosed herein. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0013] Detailed Description I. Overview As mentioned above, it becomes difficult to predict somatic variants in a biological sample when a matched normal sample is not available for analysis. To illustrate, FIG. 1 shows an example interface 100 configured to identify somatic variants in a pair of tumor / normal sequence data, according to some embodiments. The example interface 100 can include a bottom panel representing nucleic acid sequence data of a tumor sample 105 and a top panel representing nucleic acid sequence data of a normal sample 110. Grey bars may represent partially overlapping sequence reads aligned with a reference genome. Candidate variants can be highlighted within the reads using different colors. In the top panel of reads, three variants can be seen that are present in 50% to 100% of the reads. These variants can be identified as germline variants since these reads are derived from the matched normal sample. In the bottom panel of reads, the same three variants can be seen, with an additional variant present in a subset of reads (identified as a rectangle). This variant is present in the tumor sample but not in the matched normal sample, and can be identified as a somatic variant.
[0014] As shown in FIG. 1, conventional somatic variant calling techniques rely on contrasting evidence of differences between a subject's tumor sample and a matched normal sample. The lack of a matched normal sample 110 can prevent the identification of somatic variants in the tumor sample 105, which can greatly reduce the accuracy rate of conventional somatic variant calling techniques. For example, removing the matched normal sample 100 from the schematic example 100 can make it difficult to determine which of the candidate variants in the bottom panel are germline variants and which are somatic variants. The lack of a matched normal sample 110 can increase the amount of false positives (e.g., germline variants) in determining somatic variants. In some instances, the false positives caused by germline contamination (e.g., ) in the somatic variant calling output are greatly increased.
[0015] To address at least some of the deficiencies of conventional systems, the techniques of the present invention can be used to identify somatic variants in unpaired biological samples and distinguish somatic variants from germline variants. A trained machine learning model including one or more classification models can be used to predict somatic variants based on features extracted from nucleic acid sequencing data obtained from unpaired biological samples. In some examples, an additional data source (e.g., a database) can be used to predict somatic variants. For example, a highly sensitive algorithm can be used to identify candidate variants in the nucleic acid sequencing data. An attribute table can be generated, where the attribute table can include one or more features identified for each candidate variant. The trained machine learning model can be used to identify somatic variants based on the contents of the attribute table. A report identifying the somatic variants can be output. In some examples, the report includes a diagnostic report, a prognostic report, and / or a treatment recommendation.
[0016] Nucleic acid sequence data of a biological sample of a subject can be obtained. In some embodiments, the sequencing data is from a tumor sample. The sequencing can include whole exome sequencing. In some embodiments, the sequencing can include whole genome sequencing. In some embodiments, the sequencing can include shotgun sequencing. In some embodiments, the sequencing can include sequencing a selected portion of a genome or exome.
[0017] The nucleic acid sequence data can be aligned to a reference genome. As used herein, a reference genome corresponds to a nucleic acid sequence that corresponds to a representative set of genes in a certain ideal individual organism. Based on the aligned nucleic acid sequence data, a set of candidate variants in the nucleic acid sequence data can be identified. In some examples, the set of candidate variants includes one or more somatic variants and one or more germline variants. As used herein, a "somatic variant" refers to a change in DNA that occurs after conception and is not present in the germline. A somatic variant can occur in any cell of the body other than germ cells (sperm and egg cells), and therefore cannot be inherited. Furthermore, a "germline variant" refers to a genetic change in germ cells (sperm and egg cells) that is incorporated into the DNA of all cells of the offspring's body. A variant (or mutation) contained within the germline can be transmitted from parent to offspring, and is therefore heritable. In some examples, a somatic variant, instead of a germline variant, indicates the presence or level of cancer in a subject.
[0018] An attribute table (for example) can be generated, where the attribute table can include a number of features for each candidate variant. In some embodiments, the attribute table includes attributes from the sequencing data corresponding to a particular candidate variant. The attribute table can include attributes from a file including processed sequencing data. In some embodiments, the attribute table includes one or more attributes, such as: (a) pileup attributes from a BCFtools output file; (b) allele frequency data; (c) base quality data; (d) read depth data; (e) estimate of tumor cytology (which may be calculated based on B allele frequency distribution); (f) predicted germline variants; (g) predicted somatic variants; (h) copy number change data; (i) population frequency data from one or more databases; (j) data from at least one database selected from the group consisting of Cosmic, GnomAD, Dbsnp, and Mills Indels; (k) data regarding the presence of the candidate somatic variant in a problematic region of the genome; and (l) data regarding the presence of the candidate somatic variant in homopolymers.
[0019] The trained machine learning model can be used to process the set of candidate variants to identify somatic variants without nucleic acid sequencing data from a matched normal sample of the subject. In some examples, the trained machine learning model includes a gradient boosting decision tree that facilitates a significant reduction in the false positive rate corresponding to somatic variant calling. Thus, the present technology can detect somatic variants with enhanced sensitivity and specificity from unmatched biological samples compared to traditional heuristic approaches. In some embodiments, the trained machine learning model includes a two-model classification method. The machine learning model can include a filtering model that removes false positives. The machine learning model can include a rescue model that rescues false negatives. In some embodiments, the somatic variants are predicted with a precision of at least 0.5. In some embodiments, the somatic variants are predicted with a recall of at least 0.5. In some embodiments, the machine learning model includes hyperparameters that are tuned by random search. In some embodiments, the hyperparameters include a maximum depth of 5 to 100, a minimum data in leaves of 2 to 50, and at least 2 to 2048 leaves. In some embodiments, the filtering model includes a threshold of about 0.45. In some embodiments, the rescue model includes a threshold value of about 0.9995.
[0020] A report can be output identifying the somatic variants. In some embodiments, the report includes information identifying at least one diagnostic marker, at least one prognostic marker. In some embodiments, the absence of somatic variants, a treatment recommendation, a recommendation to administer a treatment to the human subject, and / or a recommendation not to administer a treatment to the human subject. In some embodiments, the recommended treatment is administered to the human subject.
[0021] Thus, embodiments of the present disclosure provide technical advantages over conventional systems by increasing the accuracy rate of somatic variants called from unpaired biological samples. Such techniques may improve the accuracy of diagnostic, prognostic, and / or treatment recommendation reports generated based on sequencing data from unpaired biological samples. Such techniques may also reduce the costs and resources required to identify somatic variants in tumors.
[0022] While various embodiments of the disclosed invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous modifications, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be used in carrying out any one of the inventions described herein.
[0023] II. Machine learning models for calling somatic variants from unpaired biospecimens A. Training a machine learning model to identify somatic variants from unpaired biospecimens A machine learning model for identifying somatic variants from unmatched biological samples can be trained using a training dataset that includes tumor samples and normal samples that match the tumor samples. For example, the training dataset can include the sequencing data obtained for 350 tumor / normal sample pairs (for example). DNA is extracted from the training samples, processed, and subjected to whole exome sequencing. Sequencing reads are subjected to quality control processing (for example, by FastQC) to provide FASTQ files. The FASTQ files are aligned with a reference genome to generate BAM files. A BCFtool is used to identify a set of candidate somatic variants for each training sample with high sensitivity. The set of candidate somatic variants includes false positives, such as germline variants.
[0024] For the set of candidate somatic variants, generate an attribute table containing multiple features (e.g., approximately 10-20 features) for each candidate variant. The attribute table can include (i) pileup attributes from the initial BCFtools output, such as allele frequency (e.g., B allele frequency), base quality, read depth, (ii) an estimate of tumor purity determined using a deep learning neural network based on the whole-exome B allele frequency distribution in the sample, (iii) whether the variant is identified as a germline variant using the GATK HaplotypeCaller, (iv) somatic copy number alteration (CNA) status for each variant site, (v) frequency of the variant in a population (e.g., in healthy human populations and / or in cancer exomes from databases such as Cosmic, GnomAD, Dbsnp, Mills Indels), (vi) presence of the variant in problematic regions, such as in homopolymers, and (vii) whether the variant is identified by standard somatic callers (run in the context of a single tumor) such as MuTect and MuTect2.
[0025] Classification labels are generated based on the presence of candidate variants in VCF files generated by MuTect or MuTect2 using default parameters with in-house reporting standards applied. Matched normal samples are examined by MuTect / MuTect2 to generate these classification labels, which identify "true" somatic variants, which are then used to evaluate model performance.
[0026] In some examples, the machine learning model is trained and tested to identify somatic variants based on the contents of the attribute table. The training dataset can be divided into a training (90%) set and a test (10%) set. In some embodiments, the trained machine learning model is trained with the training dataset to achieve one or more predetermined performance levels for estimating tumor purity. The one or more predetermined performance levels include: In some examples, the trained machine learning model can identify somatic variants with a precision of at least about 0.2, 0.25, 0.3, 0.35, 0.4, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, 0.95, or more. It is trained to predict with a precision of 0.4~0.8, 0.4~0.7, 0.4~0.6, 0.4~0.5, 0.5~1.0, 0.5~0.9, 0.5~0.8, 0.5~0.7, 0.5~0.6, 0.6~1.0, 0.6~0.9, 0.6~0.8, 0.6~0.7, 0.7~1.0, 0.7~0.9, 0.7~0.8, 0.8~1.0, 0.8~0.9, or 0.9~1.0. A recall of at least about 0.2, 0.25, 0.3, 0.35, 0.4, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, 0.95, or higher. In some examples, the trained machine learning model may detect somatic variants with a recall of at least about 0.2-1.0, 0.2-0.9, 0.2-0.8, 0.2-0.7, 0.2-0.6, 0.2-0.5, 0.2-0.4, 0.2-0.3, 0.3-1.0, 0.3-0.9, 0.3-0.8, 0.3-0.7, 0.3-0.6, 0.3-0.5, 0.3-0.4, 0.4-1.0, 0.4-0.9, It is trained to predict with a recall of 0.4-0.8, 0.4-0.7, 0.4-0.6, 0.4-0.5, 0.5-1.0, 0.5-0.9, 0.5-0.8, 0.5-0.7, 0.5-0.6, 0.6-1.0, 0.6-0.9, 0.6-0.8, 0.6-0.7, 0.7-1.0, 0.7-0.9, 0.7-0.8, 0.8-1.0, 0.8-0.9, or about 0.9-1.0. ● An Fl score (macro-average Fl classification score) of at least about 0.2, 0.25, 0.3, 0.35, 0.4, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.86, 0.87, 0.88, 0.89, 0.9, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.97, 0.98, 0.99, 0.995, or higher. In some examples, the trained machine learning model identifies somatic variants at approximately 0.2-1.0, 0.2-0.99, 0.2-0.95, 0.2-0.9, 0.2-0.8, 0.2-0.7, 0.2-0.6, 0.2-0.5, 0.2-0.4, 0.2-0.3, 0.3-1.0, 0.3-0.99, 0.3-0.95 ,0.3~0.9,0.3~0.8,0.3~0.7,0.3~0.6,0.3~0.5,0.3~0.4,0.4~1.0,0.4~0.99,0.4~0.95,0.4~0.9,0.4~0.8,0.4~0.7,0.4~0.6,0.4~0.5,0.5~1.0,0.5~0.99,0.5~0.95,0 .5~0.9,0.5~0.8,0.5~0.7,0.5~0.6,0.6~1.0,0.6~0.99,0.6~0.95,0.6~0.9,0.6~0.8,0.6~0.7,0.7~1.0,0.7~0.99,0.7~0.98,0.7~0.97,0.7~0.96,0.7~0.95,0.7~0.9, It is trained to predict with Fl-Scores between 0.7-0.8, 0.8-1.0, 0.8-0.99, 0.8-0.98, 0.8-0.97, 0.8-0.96, 0.8-0.95, 0.8-0.9, 0.9-1.0, 0.9-0.99, 0.9-0.98, 0.9-0.97, 0.9-0.96, or 0.9-0.95. a false positive rate of at most about 0.001%, 0.01%, 0.1%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 30%, 35%, 40%, or 50%; and Area Under the Curve-Receiver Operating Characteristic (AUC-ROC) of at least about 0.1, 0.2, 0.3, 0.4, 0.45, 0.5, 0.55, 0.6, 0.7, 0.8, 0.9, 0.95, 0.96, 0.97, 0.98, 0.99, 0.995, 0.999, 0.9995, 0.9999, or more. In some cases, the trained machine learning model is trained to achieve an AUC-ROC of at most about 0.8, 0.9, 0.95, 0.99, 0.995, 0.999, 0.9995, 0.9999, or less. In some cases, the trained machine learning model has a metric of approximately 0.5-1.0, 0.5-0.9995, 0.5-0.999, 0.5-0.99, 0.5-0.95, 0.5-0.9, 0.5-0.8, 0.5-0.7, 0.5-0.6, 0.6-1.0, 0.6-0.9995, 0.6-0.99, 0.6-0.95, 0.6-0.9, 0.6-0.8, 0.6-0.7, 0.7-1.0, 0.7-0.9999, 0.7-0.9995, 0.7-0.999, 0.7-0.99, 0.7-0.98, 0.7-0.97, It is trained to achieve an AUC-ROC of 0.7-0.96, 0.7-0.95, 0.7-0.9, 0.7-0.8, 0.8-1.0, 0.8-0.9999, 0.8-0.9995, 0.8-0.999, 0.8-0.99, 0.8-0.98, 0.8-0.97, 0.8-0.96, 0.8-0.95, 0.8-0.9, 0.9-1.0, 0.9-0.9999, 0.9-0.9995, 0.9-0.999, 0.9-0.99, 0.9-0.98, 0.9-0.97, 0.9-0.96, or 0.9-0.95. In some cases, the trained machine learning model is trained to achieve an AUC-ROC of about 0.5, 0.6, 0.7, 0.8, 0.85, 0.9, 0.95, 0.96, 0.97, 0.98, 0.99, 0.995, 0.997, 0.999, 0.9995, or 0.9999. In some cases, a high AUC-ROC value indicates a high probability of distinguishing true positive variants from true negative variants.
[0027] The trained machine learning model can use one or more thresholds. The threshold for the model can be selected based on (for example) maximizing the mean sample AUC of the precision-recall curve. In some cases, the filtering model can use a threshold of at least about 0.1, 0.2, 0.3, 0.4, 0.45, 0.5, 0.55, 0.6, 0.7, 0.8, 0.9, 0.95, or 0.99, or more.
[0028] B. Training a machine learning model framework with a classification model to identify somatic variants The trained machine learning model can correspond to one or more classification models. For example, the trained machine learning model can correspond to 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 models. In some embodiments, one classification model is trained and tested to identify somatic variants from an attribute table. The classification model can include a gradient boosting decision tree, which can be trained to predict somatic variants using an XGBoost framework (for example). The hyperparameters of the model can be tuned to maximize the macro-average Fl classification score.
[0029] FIG. 2 shows a graph 200 that identifies the difference in precision and recall values between the gradient boosting decision tree model and the baseline according to some embodiments. After training, the trained machine learning model can show an increase in the average Fl score compared to the baseline. The trained machine learning model can achieve a high AUC-ROC (area under the curve-receiver operating characteristic) of 0.997, indicating that it has the ability to distinguish true positive variants from true negative variants. The results in FIG. 2 show the feasibility of using the trained machine learning model to predict somatic variants from unpaired tumor sequencing data, and indicate that the model that allows for greater control of the threshold may achieve an increased accuracy rate.
[0030] C. Training a machine learning model framework with two classification models to identify somatic variants In some embodiments, the trained machine learning model corresponds to two classification models, each of which is trained and tested to identify somatic variants from the attribute table. To improve the controllability of the threshold, the somatic variant classification problem is decomposed into two sub-decomposition problems: (1) removing false positives during tumor-only calls from each variant caller, and (2) rescuing false negative candidate variants that are not present in the tumor-only calls.
[0031] FIG. 3 illustrates two classification models 300 that can be trained to identify somatic variants in unpaired biological samples, according to some embodiments. To improve the controllability of the threshold, the somatic variant classification problem is decomposed into two sub-decomposition problems: (1) removing false positives in tumor-only calls from each variant caller, and (2) rescuing false negative candidate variants that are not present in the tumor-only calls. The attribute table can therefore be split into two training datasets. In some examples, the two models are trained using a gradient boosting framework (e.g., the LightGBM framework).
[0032] The first training dataset 305 may include candidate variants identified by another variant detection algorithm in a tumor-only context (e.g., MuTect, MuTect2, etc.). A filtering model 310 can be trained to remove false positives from the first training dataset. In some examples, the first training dataset 305 includes a large portion of the training dataset, such as, for example, about 71% of tumor-normal calls.
[0033] The second training dataset 315 may include residuals of candidate variants. A rescue model 320 may be trained to rescue false negatives from the second training dataset. In some examples, the rescue model 320 is trained to distinguish between false negatives and true negatives, where false negatives correspond to those that the variant detection algorithm fails to identify.
[0034] In some examples, two classification models are trained using a gradient boosting framework (e.g., LightGBM). The classification results from both of these classification models can be combined to generate a final set of somatic variants 325. The final set of somatic variants can then be used to train the classification models 310 and 320. In some examples, training the classification models 310 and 320 includes tuning one or more hyperparameters (e.g., learning rate). During training, 300 iterations of random search are used for the following set of hyperparameters for the known classification problem: (i) maximum depth: 5 to 100, (ii) minimum data in a leaf: 3 to 50, and (iii) number of leaves: 3 to 2048 (log scale). Each iteration can train each of the classification models, followed by stratified 5-fold cross-validation. The averaged model for the 5 best-fit cross-validation models according to AUC-ROC (area under the curve - receiver operating characteristic) can be applied to the test dataset.
[0035] 4 illustrates a precision-recall curve 400 corresponding to a trained filtering model for removing false positives from a set of candidate somatic variants, according to some embodiments. As shown in FIG. 4, the precision-recall curve 400 illustrates the ability of the filtering model to remove a majority of the false positives from the dataset. Although noise due to varying positive class support is observed in the precision-recall curve, the AUC-ROC remains fairly constant.
[0036] A threshold for the filtering model 310 can be selected based on maximizing the average sample AUC of the precision-recall curve. For example, a threshold of 0.45 can be selected for the filtering model 310. In some cases, the filtering model 310 includes a threshold of at most about 0.3, 0.4, 0.45, 0.5, 0.55, 0.6, 0.7, 0.8, 0.9, 0.95, or 0.99, or less. In some cases, the filtering model 310 includes a threshold of at most about 0.2-1.0, 0.2-0.99, 0.2-0.95, 0.2-0.9, 0.2-0.8, 0.2-0.7, 0.2-0.6, 0.2-0.5, 0.2-0.4, 0.2-0.3, 0.3-1.0, 0.3-0.99, 0.3-0.95, 0. .3~0.9,0.3~0.8,0.3~0.7,0.3~0.6,0.3~0.5,0.3~0.4,0.4~1.0,0.4~0.99,0.4~0.95,0.4~0.9,0.4~0.8,0.4~0.7,0.4~0.6,0.4~0.5,0.5~1.0,0.5~0.99,0.5~0.9 5,0.5~0.9,0.5~0.8,0.5~0.7,0.5~0.6,0.6~1.0,0.6~0.99,0.6~0.95,0.6~0.9,0.6~0.8,0.6~0.7,0.7~1.0,0.7~0.99,0.7~0.98,0.7~0.97,0.7~0.96,0.7~0.95, The filtering model 310 may include a threshold value of about 0.7 to 0.9, 0.7 to 0.8, 0.8 to 1.0, 0.8 to 0.99, 0.8 to 0.98, 0.8 to 0.97, 0.8 to 0.96, 0.8 to 0.95, 0.8 to 0.9, 0.9 to 1.0, 0.9 to 0.99, 0.9 to 0.98, 0.9 to 0.97, 0.9 to 0.96, or 0.9 to 0.95. In some cases, the filtering model 310 may include a threshold value of about 0.1, 0.2, 0.3, 0.4, 0.45, 0.5, 0.55, 0.6, 0.7, 0.8, 0.9, 0.95, or 0.99. In some embodiments, the filtering model 310 may include a threshold value of about 0.4 to about 0.5. In some embodiments, the filtering model 310 includes a threshold value of about 0.45.
[0037] FIG. 5 illustrates a Shapley Additive exPlanations (SHAP) graph 500 that identifies which attributes of an attribute table influenced the output of a trained filtering model, according to some embodiments. The SHAP graph 500 illustrates graphical information that reveals the extent to which each attribute in the attribute table contributed to a false positive identification of a somatic variant in a biological sample. The SHAP graph 500 includes a left side portion 505 that identifies a number of features from the attribute table, where each column corresponds to one of a number of attributes determined for a known candidate variant. The SHAP graph 500 also includes a right side portion 510 that reveals the extent to which the known attributes contributed to a false positive identification of a somatic variant in a biological sample. In some examples, the attributes are arranged up and down based on their relative contribution to the false positive identification. For example, the attribute (gnomAD_AF) corresponding to the top column can be associated with the highest contribution to the false positive identification. In this example, gnomAD_AF may refer to the frequency of existing variants in the exomes corresponding to the combined population, where the existing variants are identified from an aggregated genomic database (e.g., gnomAD).
[0038] Figure 6 shows a precision-recall curve 600 corresponding to a trained rescue model for removing false negatives from a set of candidate somatic variants, according to some embodiments. As shown in Figure 6, the rescue model data shows that feature importance is non-linear and difficult to classify. Due to the overwhelming negative class support, precision drops off rapidly as recall of the rescue model increases.
[0039] A threshold for the filtering model 320 can be selected based on maximizing the mean sample AUC of the precision-recall curve. For example, a threshold of 0.9995 can be selected for the rescue model 320. In some cases, the rescue model 320 includes a threshold of at least about 0.1, 0.2, 0.3, 0.4, 0.45, 0.5, 0.55, 0.6, 0.7, 0.8, 0.9, 0.95, 0.96, 0.97, 0.98, 0.99, 0.995, 0.999, 0.9995, 0.9999 or more. In some cases, the remedy model 320 includes a threshold value that is at most about 0.3, 0.4, 0.45, 0.5, 0.55, 0.6, 0.7, 0.8, 0.9, 0.95, 0.99, 0.995, 0.999, 0.9995, 0.9999, or less.In some cases, the relief model 320 may be about 0.2-1.0, 0.2-0.9995, 0.2-0.99, 0.2-0.95, 0.2-0.9, 0.2-0.8, 0.2-0.7, 0.2-0.6, 0.2-0.5, 0.2-0.4, 0.2-0.3, 0.3-1.0, 0.3-0.9995, 0.3-0.99, 0.3-0.95, 0.3-0.9, 0.3-0.8, 0.3- 0.7,0.3~0.6,0.3~0.5,0.3~0.4,0.4~1.0,0.4~0.9995,0.4~0.99,0.4~0.95,0.4~0.9,0.4~0.8,0.4~0.7,0.4~0.6,0.4~0.5,0.5~1.0,0.5~0.9995,0.5~0.99,0.5~0.95,0.5~0.9,0.5~0.8,0.5~0.7,0.5~0.6, 0.6~1.0,0.6~0.9995,0.6~0.99,0.6~0.95,0.6~0.9,0.6~0.8,0.6~0.7,0.7~1.0,0.7~0.9999,0.7~0.9995,0.7~0.999,0.7~0.99,0.7~0.98,0.7~0.97,0.7~0.96,0.7~0.95,0.7~0.9,0.7~0.8,0.8~1.0,0.8~ Includes thresholds of 0.9999, 0.8~0.9995, 0.8~0.999, 0.8~0.99, 0.8~0.98, 0.8~0.97, 0.8~0.96, 0.8~0.95, 0.8~0.9, 0.9~1.0, 0.9~0.9999, 0.9~0.9995, 0.9~0.999, 0.9~0.99, 0.9~0.98, 0.9~0.97, 0.9~0.96, or 0.9~0.95. In some cases, the rescue model 320 includes a threshold value of about 0.1, 0.2, 0.3, 0.4, 0.45, 0.5, 0.55, 0.6, 0.7, 0.8, 0.9, 0.95, 0.99, 0.995, 0.999, 0.9995, or 0.9999. In some embodiments, the rescue model 320 includes a threshold value of about 0.9 to about 0.9999. In some embodiments, the rescue model 320 includes a threshold value of about 0.9995.
[0040] FIG. 7 shows a SHAP graph 700 that identifies which attributes of an attribute table influenced the output of a trained rescue model, according to some embodiments. The SHAP graph 700 shows graphical information that reveals the extent to which each attribute in the attribute table contributed to the false negative identification of a somatic variant in a biological sample. The SHAP graph 700 includes a left side 705 that identifies a number of features from the attribute table, where each column corresponds to one of a number of attributes determined for a known candidate variant. The SHAP graph 700 also includes a right side 710 that reveals the extent to which the known attributes contributed to the false negative identification of a somatic variant in a biological sample. In some examples, the attributes are arranged up and down based on their relative contribution to the false negative identification. For example, the attribute corresponding to the top column ("QA") can be associated with the highest contribution to the false negative identification. In this example, QA refers to the sum of allele quality instead of Phred, and the Phred quality score can indicate a measure of the identification quality of a nucleic acid base made by automated DNA sequencing.
[0041] The ability of the two-model classifier to predict somatic variants from unpaired tumor sequencing data can be evaluated before and after training and threshold adjustment. Baseline performance is summarized in Table 1. Macro-average statistics of precision and recall are provided for each sample set. Variance is explained by similar true positive / false positive rates per sample with varying positive class support. [Table 1]
[0042] Overall, a precision of 0.189±0.19 and a recall of 0.677±0.15 are observed as the baseline. After training and threshold adjustment, the two-model classifier achieves a precision of 0.644 and a recall of 0.634.
[0043] Figure 8 shows a comparison 800 of the performance of a machine learning model with a filtering model and a rescue model before and after training and threshold adjustment according to some embodiments. In this comparison, the machine learning model is used to predict somatic variants from unpaired tumor sequencing data. The precision and recall at baseline and after training and threshold adjustment are shown.
[0044] As shown in Figure 8, the comparison data showed that the trained machine learning model with the filtering model and the rescue model could predict somatic variants from unpaired tumor sequencing data with a high accuracy rate compared to alternative methods (e.g., MuTect and MuTect2).
[0045] III. Identification of somatic variants in unmatched biological samples A. Subjects and Samples An unmatched biological sample is obtained from a cancer patient (i.e., a tumor sample with no corresponding normal sample). The subject can be a human. The subject can be male or female. The subject can be a fetus, infant, child, adolescent, teenager, or adult. The subject can be a patient of any age. For example, the subject can be a patient less than about 10 years old. For example, the subject can be a patient at least about 0, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 years old. Often, the subject is a patient or other individual undergoing or being evaluated for a treatment regimen (cancer treatment). However, in some instances, the subject is not undergoing a treatment regimen.
[0046] In some cases, the subject may be a mammal or a non-mammal. In some cases, the subject is a mammal, such as a human, a non-human primate (e.g., ape, monkey, chimpanzee), cat, dog, rabbit, goat, horse, cow, pig, rodent, mouse, SCID mouse, rat, guinea pig, or sheep. In some methods, species variants of non-human animal models or homologs of these genes can be used. Species variants can be genes of different species that have the highest sequence homology and similar functional properties to each other. Many of these species variants of human genes can be listed in the Swiss-Prot database.
[0047] Some embodiments may include obtaining a sample from a subject, such as a human subject. Specifically, the method may include obtaining a clinical specimen from the patient. For example, blood may be drawn from the patient. Some embodiments may specifically include detecting, profiling, or quantifying molecules (e.g., nucleic acids, DNA, RNA, etc.) within the biological sample.
[0048] The sample may be a tissue sample or a bodily fluid. In some instances, the sample is an organ sample, such as a tissue sample or a biopsy. In some cases, the sample includes cancerous cells. In some cases, the sample includes cancerous cells and normal cells. In some cases, the sample is a tumor biopsy. The bodily fluid may be sweat, saliva, tears, urine, blood, menstrual, semen, and / or spinal fluid. In some cases, the sample is a blood sample. The sample may include one or more peripheral blood lymphocytes. The sample may be a whole blood sample. The blood sample may be a peripheral blood sample. In some cases, the sample includes peripheral blood mononuclear cells (PBMCs), and in some cases, the sample includes peripheral blood lymphocytes (PBLs). The sample may be a serum sample.
[0049] The sample may be obtained using any method capable of providing a sample suitable for the analytical methods described herein. The sample may be obtained by non-invasive methods such as throat swabs, buccal swabs, bronchial washing, urine collection, scraping of the skin or cervix, buccal swabs, saliva collection, fecal collection, menstrual collection, or semen collection. The sample may be obtained by minimally invasive methods such as blood collection. The sample may be obtained by venipuncture. In other examples, the sample is obtained by invasive procedures including, but not limited to, biopsy, alveolar or pulmonary lavage, or needle aspiration. Biopsy methods may include surgical biopsy, incisional biopsy, excision biopsy, punch biopsy, shave biopsy, or skin biopsy. The sample may be a formalin-fixed section. Needle aspiration methods may further include fine needle aspiration, core needle biopsy, aspiration biopsy, or large core biopsy. In some cases, multiple samples may be obtained by the methods herein to ensure a sufficient amount of biomaterial. In some instances, the sample is not obtained by biopsy. In some instances, the sample is not a kidney biopsy.
[0050] B. Generating Nucleic Acid Sequencing Data In some embodiments, the sample is processed to obtain nucleic acid sequence data. A "nucleic acid" or "nucleic acid molecule" can correspond to a polymeric form of nucleotides of any length, either ribonucleotides, deoxyribonucleotides, or peptide nucleic acids (PNAs), containing purine and pyrimidine bases, or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases. The backbone of a polynucleotide can contain sugars and phosphate groups, as may typically be found in RNA or DNA, or modified or substituted sugars and phosphate groups. A polynucleotide may contain modified nucleotides, such as methylated nucleotides and nucleotide analogs. The sequence of nucleotides may be interrupted by non-nucleotide components. Thus, the terms nucleoside, nucleotide, deoxynucleoside, and deoxynucleotide generally include analogs such as those described herein. These analogs are molecules that have some structural features in common with naturally occurring nucleosides or nucleotides, such that when incorporated into a nucleic acid or oligonucleoside sequence, they are capable of hybridizing to naturally occurring nucleic acid sequences in solution. Typically, these analogs are derived from naturally occurring nucleosides or nucleotides by substituting and / or modifying the base, ribose, or phosphodiester moiety. These modifications can be aimed at stabilizing or destabilizing hybridization or enhancing the specificity of hybridization with a desired complementary nucleic acid sequence. The nucleic acid molecule can be a DNA molecule. The nucleic acid molecule can be an RNA molecule.
[0051] DNA is extracted from tumor samples, processed, and subjected to whole exome sequencing. Sequencing reads are subjected to a quality control process (e.g., by FastQC) to provide FASTQ files. The FASTQ files are aligned to a reference genome to generate BAM files.
[0052] In some cases, sample processing includes nucleic acid sample processing and subsequent nucleic acid sample sequencing. A portion or all of the nucleic acid sample may be sequenced to provide sequence information, which may be stored or otherwise maintained in an electronic, magnetic or optical storage location. The sequence information may be analyzed with the aid of a computer processor, and the analyzed sequence information may be stored in an electronic storage location. The electronic storage location may contain a pool or collection of sequence information generated from the nucleic acid sample and analyzed sequence information. The nucleic acid sample may be collected from a subject, for example, a subject having or suspected of having cancer.
[0053] Some embodiments may include using whole genome sequencing. In some cases, whole genome sequencing is used to identify variants in an individual. In some cases, sequencing may include deep sequencing over a fraction of the genome. For example, the fraction of the genome may be at least about 50; 75; 100; 125; 150; 175; 200; 225; 250; 275; 300; 350; 400; 450; 500; 550; 600; 650; 700; 750; 800; 850; 900; 950; 1,000; 1100; 1200; 1300; 1400; 1500; 1600; 1700; In some cases, the genome may be sequenced over 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 60, 7, 8, 9, 10, 10, 15, 20, 30, 40, 50, 60, 70, 8, 9, 10 ... In some cases, the genome may be sequenced across the entire exome (e.g., whole exome sequencing). In some cases, deep sequencing may include obtaining multiple reads across a fraction of the genome. For example, obtaining multiple reads may include at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 10,000 reads across the fraction of the genome, or more than 10,000 reads.
[0054] Some embodiments may include detecting low allele fractions by deep sequencing. In some cases, deep sequencing is performed by next generation sequencing. In some cases, deep sequencing is performed by avoiding error-prone regions. In some cases, error-prone regions may include regions adjacent to sequence duplications, regions with abnormally high or low GC ratios, regions adjacent to homopolymers, dinucleotides and trinucleotides, and regions adjacent to other short repeat sequences. In some cases, error-prone regions may include regions that lead to DNA sequencing errors (e.g., polymerase slippage in homopolymers).
[0055] Some embodiments may include performing one or more sequencing reactions on one or more nucleic acid molecules in a sample. Some embodiments may include performing 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 15 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 200 or more, 300 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more sequencing reactions on one or more nucleic acid molecules in a sample. Sequencing reactions may be run simultaneously, sequentially, or in combination. The sequencing reactions may include whole genome sequencing or whole exome sequencing. The sequencing reactions may include Maxim-Gilbert, chain termination, or high throughput systems. Alternatively, or in addition, the sequencing reaction may include Helioscope™ single molecule sequencing, nanopore DNA sequencing, Lynx Therapeutics Massively Parallel Signature Sequencing (MPSS), 454 Pyrosequencing, single molecule real-time (RNAP) sequencing, Illumina (Solexa) sequencing, SOLiD sequencing, Ion Torrent™ Ion semiconductor sequencing, single molecule SMRT™ sequencing, polony sequencing, DNA nanoball sequencing, VisiGen Biotechnologies techniques, or combinations thereof. Alternatively or additionally, the sequencing reaction can include one or more sequencing platforms, including, but not limited to, single molecule real-time (SMRT™) technology, such as the Genome Analyzer IIx, HiSeq, and MiSeq offered by Illumina, the PacBio RS system and Solexa sequencers offered by Pacific Biosciences (California), and True Single Molecule Sequencing (tSMS™) technology, such as the HeliScope™ sequencer offered by Helicos Inc. (Cambridge, MA).The sequencing reaction may also include an electron microscope or a chemically sensitive field effect transistor (chemFET) array. In some embodiments, the sequencing reaction includes capillary sequencing, next generation sequencing, Sanger sequencing, sequencing by synthesis, sequencing by ligation, sequencing by hybridization, single molecule sequencing, or a combination thereof. Sequencing by synthesis may include reversible terminator sequencing, processive single molecule sequencing, sequential flow sequencing, or a combination thereof. Sequential flow sequencing may include pyrosequencing, pH-mediated sequencing, semiconductor sequencing, or a combination thereof.
[0056] Some embodiments may include performing at least one long-read sequencing reaction and at least one short-read sequencing reaction. The long-read sequencing reaction and / or the short-read sequencing reaction may be performed on at least a portion of a subset of the nucleic acid molecules. The long-read sequencing reaction and / or the short-read sequencing reaction may be performed on at least a portion of two or more subsets of the nucleic acid molecules. Both the long-read sequencing reaction and the short-read sequencing reaction may be performed on at least a portion of one or more subsets of the nucleic acid molecules.
[0057] Sequencing of one or more nucleic acid molecules or a subset thereof may be at least about 5; 10; 15; 20; 25; 30; 35; 40; 45; 50; 60; 70; 80; 90; 100; 200; 300; 400; 500; 600; 700; 800; 900; 1,000; 1500; 2,000; 2500; 3,000; 3500; 4,000; 4500; 5,000; 5500; 6,000; 6500; 7,000; 7500; 8 ,000;8500;9,000;10,000;25,000;50,000;75,000;100,000;250,000;500,000;750,000;10,000,000;25,000,000;50,000,000;100,000,000;250,000,000;500,000,000;750,000,000;1,000,000,000 or more sequencing reads.
[0058] The sequencing reaction may comprise determining at least about 50; 60; 70; 80; 90; 100; 110; 120; 130; 140; 150; 160; 170; 180; 190; 200; 210; 220; 230; 240; 250; 260; 270; 280; 290; 300; 325; 350; 375; 400; 425; 450; 475; 500; 600; 700; 800; 900; 1,000; 1500 ;2,000;2500;3,000;3500;4,000;4500;5,000;5500;6,000;6500;7,000;7500;8,000;8500;9,000;10,000;20,000;30,000;40,000;50,000;60,000;70,000;80,000;90,000;100,000 or more bases or base pairs. The sequencing reaction may comprise determining the number of sequences of at least about 50;60;70;80;90;100;110;120;130;140;150;160;170;180;190;200;210;220;230;240;250;260;270;280;290;300;325;350;375;400;425;450;475;500;600;700;800;900;1,000;1500;2 The method may include sequencing 2,000; 2500; 3,000; 3500; 4,000; 4500; 5,000; 5500; 6,000; 6500; 7,000; 7500; 8,000; 8500; 9,000; 10,000; 20,000; 30,000; 40,000; 50,000; 60,000; 70,000; 80,000; 90,000; 100,000 or more consecutive bases or base pairs.
[0059] Preferably, the sequencing techniques used in the methods of the invention generate at least 100 reads per run, at least 200 reads per run, at least 300 reads per run, at least 400 reads per run, at least 500 reads per run, at least 600 reads per run, at least 700 reads per run, at least 800 reads per run, at least 900 reads per run, at least 1000 reads per run, at least 5,000 reads per run, at least 10,000 reads per run, at least 50,000 reads per run, at least 100,000 reads per run, at least 500,000 reads per run, or at least 1,000,000 reads per run. Alternatively, the sequencing technology used in the methods of the invention produces at least 1,500,000 reads per run, at least 2,000,000 reads per run, at least 2,500,000 reads per run, at least 3,000,000 reads per run, at least 3,500,000 reads per run, at least 4,000,000 reads per run, at least 4,500,000 reads per run, or at least 5,000,000 reads per run.
[0060] Preferably, the sequencing technology used in the method of the present invention generates at least about 30 base pairs, at least about 40 base pairs, at least about 50 base pairs, at least about 60 base pairs, at least about 70 base pairs, at least about 80 base pairs, at least about 90 base pairs, at least about 100 base pairs, at least about 110, at least about 120 base pairs per read, at least about 150 base pairs, at least about 200 base pairs, at least about 250 base pairs, at least about 300 base pairs, at least about 350 base pairs, at least about 400 base pairs, at least about 450 base pairs, at least about 500 base pairs, at least about 550 base pairs, at least about 600 base pairs, at least about 700 base pairs, at least about 800 base pairs, at least about 900 base pairs, or at least about 1,000 base pairs per read. Alternatively, the sequencing technology used in the method of the present invention can generate long sequencing reads. In some examples, the sequencing technology used in the methods of the invention provides a sequencing technique that provides at least about 1,200 base pairs per read, at least about 1,500 base pairs per read, at least about 1,800 base pairs per read, at least about 2,000 base pairs per read, at least about 2,500 base pairs per read, at least about 3,000 base pairs per read, at least about 3,500 base pairs per read, at least about 4,000 base pairs per read, at least about 4,500 base pairs per read, at least about 5,000 base pairs per read, at least about 6,000 base pairs per read, at least about 7,000 base pairs per read, at least about 8,000 base pairs per read, at least about 9,000 base pairs per read, at least about 10,000 base pairs per read, at least about 11,000 base pairs per read, at least about 12,000 base pairs per read, at least about 13,000 base pairs per read, at least about 14,000 base pairs per read, at least about 15,000 base pairs per read, at least about 16,000 base pairs per read, at least about 17,000 base pairs per read, at least about 18,000 base pairs per read, at least about 20,000 base pairs per read, at least about 21,000 base pairs per read, at least about 22,000 base pairs per read, at least about 23,000 base pairs per read, at least about 24,000 base pairs per read, at least about 25,000 base pairs per read, at least about 26,000 base pairs per read, at least about 27,000 base pairs per read, at least about 28,000 base pairs per read, at least about 29,000 base pairs per read, at least about 30,000 In some embodiments, at least about 6,000 base pairs per read, at least about 7,000 base pairs per read, at least about 8,000 base pairs per read, at least about 9,000 base pairs per read, at least about 10,000 base pairs per read, 20,000 base pairs per read, 30,000 base pairs per read, 40,000 base pairs per read, 50,000 base pairs per read, 60,000 base pairs per read, 70,000 base pairs per read, 80,000 base pairs per read, 90,000 base pairs per read, or 100,000 base pairs per read can be generated.
[0061] High throughput sequencing systems may be used to detect sequenced nucleotides immediately after or as they are incorporated into the growing strand, i.e., detect sequences in real time or substantially in real time. In some cases, high throughput sequencing produces at least 1,000, at least 5,000, at least 10,000, at least 20,000, at least 30,000, at least 40,000, at least 50,000, at least 100,000, or at least 500,000 sequence reads per hour, with each read being at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 120, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, or at least 500 bases per read. Sequencing can be carried out using a nucleic acid as described herein, such as genomic DNA, cDNA derived from an RNA transcript, or RNA, as a template.
[0062] C. Identifying candidate variants The nucleic acid sequence data can be aligned with the reference genome. Based on the aligned nucleic acid sequence data, a set of candidate variants in the nucleic acid sequence data can be identified. In some examples, the set of candidate variants includes one or more somatic variants and one or more germline variants. For example, BCFtools can be used to identify a set of candidate somatic variants for each sample with high sensitivity. The set of candidate somatic variants will include false positives, such as germline variants.
[0063] For the set of candidate somatic variants, create an attribute table that includes a number of features (e.g., about 10-20 features) for each candidate variant. The attribute table can include any combination of the attributes described in Example 3. The attribute table can include a number of features for each candidate variant. Examples of features that the attribute table may include include, but are not limited to, (i) pileup attributes from the initial BCFtools output, such as allele frequency (e.g., B allele frequency), base quality, read depth, (ii) an estimate of tumor purity determined using a deep learning neural network based on the whole exome B allele frequency distribution in the sample, (iii) whether the variant is identified as a germline variant using the GATK HaplotypeCaller, (iv) somatic copy number alteration (CNA) status for each variant site, (v) the frequency of the variant in a population (e.g., in healthy human populations and / or in cancer exomes from databases such as Cosmic, GnomAD, Dbsnp, Mills Indels), (vi) the presence of the variant in problematic regions, such as in homopolymers, and (vii) whether the variant is identified by standard somatic callers (run in the context of a single tumor) such as MuTect and MuTect2.
[0064] The attribute table can include any number of features that can contribute to accurate prediction of somatic variants. For example, the attribute table can include at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 109, 109, 109, 110, 111, 112, 113, 114, 1 The number of features may be 3, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 or more features. In some cases, the attribute table is at most about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109 ... The number of features may be 3, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 or fewer features.In some embodiments, the attribute table may be about 1-100, 1-90, 1-80, 1-70, 1-60, 1-50, 1-40, 1-30, 1-20, 1-10, 1-5, 5-100, 5-90, 5-80, 5-70, 5-60, 5-50, 5-40, 5-30, 5-20, 5-10, 10-100, 10-90, 10-80, 10- The attribute table may include 70, 10-60, 10-50, 10-40, 10-30, 10-20, 15-100, 15-90, 15-80, 15-70, 15-60, 15-50, 15-40, 15-30, 15-20, 20-100, 20-90, 20-80, 20-70, 20-60, 20-50, 20-40, or 20-30 features. In some cases, the attribute table includes about 10-20 features.
[0065] In some embodiments, identifying a set of candidate variants may include identifying one or more genomic regions that comprise one or more nucleotide sequence variants. The one or more genomic regions may include one or more genomic region features. The genomic region features may include the entire genome or a portion thereof. The genomic region features may include the entire exome or a portion thereof. The genomic region features may include one or more sets of genes. The genomic region features may include one or more genes. The genomic region features may include one or more sets of regulatory regions. The genomic region features may include one or more regulatory regions. The genomic region features may include a set of polymorphisms of genes. The genomic region features may include one or more gene polymorphisms. The genomic region features may relate to GC content, complexity, and / or mappability of one or more nucleic acid molecules. The genomic region features may include one or more simple tandem repeats (STRs), unstable expanding repeats, segmental duplications, single and paired read degenerative mapping scores, GRCh37 patches, or combinations thereof. The genomic region features may include one or more low average coverage regions from whole genome sequencing (WGS), zero average coverage regions from WGS, validated compressions, or combinations thereof. The genomic region features may include one or more surrogate or non-reference sequences. The genomic region features may include one or more gene phasing and reconstructed genes. In some embodiments, the one or more genomic region features are not mutually exclusive. For example, a genomic region feature that includes an entire genome or a portion thereof can be overlapped by additional genomic region features such as an entire exome or a portion thereof, one or more genes, one or more regulatory elements, etc. Alternatively, the one or more genomic region features are mutually exclusive.For example, a genomic region that includes a non-coding portion of the entire genome will not overlap with a genomic region feature such as an exome or a portion thereof, or a coding portion of a gene. Alternatively, or in addition, one or more genomic region features are partially exclusive or partially inclusive. For example, a genomic region that includes a whole exome or a portion thereof can partially overlap with a genomic region that includes an exon portion of a gene. However, a genomic region that includes a whole exome or a portion thereof will not overlap with a genomic region that includes an intron portion of a gene. Thus, a genomic region feature that includes a gene or a portion thereof may partially exclude and / or partially include a genomic region that includes a whole exome or a portion thereof.
[0066] Some embodiments may include a nucleic acid sample or molecule comprising one or more genomic regions, at least one of which comprises a genomic region feature comprising an entire genome or a portion thereof. The entire genome or a portion thereof may comprise one or more coding portions of the genome, one or more non-coding portions of the genome, or a combination thereof. The coding portion of the genome may comprise one or more coding portions of a gene encoding one or more proteins. The one or more coding portions of the genome may comprise an entire exome or a portion thereof. Alternatively, or in addition, the one or more coding portions of the genome may comprise one or more exons. The one or more non-coding portions of the genome may comprise one or more non-coding molecules or portions thereof. The non-coding molecules may comprise one or more non-coding RNAs, one or more regulatory elements, one or more introns, one or more pseudogenes, one or more repeat sequences, one or more transposons, one or more viral elements, one or more telomeres, portions thereof, or a combination thereof. Non-coding RNA may be a functional RNA molecule that is not translated into a protein. Examples of non-coding RNA include, but are not limited to, ribosomal RNA, transfer RNA, Piwi-binding RNA, microRNA, small interfering RNA (siRNA), small hairpin RNA (shRNA), small nucleolar ribonucleic acid (snoRNA), small non-coding RNA (sncRNA) and long non-coding RNA (lncRNA). Pseudogenes may be associated with known genes and are typically no longer expressed. The repeat sequence may include one or more tandem repeats, one or more interspersed repeats, or a combination thereof. The tandem repeats may include one or more satellite DNA, one or more minisatellites, one or more microsatellites, or a combination thereof. The interspersed repeats may include one or more transposons. The transposons may be mobile genetic elements. Mobile genetic elements are often capable of changing their position within the genome. Transposons may be classified as class I transposable elements (class I TEs) or class II transposable elements (class II TEs).Class I TEs (retrotransposons) can often copy themselves in two steps, first from DNA to RNA by transcription, then from RNA back to DNA by reverse transcription. They may insert the DNA copy into the genome at a new location. Class I TEs may include one or more long terminal repeats (LTRs), one or more long interspersed repeats (LINEs), one or more short interspersed repeats (SINEs), or a combination thereof. Examples of LTRs include, but are not limited to, human endogenous retroviruses (HERVs), medium interspersed repeats 4 (MER4), and retrotransposons. Examples of LINEs include, but are not limited to, LINE1 and LINE2. SINEs include, but are not limited to, one or more Alu sequences, one or more mammalian widely interspersed repeats (MIRs), or a combination thereof. Class II TEs (e.g., DNA transposons) often do not involve an RNA intermediate. DNA transposons often excise from one site in the genome and insert into another site. Alternatively, DNA transposons are replicated and inserted into new locations in the genome. Examples of DNA transposons include, but are not limited to, MER1, MER2, and mariner. Viral elements may include one or more endogenous retroviral sequences. Telomeres are often repeated DNA regions at the ends of chromosomes.
[0067] Some embodiments may include a subset of nucleic acid samples or nucleic acid molecules comprising one or more genomic regions, at least one of which comprises a genomic region feature comprising an entire exome or a portion thereof. An exome is often a portion of a genome formed by exons. An exome may be formed by untranslated regions (UTRs), splice sites and / or intronic regions. An entire exome or a protein thereof may comprise one or more exons of a protein-coding gene. An entire exome or a protein thereof may comprise one or more untranslated regions (UTRs), splice sites and / or introns.
[0068] Some embodiments may include a nucleic acid sample or molecule comprising one or more genomic regions, at least one of the one or more genomic regions comprising a genomic region feature comprising a gene or a portion thereof. A gene typically comprises a stretch of nucleic acid encoding a polypeptide or functional RNA. A gene may comprise one or more exons, one or more introns, one or more untranslated regions (UTRs), or a combination thereof. Exons are often coding sections of a gene that are transcribed into a pre-mRNA sequence and within the final mature RNA product of the gene. Introns are often non-coding sections of a gene that are transcribed into a pre-mRNA sequence and removed by RNA splicing. A UTR may refer to the sections on either side of a coding sequence of a strand of mRNA. A UTR located 5' of a coding sequence may be referred to as a 5'UTR (or leader sequence). A UTR located 3' of a coding sequence may be referred to as a 3'UTR (or trailer sequence). A UTR may comprise one or more elements for regulating gene expression. Elements such as regulatory elements may be located in the 5'UTR. Regulatory elements such as polyadenylation signals, protein binding sites, and miRNA binding sites may be located in the 3'UTR. Protein binding sites located in the 3'UTR may include, but are not limited to, selenocysteine insertion sequence (SECIS) elements and AU-rich elements (AREs). SECIS elements may direct ribosomes to translate the codon UGA as selenocysteine rather than a stop codon. AREs are often stretches consisting mainly of adenine and uracil nucleotides, which may affect mRNA stability.
[0069] Some embodiments may include a subset of nucleic acid samples or nucleic acid molecules comprising one or more genomic regions, at least one of the one or more genomic regions comprising a genomic region feature comprising a set of genes. The set of genes may include, but is not limited to, Mendel DB genes, Human Gene Mutation Database (HGMD) genes, Cancer Gene Census genes, Online Mendelian Inheritance in Man (OMIM) Mendelian genes, HGMD Mendelian genes, and human leukocyte antigen (HLA) genes. The set of genes may include one or more known Mendelian traits, one or more known disease traits, one or more drug traits, one or more biomedically interpretable variants, or combinations thereof. The Mendelian traits may be regulated by a single genetic locus and may exhibit a Mendelian inheritance pattern. A set of genes with known Mendelian traits may include one or more genes encoding Mendelian traits including, but not limited to, the ability to taste phenylthiocarbamide (dominant), the ability to smell hydrogen cyanide (bitter almond-like) (recessive), albinism (recessive), brachydactyly (short fingers and toes), and wet (dominant) or dry (recessive) ear wax. A disease trait causes a disease or increases the risk of a disease and can be inherited in a Mendelian or complex pattern. A set of genes with known disease traits may include one or more genes encoding disease traits including, but not limited to, cystic fibrosis, hemophilia, and Lynch syndrome. A drug trait can alter the metabolism, optimal dosage, adverse reactions, and side effects of one or more drugs or groups of drugs. A set of genes with known drug traits may include one or more genes encoding drug traits including, but not limited to, CYP2D6, UGT1A1, and ADRB1. Biomedically interpretable variants may be polymorphisms in genes that are associated with a disease or condition.The set of genes with known biomedically interpretable variants may include one or more genes encoding biomedically interpretable variants including, but not limited to, cystic fibrosis (CF) mutations, muscular dystrophy mutations, p53 mutations, Rb mutations, cell cycle control, receptors, and kinases. Alternatively, or in addition, the set of genes with known biomedically interpretable variants may include one or more genes associated with Huntington's disease, cancer, cystic fibrosis, muscular dystrophy (e.g., Duchenne muscular dystrophy).
[0070] Some embodiments may include a nucleic acid sample or molecule comprising one or more genomic regions, at least one of the one or more genomic regions comprising a genomic region feature comprising a regulatory element or a portion thereof. The regulatory element may be a cis-regulatory element or a trans-regulatory element. The cis-regulatory element may be a sequence that controls the transcription of a nearby gene. The cis-regulatory element may be located in a 5' or 3' untranslated region (UTR) or within an intron. The trans-regulatory element may control the transcription of a distant gene. The regulatory element may include one or more promoters, one or more enhancers, or a combination thereof. A promoter may promote the transcription of a particular gene and may be found upstream of a coding region. An enhancer may exert a distant effect on the transcription level of a gene.
[0071] Some embodiments may include a subset of nucleic acid samples or molecules that include one or more genomic regions, where at least one of the one or more genomic regions includes a genomic region feature that includes a polymorphism or a portion thereof. A polymorphism generally refers to a variation in a genotype. A polymorphism can be a germline variant or a somatic variant. A polymorphism may include one or more base changes, insertions, repeats, or deletions of one or more bases. Copy number variations (CNVs), transversions, and other rearrangements are also forms of genetic variation. Polymorphic markers include restriction fragment length polymorphisms, variable number tandem repeats (VNTRs), hypervariable regions, minisatellites, dinucleotide repeats, trinucleotide repeats, tetranucleotide repeats, simple sequence repeats, and insertion elements such as Alu. The allelic form that occurs most frequently in a selected population is sometimes referred to as the wild type. Diploid organisms may be homozygous or heterozygous for an allelic form. Diallelic polymorphisms have two forms. Triallelic polymorphisms have three forms. Single nucleotide polymorphisms (SNPs) are one form of polymorphism. In some embodiments, the one or more polymorphisms include one or more single nucleotide variations, indels, small insertions, small deletions, structural variant junctions, variable length tandem repeats, flanking sequences, or combinations thereof. The one or more polymorphisms may be located in coding and / or non-coding regions. The one or more polymorphisms may be located within, around, or near genes, exons, introns, splice sites, untranslated regions, or combinations thereof. The one or more polymorphisms may span at least genes, exons, introns, and untranslated regions.
[0072] Some embodiments may include a nucleic acid sample or molecule comprising one or more genomic regions, where at least one of the one or more genomic regions comprises a genomic region feature comprising one or more simple tandem repeats (STRs), unstable expanding repeats, segmental duplications, single and paired read degenerative mapping scores, GRCh37 patches, or a combination thereof. The one or more STRs may comprise one or more homopolymers, one or more dinucleotide repeats, one or more trinucleotide repeats, or a combination thereof. The one or more homopolymers may be about 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more bases or base pairs. The dinucleotide repeats and / or trinucleotide repeats may be 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50 or more bases or base pairs. Single and paired read degenerative mapping scores may be based on or derived from 100-mer GEM alignability from ENCODE / CRG (Guigo), 75-mer GEM alignability from ENCODE / CRG (Guigo), 100 base pair box car average for signal mappability, max of locus and possible pairs for paired read score, or combinations thereof. The genomic region features may include one or more low average coverage regions from whole genome sequencing (WGS), zero average coverage regions from WGS, validated compressions, or combinations thereof.The WGS-derived low average coverage regions may include regions generated from Illumina v3 chemistry, regions below the first percentile of a Poisson distribution based on average coverage, or combinations thereof. The WGS-derived zero average coverage regions may include regions generated from Illumina v3 chemistry. The validated compressions may include regions of high mapped depth, regions with two or more observed haplotypes, regions expected to be missing repeats in a reference, or combinations thereof. The genomic region features may include one or more alternative or non-reference sequences. The one or more alternative or non-reference sequences may include known structural variant junctions, known insertions, known deletions, alternative haplotypes, or combinations thereof. The genomic region features may include one or more gene phasing and reconstructed genes. Examples of phasing and remodeling genes include, but are not limited to, one or more of major histocompatibility complex, blood group, and amylase gene family. The one or more major histocompatibility complex may include one or more of HLA class I, HLA class II, or a combination thereof. The one or more HLA class I may include HLA-A, HLA-B, HLA-C, or a combination thereof. The one or more HLA class II may include HLA-DP, HLA-DM, HLA-DOA, HLA-DOB, HLA-DQ, HLA-DR, or a combination thereof. The blood group genes may include ABO, RHD, RHCE, or a combination thereof.
[0073] Some embodiments may include a nucleic acid sample or molecule comprising one or more genomic regions, at least one of the one or more genomic regions comprising a genomic region feature related to the GC content of one or more nucleic acid molecules. GC content may refer to the GC content of the nucleic acid molecule. Alternatively, GC content may refer to the GC content of one or more nucleic acid molecules and may be referred to as the average GC content. As used herein, the terms "GC content" and "average GC content" may be used interchangeably. The GC content of the genomic region may be a high GC content. Typically, a high GC content refers to a GC content of about 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97% or more. In some embodiments, a high GC content may refer to a GC content of about 70% or more. The GC content of the genomic region may be a low GC content. Typically, low GC content refers to a GC content of about 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 5%, 2% or less.
[0074] Some embodiments may include a nucleic acid sample or molecule comprising one or more genomic regions, at least one of the one or more genomic regions comprising a genomic region characteristic associated with the complexity of one or more nucleic acid molecules. The complexity of a nucleic acid molecule may refer to the randomness of the nucleotide sequence. Low complexity may refer to the pattern, repetition and / or depletion of one or more species of nucleotides in the sequence.
[0075] Some embodiments may include a nucleic acid sample or molecule comprising one or more genomic regions, at least one of the one or more genomic regions comprising genomic region features related to the mappability of one or more nucleic acid molecules. The mappability of a nucleic acid molecule may refer to the uniqueness of the alignment of the nucleic acid molecule to a reference sequence. A nucleic acid molecule with low mappability may have low alignment to a reference sequence.
[0076] D. Predicting whether a candidate variant is a somatic variant The two-model classification method is used to predict somatic variants from an attribute table. For example, the attribute table can be subdivided into two datasets as shown in FIG. 4 and processed using a trained model, such as the model described in Example 3. The first dataset can include candidate somatic variants identified by one or more bioinformatics tools. The first model can be applied to remove false positives from this dataset. The second dataset can include the remainder of the candidate variants, including false negatives and true negatives. The second model can be applied to rescue false negatives from this dataset. This method can predict somatic variants with an acceptable accuracy rate, despite the lack of corresponding normal samples.
[0077] To improve the controllability of the threshold, the somatic variant classification problem is decomposed into two sub-decomposition problems: (1) removing false positives during tumor-only calls from each variant caller, and (2) rescuing false negative candidate variants that are not present in the tumor-only calls. The attribute table is subdivided into two datasets. The first dataset contains the candidate variants identified by MuTect and MuTect2 (tumor-only context). A first model is trained to remove false positives from this dataset. The second dataset contains the rest of the candidate variants. A second model is trained to rescue false negatives from this dataset. The models are trained using Microsoft's LightGBM framework (LGBM). The classification results from both of these models are then combined to create the final set of candidate variants.
[0078] E. Generate a report identifying candidate variants One or more reports can be generated (diagnostic and / or prognostic reports) that include some or all of the predicted candidate variants. Based on the predicted candidate variants and / or the reports, one or more treatments can be administered or not administered to the patient. For example, the predicted candidate variants can be compared to one or more databases of known cancer mutations to diagnose or characterize cancer. Variants associated with responsiveness or non-responsiveness to a particular cancer treatment can be identified and treatment recommendations can be provided. The cancer can be treated based on the recommendations.
[0079] IV. Process for Somatic Variants from Unmatched Biological Samples FIG. 9 includes a flowchart 900 illustrating an example of a method for somatic variant calling from unmatched biological samples, according to some embodiments. The operations described in the flowchart 900 may be performed by a computer system running a trained machine learning model, including, for example, a filtering model and a rescue model. Although the flowchart 900 may describe the operations as a sequential process, in various embodiments, many of the operations may be performed in parallel or simultaneously. Furthermore, the order of the operations may be rearranged. The operations may have additional steps not shown in the figures. Furthermore, the method embodiments may be performed by hardware, software, firmware, middleware, microcode, hardware description languages, or combinations thereof. When executed by the software, firmware, middleware, or microcode, the program code or code segments for performing the associated tasks may be stored in a computer-readable medium, such as a storage medium.
[0080] In operation 910, the computer system obtains nucleic acid sequence data of a biological sample of the subject. The nucleic acid sequence data can be generated by sequencing a plurality of nucleic acid molecules of a tumor sample. In some embodiments, the tumor sample is from a human subject. The sequencing can include whole exome sequencing. In some embodiments, the sequencing can include whole genome sequencing. In some embodiments, the sequencing includes shotgun sequencing. In some embodiments, the sequencing includes sequencing a selected portion of a genome or exome.
[0081] In operation 920, the computer system aligns the nucleic acid sequence data to the reference genome. For example, a FASTQ file, corresponding to the nucleic acid sequence data, can be aligned to the reference genome to create one or more BAM files.
[0082] In operation 930, the computer system identifies a set of candidate variants in the nucleic acid sequence data based on the aligned nucleic acid sequence data. In some examples, the set of candidate variants includes one or more somatic variants and one or more germline variants. Somatic variants refer to DNA changes that occur after conception and are not present in the germline. Germline variants refer to genetic changes in germ cells (sperm and egg cells) that are incorporated into the DNA of all cells of the body of an offspring. In some examples, somatic variants, instead of germline variants, are indicative of the presence or level of cancer in a subject.
[0083] An attribute table can be created, where the attribute table can include a number of features for each candidate variant. In some embodiments, the attribute table includes attributes from the sequencing data corresponding to a particular candidate variant. The attribute table can include attributes from a file including processed sequencing data. In some embodiments, the attribute table includes one or more attributes, such as: (a) pileup attributes from a BCFtools output file; (b) allele frequency data; (c) base quality data; (d) read depth data; (e) tumor cytology estimates (which may be calculated based on B allele frequency distributions); (f) predicted germline variants; (g) predicted somatic variants; (h) copy number alteration data; (i) population frequency data from one or more databases; (j) data from at least one database selected from the group consisting of Cosmic, GnomAD, Dbsnp, and Mills Indels; (k) data regarding the presence of the candidate somatic variant in problematic regions of the genome; and (l) data regarding the presence of the candidate somatic variant in homopolymers.
[0084] In operation 940, the computer system processes the set of candidate variants using the trained machine learning model to identify somatic variants without using nucleic acid sequence data of the subject's corresponding biological sample. In some examples, the trained machine learning model includes a gradient boosting decision tree that facilitates a significant reduction in the false positive rate corresponding to the somatic variant call. In some embodiments, the trained machine learning model includes a two-model classification method. The trained machine learning model may include a filtering model that removes false positives. The trained machine learning model may also include a rescue model that rescues false negatives. In some embodiments, the attribute table includes attributes from the sequencing data.
[0085] In operation 950, the computer system outputs a report identifying the somatic variants. In some embodiments, the report includes information identifying at least one diagnostic marker, at least one prognostic marker, in some embodiments, the absence of a somatic variant, a treatment recommendation, a recommendation to administer a treatment to the human subject, and / or a recommendation not to administer a treatment to the human subject. In some embodiments, the recommended treatment is administered to the human subject. Process 900 then ends.
[0086] V. Further Considerations A. Search technology Some embodiments may include one or more labels. One or more labels may be attached to one or more capture probes, nucleic acid molecules, beads, primers, or combinations thereof. Examples of labels include, but are not limited to, detectable labels such as radioisotopes, fluorophores, chemiluminophores, chromophores, lumiphores, enzymes, colloidal particles, and fluorescent microparticles, quantum dots, and antigens, antibodies, haptens, avidin / streptavidin, biotin, haptens, enzyme cofactors / substrates, one or more members of a quenching system, chromogens, haptens, magnetic particles, substances exhibiting nonlinear optics, semiconductor nanocrystals, metal nanoparticles, enzymes, aptamers, and one or more members of a binding pair.
[0087] Some embodiments may include one or more capture probes, multiple capture probes, or one or more capture probe sets. Typically, the capture probe comprises a nucleic acid binding site. The capture probe may further comprise one or more linkers. The capture probe may further comprise one or more labels. The one or more linkers may attach one or more labels to the nucleic acid binding site.
[0088] The capture probe may hybridize to one or more nucleic acid molecules in the sample. The capture probe may hybridize to one or more genomic regions. The capture probe may hybridize to one or more genomic regions within, surrounding, near or spanning one or more genes, exons, introns, UTRs, or combinations thereof. The capture probe may hybridize to one or more genomic regions spanning one or more genes, exons, introns, UTRs, or combinations thereof. The capture probe may hybridize to one or more known indels. The capture probe may hybridize to one or more known structural variants.
[0089] Some embodiments may include one or more capture probes or capture probe sets: 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more. The one or more capture probes or capture probe sets may be different, similar, identical, or a combination thereof.
[0090] The one or more capture probes may comprise a nucleic acid binding site that hybridizes to at least a portion of one or more nucleic acid molecules or variants or derivatives thereof or a subset of nucleic acid molecules in a sample. The capture probes may comprise a nucleic acid binding site that hybridizes to one or more genomic regions. The capture probes may hybridize to different, similar, and / or identical genomic regions. The one or more capture probes may be at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99% or more complementary to one or more nucleic acid molecules or variants or derivatives thereof.
[0091] The capture probe may comprise one or more nucleotides. The capture probe may comprise 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more nucleotides. The capture probe may comprise about 100 nucleotides. The capture probe may comprise from about 10 to about 500 nucleotides, from about 20 to about 450 nucleotides, from about 30 to about 400 nucleotides, from about 40 to about 350 nucleotides, from about 50 to about 300 nucleotides, from about 60 to about 250 nucleotides, from about 70 to about 200 nucleotides, or from about 80 to about 150 nucleotides. In some embodiments, the capture probe may comprise from about 80 nucleotides to about 100 nucleotides.
[0092] A plurality of capture probes or a capture probe set may include two or more capture probes with identical, similar, and / or different nucleic acid binding site sequences, linkers, and / or labels. For example, two or more capture probes include identical nucleic acid binding sites. In another example, two or more capture probes include similar nucleic acid binding sites. In yet another example, two or more capture probes include different nucleic acid binding sites. Two or more capture probes may further include one or more linkers. Two or more capture probes may further include different linkers. Two or more capture probes may further include similar linkers. Two or more capture probes may further include identical linkers. Two or more capture probes may further include one or more labels. Two or more capture probes may further include different labels. Two or more capture probes may further include similar labels. Two or more capture probes may further include identical labels.
[0093] B. Assays and Amplification Techniques Some embodiments may include performing one or more assays on a sample comprising one or more nucleic acid molecules. Producing two or more subsets of nucleic acid molecules may include performing one or more assays. An assay may be performed on a subset of nucleic acid molecules from the sample. An assay may be performed on one or more nucleic acid molecules from the sample. An assay may be performed on at least a portion of the subset of nucleic acid molecules. An assay may include one or more techniques, reagents, capture probes, primers, labels, and / or components for detection, quantification, and / or analysis of one or more nucleic acid molecules.
[0094] The assay may include, but is not limited to, sequencing, amplification, hybridization, enrichment, separation, elution, fragmentation, detection, quantification of one or more nucleic acid molecules. The assay may include a method of preparing one or more nucleic acid molecules.
[0095] Some embodiments may include performing one or more amplification reactions on one or more nucleic acid molecules in a sample. The term "amplification" refers to any process that creates at least one copy of a nucleic acid molecule. The terms "amplification product" and "amplified nucleic acid molecule" refer to copies of a nucleic acid molecule and can be used interchangeably. The amplification reaction can include a PCR-based method, a non-PCR-based method, or a combination thereof. Examples of non-PCR-based methods include, but are not limited to, multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, rolling circle amplification, or circle-to-circle amplification. PCR-based methods can include, but are not limited to, PCR, HD-PCR, next-generation PCR (Next Gen PCR), digital RTA, or any combination thereof. Additional PCR methods include, but are not limited to, linear amplification, allele-specific PCR, Alu PCR Alu PCR, assembly PCR, asymmetric PCR, droplet PCR, emulsion PCR, helicase-dependent amplification HDA, hot start PCR, inverse PCR, exponential-after-linear (LATE)-PCR (linear-after-the-exponential (LATE)-PCR), long PCR, multiplex PCR, nested PCR, hemi-nested PCR, quantitative PCR, RT-PCR, real-time PCR, single-cell PCR, and touchdown PCR.
[0096] Some embodiments may include performing one or more hybridization reactions on one or more nucleic acid molecules in a sample. The hybridization reaction may include hybridization of one or more capture probes to one or more nucleic acid molecules or a subset of nucleic acid molecules in a sample. The hybridization reaction may include hybridization of one or more sets of capture probes to one or more nucleic acid molecules or a subset of nucleic acid molecules in a sample. The hybridization reaction may include one or more hybridization arrays, multiplex hybridization reactions, hybridization chain reactions, isothermal hybridization reactions, nucleic acid hybridization reactions, or combinations thereof. The one or more hybridization arrays may include hybridization array genotyping, hybridization array proportional sensing, DNA hybridization arrays, macroarrays, microarrays, high density oligonucleotide arrays, genomic hybridization arrays, comparative hybridization arrays, or combinations thereof. A hybridization reaction may include one or more capture probes, one or more beads, one or more labels, one or more subsets of nucleic acid molecules, one or more nucleic acid samples, one or more reagents, one or more wash buffers, one or more elution buffers, one or more hybridization buffers, one or more hybridization chambers, one or more incubators, one or more separators, or combinations thereof.
[0097] Some embodiments may include performing one or more enrichment reactions on one or more nucleic acid molecules in a sample. The enrichment reaction may include contacting the sample with one or more beads or bead sets. The enrichment reaction may include differential amplification of two or more subsets of nucleic acid molecules based on one or more genomic region features. For example, the enrichment reaction includes differential amplification of two or more subsets of nucleic acid molecules based on GC content. Alternatively, or in addition, the enrichment reaction includes differential amplification of two or more subsets of nucleic acid molecules based on methylation sites. The enrichment reaction may include one or more hybridization reactions. The enrichment reaction may further include separation and / or purification of one or more hybridized nucleic acid molecules, one or more bead-bound nucleic acid molecules, one or more free nucleic acid molecules (e.g., nucleic acid molecules without capture probes, nucleic acid molecules without beads), one or more labeled nucleic acid molecules, one or more unlabeled nucleic acid molecules, one or more amplification products, one or more unamplified nucleic acid molecules, or combinations thereof. Alternatively, or in addition, the enrichment reaction may include enriching one or more cell types in the sample. One or more cell types may be enriched by flow cytometry.
[0098] The enrichment reaction or reactions may produce one or more enriched nucleic acid molecules. The enriched nucleic acid molecules may include nucleic acid molecules or variants or derivatives thereof. For example, the enriched nucleic acid molecules include one or more hybridized nucleic acid molecules, one or more bead-bound nucleic acid molecules, one or more free nucleic acid molecules (e.g., nucleic acid molecules without capture probes, nucleic acid molecules without beads), one or more labeled nucleic acid molecules, one or more unlabeled nucleic acid molecules, one or more amplification products, one or more unamplified nucleic acid molecules, or combinations thereof. The enriched nucleic acid molecules may be distinguished from non-enriched nucleic acids by GC content, molecular size, genomic region, genomic region features, or combinations thereof. The enriched nucleic acid molecules may originate from one or more assays, supernatants, eluates, or combinations thereof. The enriched nucleic acid molecules may differ from the non-enriched nucleic acid molecules in average size, average GC content, genomic region, or combinations thereof.
[0099] Some embodiments may include performing one or more separation or purification reactions on one or more nucleic acid molecules in a sample. The separation or purification reaction may include contacting the sample with one or more beads or bead sets. The separation or purification reaction may include one or more hybridization reactions, enrichment reactions, amplification reactions, sequencing reactions, or combinations thereof. The separation or purification reaction may include the use of one or more separators. The one or more separators may include a magnetic separator. The separation or purification reaction may include separating bead-bound nucleic acid molecules from bead-free nucleic acid molecules. The separation or purification reaction may include separating capture probe-hybridized nucleic acid molecules from capture probe-free nucleic acid molecules. The separation or purification reaction may include separating a first subset of nucleic acid molecules from a second subset of nucleic acid molecules, the first subset of nucleic acid molecules differing from the second subset of nucleic acid molecules in average size, average GC content, genomic region, or combinations thereof.
[0100] Some embodiments may include performing one or more elution reactions on one or more nucleic acid molecules in the sample. The elution reaction may include contacting the sample with one or more beads or bead sets. The elution reaction may include separating bead-bound nucleic acid molecules from bead-free nucleic acid molecules. The elution reaction may include separating capture probe-hybridized nucleic acid molecules from capture probe-free nucleic acid molecules. The elution reaction may include separating a first subset of nucleic acid molecules from a second subset of nucleic acid molecules, the first subset of nucleic acid molecules differing from the second subset of nucleic acid molecules in average size, average GC content, genomic region, or combinations thereof.
[0101] Some embodiments may include one or more fragmentation reactions. The fragmentation reaction may include fragmenting one or more nucleic acid molecules or a subset of nucleic acid molecules in a sample to generate one or more fragmented nucleic acid molecules. The one or more nucleic acid molecules may be fragmented by sonication, needle shearing, nebulization, shearing (e.g., acoustic shearing, mechanical shearing, point-sink shearing, etc.), passage through a French pressure cell, or enzymatic digestion. Enzymatic digestion may include nuclease digestion (e.g., micrococcal nuclease digestion, endonuclease, exonuclease, ribonuclease H (RNAse H) or deoxyribonuclease I (DNase I). Fragmentation of the one or more nucleic acid molecules may produce fragments of about 100 base pairs to about 2000 base pairs, about 200 base pairs to about 1500 base pairs, about 200 base pairs to about 1000 base pairs, about 200 base pairs to about 500 base pairs, about 500 base pairs to about 1500 base pairs, about 500 base pairs to about 1000 base pairs in size. The one or more fragmentation reactions may produce fragments of about 50 base pairs to about 1000 base pairs in size. The one or more fragmentation reactions may produce fragments that are about 100 base pairs, 150 base pairs, 200 base pairs, 250 base pairs, 300 base pairs, 350 base pairs, 400 base pairs, 450 base pairs, 500 base pairs, 550 base pairs, 600 base pairs, 650 base pairs, 700 base pairs, 750 base pairs, 800 base pairs, 850 base pairs, 900 base pairs, 950 base pairs, 1000 base pairs, or more in size.
[0102] Fragmenting one or more nucleic acid molecules may include mechanically shearing one or more nucleic acid molecules in a sample for a period of time. The fragmentation reaction may occur for at least about 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500 seconds, or more.
[0103] Fragmenting the one or more nucleic acid molecules may include contacting the nucleic acid sample with one or more beads. Fragmenting the one or more nucleic acid molecules may include contacting the nucleic acid sample with a plurality of beads, wherein the ratio of the volume of the plurality of beads to the volume of the nucleic acid sample is about 0.10, 0.20, 0.30, 0.40, 0.50, 0.60, 0.70, 0.80, 0.90, 1.00, 1.10, 1.20, 1.30, 1.40, 1.50, 1.60, 1.70, 1.80, 1.90, 2.00 or more. Fragmenting the one or more nucleic acid molecules may include contacting the nucleic acid sample with a plurality of beads, wherein the ratio of the volume of the plurality of beads to the volume of the nucleic acid is about 2.00, 1.90, 1.80, 1.70, 1.60, 1.50, 1.40, 1.30, 1.20, 1.10, 1.00, 0.90, 0.80, 0.70, 0.60, 0.50, 0.40, 0.30, 0.20, 0.10, 0.05, 0.04, 0.03, 0.02, 0.01 or less.
[0104] Some embodiments may include performing one or more detection reactions on one or more nucleic acid molecules in the sample. The detection reaction may include one or more sequencing reactions. Alternatively, performing the detection reaction includes optical sensing, electrical sensing, or a combination thereof. The optical sensing may include optical sensing of photoluminescence photon emission, fluorescence photon emission, pyrophosphate photon emission, chemiluminescence photon emission, or a combination thereof. The electrical sensing may include electrical sensing of ion concentration, ion current modulation, nucleotide electric field, nucleotide tunneling current, or a combination thereof.
[0105] Some embodiments may include performing one or more quantification reactions on one or more nucleic acid molecules in the sample. The quantification reactions may include sequencing, PCR, qPCR, digital PCR, or a combination thereof.
[0106] Some embodiments may include one or more samples. Some embodiments may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more samples. The samples may be derived from a subject. The two or more samples may be derived from a single subject. The two or more samples may be derived from 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more different subjects. The subject may be a mammal, reptile, amphibian, bird, and fish. The mammal may be a human, ape, orangutan, monkey, chimpanzee, cow, pig, horse, rodent, bird, reptile, dog, cat, or other animal. The reptile may be a lizard, snake, alligator, turtle, crocodile, and turtle. The amphibian may be a toad, frog, newt, and salamander. Examples of birds include, but are not limited to, ducks, geese, penguins, ostriches, and owls. Examples of fish include, but are not limited to, catfish, eels, sharks, and swordfish. Preferably, the subject is a human. The subject may be suffering from a disease or condition (e.g., cancer).
[0107] Two or more samples may be collected over 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 or more time points. The time points may occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 hours or more. The time points may occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 days or more. The time points may occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 weeks or more. The time points may occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 months or more. The time points may occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 years or more.
[0108] The sample may be from a body fluid, a cell, skin, a tissue, an organ, or a combination thereof. The sample may be blood, plasma, a blood fraction, saliva, sputum, urine, semen, vaginal fluid, cerebrospinal fluid, stool, a cell or tissue biopsy. The sample may be from the adrenal gland, appendix, bladder, brain, ear, esophagus, eye, gallbladder, heart, kidney, large intestine, liver, lung, mouth, muscle, nose, pancreas, parathyroid gland, pineal gland, pituitary gland, skin, small intestine, spleen, stomach, thymus, thyroid gland, trachea, uterus, vermiform appendix, cornea, skin, heart valve, artery, or vein.
[0109] The sample may include one or more nucleic acid molecules. The nucleic acid molecules may be DNA molecules, RNA molecules (e.g., messenger RNA (mRNA), complementary RNA (cRNA) or microRNA (miRNA)), and DNA / RNA hybrids. Examples of DNA molecules include, but are not limited to, double-stranded DNA, single-stranded DNA, single-stranded DNA hairpins, complementary DNA (cDNA), and genomic DNA. The nucleic acid may be an RNA molecule, such as double-stranded RNA, single-stranded RNA, non-coding RNA (ncRNA), RNA hairpins, and messenger RNA (mRNA). Examples of non-coding RNA (ncRNA) include, but are not limited to, small interfering RNA (siRNA), microRNA (miRNA), small nucleolar ribonucleic acid (snoRNA), Piwi-binding RNA (piRNA), transcription initiator RNA (tiRNA), PASR, TASR, aTASR, TSSa-RNA, small nuclear RNA (snRNA), RE-RNA, uaRNA, x-ncRNA, hYRNA, usRNA, snaR, and vtRNA.
[0110] Some embodiments may include one or more containers. Some embodiments may include 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more containers. The one or more containers may be different, similar, identical, or a combination thereof. Examples of vessels include, but are not limited to, plates, microplates, PCR plates, wells, microwells, tubes, Eppendorf tubes, vials, arrays, microarrays, and chips.
[0111] Some embodiments may include one or more reagents. Some embodiments may include 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more reagents. One or more reagents may be different, similar, identical, or a combination thereof. A reagent may improve the efficiency of one or more assays. The reagent may enhance the stability of the nucleic acid molecule or variant or derivative thereof. The reagent may include, but is not limited to, enzymes, proteases, nucleases, molecules, polymerases, reverse transcriptases, ligases, and chemicals. Some embodiments may include performing an assay that includes one or more antioxidants. Generally, an antioxidant is a molecule that inhibits the oxidation of another molecule. Examples of antioxidants include, but are not limited to, ascorbic acid (e.g., vitamin C), glutathione, lipoic acid, uric acid, carotene, α-tocopherol (e.g., vitamin E), ubiquinol (e.g., coenzyme Q), and vitamin A.
[0112] Some embodiments may include one or more buffers or solutions. The one or more buffers or solutions may be different, similar, identical, or a combination thereof. The buffers or solutions may improve the efficiency of one or more assays. The buffers or solutions may enhance the stability of the nucleic acid molecules or variants or derivatives thereof. The buffers or solutions may include, but are not limited to, wash buffers, elution buffers, and hybridization buffers.
[0113] Some embodiments may include one or more beads, multiple beads, or one or more bead sets. Some embodiments may include one or more beads or bead sets of 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more. One or more beads or bead sets may be different, similar, identical, or a combination thereof. The beads may be magnetic, antibody-coated, protein A-crosslinked, protein G-crosslinked, streptavidin-coated, oligonucleotide-linked, silica-coated, or combinations thereof. Examples of beads include, but are not limited to, Ampur beads, Ampur XP beads, streptavidin beads, agarose beads, magnetic beads, Dynabeads®, MACS® microbeads, antibody-linked beads (e.g., anti-immunoglobulin microbeads), protein A-linked beads, protein G-linked beads, protein A / G-linked beads, protein L-linked beads, oligo dT-linked beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorescent dye microbeads, and BcMag™ carboxy-terminated magnetic beads. In some embodiments, one or more beads comprise one or more Ampur beads. Alternatively, or in addition, one or more beads comprise Ampur XP beads.
[0114] Some embodiments may include one or more primers, multiple primers, or one or more primer sets. The primers may further include one or more linkers. The primers may further include one or more labels. The primers may be used in one or more assays. For example, the primers are used in one or more sequencing reactions, amplification reactions, or a combination thereof. Some embodiments may comprise one or more primers or primer sets: 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more. The primers may comprise about 100 nucleotides. The primers may comprise about 10 to about 500 nucleotides, about 20 to about 450 nucleotides, about 30 to about 400 nucleotides, about 40 to about 350 nucleotides, about 50 to about 300 nucleotides, about 60 to about 250 nucleotides, about 70 to about 200 nucleotides, or about 80 to about 150 nucleotides. In some embodiments, the primers comprise about 80 nucleotides to about 100 nucleotides. The one or more primers or primer sets may be different, similar, identical, or a combination thereof.
[0115] The primers may hybridize to at least a portion of one or more nucleic acid molecules or variants or derivatives thereof or a subset of nucleic acid molecules in a sample. The primers may hybridize to one or more genomic regions. The primers may hybridize to different, similar, and / or identical genomic regions. The one or more primers may be at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99% or more complementary to one or more nucleic acid molecules or variants or derivatives thereof.
[0116] The primer may comprise one or more nucleotides. The primer may comprise 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more nucleotides. The primer may comprise about 100 nucleotides. The primer may comprise from about 10 to about 500 nucleotides, from about 20 to about 450 nucleotides, from about 30 to about 400 nucleotides, from about 40 to about 350 nucleotides, from about 50 to about 300 nucleotides, from about 60 to about 250 nucleotides, from about 70 to about 200 nucleotides, or from about 80 to about 150 nucleotides. In some embodiments, the primer comprises from about 80 nucleotides to about 100 nucleotides.
[0117] A plurality of primers or primer sets may include two or more primers with identical, similar, and / or different sequences, linkers, and / or labels. For example, two or more primers include identical sequences. In another example, two or more primers include similar sequences. In yet another example, two or more primers include different sequences. Two or more primers may further include one or more linkers. Two or more primers may further include different linkers. Two or more primers may further include similar linkers. Two or more primers may further include identical linkers. Two or more primers may further include one or more labels. Two or more primers may further include different labels. Two or more primers may further include similar labels. Two or more primers may further include identical labels.
[0118] The capture probe, primer, label, and / or bead may comprise one or more nucleotides that may include RNA, DNA, a mixture of DNA and RNA residues, or modified analogs thereof, such as 2'-0Me or 2'-fluoro (2'-F), locked nucleic acid (LNA), or abasic sites.
[0119] Some embodiments may include one or more labels. Some embodiments may include one or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more labels. The one or more labels may be different, similar, identical, or a combination thereof.
[0120] Examples of labels include, but are not limited to, chemical, biochemical, biological, chromogenic, enzymatic, fluorescent, and luminescent labels that are well known in the art. Labels include dyes, photocrosslinkers, cytotoxic compounds, drugs, affinity labels, photoaffinity labels, reactive compounds, antibodies or antibody fragments, biomaterials, nanoparticles, spin labels, fluorophores, metal-containing moieties, radioactive moieties, novel functional groups, groups that interact covalently or non-covalently with other molecules, photocaged moieties, actinic excitable moieties, ligands, photoisomerizable moieties, biotin, biotin analogs, moieties incorporating heavy atoms, chemically cleavable groups, photocleavable groups, redox-active agents, isotopically labeled moieties, biophysical probes, phosphorescent groups, chemiluminescent groups, electron dense groups, magnetic groups, intercalating groups, chromophores, energy transfer agents, bioactive agents, detectable labels, or combinations thereof.
[0121] The label may be a chemical label. Examples of chemical labels may include, but are not limited to, biotin and radioisotopes (e.g., iodine, carbon, phosphate, hydrogen).
[0122] The methods, kits, and compositions disclosed herein may include biolabels, including, but not limited to, metabolic labels, including bioorthogonal azide-modified amino acids, sugars, and other compounds.
[0123] The methods, kits, and compositions disclosed herein may include an enzyme label. The enzyme label may include, but is not limited to, horseradish peroxidase (HRP), alkaline phosphatase (AP), glucose oxidase, and O-galactosidase. The enzyme label may be luciferase.
[0124] The methods, kits, and compositions disclosed herein may include a fluorescent label. The fluorescent label may include an organic dye (e.g., FITC), a biological fluorophore (e.g., green fluorescent protein), or a quantum dot. A non-limiting list of fluorescent labels includes fluorescein isothiocyanate (FITC), DyLight Fluor, fluorescein, rhodamine (tetramethylrhodamine isothiocyanate, TRITC), coumarin, lucifer yellow, and BODIPY. The label may be a fluorophore. Exemplary fluorophores include, but are not limited to, indocarbocyanine (C3), indodicarbocyanine (C5), Cy3, Cy3.5, Cy5, Cy5.5, Cy7, Texas Red, Pacific Blue, Oregon Green 488, Alexa Fluor®-355, Alexa Fluor 488, Alexa Fluor 532, Alexa Fluor 546, Alexa Fluor-555, Alexa Fluor 568, Alexa Fluor 594, Alexa Fluor 647, Alexa Fluor 660, Alexa Fluor 680, JOE, Lissamine, Rhodamine. Fluorescent labels include Green, BODIPY, fluorescein isothiocyanate (FITC), carboxyfluorescein (FAM), phycoerythrin, rhodamine, dichlororhodamine (dRhodamine), carboxytetramethylrhodamine (TAMRA), carboxy-X-rhodamine (ROX™), LIZ™, VICTM NED™ PET™, SYBR, PicoGreen, RiboGreen, and the like. The fluorescent label may be green fluorescent protein (GFP), red fluorescent protein (RFP), yellow fluorescent protein, phycobiliproteins (e.g., allophycocyanin, phycocyanin, phycoerythrin, and phycoerythrocyanin).
[0125] Some embodiments may include one or more linkers. Some embodiments may include one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, twenty or more, thirty or more, forty or more, fifty or more, sixty or more, seventy or more, eighty or more, ninety or more, one hundred or more, one hundred or more, twenty-five or more, fifty or more, one hundred or more, seventy-five or more, three hundred or more, three hundred or more, four hundred or more, five hundred or more, six hundred or more, seven hundred or more, eight hundred or more, nine hundred or more, or one thousand or more linkers. The linkers may be different, similar, identical, or a combination thereof.
[0126] Suitable linkers include any chemical or biological compound that can be attached to the label, primer, and / or capture probe disclosed herein. When the linker is attached to both the label and the primer or capture probe, a suitable linker will be able to sufficiently separate the label and the primer or capture probe. A suitable linker will not significantly interfere with the ability of the primer and / or capture probe to hybridize to the nucleic acid molecule, a portion thereof, or a variant or derivative thereof. A suitable linker will not significantly interfere with the ability to detect the label. The linker may be rigid. The linker may be flexible. The linker may be semi-rigid. The linker may be proteolytically stable (e.g., resistant to proteolytic cleavage). The linker may be proteolytically unstable (e.g., susceptible to proteolytic cleavage). The linker may be helical. The linker may be non-helical. The linker may be coiled. The linker may be 3-stranded. The linker may include a rotated conformation. The linker may be single chain. The linker may be long chain. The linker may be short chain. The linker may be at least about 5 residues, at least about 10 residues, at least about 15 residues, at least about 20 residues, at least about 25 residues, at least about 30 residues, at least about 40 residues, or more.
[0127] Examples of linkers include, but are not limited to, hydrazones, disulfides, thioethers, and peptide linkers. The linker may be a peptide linker. The peptide linker may include a proline residue. The peptide linker may include arginine, phenylalenine, threonine, glutamine, glutamic acid, or any combination thereof. The linker may be a heterobifunctional crosslinker.
[0128] Some embodiments may include performing 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 20 or more, 25 or more, 30 or more, 35 or more, 40 or more, 45 or more, or 50 or more assays on a sample containing one or more nucleic acid molecules. The two or more assays may be different, similar, identical, or a combination thereof. For example, some embodiments include performing two or more sequencing reactions. In another example, some embodiments include performing two or more assays, where at least one of the two or more assays includes a sequencing reaction. In yet another example, some embodiments include performing two or more assays, where at least two of the two or more assays include a sequencing reaction and a hybridization reaction. The two or more assays may be performed sequentially, simultaneously, or a combination thereof. For example, two or more sequencing reactions may be performed simultaneously. In another example, some embodiments include performing a hybridization reaction followed by a sequencing reaction. In yet another example, some embodiments include performing two or more hybridization reactions simultaneously followed by two or more sequencing reactions simultaneously. Two or more assays may be performed by one or more devices. For example, two or more amplification reactions may be performed by a PCR device. In another example, two or more sequencing reactions may be performed by two or more sequencers.
[0129] C. Equipment Some embodiments may include one or more devices. Some embodiments may include one or more assays that include one or more devices. Some embodiments may include the use of one or more devices to perform one or more steps or assays. Some embodiments may include the use of one or more devices in one or more steps or assays. For example, performing a sequencing reaction may include one or more sequencers. In another example, generating a subset of nucleic acid molecules may include the use of one or more magnetic separators. In yet another example, one or more processors may be used in the analysis of one or more nucleic acid samples. Exemplary devices include, but are not limited to, sequencers, thermocyclers, real-time PCR instruments, magnetic separators, transmission devices, hybridization chambers, electrophoresis devices, centrifuges, microscopes, imaging devices, fluorometers, luminometers, plate readers, computers, processors, and bioanalyzers.
[0130] Some embodiments may include one or more sequencers. The one or more sequencers may include one or more HiSeq, MiSeq, HiScan, Genome Analyzer IIx, SOLiD sequencer, Ion Torrent PGM, 454 GS Junior, Pac Bio RS, or a combination thereof. The one or more sequencers may include one or more sequencing platforms. The one or more sequencing platforms may include 454 GS FLX by Life Technologies / Roche, Genome Analyzer by Solexa / Illumina, SOLiD by Applied Biosystems, CGA Platform by Complete Genomics, PacBio RS by Pacific Biosciences, or a combination thereof.
[0131] Some embodiments may include one or more thermocyclers. The one or more thermocyclers may be used to amplify one or more nucleic acid molecules. Some embodiments may include one or more real-time PCR instruments. The one or more real-time PCR instruments may include a thermal cycler and a fluorometer. The one or more thermocyclers may be used to amplify and detect one or more nucleic acid molecules.
[0132] Some embodiments may include one or more magnetic separators. The one or more magnetic separators may be used to separate the paramagnetic and ferromagnetic particles from the suspension. The one or more magnetic separators may include one or more LifeStep™ biomagnetic separators, SPHERO™ FlexiMag separators, SPHERO™ MicroMag separators, SPHERO™ HandiMag separators, SPHERO™ MiniTube Mag separators, SPHERO™ UltraMag separators, DynaMag™ magnets, DynaMag™ 2 magnets, or combinations thereof.
[0133] Some embodiments may include one or more bioanalyzers. Generally, a bioanalyzer is a microchip-based capillary electrophoresis device that can analyze RNA, DNA, and proteins. The one or more bioanalyzers may include an Agilent 2100 Bioanalyzer.
[0134] Some embodiments may include one or more processors. The one or more processors may analyze, collect, store, sort, combine, evaluate, or otherwise process one or more data and / or results of one or more assays, one or more data and / or results based on or derived from one or more assays, one or more outputs of one or more assays, one or more outputs based on or derived from one or more assays, one or more outputs from one or more data and / or results, one or more outputs based on or derived from one or more data and / or results, or a combination thereof. The one or more processors may transmit one or more data, results, or outputs of one or more assays, one or more data, results, or outputs based on or derived from one or more assays, one or more outputs of one or more data or results, one or more outputs based on or derived from one or more data or results, or a combination thereof. The one or more processors may receive and / or store requests from a user. The one or more processors may generate or create one or more data, results, outputs. The one or more processors may generate or create one or more biomedical reports. The one or more processors may transmit one or more biomedical reports. The one or more processors may analyze, collect, store, sort, combine, evaluate, or otherwise process information from one or more databases, one or more data or results, one or more outputs, or a combination thereof. The one or more processors may analyze, collect, store, sort, combine, evaluate, or otherwise process information from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30 or more databases. The one or more processors may send one or more requests, data, results, output and / or information to one or more users, processors, computers, computer systems, memory locations, devices, databases, or combinations thereof.The one or more processors may receive one or more requests, data, results, output and / or information from one or more users, processors, computers, computer systems, memory locations, devices, databases, or combinations thereof. The one or more processors may retrieve one or more requests, data, results, output and / or information from one or more users, processors, computers, computer systems, memory locations, devices, databases, or combinations thereof.
[0135] Some embodiments may include one or more memory locations. The one or more memory locations may store information, data, results, output, requests, or combinations thereof. The one or more memory locations may receive information, data, results, output, requests, or combinations thereof from one or more users, processors, computers, computer systems, devices, or combinations thereof.
[0136] The methods described herein can be performed with the aid of one or more computers and / or computer systems. A computer or computer system may include electronic storage locations (e.g., databases, memory) having machine-executable code for performing the methods provided herein, and one or more processors for executing the machine-executable code.
[0137] The code may be pre-compiled and configured for use with a machine having a processor adapted to execute the code, or it may be compiled during run-time. The code may be provided in a programming language that can be selected to allow the code to be executed in a pre-compiled or compiled manner.
[0138] The one or more computers and / or computer systems may analyze, collect, store, sort, combine, evaluate, or otherwise process one or more data and / or results of one or more assays, one or more data and / or results based on or derived from one or more assays, one or more outputs of one or more assays, one or more outputs based on or derived from one or more assays, one or more outputs from one or more data and / or results, one or more outputs based on or derived from one or more data and / or results, or a combination thereof. The one or more computers and / or computer systems may transmit one or more data, results, or outputs of one or more assays, one or more data, results, or outputs based on or derived from one or more assays, one or more outputs of one or more data or results, one or more outputs based on or derived from one or more data or results, or a combination thereof. The one or more computers and / or computer systems may receive and / or store requests from a user. One or more computers and / or computer systems may generate or create one or more data, results, outputs. One or more computers and / or computer systems may generate or create one or more biomedical reports. One or more computers and / or computer systems may transmit one or more biomedical reports. One or more computers and / or computer systems may analyze, collect, store, filter, combine, evaluate, or otherwise process information from one or more databases, one or more data or results, one or more outputs, or combinations thereof. One or more computers and / or computer systems may analyze, collect, store, filter, combine, evaluate, or otherwise process information from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, or more databases.One or more computers and / or computer systems may send one or more requests, data, results, output and / or information to one or more users, processors, computers, computer systems, memory locations, devices, or combinations thereof. One or more computers and / or computer systems may receive one or more requests, data, results, output and / or information from one or more users, processors, computers, computer systems, memory locations, devices, or combinations thereof. One or more computers and / or computer systems may retrieve one or more requests, data, results, output and / or information from one or more users, processors, computers, computer systems, memory locations, devices, databases, or combinations thereof.
[0139] D. Database Some embodiments may include one or more databases. Some embodiments may include at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, or more databases. The databases may include genomic, proteomic, pharmacogenomic, biomedical, and scientific databases. The databases may be publicly available databases. Alternatively, or in addition, the databases may be proprietary databases. The databases may be commercial databases. Databases include, but are not limited to, Cosmic, GnomAD, Dbsnp, Mills Indels, MendelDB, PharmGKB, Varimed, Regulome, curated BreakSeq junctions, Online Mendelian Inheritance in Man (OMIM), Human Gene Variation Database (HGMD), NCBI db SNP, NCBI RefSeq, GENCODE, GO (Gene Ontology), and Kyoto Encyclopedia of Genes and Genomes (KEGG).
[0140] Some embodiments may include analyzing one or more databases. Some embodiments may include analyzing at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30 or more databases. Analyzing the one or more databases may include one or more algorithms, computers, processors, memory locations, devices, or combinations thereof.
[0141] Some embodiments may include identifying one or more nucleic acid regions based on data and / or information from one or more databases. Some embodiments may include identifying one or more sets of nucleic acid regions based on data and / or information from one or more databases. Some embodiments may include identifying one or more nucleic acid regions and / or sets of nucleic acid regions based on data and / or information from at least two or more databases. Some embodiments may include identifying one or more nucleic acid regions and / or sets of nucleic acid regions based on data and / or information from at least three or more databases. Some embodiments may include identifying one or more nucleic acid regions and / or sets of nucleic acid regions based on data and / or information from at least about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30 or more databases.
[0142] Some embodiments may include analyzing one or more results based on data and / or information from one or more databases. Some embodiments may include analyzing one or more sets of results based on data and / or information from one or more databases. Some embodiments may include analyzing one or more aggregated results based on data and / or information from one or more databases. Some embodiments may include analyzing one or more results, sets of results, and / or aggregated results based on data and / or information from at least about two or more databases. Some embodiments may include analyzing one or more results, sets of results, and / or aggregated results based on data and / or information from at least about three or more databases. Some embodiments may include analyzing one or more results, sets of results, and / or aggregated results based on data and / or information from at least about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30 or more databases.
[0143] Some embodiments may include comparing one or more results based on data and / or information from one or more databases. Some embodiments may include comparing one or more sets of results based on data and / or information from one or more databases. Some embodiments may include comparing one or more aggregated results based on data and / or information from one or more databases. Some embodiments may include comparing one or more results, sets of results, and / or aggregated results based on at least about two or more database data and / or information. Some embodiments may include comparing one or more results, sets of results, and / or aggregated results based on data and / or information from at least about three or more database data. Some embodiments may include comparing one or more results, sets of results, and / or aggregated results based on data and / or information from at least about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30 or more databases.
[0144] Some embodiments may include data and / or information from one or more databases, biomedical databases, genomic databases, biomedical reports, disease reports, case-control analyses, and rare variant discovery analyses based on one or more assays, one or more data or results, one or more outputs based on or derived from one or more assays, one or more outputs based on or derived from one or more data or results, or combinations thereof.
[0145] E. Analysis Some embodiments may include one or more data, one or more data sets, one or more summed data, one or more summed data sets, one or more results, one or more sets of results, one or more summed results, or a combination thereof. The data and / or results may be based on one or more assays, one or more databases, or a combination thereof. Some embodiments may include analysis of one or more data, one or more data sets, one or more summed data, one or more summed data sets, one or more results, one or more sets of results, one or more summed results, or a combination thereof. Some embodiments may include processing of one or more data, one or more data sets, one or more summed data, one or more summed data sets, one or more results, one or more sets of results, one or more summed results, or a combination thereof.
[0146] Some embodiments may include at least one analysis and at least one processing of one or more data, one or more data sets, one or more summed data, one or more summed data sets, one or more results, one or more sets of results, one or more summed results, or combinations thereof.Some embodiments may include one or more analyses and one or more processing of one or more data, one or more data sets, one or more summed data, one or more summed data sets, one or more results, one or more sets of results, one or more summed results, or combinations thereof. Some embodiments may include at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 or more different analyses of one or more data, one or more datasets, one or more summed data, one or more summed datasets, one or more results, one or more sets of results, one or more summed results, or combinations thereof. Some embodiments may include at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 or more different processes of one or more data, one or more data sets, one or more summed data, one or more summed data sets, one or more results, one or more sets of results, one or more summed results, or combinations thereof. The one or more analyses and / or one or more processes may occur simultaneously, over time, or combinations thereof.
[0147] The one or more analyses and / or one or more treatments may occur at more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 time points. The time points may occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 hours or more. The time points may occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 days or more. The time points may occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 weeks or more. The time points may occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 months or more. The time points may occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 years or more.
[0148] Some embodiments may include one or more data. The one or more data may include one or more raw data based on or derived from one or more assays. The one or more data may include one or more raw data based on or derived from one or more databases. The one or more data may include at least partially analyzed data based on or derived from the one or more raw data. The one or more data may include at least partially processed data based on or derived from the one or more raw data. The one or more data may include fully analyzed data based on or derived from the one or more raw data. The one or more data may include fully processed data based on or derived from the one or more raw data. The data may include sequencing read data or expression data. The data may include biomedical, scientific, pharmacological, and / or genetic information.
[0149] Some embodiments may include one or more summed data. The one or more summed data may include two or more data. The one or more summed data may include two or more data sets. The one or more summed data may include one or more raw data based on or derived from one or more assays. The one or more summed data may include one or more raw data based on or derived from one or more databases. The one or more summed data may include at least partially analyzed data based on or derived from one or more raw data. The one or more summed data may include at least partially processed data based on or derived from one or more raw data. The one or more summed data may include fully analyzed data based on or derived from one or more raw data. The one or more summed data may include fully processed data based on or derived from one or more raw data. The one or more summed data may include sequencing read data or expression data. The one or more aggregated data may include biomedical, chemical, pharmacological, and / or genetic information.
[0150] Some embodiments may include one or more datasets. The one or more datasets may include one or more data. The one or more datasets may include one or more summed data. The one or more datasets may include one or more raw data based on or derived from one or more assays. The one or more datasets may include one or more raw data based on or derived from one or more databases. The one or more datasets may include at least partially analyzed data based on or derived from one or more raw data. The one or more datasets may include at least partially processed data based on or derived from one or more raw data. The one or more datasets may include fully analyzed data based on or derived from one or more raw data. The one or more datasets may include fully processed data based on or derived from one or more raw data. The datasets may include sequencing read data or expression data. The datasets may include biomedical, scientific, pharmacological, and / or genetic information.
[0151] Some embodiments may include one or more summed datasets. The one or more summed datasets may include two or more data. The one or more summed datasets may include two or more summed data. The one or more summed datasets may include two or more datasets. The one or more summed datasets may include one or more raw data based on or derived from one or more assays. The one or more summed datasets may include one or more raw data based on or derived from one or more databases. The one or more summed datasets may include at least partially analyzed data based on or derived from one or more raw data. The one or more summed datasets may include at least partially processed data based on or derived from one or more raw data. The one or more summed datasets may include fully analyzed data based on or derived from one or more raw data. The one or more summed datasets may include fully processed data based on or derived from one or more raw data. Some embodiments may further include further processing and / or analysis of the combined dataset. One or more combined datasets may include sequencing read data or expression data. One or more combined datasets may include biomedical, chemical, pharmacological, and / or genetic information.
[0152] Some embodiments may include one or more results. The one or more results may include one or more data, datasets, summed data, and / or summed datasets. The one or more results may be based on or derived from one or more data, datasets, summed data, and / or summed datasets. The one or more results may be generated from one or more assays. The one or more results may be based on or derived from one or more assays. The one or more results may be based on or derived from one or more databases. The one or more results may include at least partially analyzed results based on or derived from one or more data, datasets, summed data, and / or summed datasets. The one or more results may include at least partially processed results based on or derived from one or more data, datasets, summed data, and / or summed datasets. The one or more results may include fully analyzed results based on or derived from one or more data, datasets, summed data, and / or summed datasets. The one or more results may include one or more data, datasets, aggregated data, and / or fully processed results based on or derived from the aggregated dataset. The results may include sequencing read data or expression data. The results may include biomedical, scientific, pharmacological, and / or genetic information.
[0153] Some embodiments may include one or more sets of results. The one or more sets of results may include one or more data, data sets, summed data, and / or summed datasets. The one or more sets of results may be based on or derived from one or more data, data sets, summed data, and / or summed datasets. The one or more sets of results may be generated from one or more assays. The one or more sets of results may be based on or derived from one or more assays. The one or more sets of results may be based on or derived from one or more databases. The one or more sets of results may include at least partially analyzed sets of results based on or derived from one or more data, data sets, summed data, and / or summed datasets. The one or more sets of results may include at least partially processed sets of results based on or derived from one or more data, data sets, summed data, and / or summed datasets. The one or more sets of results may include fully analyzed set results based on or derived from one or more data, data sets, summed data, and / or summed datasets. The one or more sets of results may include fully processed set results based on or derived from one or more data, data sets, summed data, and / or summed datasets. The set results may include sequencing read data or expression data. The set results may include biomedical, scientific, pharmacological, and / or genetic information.
[0154] Some embodiments may include one or more summed results. The summed results may include one or more results, sets of results, and / or summed sets of results. The summed results may be based on or derived from one or more results, sets of results, and / or summed sets of results. The one or more summed results may include one or more data, data sets, summed data, and / or summed datasets. The one or more summed results may be based on or derived from one or more data, data sets, summed data, and / or summed datasets. The one or more summed results may be generated from one or more assays. The one or more summed results may be based on or derived from one or more assays. The one or more summed results may be based on or derived from one or more databases. The one or more summed results may include at least partially analyzed summed results based on or derived from one or more data, data sets, summed data, and / or summed datasets. The one or more summed results may include at least partially processed summed results based on or derived from one or more data, datasets, summed data, and / or summed datasets. The one or more summed results may include fully analyzed summed results based on or derived from one or more data, datasets, summed data, and / or summed datasets. The one or more summed results may include fully processed summed results based on or derived from one or more data, datasets, summed data, and / or summed datasets. The summed results may include sequencing read data or expression data. The summed results may include biomedical, scientific, pharmacological, and / or genetic information.
[0155] Some embodiments may include a summed set result. The summed set result may include one or more results, set results, and / or summed set results. The summed set result may be based on or derived from one or more results, set results, and / or summed set results. The one or more summed set results may include one or more data, datasets, summed data, and / or summed datasets. The one or more summed set results may be based on or derived from one or more data, datasets, summed data, and / or summed datasets. The one or more summed set results may be generated from one or more assays. The one or more summed set results may be based on or derived from one or more assays. The one or more summed set results may be based on or derived from one or more databases. The one or more summed set results may include at least partially analyzed summed set results based on or derived from one or more data, datasets, summed data, and / or summed datasets. The one or more summed set results may include at least partially processed summed set results based on or derived from one or more data, datasets, summed data, and / or summed datasets. The one or more summed set results may include fully analyzed summed set results based on or derived from one or more data, datasets, summed data, and / or summed datasets. The one or more summed set results may include fully processed summed set results based on or derived from one or more data, datasets, summed data, and / or summed datasets. The summed set results may include sequencing read data or expression data. The summed set results may include biomedical, scientific, pharmacological, and / or genetic information.
[0156] Some embodiments may include one or more outputs, set outputs, summed outputs, and / or summed set outputs. The methods, libraries, kits, and systems herein may include generating one or more outputs, set outputs, summed outputs, and / or summed set outputs. A set output may include one or more outputs, one or more summed outputs, or a combination thereof. A summed output may include one or more outputs, one or more set outputs, one or more summed set outputs, and / or a combination thereof. A summed set output may include one or more outputs, one or more set outputs, one or more summed outputs, or a combination thereof. One or more outputs, set outputs, summed outputs, and / or summed set outputs may be based on or derived from one or more data, one or more datasets, one or more summed data, one or more summed datasets, one or more results, one or more set results, one or more summed results, or a combination thereof. The one or more outputs, set outputs, summed outputs, and / or summed set outputs may be based on or derived from one or more databases. The one or more outputs, set outputs, summed outputs, and / or summed set outputs may include one or more biomedical reports, biomedical outputs, rare variant outputs, pharmacogenetic outputs, population study outputs, case-control outputs, biomedical databases, genomic databases, disease databases, net content.
[0157] Some embodiments may include one or more biomedical outputs, one or more sets of biomedical outputs, one or more summed biomedical outputs, one or more summed sets of biomedical outputs. The methods, libraries, kits, and systems herein may include producing one or more biomedical outputs, one or more sets of biomedical outputs, one or more summed biomedical outputs, one or more summed sets of biomedical outputs. The set of biomedical outputs may include one or more biomedical outputs, one or more summed biomedical outputs, or a combination thereof. The summed biomedical outputs may include one or more biomedical outputs, one or more sets of biomedical outputs, one or more summed sets of biomedical outputs, or a combination thereof. The summed set of biomedical outputs may include one or more biomedical outputs, one or more sets of biomedical outputs, one or more summed biomedical outputs, or a combination thereof. The one or more biomedical outputs, the one or more sets of biomedical outputs, the one or more summed biomedical outputs, the one or more summed sets of biomedical outputs may be based on or derived from one or more data, one or more datasets, one or more summed data, one or more summed datasets, one or more results, one or more sets of results, one or more summed results, one or more outputs, one or more sets of outputs, one or more summed outputs, one or more sets of summed outputs, or combinations thereof. The one or more biomedical outputs may include biomedical information of the subject. The biomedical output of the subject may predict, diagnose, and / or prognosticate one or more biomedical features. The one or more biomedical features may include symptoms of a disease or condition, genetic risk of a disease or condition, reproductive risk, genetic risk to a fetus, risk of adverse drug reactions, efficacy of drug therapy, prediction of optimal drug dosage, transplantation tolerance, or combinations thereof.
[0158] Some embodiments may include one or more biomedical reports. The methods, libraries, kits, and systems herein may include generating one or more biomedical reports. The one or more biomedical reports may be based on or derived from one or more data, one or more datasets, one or more summed data, one or more summed datasets, one or more results, one or more sets of results, one or more summed results, one or more outputs, one or more sets of outputs, one or more summed outputs, one or more sets of summed outputs, one or more biomedical outputs, one or more sets of biomedical outputs, summed biomedical outputs, one or more sets of biomedical outputs, or combinations thereof. The biomedical report may predict, diagnose, and / or prognose one or more biomedical characteristics. The one or more biomedical characteristics may include symptoms of a disease or condition, genetic risk for a disease or condition, reproductive risk, genetic risk to a fetus, risk of an adverse drug reaction, efficacy of drug therapy, prediction of optimal drug dosage, transplantation tolerance, or a combination thereof.
[0159] Some embodiments may also include transmitting one or more data, information, results, outputs, reports, or combinations thereof. For example, transmitting data / information based on or derived from one or more assays to another device and / or instrument. In another example, transmitting data, results, outputs, biomedical outputs, biomedical reports, or combinations thereof to another device and / or instrument. Information obtained from an algorithm may also be transmitted to another device and / or instrument. Information based on analysis of one or more databases may be transmitted to another device and / or instrument. The transmitting of data / information may include transmitting data / information from a first source to a second source. The first and second sources may be in the same approximate location (e.g., in the same room, building, block, campus, etc.). Alternatively, the first and second sources may be in multiple locations (e.g., in multiple cities, states, countries, continents, etc.). The data, results, outputs, biomedical outputs, biomedical reports may be transmitted to a patient and / or a healthcare provider.
[0160] The sending may be based on an analysis of one or more data, results, information, databases, outputs, reports, or combinations thereof. For example, the sending of a second report is based on an analysis of a first report. Alternatively, the sending of a report is based on an analysis of one or more data or results. The sending may be based on receiving one or more requests. For example, the sending of a report may be based on receiving a request from a user (e.g., a patient, a healthcare provider, an individual, etc.).
[0161] The transmission of data / information may include digital transmission or analog transmission. Digital transmission may include the physical transmission of data (digital bit stream) over point-to-point or point-to-multipoint communication lines. Examples of such lines are copper wire, optical fiber, wireless communication lines, and recording media. Data may be written as electromagnetic signals, such as electrical voltages, radio waves, microwaves, or infrared signals.
[0162] Analog transmission may include the transmission of a continuously variable analog signal. A message may be represented by either a series of pulses according to a line code (baseband transmission) or by a finite set of continuously variable waveforms (bandwidth transmission) using digital modulation methods. Bandwidth modulation and corresponding demodulation (also known as detection) may be performed by modem equipment. According to the most common definition of a digital signal, both baseband signals and bandwidth signals representing bit streams are considered digital transmissions, while another definition considers only baseband signals to be digital and bandwidth transmission of digital data to be a form of digital-to-analog conversion.
[0163] Some embodiments may include one or more sample identifiers. Sample identifiers may include labels, bar codes, and other indicia that can be associated with one or more samples and / or subsets of nucleic acid molecules. Some embodiments may include one or more processors, one or more memory locations, one or more computers, one or more monitors, one or more computer software, one or more algorithms for associating data, results for samples, outputs, biomedical outputs, and / or biomedical reports.
[0164] Some embodiments may include a processor that correlates expression levels of one or more nucleic acid molecules with a prognosis of disease outcome. Some embodiments may include one or more of a variety of correlation techniques, including look-up tables, algorithms, multivariate models, and linear or non-linear combinations of expression models or algorithms. The expression levels may be converted into one or more likelihood scores to indicate the likelihood that the patient providing the sample will exhibit a particular disease outcome. The method and / or algorithm may be provided in a machine-readable format, and may further optionally prescribe a treatment modality for a patient or patient classification.
[0165] In some cases, the methods and systems described herein are used to generate output including detection and / or quantification of genomic DNA regions, such as regions containing DNA polymorphisms (e.g., germline variants or somatic variants). In some cases, the detection of one or more genomic regions is based on one or more algorithms, depending on the source of data input or databases described elsewhere herein. Each of the one or more algorithms can be used to receive, combine, and generate data including detection of genomic regions (i.e., polymorphisms). In some embodiments, the methods and systems can include detection of genomic regions based on one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more algorithms. The algorithms can be machine learning algorithms, computer-implemented algorithms, machine-implemented algorithms, automated algorithms, and the like.
[0166] The data generated for each nucleic acid sample can be analyzed using feature selection techniques, including filter techniques that assess the relevance of features by testing intrinsic properties of the data, wrapper methods that embed model hypotheses within the feature subset search, and embedding techniques where the search for an optimal set of features is built into an algorithm or model.
[0167] In some cases, the detection of one or more genomic regions is based on one or more statistical models. Statistical models or filtering techniques useful in the methods of the present invention include: (1) parametric methods such as two-sample t-test, ANOVA analysis, Bayesian framework, and use of gamma distribution model; (2) Wilcoxon rank sum test, between-class and within-class sum of squares test, rank product method, random permutation method, or TNoM, which includes setting a threshold point for the fold change of expression difference between two data sets, and then finding the threshold point for each gene that minimizes the number of misclassifications; and (3) multivariate methods such as bivariate methods, correlation-based feature selection (CFS), minimum redundancy maximum relevance (MRMR), Markov blanket filter, Markov model, hidden Markov model (HMM), and uncorrelated shrunken centroid method. In some cases, the hidden Markov model (HMM) is given an internal state, and the internal state is set according to the total copy number of chromosomes in the first or second nucleic acid sample. In one example, for chromosome diploids, the internal states of the HMM can be homozygous deletion (locally zero copies), heterozygous deletion (locally one copy), normal (locally two copies), duplication (two or more copies), and reference gap (present as a state to distinguish gaps from homozygous deletions). In another example, for haploid chromosomes (e.g., X or Y in males), the internal states of the HMM can be homozygous deletion (locally zero copies), normal (locally two copies), duplication (two or more copies), and reference gap (present as a state to distinguish gaps from homozygous deletions). For example, for haploid chromosomes, there may be no heterozygous deletion state available. In another example, for trisomy and / or tetrasomy, further intermediate HMM states may have additional intermediate states, which can account for various CNV possibilities. In another embodiment, a hidden Markov model is used to filter the output by testing the measured insertion size of the reads close to the limit points of the detected features.
[0168] Other models or algorithms useful in the methods of the present invention include sequential search methods, genetic algorithms, distribution estimation algorithms, random forest algorithms, support vector machine weight vector algorithms, logistic regression weight algorithms, and the like. Bioinformatics. 2007 Oct 1; 23(19):2507-17 outlines the advantages of the above algorithms or models used in data analysis. Exemplary algorithms include, but are not limited to, principal component analysis algorithms, partial least squares, independent component analysis algorithms, methods that directly handle large numbers of variables, such as statistical models, and methods that reduce the number of variables, such as methods based on machine learning techniques. Statistical models include penalized logistic regression, predictive analysis of microarrays (PAM), methods based on shrunken centroids, support vector machine analysis, and regularized linear discriminant analysis. Machine learning includes fully connected neural networks, convolutional neural networks, 1D convolutional neural networks, 2D convolutional neural networks, gradient boosting decision trees (e.g., XGBoost framework, LightGBM framework), bagging methods, boosting methods, random forest algorithms, and combinations thereof. Cancer Inform. 2008; 6: 77-97 reviews the above techniques used to analyze data. In some embodiments, the trained machine learning model includes gradient boosting decision trees (e.g., LightGBM framework). In some embodiments, the trained machine learning model includes a convolutional neural network (e.g., a 1D convolutional neural network or a 2D convolutional neural network). In some embodiments, the trained machine learning model includes a fully connected neural network.
[0169] Machine learning can include deep learning, which can be used to capture the internal structure of increasingly large, high-dimensional data sets (e.g., nucleic acid sequencing data). Deep models can enable the discovery of high-level features, improving performance over traditional models, increasing interpretability, and providing further understanding of the structure of biological data.
[0170] The trained machine learning model can include a fully connected neural network. The fully connected neural network can include a series of fully connected layers. Each output dimension can depend on each input dimension. The fully connected neural network can be a feedforward network.
[0171] The trained machine learning model may include a convolutional neural network. A convolutional neural network may rely on local connections and shared weights between units to obtain translation invariant descriptors, followed by feature pooling (subsampling). A basic convolutional neural network architecture may include one convolutional and pooling layer, optionally with a fully connected layer for supervised prediction. In practice, a convolutional neural network may consist of multiple (e.g., more than 10) convolutional and pooling layers to successfully model the input space. In some cases, a convolutional neural network requires a large data set to train successfully. In some cases, a convolutional neural network may use fewer parameters than a fully connected neural network by computing convolutions in small regions of the input space and sharing parameters between regions. The convolutional neural network may be a one-dimensional (1D) convolutional neural network. The convolutional neural network may be a two-dimensional (2D) convolutional neural network. In some embodiments, the convolutional neural network includes three or more dimensions.
[0172] The trained machine learning model can include gradient boosting decision trees. Gradient boosting is a machine learning technique that can be used for regression and classification problems to create weak predictive models, such as ensemble-type predictive models, such as decision trees. Gradient boosting decision trees can include, for example, the XGBoost framework or the LightGBM framework.
[0173] Machine learning models can include hyperparameters. Hyperparameters are configurations external to the model and their values cannot be estimated from data. Hyperparameters can be tuned, for example, tuned for a known predictive model problem. In some cases, hyperparameters are used in the process to help estimate model parameters. In some cases, hyperparameters can be specified by a physician. In some cases, hyperparameters can be set using heuristic techniques.
[0174] In some embodiments, the HMM-based detection algorithm can "segment" large or very large CNVs. In some cases, there may be small detection gaps along the length of the true CNV due to variations in coverage signal. In one example, a 1 megabase pair (Mbp) deletion may be detected as a few separate nominal detections with a small gap between them. To mitigate this, a merge operation can be used that identifies a pair of adjacent detections separated by a gap smaller than either of the two flanking detections. The merge operation then measures the median coverage level of the gap. If the median coverage exceeds a predefined threshold, the two detections are merged into a single large detection that spans the two original detections (including the enclosed detection gap). In one example, the true feature spans both detections, and the gap is a statistical artifact. Using real sequencing data of a sample known to have a large CNV, this merge operation may allow for a fairly good fidelity in terms of the true features of the CNV.
[0175] The methods and systems provided herein may further include the use of feature selection algorithms provided herein. In some embodiments of the invention, feature selection is effected by use of the LIMMA software package (Smyth, GK (2005). Limma: Linear Models for Microarray Data. In: Bioinformatics and Computational Biology Solutions using R and Bioconductor, R. Gentleman, V. Carey, S. Dudoit, R. Irizarry, W. Huber (eds.), Springer, New York, pages 397-420).
[0176] In some embodiments of the present invention, diagonal linear discriminant analysis, k-nearest neighbor algorithm, support vector machine (SVM) algorithm, linear support vector machine, random forest algorithm, or probability modeling method or combinations thereof are provided for detecting one or more genomic regions. In some embodiments, the identified markers that distinguish samples (e.g., diseased vs. normal) or genomic regions (e.g., copy number variation vs. normal) are selected based on the statistical significance of the difference in expression levels between the classes of interest. In some cases, the statistical significance is adjusted by providing Benjamini Hochberg or another correction for false discovery rate (FDR).
[0177] In some cases, the algorithms may be supplemented with meta-analysis techniques such as those described by Fishel and Kaufman et al. 2007 Bioinformatics 23(13): 1599-606. In some cases, the algorithms may be supplemented with meta-analysis techniques such as reproducibility analysis. In some cases, the reproducibility analysis selects markers that appear in at least one predictive expression product marker set.
[0178] Statistical evaluation of the detection of genomic regions may provide one or more quantitative values that are indicative of one or more of the following: the likelihood that the diagnosis is accurate; the likelihood of a disorder, disease, condition, and the like; the likelihood of a particular disorder, disease, condition; and the likelihood that a particular therapeutic intervention will be successful. Thus, the raw data does not need to be understood by a physician who is likely not trained in genetics or molecular biology. Rather, the data is presented directly to the physician in the form of a quantitative value that guides patient care. The results can be statistically evaluated using many methods known in the art, including, but not limited to: Student's t-test, two-tailed t-test, Pearson's rank sum analysis, hidden Markov model analysis, qq plot analysis, principal component analysis, one-way ANOVA, two-way ANOVA, LIMMA, and the like.
[0179] F. Diseases or Conditions Some embodiments include predicting, diagnosing, and / or prognosing a disease or condition state or outcome in a subject based on one or more biomedical outputs. Predicting, diagnosing, and / or prognosing a disease state or outcome in a subject may include diagnosing a disease or condition, identifying a disease or condition, determining a disease or condition status, assessing risk of a disease or condition, assessing risk of disease recurrence, assessing efficacy of a drug, assessing risk of adverse drug reactions, predicting optimal drug dosage, predicting drug resistance, or a combination thereof.
[0180] The samples disclosed herein may be from subjects suffering from cancer. The samples may include malignant tissue, benign tissue, or a combination thereof. The cancer may be recurrent and / or refractory cancer. Examples of cancer include, but are not limited to, sarcoma, carcinoma, lymphoma, or leukemia. In some cases, a sample containing cancer tissue is obtained, but a matched normal sample is not obtained. In some cases, a matched normal sample is not available. In some cases, a matched normal sample is obtained (e.g., for training and testing the models disclosed herein).
[0181] Sarcomas are cancers of bone, cartilage, fat, muscle, blood vessels, or other connective or supporting tissues. Sarcomas include, but are not limited to, bone cancer, fibrosarcoma, chondrosarcoma, Ewing's sarcoma, malignant hemangioendothelioma, malignant schwannoma, bilateral vestibular schwannoma, osteosarcoma, soft tissue sarcomas (e.g., alveolar soft part sarcoma, angiosarcoma, cystosarcoma phylloides, dermatofibrosarcoma, desmoid tumor, epithelioid sarcoma, extraskeletal osteosarcoma, fibrosarcoma, hemangiopericytoma, angiosarcoma, Kaposi's sarcoma, leiomyosarcoma, liposarcoma, lymphangiosarcoma, lymphosarcoma, malignant fibrous histiocytoma, neurofibrosarcoma, rhabdomyosarcoma, and synovial sarcoma.
[0182] A carcinoma is a cancer that begins in epithelial cells, which are cells that cover the surface of the body, produce hormones, and make up glands. Non-limiting examples of carcinoma include breast cancer, pancreatic cancer, lung cancer, colon cancer, colorectal cancer, rectal cancer, kidney cancer, bladder cancer, stomach cancer, prostate cancer, liver cancer, ovarian cancer, brain cancer, vaginal cancer, vulvar cancer, uterine cancer, oral cancer, penile cancer, testicular cancer, esophageal cancer, skin cancer, fallopian tube cancer, head and neck cancer, gastrointestinal stromal cancer, adenocarcinoma, cutaneous or intraocular melanoma, anal region cancer, small intestine cancer, endocrine system cancer, thyroid cancer, parathyroid cancer, adrenal gland cancer, urethra cancer, renal pelvis cancer, ureter cancer, endometrial cancer, cervical cancer, pituitary gland cancer, central nervous system (CNS) tumor, primary CNS lymphoma, brain stem glioma, and spinal axis tumor. The cancer may be basal cell carcinoma, squamous cell carcinoma, melanoma, non-melanoma cancer, or actinic (solar) keratosis.
[0183] The cancer may be lung cancer. Lung cancer may begin in the airways that branch off from the trachea and supply the small air sacs of the lungs (bronchi) or lungs (alveoli). Lung cancer includes non-small cell lung cancer (NSCLC), small cell lung cancer, and mesothelioma. Examples of NSCLC include squamous cell carcinoma, adenocarcinoma, and large cell carcinoma. Mesothelioma may be a cancerous tumor of the lining of the lung and chest cavity (pleura) or the lining of the abdomen (peritoneum). Mesothelioma may be due to asbestos exposure. The cancer may be brain cancer, such as glioblastoma.
[0184] The cancer may be a central nervous system (CNS) tumor. CNS tumors may be classified as gliomas or nongliomas. Gliomas may be malignant gliomas, high-grade gliomas, or diffuse intrinsic pontine gliomas. Examples of gliomas include astrocytomas, oligodendrogliomas (or mixed oligodendroglioma and astrocytoma elements), and ependymomas. Astrocytomas include, but are not limited to, low-grade astrocytomas, anaplastic astrocytomas, glioblastoma multiforme, pilocytic astrocytomas, pleomorphic xanthoastrocytomas, and subependymal giant cell astrocytomas. Oligodendrogliomas include low-grade oligodendrogliomas (or oligoastrocytomas) and anaplastic oligodendriogliomas. Nongliomas include meningiomas, pituitary adenomas, primary CNS lymphomas, and medulloblastomas. The cancer may be a meningioma.
[0185] The leukemia may be acute lymphocytic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, or chronic myelogenous leukemia. Further types of leukemia include hairy cell leukemia, chronic myelomonocytic leukemia, and juvenile myelomonocytic leukemia.
[0186] Lymphoma is a cancer of lymphocytes and may arise from either B or T lymphocytes. The two major types of lymphoma are Hodgkin's lymphoma, formerly known as Hodgkin's disease, and non-Hodgkin's lymphoma. Hodgkin's lymphoma is characterized by the presence of Reed-Sternberg cells. Non-Hodgkin's lymphoma is any lymphoma that is not Hodgkin's lymphoma. Non-Hodgkin's lymphoma can be low-grade lymphoma and intermediate-grade lymphoma. Non-Hodgkin's lymphomas include, but are not limited to, diffuse large B-cell lymphoma, follicular lymphoma, mucosa-associated lymphatic tissue lymphoma (MALT), small cell lymphocytic lymphoma, mantle cell lymphoma, Burkitt lymphoma, mediastinal large B-cell lymphoma, Waldenstrom's hypergammaglobulinemia, nodal marginal zone B-cell lymphoma (NMZL), splenic marginal zone lymphoma (SMZL), extranodal marginal zone B-cell lymphoma, intravascular large B-cell lymphoma, primary effusion lymphoma, and lymphomatoid granulomatosis.
[0187] Some embodiments may include treating and / or preventing a disease or condition in a subject based on the one or more biomedical outputs. The one or more biomedical outputs may recommend one or more treatments. The one or more biomedical outputs may suggest, select, prescribe, recommend or otherwise determine a course of treatment and / or prevention of a disease or condition. The one or more biomedical outputs may recommend modifying or continuing one or more treatments. Modifying one or more treatments may include administering, initiating, decreasing, increasing, and / or terminating one or more treatments. The one or more treatments may include an anti-cancer therapy, an anti-viral therapy, an anti-bacterial therapy, an anti-fungal therapy, an immunosuppressive therapy, or a combination thereof. The one or more treatments may treat, alleviate or prevent one or more diseases or symptoms.
[0188] Examples of anti-cancer therapies include, but are not limited to, surgery, chemotherapy, radiation therapy, immunotherapy / biotherapy, photodynamic therapy. Anti-cancer therapies may include chemotherapy, monoclonal antibodies (e.g., rituximab, trastuzumab), cancer vaccines (e.g., therapeutic vaccines, prophylactic vaccines), gene therapy, or a combination thereof.
[0189] G. Systems, Kits, and Libraries Certain embodiments may be implemented by means of a system, a kit, a library, or a combination thereof. The methods of the present invention may include one or more systems. The system may be implemented by means of a kit, a library, or both. The system may include one or more components for performing any of the methods or any of the steps of some embodiments. For example, the system may include one or more kits, devices, libraries, or a combination thereof. The system may include one or more sequencers, processors, memory locations, computers, computer systems, or a combination thereof. The system may include a transmission device.
[0190] The kits may include various reagents for carrying out various operations disclosed herein, including sample processing and / or analytical operations. The kits may include instructions for carrying out at least some of the operations disclosed herein. The kits may include one or more capture probes, one or more beads, one or more labels, one or more linkers, one or more devices, one or more reagents, one or more buffers, one or more samples, one or more databases, or combinations thereof.
[0191] The library may include one or more capture probes. The library may include one or more subsets of nucleic acid molecules. The library may include one or more databases. The library may be created or generated from any of the methods, kits, or systems disclosed herein. A database library may be created from one or more databases. A method of creating one or more libraries may include (a) integrating information from one or more databases to create an integrated dataset, (b) analyzing the integrated dataset, and (c) creating one or more database libraries from the integrated dataset.
[0192] From the foregoing, it should be understood that while specific implementations have been illustrated and described, various modifications may be made thereto and are contemplated herein. An embodiment of one aspect may be combined with or modified by another aspect. The present invention is not intended to be limited by the specific examples provided herein. The present invention has been described with reference to the foregoing specification, but the description and drawings of the embodiments of the present invention herein are not intended to be construed in a limiting sense. Furthermore, it should be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative proportions set forth herein, which depend upon a variety of conditions and variables. Various modifications in form and details of the embodiments of the present invention will be apparent to those skilled in the art. Thus, the present invention is also intended to cover any such modifications, variations, and equivalents.
[0193] VI. Computing Environment FIG. 10 illustrates an example of a computer system 1000 for implementing some of the embodiments disclosed herein. The computer system 1000 may have a distributed architecture, where some of the components (e.g., memory and processor) are part of an end-user device and some other similar components (e.g., memory and processor) are part of a computer server. The computer system 1000 includes at least a processor 1002, a memory 1004, a storage device 1006, an input / output (I / O) peripherals 1008, a communication peripherals 1010, and an interface bus 1012. The interface bus 1012 is configured to communicate, transfer, and transfer data, control, and commands between various components of the computer system 1000. The processor 1002 may include one or more processing units, such as a CPU, a GPU, a TPU, a systolic array, or a SIMD processor. The memory 1004 and storage 1006 include computer readable storage media such as RAM, ROM, Electrically Erasable Programmable Read Only Memory (EEPROM), hard drives, CD-ROM, optical storage, magnetic storage, electronic non-volatile computer storage devices such as flash memory, and other tangible recording media. Any of such computer readable storage media may be configured to store instructions or program code embodying aspects. The memory 1004 and storage 1006 also include computer readable signal media. A computer readable signal medium includes a propagated data signal having computer readable program code embodied therein. Such a propagated signal may take any of a variety of forms, including, but not limited to, an electromagnetic signal, an optical signal, or a combination thereof. A computer readable signal medium includes any computer readable medium that can communicate, propagate, or carry a program for use with the computer system 1000, other than a computer readable storage medium.
[0194] Additionally, the memory 1004 includes an operating system, programs, and applications. The processor 1002 is configured to execute stored instructions and includes, for example, logic processing units, microprocessors, digital signal processors, and other processors. The memory 1004 and / or the processor 1002 can be virtualized and provided within another computing system, such as, for example, a cloud network or a data center. The I / O peripherals 1008 include user interfaces such as keyboards, screens (e.g., touch screens), microphones, speakers, other input / output devices, and computing components such as graphics processing units, serial ports, parallel ports, universal serial buses, and other input / output peripherals. The I / O peripherals 1008 are connected to the processor 1002 by any of the ports coupled to the interface bus 1012. The communication peripherals 1010 are configured to facilitate communication between the computer system 1000 and other computing devices over a communication network and include, for example, network interface controllers, modems, wireless and wired interface cards, antennas, and other communication peripherals.
[0195] Although the present subject matter has been described in detail with respect to certain embodiments thereof, those skilled in the art, upon understanding the foregoing, can easily make alternatives, modifications, and equivalents of such embodiments. Thus, the present disclosure is presented for purposes of illustration rather than limitation, and should be understood not to exclude such improvements, modifications, and / or additional inclusions to the present subject matter, as would be readily understood by one skilled in the art. On the contrary, the methods and systems described herein may be embodied in other various forms, and furthermore, various omissions, substitutions, and changes in the forms of the methods and systems described herein may be made without departing from the spirit of the present disclosure. The appended claims and their equivalents are intended to cover such forms and modifications as fall within the scope and spirit of the present disclosure.
[0196] Unless specifically stated otherwise, discussions throughout this specification utilizing terms such as "processing," "computing," "calculating," "determining," and "identifying" or the like refer to operations or processes of a computing device, such as one or more computers or one or more similar electronic computing devices, that manipulate or transform data represented as physical, electronic, or magnetic quantities within the memories, registers, or other information storage devices, transmission devices, or display devices of the computing platform.
[0197] The system or systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device may include any suitable arrangement of components that provide a result that is determinative of one or more inputs. Suitable computing devices include computing systems that use general-purpose microprocessors that access stored software that programs or configures the computing device from a general-purpose computing device to a specialized computing device that executes one or more embodiments of the present subject matter. Any suitable programming, scripting, or other type of language or combination of languages may be used to implement the teachings contained herein in the software used to program or configure the computing device.
[0198] Some embodiments may be implemented in operation of such a computing device. The order of the blocks presented in the above examples may be varied, e.g., the blocks may be rearranged, combined, and / or decomposed into sub-blocks, etc. Certain blocks or processes may be performed in parallel.
[0199] In particular, conditional language used herein, such as "can," "could," "might," "may," "eg," and the like, unless specifically stated otherwise, is intended to be understood within the context in which it is otherwise used and generally to express that a particular example includes certain features, elements, and / or steps, but not other examples. Thus, such conditional language is generally intended to mean that features, elements, and / or steps do not require one or more examples at all, or that one or more examples do not necessarily include logic for determining whether those features, elements, and / or steps are included or performed in any particular example, with or without author input or instruction.
[0200] The terms "including," "including," "having," and the like, mean the same thing and are used in an open-ended, inclusive manner and do not exclude additional elements, features, acts, operations, etc. The term "or" is also used in an inclusive sense (and not an exclusive sense), such as when used in conjunction with a list of elements, where the term "or" means one, some, or all of the elements in the list. Use of "adapted to" or "configured to" herein means open and inclusive language that does not exclude apparatus adapted or configured to perform additional tasks or steps. Moreover, use of "based on" means open and inclusive language in that a process, step, calculation, or other act that is "based on" one or more recited conditions or values may in fact be based on additional conditions or values beyond those recited. Similarly, the use of "based at least in part on" is intended to be open and inclusive in that a process, step, calculation, or other act that is "based at least in part on" one or more recited conditions or values may in fact be based on additional conditions or values beyond those recited. The headings, tables, and numbering contained herein are for ease of explanation only and are not intended to be limiting.
[0201] The various features and processes described above may be used independently of one another or may be combined in various ways. All possible combinations and sub-combinations are intended to be within the scope of the present disclosure. In addition, certain method or process blocks may be omitted in some implementations. The methods and processes described herein are also not limited to any order, and the blocks or states associated therewith may be performed in other orders as appropriate. For example, the described blocks or states may be performed in orders other than those specifically disclosed, or multiple blocks or states may be combined into a single block or state. Example blocks or states may be performed sequentially, in parallel, or in some other manner. Blocks or states may be added or removed from the disclosed embodiments. Similarly, the example systems and components described herein may be configured differently than as described. For example, elements may be added, removed, or rearranged in the disclosed embodiments.
Claims
1. Obtaining nucleic acid sequence data for a biological sample of a subject, wherein a reference biological sample of the subject corresponding to the biological sample of the subject is not available and the reference biological sample contains only non-tumor cells; aligning the nucleic acid sequence data to a reference genome; identifying a set of candidate variants in the nucleic acid sequence data based on the aligned nucleic acid sequence data of the biological sample using a sensitive algorithm, wherein the set of candidate variants includes one or more somatic variants and one or more germline variants. generating an attribute table comprising one or more features identified for each of a set of candidate variants in the nucleic acid sequence data, wherein the attribute table comprises a first training dataset and a second training dataset of a trained machine learning model, the first training dataset comprising candidate variants identified using a variant detection algorithm in the context of a tumor only, and the second training dataset comprising the remainder of the candidate variants; processing the set of candidate variants with the trained machine learning model without nucleic acid sequence data of a reference biological sample of the subject to identify the somatic variants based on the generated attribute table, wherein the trained machine learning model includes both a filtering model for removing false positives from the first training dataset and a rescue model trained to distinguish between false negatives and true negatives from the second training dataset; and outputting a report identifying said somatic variants; A computer-implemented method comprising:
2. The method of claim 1 , wherein the biological sample is a tumor sample from the subject.
3. The method of claim 2 , wherein the subject is a human subject.
4. The method of claim 1 , wherein the trained machine learning model comprises a gradient boosting decision tree.
5. The method of claim 1 , wherein the trained machine learning models include two classification models.
6. 2. The method of claim 1, wherein the trained machine learning model is further trained using training data corresponding to a set of matched tumor-normal pairs.
7. The method of claim 1 , wherein the trained machine learning model is trained by tuning one or more hyperparameters via random search.
8. The method of claim 1 , wherein the report identifies at least one biomarker.
9. The method of claim 1 , wherein the report identifies at least one prognostic marker.
10. 10. The method of claim 1, wherein the report identifies the presence or absence of the one or more somatic variants.
11. The method of claim 1 , wherein the report identifies a treatment recommendation.
12. The method of claim 11 , wherein the treatment recommendation comprises a recommendation to administer a treatment to a human subject.
13. 2. The method of claim 1, wherein generating the attribute table comprises determining one or more features selected from the group including pile-up attributes resulting from high sensitivity algorithms, allele frequency data, base quality data, read depth data, estimates of tumor purity, predicted germline variants, predicted somatic variants, copy number alteration data, population frequency data from external databases, data from at least one of Cosmic, GnomAD, Dbsnp, and Mills Indels, data regarding the presence of candidate somatic variants in problematic regions of the genome, and data regarding the presence of candidate somatic variants in homopolymers.
Citation Information
Patent Citations
No-contrast somatic cell mutation detecting method and device
CN109903811A
Machine learning system and method for somatic mutation discovery
US20190189242A1
Methods for classifying somatic variations
WO2018064547A1