Somatic variant calling from unmatched biological sample

A machine learning model processes nucleic acid sequence data from unmatched biological samples to accurately identify somatic variants, addressing the challenge of absent normal samples by reducing false positives and enhancing diagnostic and treatment reports.

JP2025109805APending Publication Date: 2025-07-25PERSONALIS INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025078607
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-11-05
Filing Date
2025-05-09
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Conventional somatic variant calling techniques struggle to accurately identify somatic variants in biological samples when corresponding normal samples are unavailable, leading to increased false positives and reduced accuracy.

Method used

A method using a trained machine learning model, such as a gradient boosting decision tree, processes nucleic acid sequence data from unmatched biological samples, generating an attribute table with features like pileup attributes and allele frequency, to distinguish somatic variants from germline variants without relying on normal control samples.

Benefits of technology

Enhances the accuracy of somatic variant identification, reducing false positives and improving diagnostic, prognostic, and treatment recommendation reports based on sequencing data from unmatched biological samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025109805000002
    Figure 2025109805000002
  • Figure 2025109805000003
    Figure 2025109805000003
  • Figure 2025109805000004
    Figure 2025109805000004
Patent Text Reader

Abstract

To provide methods for somatic variant calling from an unmatched biological sample.SOLUTION: The method can include obtaining nucleic acid sequence data corresponding to a biological sample of a subject. The method can also include aligning the nucleic acid sequence data to a reference genome. The method can also include identifying, based on the aligned nucleic acid sequence data, a set of candidate variants in the nucleic acid sequence data. The set of candidate variants may include one or more somatic variants and one or more germline variants. The method can also include, without using a nucleic acid sequence data from a matching biological sample of the subject, processing the set of candidate variants using a trained machine-learning model to identify the somatic variants. The method can also include outputting a report that identifies the somatic variants.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to related applications This application claims priority to U.S. Provisional Patent Application No. 62 / 931,100, filed on November 5, 2019, which is hereby incorporated by reference in its entirety for all purposes.

[0002] The present disclosure generally relates to systems and methods for identifying somatic variants in biological samples. More specifically, but not by way of limitation, the present disclosure relates to identifying somatic variants in biological samples by filtering false positives from a set of candidate variants detected using a trained machine learning model.

Background Art

[0003] Background of the Invention Somatic variants of DNA sequences can exhibit one or more mutations that contribute to the development of cancer. For many analyses of tumor samples, identifying somatic variants promotes cancer diagnosis, prognosis, treatment decision - making, and improvement of treatment efficiency. To identify somatic variants in biological samples, it is possible to distinguish germline sequence variants and somatic variants. Conventional somatic variant calling techniques rely heavily on contrasting evidence about differences between tumor samples and corresponding normal samples. However, there are some cases where corresponding normal samples for analysis are not available.

[0004] Therefore, there is a need to accurately identify somatic variants in biological samples and distinguish somatic variants from germline variants without relying on normal control samples.

Summary of the Invention

[0005] In some embodiments, a method for identifying somatic variants from a biological sample is provided. The method can include obtaining nucleic acid sequence data corresponding to a biological sample of a subject. The method can also include aligning the nucleic acid sequence data to a reference genome (generated based on samples from other subjects). The method can also include identifying a set of candidate variants in the nucleic acid sequence data based on the aligned nucleic acid sequence data. In some examples, the set of candidate variants includes one or more somatic variants and one or more germline variants.

[0006] The method can also include processing the set of candidate variants using a trained machine learning model to identify the somatic variant without using nucleic sequencing data of the corresponding biological sample of the subject. The corresponding biological sample of the subject indicates the absence of a tumor. The method can also output a report identifying the somatic variant.

[0007] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of the one or more methods disclosed herein.

[0008] In some embodiments, a computer program product is provided that is tangibly embodied on a non-transitory machine-readable storage medium and includes instructions configured to cause one or more data processors to perform some or all of the one or more methods disclosed herein.

[0009] Some embodiments of the present disclosure include a system that includes one or more data processors. In some embodiments, the system, when executed by the one or more data processors, contains a non-transitory computer-readable storage medium containing instructions that cause the one or more data processors to execute some or all of one or more of the methods disclosed herein and / or some or all of one or more of the processes. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium and including instructions configured to cause one or more data processors to execute some or all of one or more of the methods disclosed herein and / or some or all of one or more of the processes.

[0010] The terms and expressions used are not limiting but are used as terms of explanation, and such terms and expressions are not intended to exclude any equivalents of the features shown and described or parts thereof, but it is understood that various improvements are possible within the scope of the claimed invention. Accordingly, the claimed invention is specifically disclosed by embodiments and optional features, but improvements and variations of the concepts disclosed herein may be made by those skilled in the art, and such improvements and variations are to be considered within the scope of the invention defined by the appended claims.

Brief Description of the Drawings

[0011] The features, embodiments, and advantages of the present disclosure will be better understood when the following detailed description of the invention is read with reference to the following drawings. The present invention or the application documents include at least one drawing executed in color. A copy of this patent or patent application publication including the color drawing will be provided by the Patent Office upon request and payment of the necessary fees.

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Best Mode for Carrying Out the Invention

[0013] Detailed Description I. Overview As described above, it becomes difficult to predict somatic variants in a biological sample when the corresponding normal sample is not available for analysis. To illustrate, FIG. 1 shows an example interface 100 configured to identify somatic variants in a pair of tumor / normal sequence data according to some embodiments. The example interface 100 can include a lower panel representing nucleic acid sequence data of a tumor sample 105 and an upper panel representing nucleic acid sequence data of a normal sample 110. The gray bars may represent partially overlapping sequence reads aligned to the reference genome. Candidate variants can be highlighted within the reads using different colors. In the upper panel of the reads, three variants present from 50% to 100% of the reads can be seen. Since these reads are from the corresponding normal sample, these variants can be identified as germline variants. In the lower panel of the reads, the same three variants can be confirmed, and additional variants (identified by the squares) are present in a subset of the reads. This variant is present in the tumor sample but not in the corresponding normal sample, and thus can be identified as a somatic variant.

[0014] As shown in FIG. 1, conventional somatic variant calling techniques rely on contrasting evidence about the differences between a subject's tumor sample and the corresponding normal sample. In the absence of the corresponding normal sample 110, the identification of somatic variants in the tumor sample 105 is hindered, and thereby the accuracy rate of conventional somatic variant calling techniques can be significantly reduced. For example, removing the corresponding normal sample 100 from the schematic example 100 may make it difficult to determine which of the candidate variants in the lower panel are germline variants and which are somatic variants. In the absence of the corresponding normal sample 110, the amount of false positives (e.g., germline variants) may increase in determining somatic variants. In some examples, the false positives caused by, for example, germline contamination in the somatic variant calling output increase significantly.

[0015] To address at least the above deficiencies of conventional systems, the techniques of the present invention can be used to identify somatic variants in an unmatched biological sample and to distinguish somatic variants from germline variants. A trained machine learning model including one or more classification models can be used to predict somatic variants based on features extracted from nucleic acid sequencing data obtained from an unmatched biological sample. In some examples, additional data sources (e.g., databases) are used to predict somatic variants. For example, a high-sensitivity algorithm can be used to identify candidate variants in nucleic acid sequencing data. An attribute table can be generated, which may include one or more features identified for each candidate variant. A trained machine learning model can be used to identify somatic variants based on the contents of the attribute table. A report identifying the somatic variants can be output. In some examples, this report includes a diagnostic report, a prognostic report, and / or a treatment recommendation.

[0016] Nucleic acid sequence data of a subject's biological sample can be obtained. In some embodiments, this sequencing data is derived from a tumor sample. Sequencing can include whole exome sequencing. In some embodiments, this sequencing can include whole genome sequencing. In some embodiments, this sequencing includes shotgun sequencing. In some embodiments, this sequencing includes sequencing a selected portion of the genome or exome.

[0017] Nucleic acid sequence data can be aligned to a reference genome. As used herein, a reference genome corresponds to a nucleic acid sequence corresponding to a representative example of a set of genes in an ideal individual organism of a species. Based on the aligned nucleic acid sequence data, a set of candidate variants in the nucleic acid sequence data can be identified. In some examples, the set of candidate variants includes one or more somatic variants and one or more germline variants. As used herein, "somatic variant" refers to a DNA change that occurs after conception and is not present in the germline. Somatic variants can occur in any of the body's cells other than germ cells (sperm and egg cells) and thus cannot be inherited. Further, "germline variant" refers to a genetic change in germ cells (sperm and egg cells) that is incorporated into the DNA of all cells in the body of the offspring. Variants (or mutations) contained within the germline can be passed from parent to offspring and are thus hereditary. In some examples, somatic variants indicate the presence or level of cancer in a subject instead of germline variants.

[0018] An attribute table (for example) can be generated, where the attribute table can include many features for each candidate variant. In some embodiments, the attribute table includes attributes derived from sequencing data corresponding to a particular candidate variant. The attribute table can include attributes derived from a file containing processed sequencing data. In some embodiments, the attribute table includes one or more attributes such as: (a) pileup attributes from a BCFtools output file; (b) allele frequency data; (c) base quality data; (d) read depth data; (e) an estimate of tumor cellularity (which may be calculated based on the B allele frequency distribution); (f) predicted germline variants; (g) predicted somatic variants; (h) copy number variation data; (i) population frequency data from one or more databases; (j) data from at least one database selected from the group consisting of Cosmic, GnomAD, Dbsnp, and Mills Indels; (k) data regarding the presence of candidate somatic variants in problematic regions of the genome, and (l) data regarding the presence of candidate somatic variants in homopolymers.

[0019] A set of candidate variants can be processed using a trained machine learning model to identify somatic variants without using nucleic acid sequencing data from corresponding normal samples of interest. In some examples, the trained machine learning model includes a gradient boosting decision tree that promotes a significant reduction in the false positive rate corresponding to somatic variant calls. Thus, the present technique can detect somatic variants from unmatched biological samples with enhanced sensitivity and specificity compared to conventional discovery-based approaches. In some embodiments, the trained machine learning model includes a two-model classification method. The machine learning model may include a filtering model that removes false positives. The machine learning model may include a remedial model that remedies false negatives. In some embodiments, the somatic variant is predicted with a precision of at least 0.5. In some embodiments, the somatic variant is predicted with a recall of at least 0.5. In some embodiments, the machine learning model includes hyperparameters that are tuned by random search. In some embodiments, the hyperparameters include a maximum depth of 5 to 100, a minimum data within a leaf of 2 to 50, and at least 2 to 2048 leaves. In some embodiments, the filtering model includes a threshold of about 0.45. In some embodiments, the remedial model includes a threshold of about 0.9995.

[0020] A report identifying the somatic variants can be output. In some embodiments, the report includes information identifying at least one diagnostic marker and at least one prognostic marker. In some embodiments, the absence of somatic variants, treatment recommendations, recommendations to treat a human subject, and / or recommendations not to treat a human subject. In some embodiments, the recommended treatment is performed on the human subject.

[0021] Accordingly, embodiments of the present disclosure provide a technical advantage over conventional systems by increasing the accuracy rate of somatic variants called from unmatched biological samples. Such techniques may improve the accuracy of diagnostic, prognostic, and / or treatment recommendation reports generated based on sequencing data from unmatched biological samples. Such techniques may also reduce the costs and resources required for identifying somatic variants in tumors.

[0022] While various embodiments of the invention of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be used in practicing any one of the inventions described herein.

[0023] II. Machine Learning Model for Somatic Variant Calling from Unmatched Biological Samples A. Training a Machine Learning Model for Identifying Somatic Variants from Unmatched Biological Samples A machine learning model for identifying somatic variants from unmatched biological samples can be trained using a training dataset that includes tumor samples and normal samples corresponding to the tumor samples. For example, the training dataset may include sequencing data obtained for 350 tumor / normal sample pairs (e.g.). DNA is extracted from the training samples, processed, and subjected to whole exome sequencing. The sequencing reads are subjected to a quality control process (e.g., by FastQC) to provide FASTQ files. The FASTQ files are aligned to a reference genome to generate BAM files. BCFtool is used to identify a set of candidate somatic variants for each training sample with high sensitivity. The set of candidate somatic variants may include false positives such as germline variants (e.g.).

[0024] For a set of candidate somatic variants, an attribute table is generated that includes multiple features (e.g., about 10 - 20 features) for each candidate variant. The attribute table includes (i) pileup attributes derived from initial BCFtools output such as allele frequency (e.g., B allele frequency), base quality, read depth, (ii) an estimate of tumor purity determined using a deep learning neural network based on the B allele frequency distribution of all exomes in the sample, (iii) whether the variant is identified as a germline variant using GATK HaplotypeCaller, (iv) the somatic copy number alteration (CNA) state for each variant site, (v) the frequency of the variant in the population (e.g., in a healthy human population and / or in cancer exomes from databases such as Cosmic, GnomAD, Dbsnp, Mills Indels), (vi) the presence of the variant in problematic regions such as in homopolymers, and (vii) whether the variant is identified by standard somatic callers such as MuTect and MuTect2 (running in the context of a single tumor).

[0025] Classification labels are created based on the presence of candidate variants in the VCF file generated by MuTect or MuTect2 using default parameters having applied in - house reporting criteria. The corresponding normal samples are considered by MuTect / MuTect2 to generate these classification labels, thereby identifying "true" somatic variants, and this is used to evaluate model performance.

[0026] In some examples, this machine - learning model is trained and tested to identify somatic variants based on the content of the attribute table. The training dataset can be split into a training (90%) set and a test (10%) set. In some embodiments, the trained machine - learning model is trained by the training dataset to achieve one or more predetermined performance levels for estimating tumor purity. The one or more predetermined performance levels include the following. ● A concordance rate of at least about 0.2, 0.25, 0.3, 0.35, 0.4, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, 0.95, or more. In some examples, the trained machine learning model predicts somatic variants with a concordance rate of about 0.2 to 1.0, 0.2 to 0.9, 0.2 to 0.8, 0.2 to 0.7, 0.2 to 0.6, 0.2 to 0.5, 0.2 to 0.4, 0.2 to 0.3, 0.3 to 1.0, 0.3 to 0.9, 0.3 to 0.8, 0.3 to 0.7, 0.3 to 0.6, 0.3 to 0.5, 0.3 to 0.4, 0.4 to 1.0, 0.4 to 0.9, 0.4 to 0.8, 0.4 to 0.7, 0.4 to 0.6, 0.4 to 0.5, 0.5 to 1.0, 0.5 to 0.9, 0.5 to 0.8, 0.5 to 0.7, 0.5 to 0.6, 0.6 to 1.0, 0.6 to 0.9, 0.6 to 0.8, 0.6 to 0.7, 0.7 to 1.0, 0.7 to 0.9, 0.7 to 0.8, 0.8 to 1.0, 0.8 to 0.9, or 0.9 to 1.0. ● A reproducibility rate of at least about 0.2, 0.25, 0.3, 0.35, 0.4, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.9, 0.95, or more. In some examples, the trained machine learning model predicts somatic variants with a reproducibility rate of about 0.2 to 1.0, 0.2 to 0.9, 0.2 to 0.8, 0.2 to 0.7, 0.2 to 0.6, 0.2 to 0.5, 0.2 to 0.4, 0.2 to 0.3, 0.3 to 1.0, 0.3 to 0.9, 0.3 to 0.8, 0.3 to 0.7, 0.3 to 0.6, 0.3 to 0.5, 0.3 to 0.4, 0.4 to 1.0, 0.4 to 0.9, 0.4 to 0.8, 0.4 to 0.7, 0.4 to 0.6, 0.4 to 0.5, 0.5 to 1.0, 0.5 to 0.9, 0.5 to 0.8, 0.5 to 0.7, 0.5 to 0.6, 0.6 to 1.0, 0.6 to 0.9, 0.6 to 0.8, 0.6 to 0.7, 0.7 to 1.0, 0.7 to 0.9, 0.7 to 0.8, 0.8 to 1.0, 0.8 to 0.9, or about 0.9 to 1.0. ● An Fl score (macro-average Fl classification score) of at least about 0.2, 0.25, 0.3, 0.35, 0.4, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7, 0.75, 0.8, 0.85, 0.86, 0.87, 0.88, 0.89, 0.9, 0.91, 0.92, 0.93, 0.94, 0.95, 0.96, 0.97, 0.98, 0.99, 0.995, or more. In some examples, the trained machine learning model predicts somatic variants with an Fl score of about 0.2 - 1.0, 0.2 - 0.99, 0.2 - 0.95, 0.2 - 0.9, 0.2 - 0.8, 0.2 - 0.7, 0.2 - 0.6, 0.2 - 0.5, 0.2 - 0.4, 0.2 - 0.3, 0.3 - 1.0, 0.3 - 0.99, 0.3 - 0.95, 0.3 - 0.9, 0.3 - 0.8, 0.3 - 0.7, 0.3 - 0.6, 0.3 - 0.5, 0.3 - 0.4, 0.4 - 1.0, 0.4 - 0.99, 0.4 - 0.95, 0.4 - 0.9, 0.4 - 0.8, 0.4 - 0.7, 0.4 - 0.6, 0.4 - 0.5, 0.5 - 1.0, 0.5 - 0.99, 0.5 - 0.95, 0.5 - 0.9, 0.5 - 0.8, 0.5 - 0.7, 0.5 - 0.6, 0.6 - 1.0, 0.6 - 0.99, 0.6 - 0.95, 0.6 - 0.9, 0.6 - 0.8, 0.6 - 0.7, 0.7 - 1.0, 0.7 - 0.99, 0.7 - 0.98, 0.7 - 0.97, 0.7 - 0.96, 0.7 - 0.95, 0.7 - 0.9, 0.7 - 0.8, 0.8 - 1.0, 0.8 - 0.99, 0.8 - 0.98, 0.8 - 0.97, 0.8 - 0.96, 0.8 - 0.95, 0.8 - 0.9, 0.9 - 1.0, 0.9 - 0.99, 0.9 - 0.98, 0.9 - 0.97, 0.9 - 0.96, or 0.9 - 0.95. ● A false positive rate of at most about 0.001%, 0.01%, 0.1%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 30%, 35%, 40%, or 50%, and ● At least about 0.1, 0.2, 0.3, 0.4, 0.45, 0.5, 0.55, 0.6, 0.7, 0.8, 0.9, 0.95, 0.96, 0.97, 0.98, 0.99, 0.995, 0.999, 0.9995, 0.9999, or more area under the curve - receiver operating characteristic (AUC-ROC). In some cases, the trained machine learning model is trained to achieve an AUC-ROC of at most about 0.8, 0.9, 0.95, 0.99, 0.995, 0.999, 0.9995, 0.9999 or less. In some cases, the trained machine learning model is about 0.5 - 1.0, 0.5 - 0.9995, 0.5 - 0.999, 0.5 - 0.99, 0.5 - 0.95, 0.5 - 0.9, 0.5 - 0.8, 0.5 - 0.7, 0.5 - 0.6, 0.6 - 1.0, 0.6 - 0.9995, 0.6 - 0.99, 0.6 - 0.95, 0.6 - 0.9, 0.6 - 0.8, 0.6 - 0.7, 0.7 - 1.0, 0.7 - 0.9999, 0.7 - 0.9995, 0.7 - 0.999, 0.7 - 0.99, 0.7 - 0.98, 0.7 - 0.97, 0.7 - 0.96, 0.7 - 0.95, 0.7 - 0.9, 0.7 - 0.8, 0.8 - 1.0, 0.8 - 0.9999, 0.8 - 0.9995, 0.8 - 0.999, 0.8 - 0.99, 0.8 - 0.98, 0.8 - 0.97, 0.8 - 0.96, 0.8 - 0.95, 0.8 - 0.9, 0.9 - 1.0, 0.9 - 0.9999, 0.9 - 0.9995, 0.9 - 0.999, 0.9 - 0.99, 0.9 - 0.98, 0.9 - 0.97, 0.9 - 0.96, or 0.9 - 0.95 AUC-ROC. In some cases, the trained machine learning model is trained to achieve an AUC-ROC of about 0.5, 0.6, 0.7, 0.8, 0.85, 0.9, 0.95, 0.96, 0.97, 0.98, 0.99, 0.995, 0.997, 0.999, 0.9995, or 0.9999. In some cases, a high AUC-ROC value indicates a high likelihood of distinguishing true positive variants from true negative variants.

[0027] A trained machine learning model can use one or more thresholds. The threshold for the model can be selected based on, for example, maximizing the average sample AUC of the precision-recall curve. In some cases, in a filtering model, a threshold of at least about 0.1, 0.2, 0.3, 0.4, 0.45, 0.5, 0.55, 0.6, 0.7, 0.8, 0.9, 0.95, or 0.99, or more can be used.

[0028] B. Training a machine learning model framework with a classification model to identify somatic variants A trained machine learning model can correspond to one or more classification models. For example, a trained machine learning model can correspond to 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 models. In some embodiments, one classification model is trained and tested to identify somatic variants from an attribute table. The classification model can include a gradient boosting decision tree, which can be trained to predict somatic variants using, for example, the XGBoost framework. The hyperparameters of the model can be tuned to maximize the macro-average F1 classification score.

[0029] FIG. 2 shows a graph 200 identifying differences in precision and recall values between a gradient boosting decision tree model and a baseline according to some embodiments. After training, the trained machine learning model can show an increase in the average F1 score compared to the baseline. The trained machine learning model can achieve a high AUC-ROC (area under the curve - receiver operating characteristic) of 0.997, indicating the ability to distinguish true positive variants from true negative variants. From the results of FIG. 2, it is shown that it is feasible to predict somatic variants from unmatched tumor sequencing data using the trained machine learning model, and it can be seen that an increase in the accuracy rate can be achieved by a model that can enhance the controllability of the threshold.

[0030] C. Training a machine learning model framework with two classification models to identify somatic variants In some embodiments, the trained machine learning model corresponds to two classification models, each of which is trained and tested to identify somatic variants from an attribute table. To enhance the controllability of the threshold, the classification problem of somatic variants is decomposed into two sub-problems: (1) removing false positives during tumor-only calls from each variant caller, and (2) salvaging false negative candidate variants that do not exist in tumor-only calls.

[0031] FIG. 3 shows two classification models 300 that can be trained to identify somatic variants of non-corresponding biological samples according to some embodiments. To enhance the controllability of the threshold, the classification problem of somatic variants is decomposed into two sub-problems: (1) removing false positives during tumor-only calls from each variant caller, and (2) salvaging false negative candidate variants that do not exist in tumor-only calls. The attribute table can thus be split into two training datasets. In some examples, these two models are trained using a gradient boosting framework (e.g., the LightGBM framework).

[0032] The first training dataset 305 may include candidate variants identified by another variant detection algorithm in a tumor-only context (e.g., MuTect, MuTect2, etc.). A filtering model 310 can be trained to remove false positives from the first training dataset. In some examples, the first training dataset 305 includes a majority of the training dataset, such as about 71% of tumor-normal calls.

[0033] The second training dataset 315 may include the remainder of the candidate variants. A remedial model 320 can be trained to remedy false negatives from the second training dataset. In some examples, the remedial model 320 is trained to distinguish false negatives from true negatives, where false negatives correspond to those that the variant detection algorithm was unable to identify.

[0034] In some examples, two classification models are trained using a gradient boosting framework (such as LightGBM). The classification results from both of these classification models can be combined to create a final set 325 of somatic variants. The final set of somatic variants can then be used to train the classification models 310 and 320. In some examples, training the classification models 310 and 320 includes tuning one or more hyperparameters (such as the learning rate). During training, 300 repetitions of random search are used on the following set of hyperparameters for a known classification problem: (i) maximum depth: 5 - 100, (ii) minimum data within a leaf: 3 - 50, and (iii) number of leaves: 3 - 2048 (logarithmic scale). Each repetition can train each of the classification models, after which stratified 5 - fold cross - validation can be performed. A model averaged over the 5 best - fitting cross - validation models according to AUC - ROC (area under the curve - receiver operating characteristic) can be applied to the test dataset.

[0035] Figure 4 shows a precision - recall curve 400 corresponding to a trained filtering model for removing false positives from a set of candidate somatic variants, according to some embodiments. As shown in Figure 4, the precision - recall curve 400 shows the ability of the filtering model to remove most of the false positives from the dataset. In the precision - recall curve, noise due to varying positive class support is observed, but the AUC - ROC remains fairly constant.

[0036] The threshold of the filtering model 310 can be selected based on maximizing the average sample AUC of the precision-recall curve. For example, a threshold of 0.45 can be selected for the filtering model 310. In some cases, the filtering model 310 includes a threshold of at most about 0.3, 0.4, 0.45, 0.5, 0.55, 0.6, 0.7, 0.8, 0.9, 0.95, or 0.99, or less. In some cases, the filtering model 310 includes a threshold of about 0.2 to 1.0, 0.2 to 0.99, 0.2 to 0.95, 0.2 to 0.9, 0.2 to 0.8, 0.2 to 0.7, 0.2 to 0.6, 0.2 to 0.5, 0.2 to 0.4, 0.2 to 0.3, 0.3 to 1.0, 0.3 to 0.99, 0.3 to 0.95, 0.3 to 0.9, 0.3 to 0.8, 0.3 to 0.7, 0.3 to 0.6, 0.3 to 0.5, 0.3 to 0.4, 0.4 to 1.0, 0.4 to 0.99, 0.4 to 0.95, 0.4 to 0.9, 0.4 to 0.8, 0.4 to 0.7, 0.4 to 0.6, 0.4 to 0.5, 0.5 to 1.0, 0.5 to 0.99, 0.5 to 0.95, 0.5 to 0.9, 0.5 to 0.8, 0.5 to 0.7, 0.5 to 0.6, 0.6 to 1.0, 0.6 to 0.99, 0.6 to 0.95, 0.6 to 0.9, 0.6 to 0.8, 0.6 to 0.7, 0.7 to 1.0, 0.7 to 0.99, 0.7 to 0.98, 0.7 to 0.97, 0.7 to 0.96, 0.7 to 0.95, 0.7 to 0.9, 0.7 to 0.8, 0.8 to 1.0, 0.8 to 0.99, 0.8 to 0.98, 0.8 to 0.97, 0.8 to 0.96, 0.8 to 0.95, 0.8 to 0.9, 0.9 to 1.0, 0.9 to 0.99, 0.9 to 0.98, 0.9 to 0.97, 0.9 to 0.96, or 0.9 to 0.95. In some cases, the filtering model 310 includes a threshold of about 0.1, 0.2, 0.3, 0.4, 0.45, 0.5, 0.55, 0.6, 0.7, 0.8, 0.9, 0.95, or 0.99. In some embodiments, the filtering model 310 includes a threshold of about 0.4 to about 0.5. In some embodiments, the filtering model 310 includes a threshold of about 0.45.

[0037] FIG. 5 shows a Shapley Additive exPlanations (SHAP) graph 500 that identifies which attributes of an attribute table affected the output of a trained filtering model according to some embodiments. The SHAP graph 500 shows graph information that reveals the extent to which each attribute in the attribute table contributed to the identification of false positives of somatic variants in a biological sample. The SHAP graph 500 includes a left side portion 505 that identifies a plurality of features derived from the attribute table, where each column corresponds to one of a plurality of attributes determined for known candidate variants. The SHAP graph 500 also includes a right side portion 510 that reveals the extent to which the known attributes contributed to the identification of false positives of somatic variants in a biological sample. In some examples, these attributes are arranged vertically based on their relative contribution to the identification of false positives. For example, the attribute (gnomAD_AF) corresponding to the upper column can be associated with the highest contribution to the identification of false positives. In this example, gnomAD_AF may refer to the frequency of variants existing in the exome corresponding to the aggregated population, and the existing variants are identified from an aggregated genomic database (e.g., gnomAD).

[0038] FIG. 6 shows a precision-recall curve 600 corresponding to a trained rescue model for removing false negatives from a set of candidate somatic variants according to some embodiments. As shown in FIG. 6, it can be seen from the rescue model data that the feature importance is non-linear and classification is difficult. Due to the overwhelming negative class support, the precision rapidly decreases as the recall of the rescue model increases.

[0039] The threshold of the filtering model 320 can be selected based on maximizing the average sample AUC of the precision-recall curve. For example, a threshold of 0.9995 can be selected for the salvage model 320. In some cases, the salvage model 320 includes a threshold of at least about 0.1, 0.2, 0.3, 0.4, 0.45, 0.5, 0.55, 0.6, 0.7, 0.8, 0.9, 0.95, 0.96, 0.97, 0.98, 0.99, 0.995, 0.999, 0.9995, 0.9999 or more. In some cases, the salvage model 320 includes a threshold of at most about 0.3, 0.4, 0.45, 0.5, 0.55, 0.6, 0.7, 0.8, 0.9, 0.95, 0.99, 0.995, 0.999, 0.9995, 0.9999, or less.In some cases, the relief model 320 includes thresholds of about 0.2 to 1.0, 0.2 to 0.9995, 0.2 to 0.99, 0.2 to 0.95, 0.2 to 0.9, 0.2 to 0.8, 0.2 to 0.7, 0.2 to 0.6, 0.2 to 0.5, 0.2 to 0.4, 0.2 to 0.3, 0.3 to 1.0, 0.3 to 0.9995, 0.3 to 0.99, 0.3 to 0.95, 0.3 to 0.9, 0.3 to 0.8, 0.3 to 0.7, 0.3 to 0.6, 0.3 to 0.5, 0.3 to 0.4, 0.4 to 1.0, 0.4 to 0.9995, 0.4 to 0.99, 0.4 to 0.95, 0.4 to 0.9, 0.4 to 0.8, 0.4 to 0.7, 0.4 to 0.6, 0.4 to 0.5, 0.5 to 1.0, 0.5 to 0.9995, 0.5 to 0.99, 0.5 to 0.95, 0.5 to 0.9, 0.5 to 0.8, 0.5 to 0.7, 0.5 to 0.6, 0.6 to 1.0, 0.6 to 0.9995, 0.6 to 0.99, 0.6 to 0.95, 0.6 to 0.9, 0.6 to 0.8, 0.6 to 0.7, 0.7 to 1.0, 0.7 to 0.9999, 0.7 to 0.9995, 0.7 to 0.999, 0.7 to 0.99, 0.7 to 0.98, 0.7 to 0.97, 0.7 to 0.96, 0.7 to 0.95, 0.7 to 0.9, 0.7 to 0.8, 0.8 to 1.0, 0.8 to 0.9999, 0.8 to 0.9995, 0.8 to 0.999, 0.8 to 0.99, 0.8 to 0.98, 0.8 to 0.97, 0.8 to 0.96, 0.8 to 0.95, 0.8 to 0.9, 0.9 to 1.0, 0.9 to 0.9999, 0.9 to 0.9995, 0.9 to 0.999, 0.9 to 0.99, 0.9 to 0.98, 0.9 to 0.97, 0.9 to 0.96, or 0.9 to 0.95. In some cases, the relief model 320 includes thresholds of about 0.1, 0.2, 0.3, 0.4, 0.45, 0.5, 0.55, 0.6, 0.7, 0.8, 0.9, 0.95, 0.99, 0.995, 0.999, 0.9995, or 0.9999. In some embodiments, the relief model 320 includes a threshold of about 0.9 to about 0.9999. In some embodiments, the relief model 320 includes a threshold of about 0.9995.

[0040] FIG. 7 shows a SHAP graph 700 that identifies which attributes of an attribute table affected the output of a trained relief model according to some embodiments. The SHAP graph 700 shows graph information that reveals the extent to which each attribute in the attribute table contributed to the identification of false negatives of somatic variants in a biological sample. The SHAP graph 700 includes a left side portion 705 that identifies a plurality of features derived from the attribute table, with each column corresponding to one of a plurality of attributes determined for a known candidate variant. The SHAP graph 700 also includes a right side portion 710 that reveals the extent to which a known attribute contributed to the identification of false negatives of somatic variants in a biological sample. In some examples, these attributes are arranged vertically based on their relative contribution to the identification of false negatives. For example, the attribute corresponding to the upper column ("QA") can be associated with the highest contribution to the identification of false negatives. In this example, QA refers to the sum of allele qualities instead of Phred, and the Phred quality score can indicate the magnitude of the identification quality of nucleobases created by automated DNA sequencing.

[0041] The ability of the two-model classification method to predict paired tumor sequencing data from somatic variants can be evaluated before and after training and threshold adjustment. The baseline performance is summarized in Table 1. Macro-average statistics of precision and recall are provided for each sample set. The variance is explained by the similar true positive rate / false positive rate per sample while varying the positive class support. [Table 1]

[0042] Overall, a precision of 0.189 ± 0.19 and a recall of 0.677 ± 0.15 are observed as the baseline. After training and threshold adjustment, the two-model classification method reaches a precision of 0.644 and a recall of 0.634.

[0043] Figure 8 shows a comparison 800 of the performance of a machine learning model having a filtering model and a rescue model before and after training and threshold adjustment, according to some embodiments. In this comparison, the machine learning model was used to predict somatic variants from unmatched tumor sequencing data. The baseline as well as the precision and recall after training and threshold adjustment are shown.

[0044] As shown in Figure 8, it was found that the trained machine learning model having a filtering model and a rescue model can predict somatic variants from unmatched tumor sequencing data with a high precision compared to alternative methods (e.g., MuTect and MuTect2).

[0045] III. Identification of Somatic Variants in Unmatched Biological Samples A. Subjects and Samples An unmatched biological sample is obtained from a cancer patient (i.e., a tumor sample without a corresponding normal sample). The subject can be human. The subject can be male or female. The subject can be a fetus, infant, child, adolescent, teenager, or adult. The subject can be a patient of any age. For example, the subject can be a patient less than about 10 years old. For example, the subject can be a patient at least about 0, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 years old. Commonly, the subject is a patient or other individual who is undergoing a treatment plan or being evaluated for a treatment plan (cancer treatment). However, in some instances, the subject is not undergoing a treatment plan.

[0046] In some cases, the subject may be a mammal or a non-mammal. In some cases, the subject is a mammal such as a human, non-human primate (e.g., ape, monkey, chimpanzee), cat, dog, rabbit, goat, horse, female cow, pig, rodent, mouse, SCID mouse, rat, guinea pig, or sheep. In some methods, variants of the species of the non-human animal model or homologs of these genes can be used. The variants of the species may be genes of different species with the highest sequence homology and similar functional characteristics to each other. Many of the human genes of such species variants may be described in the Swiss-Prot database.

[0047] Some embodiments may include obtaining a sample from a subject, such as a human subject. Specifically, the method may include obtaining a clinical specimen from a patient. For example, blood may be drawn from the patient. Some embodiments may specifically include detecting, profiling, or quantifying molecules (e.g., nucleic acids, DNA, RNA, etc.) within the biological sample.

[0048] The sample may be a tissue sample or a body fluid. In some examples, the sample is an organ sample such as a tissue sample or a biopsy material. In some cases, the sample contains cancerous cells. In some cases, the sample contains cancerous cells and normal cells. In some cases, the sample is a tumor biopsy material. The body fluid may be sweat, saliva, tears, urine, blood, menstrual fluid, semen, and / or cerebrospinal fluid. In some cases, the sample is a blood sample. The sample may contain one or more peripheral blood lymphocytes. The sample may be a whole blood sample. The blood sample may be a peripheral blood sample. In some cases, the sample contains peripheral blood mononuclear cells (PBMCs), and in some cases, the sample contains peripheral blood lymphocytes (PBLs). The sample may be a serum sample.

[0049] The sample can be obtained using any method capable of providing a sample suitable for the analysis methods described herein. The sample may be obtained by non-invasive methods such as throat swabbing, buccal swabbing, bronchial lavage, urine collection, scraping of the skin or cervix, cheek swabbing, saliva collection, fecal collection, menstrual collection, or semen collection. The sample may be obtained by minimally invasive methods such as blood collection. The sample may be obtained by venipuncture. In other instances, the sample is obtained by invasive procedures including, but not limited to, biopsy, alveolar lavage or lung lavage, or needle aspiration biopsy. Biopsy methods may include surgical biopsy, incisional biopsy, excisional biopsy, punch biopsy, shave biopsy, or skin biopsy. The sample may be a formalin-fixed section. The needle aspiration biopsy method may further include fine needle aspiration, core needle biopsy, aspiration biopsy, or large core biopsy. In some cases, multiple samples may be obtained by the methods herein to ensure a sufficient amount of biological material. In some examples, the sample is not obtained by biopsy. In some examples, the sample is not renal biopsy material.

[0050] B. Generating nucleic acid sequencing data In some embodiments, a sample is processed to obtain nucleic acid sequence data. A "nucleic acid" or "nucleic acid molecule" can correspond to a polymeric form of nucleotides of any length, including ribonucleotides, deoxyribonucleotides, or peptide nucleic acids (PNAs), which contain purine and pyrimidine bases, or other natural, chemically modified, biochemically modified, unnatural, or derivatized nucleotide bases. The backbone of the polynucleotide can contain sugars and phosphate groups, such as can typically be found in RNA or DNA, or modified or substituted sugars and phosphate groups. The polynucleotide may contain modified nucleotides such as methylated nucleotides and nucleotide analogs. The nucleotide sequence may be interrupted by non-nucleotide components. Thus, the terms nucleoside, nucleotide, deoxynucleoside, and deoxynucleotide generally include analogs such as those described herein. These analogs are molecules that have some structural features in common with natural nucleosides or nucleotides such that they can hybridize with a naturally occurring nucleic acid sequence in solution when incorporated into a nucleic acid or oligonucleoside sequence. Typically, these analogs are derived from natural nucleosides or nucleotides by substituting and / or modifying the base, ribose, or phosphodiester moiety. These alterations can serve the purpose of stabilizing or destabilizing hybridization or enhancing the specificity of hybridization with a complementary nucleic acid sequence as desired. The nucleic acid molecule can be a DNA molecule. The nucleic acid molecule can be an RNA molecule.

[0051] DNA is extracted from a tumor sample, processed, and subjected to whole exome sequencing. The sequencing reads are subjected to quality control processing (e.g., by FastQC) to provide a FASTQ file. The FASTQ file is aligned to a reference genome to create a BAM file.

[0052] In some cases, sample processing includes nucleic acid sample processing and subsequent nucleic acid sample sequencing. Some or all of the nucleic acid sample may be sequenced to provide sequence information, which may be stored or otherwise maintained in an electronic, magnetic, or optical storage location. The sequence information may be analyzed with the aid of a computer processor, and the analyzed sequence information may be stored in an electronic storage location. The electronic storage location may include a pool or collection of sequence information created from and analyzed nucleic acid samples. The nucleic acid sample may be recovered from a subject, such as a subject having or suspected of having cancer.

[0053] Some embodiments may include using whole genome sequencing. In some cases, whole genome sequencing is used to identify an individual's variants. In some cases, the sequencing can include deep sequencing across a fraction of the genome. For example, the fraction of the genome can be at least about 50; 75; 100; 125; 150; 175; 200; 225; 250; 275; 300; 350; 400; 450; 500; 550; 600; 650; 700; 750; 800; 850; 900; 950; 1,000; 1100; 1200; 1300; 1400; 1500; 1600; 1700; 1800; 1900; 2,000; 3,000; 4,000; 5,000; 6,000; 7,000; 8,000; 9,000; 10,000; 15,000; 20,000; 30,000; 40,000; 50,000; 60,000; 70,000; 80,000; 90,000; 100,000, or more bases or base pairs. In some cases, the genome may be sequenced across 1 million, 2 million, 3 million, 4 million, 5 million, 6 million, 7 million, 8 million, 9 million, 10 million or more than 10 million bases or base pairs. In some cases, the genome may be sequenced across the entire exome (e.g., whole exome sequencing). In some cases, deep sequencing may include obtaining multiple reads across a fraction of the genome. For example, obtaining multiple reads can include at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 10,000 reads, or more than 10,000 reads across a fraction of the genome.

[0054] Some embodiments may include detecting low allele fractions by deep sequencing. In some cases, deep sequencing is performed by next-generation sequencing. In some cases, deep sequencing is performed by avoiding error-prone regions. In some cases, error-prone regions may include regions proximate to sequence duplications, regions with unusually high or low GC content, homopolymers, regions proximate to dinucleotides and trinucleotides, and regions proximate to other short repetitive sequences. In some cases, error-prone regions may include regions leading to DNA sequencing errors (e.g., polymerase translational slippage in homopolymers).

[0055] Some embodiments may include performing one or more sequencing reactions on one or more nucleic acid molecules in a sample. Some embodiments may include performing one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, fifteen or more, twenty or more, thirty or more, forty or more, fifty or more, sixty or more, seventy or more, eighty or more, ninety or more, one hundred or more, two hundred or more, three hundred or more, four hundred or more, five hundred or more, six hundred or more, seven hundred or more, eight hundred or more, nine hundred or more, or one thousand or more sequencing reactions on one or more nucleic acid molecules in a sample. The sequencing reactions may be run simultaneously with this, over time with this, or in combinations of these. The sequencing reactions may include whole genome sequencing or whole exome sequencing. The sequencing reactions may include the Maxim-Gilbert method, chain termination reactions, or high-throughput systems. Alternatively, or additionally, the sequencing reactions may include, but are not limited to, Helioscope™ single molecule sequencing, nanopore DNA sequencing, Lynx Therapeutics Massively Parallel Signature Sequencing method (MPSS), 454 pyrosequencing, single molecule real-time (RNAP) sequencing, Illumina (Solexa) sequencing, SOLiD sequencing, Ion Torrent™ Ion semiconductor sequencing, single molecule SMRT™ sequencing, polony sequencing, DNA nanoball sequencing, the method of VisiGen Biotechnologies, or combinations of these. Alternatively or additionally, the sequencing reactions may include one or more sequencing platforms including, but not limited to, the Genome Analyzer IIx, HiSeq, and MiSeq provided by Illumina, single molecule real-time (SMRT™) technologies such as the PacBio RS system and the Solexa sequencer provided by Pacific Biosciences (California), True Single Molecule Sequencing (tSMSTM) technologies such as the HeliScope™ sequencer provided by Helicos Inc. (Cambridge, MA).The array determination reaction may also include an electron microscope or a chemFET (chemical field effect transistor) array. In some embodiments, the array determination reaction includes capillary sequencing, next-generation sequencing, Sanger sequencing, sequencing by synthesis, sequencing by ligation, sequencing by hybridization, single-molecule sequencing, or combinations thereof. Sequencing by synthesis may include reversible terminator sequencing, processive single molecule sequencing, sequential flow sequencing, or combinations thereof. Sequential flow sequencing may include pyrosequencing, pH-mediated sequencing, semiconductor sequencing, or combinations thereof.

[0056] Some embodiments may include performing at least one long-read array determination reaction and at least one short-read array determination reaction. The long-read array determination reaction and / or the short-read array determination reaction may be performed on at least a portion of a subset of nucleic acid molecules. The long-read array determination reaction and / or the short-read array determination reaction may be performed on at least a portion of two or more subsets of nucleic acid molecules. Both the long-read array determination reaction and the short-read array determination reaction may be performed on at least a portion of one or more subsets of nucleic acid molecules.

[0057] Sequencing of one or more nucleic acid molecules or subsets thereof may include at least about 5; 10; 15; 20; 25; 30; 35; 40; 45; 50; 60; 70; 80; 90; 100; 200; 300; 400; 500; 600; 700; 800; 900; 1,000; 1500; 2,000; 2500; 3,000; 3500; 4,000; 4500; 5,000; 5500; 6,000; 6500; 7,000; 7500; 8,000; 8500; 9,000; 10,000; 25,000; 50,000; 75,000; 100,000; 250,000; 500,000; 750,000; 10,000,000; 25,000,000; 50,000,000; 100,000,000; 250,000,000; 500,000,000; 750,000,000; 1,000,000,000 or more sequencing reads.

[0058] The sequencing reaction may include sequencing at least about 50; 60; 70; 80; 90; 100; 110; 120; 130; 140; 150; 160; 170; 180; 190; 200; 210; 220; 230; 240; 250; 260; 270; 280; 290; 300; 325; 350; 375; 400; 425; 450; 475; 500; 600; 700; 800; 900; 1,000; 1500; 2,000; 2500; 3,000; 3500; 4,000; 4500; 5,000; 5500; 6,000; 6500; 7,000; 7500; 8,000; 8500; 9,000; 10,000; 20,000; 30,000; 40,000; 50,000; 60,000; 70,000; 80,000; 90,000; 100,000 or more bases or base pairs of one or more nucleic acid molecules. The sequencing reaction may include sequencing at least about 50; 60; 70; 80; 90; 100; 110; 120; 130; 140; 150; 160; 170; 180; 190; 200; 210; 220; 230; 240; 250; 260; 270; 280; 290; 300; 325; 350; 375; 400; 425; 450; 475; 500; 600; 700; 800; 900; 1,000; 1500; 2,000; 2500; 3,000; 3500; 4,000; 4500; 5,000; 5500; 6,000; 6500; 7,000; 7500; 8,000; 8500; 9,000; 10,000; 20,000; 30,000; 40,000; 50,000; 60,000; 70,000; 80,000; 90,000; 100,000 or more consecutive bases or base pairs of one or more nucleic acid molecules.

[0059] Preferably, the sequencing technology used in the method of the present invention generates at least 100 reads per run, at least 200 reads per run, at least 300 reads per run, at least 400 reads per run, at least 500 reads per run, at least 600 reads per run, at least 700 reads per run, at least 800 reads per run, at least 900 reads per run, at least 1000 reads per run, at least 5,000 reads per run, at least 10,000 reads per run, at least 50,000 reads per run, at least 100,000 reads per run, at least 500,000 reads per run, or at least 1,000,000 reads per run. Alternatively, the sequencing technology used in the method of the present invention generates at least 1,500,000 reads per run, at least 2,000,000 reads per run, at least 2,500,000 reads per run, at least 3,000,000 reads per run, at least 3,500,000 reads per run, at least 4,000,000 reads per run, at least 4,500,000 reads per run, or at least 5,000,000 reads per run.

[0060] Preferably, the sequencing technology used in the method of the present invention generates at least about 30 base pairs, at least about 40 base pairs, at least about 50 base pairs, at least about 60 base pairs, at least about 70 base pairs, at least about 80 base pairs, at least about 90 base pairs, at least about 100 base pairs, at least about 110, at least about 120 base pairs per read, at least about 150 base pairs, at least about 200 base pairs, at least about 250 base pairs, at least about 300 base pairs, at least about 350 base pairs, at least about 400 base pairs, at least about 450 base pairs, at least about 500 base pairs, at least about 550 base pairs, at least about 600 base pairs, at least about 700 base pairs, at least about 800 base pairs, at least about 900 base pairs, or at least about 1,000 base pairs per read. Alternatively, the sequencing technology used in the method of the present invention can generate long sequencing reads. In some examples, the sequencing technology used in the method of the present invention generates at least about 1,200 base pairs per read, at least about 1,500 base pairs per read, at least about 1,800 base pairs per read, at least about 2,000 base pairs per read, at least about 2,500 base pairs per read, at least about 3,000 base pairs per read, at least about 3,500 base pairs per read, at least about 4,000 base pairs per read, at least about 4,500 base pairs per read, at least about 5,000 base pairs per read, at least about 6,000 base pairs per read, at least about 7,000 base pairs per read, at least about 8,000 base pairs per read, at least about 9,000 base pairs per read, at least about 10,000 base pairs per read, 20,000 base pairs per read, 30,000 base pairs per read, 40,000 base pairs per read, 50,000 base pairs per read, 60,000 base pairs per read, 70,000 base pairs per read, 80,000 base pairs per read, 90,000 base pairs per read, or 100,000 base pairs per read.

[0061] Using a high-throughput sequencing system, the sequenced nucleotides can be detected immediately after or when incorporated into the growing strand, i.e., the sequence can be detected in real-time or substantially in real-time. In some cases, high-throughput sequencing results in each read being at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 120, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, or at least 500 bases long, and at least 1,000, at least 5,000, at least 10,000, at least 20,000, at least 30,000, at least 40,000, at least 50,000, at least 100,000 or at least 500,000 sequence reads per unit time. Sequencing can be performed using as a template a nucleic acid as described herein, such as genomic DNA, cDNA or RNA derived from an RNA transcript.

[0062] C. Identifying candidate variants The nucleic acid sequence data can be aligned with a reference genome. Based on the aligned nucleic acid sequence data, a set of candidate variants in the nucleic acid sequence data can be identified. In some examples, the set of candidate variants includes one or more somatic variants and one or more germline variants. For example, BCFtools can be used to identify with high sensitivity a set of candidate somatic variants for each sample. The set of candidate somatic variants will include false positives, e.g., germline variants.

[0063] For a set of candidate somatic variants, create an attribute table that includes many features (e.g., about 10-20 features) for each candidate variant. The attribute table can include any combination of the attributes described in Example 3. The attribute table can include many features for each candidate variant. Examples of features that the attribute table can include include, but are not limited to, (i) pileup attributes derived from the initial BCFtools output such as allele frequency (e.g., B allele frequency), base quality, read depth, (ii) an estimate of tumor purity determined using a deep learning neural network based on the overall exome B allele frequency distribution in the sample, (iii) whether the variant is identified as a germline variant using GATK HaplotypeCaller, (iv) the somatic copy number alteration (CNA) state for each variant site, (v) the frequency of the variant in the population (e.g., in a healthy human population and / or in cancer exomes from databases such as Cosmic, GnomAD, Dbsnp, Mills Indels), (vi) the presence of the variant in a problematic region such as in a homopolymer, and (vii) whether the variant is identified by standard somatic callers such as MuTect and MuTect2 (running in the context of a single tumor).

[0064] The attribute table can include any number of features that can contribute to the accurate prediction of somatic cell variants. For example, the attribute table can include at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 features, or more. In some cases, the attribute table can include at most about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 features or fewer.In some embodiments, the attribute table can include about 1 to 100, 1 to 90, 1 to 80, 1 to 70, 1 to 60, 1 to 50, 1 to 40, 1 to 30, 1 to 20, 1 to 10, 1 to 5, 5 to 100, 5 to 90, 5 to 80, 5 to 70, 5 to 60, 5 to 50, 5 to 40, 5 to 30, 5 to 20, 5 to 10, 10 to 100, 10 to 90, 10 to 80, 10 to 70, 10 to 60, 10 to 50, 10 to 40, 10 to 30, 10 to 20, 15 to 100, 15 to 90, 15 to 80, 15 to 70, 15 to 60, 15 to 50, 15 to 40, 15 to 30, 15 to 20, 20 to 100, 20 to 90, 20 to 80, 20 to 70, 20 to 60, 20 to 50, 20 to 40, or 20 to 30 features. In some cases, the attribute table includes about 10 to 20 features.

[0065] In some embodiments, identifying a set of candidate variants may include identifying one or more genomic regions that include one or more nucleotide sequence variants. The one or more genomic regions may include one or more genomic region features. The genomic region features may include the entire genome or a portion thereof. The genomic region features may include the entire exome or a portion thereof. The genomic region features may include one or more sets of genes. The genomic region features may include one or more genes. The genomic region features may include one or more sets of regulatory regions. The genomic region features may include one or more regulatory regions. The genomic region features may include gene set polymorphisms. The genomic region features may include one or more gene polymorphisms. The genomic region features may relate to the GC content, complexity, and / or mappability of one or more nucleic acid molecules. The genomic region features may include one or more simple tandem repeats (STRs), unstable expanding repeats, segmental duplications, single and paired read degenerative mapping scores, GRCh37 patches, or combinations thereof. The genomic region features may include one or more low mean coverage regions from whole genome sequencing (WGS), zero mean coverage regions from WGS, validated compressions, or combinations thereof. The genomic region features may include one or more proxy or non-reference sequences. The genomic region features may include one or more gene phasing and reconstructed genes. In some aspects, the one or more genomic region features are not mutually exclusive. For example, a genomic region feature that includes the entire genome or a portion thereof may overlap with additional genomic region features such as the entire exome or a portion thereof, one or more genes, one or more regulatory elements, etc. Alternatively, the one or more genomic region features are mutually exclusive.For example, a genomic region that includes the non-coding part of the entire genome will not overlap with genomic region features such as the exon of a gene or a part thereof, or the coding part. Alternatively, or on top of that, one or more genomic region features are partially exclusive or partially inclusive. For example, a genomic region that includes the entire exon or a part thereof can partially overlap with a genomic region that includes the exon part of a gene. However, a genomic region that includes the entire exon or a part thereof will not overlap with a genomic region that includes the intron part of a gene. Therefore, a genomic region feature that includes a gene or a part thereof may partially exclude and / or partially include a genomic region that includes the entire exon or a part thereof.

[0066] Some embodiments may include a nucleic acid sample or molecule that includes one or more genomic regions, at least one of the one or more genomic regions including a genomic region feature that includes the entire genome or a portion thereof. The entire genome or a portion thereof may include one or more coding portions of the genome, one or more non-coding portions of the genome, or a combination thereof. The coding portion of the genome may include one or more coding portions of genes that encode one or more proteins. One or more coding portions of the genome may include the entire exome or a portion thereof. Alternatively, or in addition, one or more coding portions of the genome may include one or more exons. One or more non-coding portions of the genome may include one or more non-coding molecules or portions thereof. Non-coding molecules may include one or more non-coding RNAs, one or more regulatory elements, one or more introns, one or more pseudogenes, one or more repetitive sequences, one or more transposons, one or more viral elements, one or more telomeres, portions thereof, or combinations thereof. Non-coding RNAs may be functional RNA molecules that are not translated into proteins. Examples of non-coding RNAs include, but are not limited to, ribosomal RNA, transfer RNA, Piwi-interacting RNA, microRNA, small interfering RNA (siRNA), small hairpin RNA (shRNA), small nucleolar RNA (snoRNA), small non-coding RNA (sncRNA), and long non-coding RNA (lncRNA). Pseudogenes may be related to known genes and typically are no longer expressed. Repetitive sequences may include one or more tandem repeats, one or more interspersed repeats, or combinations thereof. Tandem repeats may include one or more satellite DNAs, one or more minisatellites, one or more microsatellites, or combinations thereof. Interspersed repeats may include one or more transposons. Transposons may be mobile genetic elements. Mobile genetic elements are often able to change their position within the genome. Transposons may be classified as class I transposable elements (class I TEs) or class II transposable elements (class II TEs).Class I TEs (retrotransposons) can often copy themselves in two steps: first, from DNA to RNA by transcription, and then back from RNA to DNA by reverse transcription. The DNA copy may be inserted into the genome at a new location. Class I TEs may include one or more long terminal repeats (LTRs), one or more long interspersed nuclear elements (LINEs), one or more short interspersed nuclear elements (SINEs), or combinations thereof. Examples of LTRs include, but are not limited to, human endogenous retroviruses (HERVs), medium reiterated repeats 4 (MER4), and retrotransposons. Examples of LINEs include, but are not limited to, LINE1 and LINE2. SINEs include one or more Alu sequences, one or more mammalian-wide interspersed repeats (MIRs), or combinations thereof. Class II TEs (e.g., DNA transposons) often do not involve an RNA intermediate. DNA transposons are often cut from one site in the genome and inserted into another site. Alternatively, DNA transposons are replicated and inserted into a new location in the genome. Examples of DNA transposons include, but are not limited to, MER1, MER2, and mariner. Viral elements may include one or more endogenous retroviral sequences. Telomeres are often repetitive DNA regions at the ends of chromosomes.

[0067] Some embodiments may include a nucleic acid sample or a subset of nucleic acid molecules that includes one or more genomic regions, wherein at least one of the one or more genomic regions includes genomic region characteristics that include the entire exosome or a portion thereof. The exosome is often a portion of the genome formed by exons. The exosome may be formed by untranslated regions (UTRs), splice sites, and / or intron regions. The entire exosome or its proteins may include one or more exons of a protein-coding gene. The entire exosome or its proteins may include one or more untranslated regions (UTRs), splice sites, and / or introns.

[0068] Some embodiments may include a nucleic acid sample or molecule comprising one or more genomic regions, wherein at least one of the one or more genomic regions comprises a genomic region feature that includes a gene or a portion thereof. A gene typically comprises a stretch of nucleic acid that encodes a polypeptide or functional RNA. A gene may include one or more exons, one or more introns, one or more untranslated regions (UTRs), or combinations thereof. Exons are often the coding sections of a gene that are transcribed into the mRNA sequence precursor and are within the final mature RNA product of the gene. Introns are often the non-coding sections of a gene that are transcribed into the mRNA sequence precursor and are removed by RNA splicing. A UTR may refer to sections on each side of the coding sequence of an mRNA strand. The UTR located on the 5’ side of the coding sequence may be referred to as the 5’UTR (or leader sequence). The UTR located on the 3’ side of the coding sequence may be referred to as the 3’UTR (or trailer sequence). A UTR may include one or more elements for regulating gene expression. Elements such as regulatory elements may be located in the 5’UTR. Regulatory elements such as polyadenylation signals, protein binding sites, and miRNA binding sites may be located in the 3’UTR. Protein binding sites located in the 3’UTR may include, but are not limited to, the selenocysteine insertion sequence (SECIS) element and the AU rich element (ARE). The SECIS element may instruct the ribosome to translate the codon UGA as selenocysteine rather than as a stop codon. The ARE is often a stretch consisting mainly of adenine and uracil nucleotides, which may affect the stability of the mRNA.

[0069] Some embodiments may include a nucleic acid sample or a subset of nucleic acid molecules that includes one or more genomic regions, and at least one of the one or more genomic regions includes a genomic region feature that includes a set of genes. The set of genes may include, but is not limited to, Mendel DB genes, Human Gene Mutation Database (HGMD) genes, Cancer Gene Census genes, Online Mendelian Inheritance in Man (OMIM) Mendelian genes, HGMD Mendelian genes, and human leukocyte antigen (HLA) genes. The set of genes may include one or more known Mendelian traits, one or more known disease characteristics, one or more drug traits, one or more biomedically interpretable variants, or combinations thereof. The Mendelian trait may be regulated by a single locus and may exhibit a Mendelian inheritance pattern. The set of genes having known Mendelian traits may include, but is not limited to, one or more genes encoding Mendelian traits including the ability to taste phenylthiocarbamide (dominant), the ability to smell hydrogen cyanide (like bitter almonds) (recessive), albinism (recessive), brachydactyly (short fingers and toes), and wet (dominant) or dry (recessive) earwax. Disease characteristics may cause or increase the risk of a disease and may be inherited in a Mendelian or complex pattern. The set of genes having known disease characteristics may include, but is not limited to, one or more genes encoding disease characteristics including cystic fibrosis, hemophilia, and Lynch syndrome. A drug trait can alter the metabolism, optimal dosage, adverse reactions, and side effects of one or more drugs or drug groups. The set of genes having known drug traits may include, but is not limited to, one or more genes encoding drug traits including CYP2D6, UGT1A1, and ADRB1. Biomedically interpretable variants may be polymorphisms of genes associated with a disease or a symptom.A set of genes having known biomedically interpretable variants may include one or more genes encoding biomedically interpretable variants including, but not limited to, cystic fibrosis (CF) variants, muscular dystrophy variants, p53 variants, Rb variants, cell cycle control, receptors, and kinases. Alternatively, or in addition, a set of genes having known biomedically interpretable variants may include one or more genes associated with Huntington's disease, cancer, cystic fibrosis, muscular dystrophy (e.g., Duchenne muscular dystrophy).

[0070] Some embodiments may include a nucleic acid sample or molecule comprising one or more genomic regions, at least one of the one or more genomic regions comprising a genomic region feature comprising a regulatory element or a portion thereof. The regulatory element may be a cis-regulatory element or a trans-regulatory element. The cis-regulatory element may be a sequence that controls the transcription of nearby genes. The cis-regulatory element may be located in the 5' or 3' untranslated region (UTR), or within an intron. The trans-regulatory element can control the transcription of distant genes. The regulatory element may include one or more promoters, one or more enhancers, or a combination thereof. The promoter may facilitate the transcription of a particular gene and can be found upstream of the coding region. The enhancer can exert a remote effect on the transcription level of a gene.

[0071] Some embodiments may include a nucleic acid sample or a subset of nucleic acid molecules that includes one or more genomic regions, at least one of the one or more genomic regions including a genomic region feature that includes a genetic polymorphism or a part thereof. A genetic polymorphism generally refers to a variation in genotype. A genetic polymorphism can be a germline variant or a somatic variant. A genetic polymorphism may include a change, insertion, repeat, or deletion of one or more bases. Copy number polymorphisms (CNVs), transversions, and other rearrangements are also forms of genetic variation. Genetic polymorphism markers include restriction fragment length polymorphisms, variable number tandem repeats (VNTRs), hypervariable regions, minisatellites, dinucleotide repeats, trinucleotide repeats, tetranucleotide repeats, simple sequence repeats, and insertion elements such as Alu. The allele form that occurs most frequently in a selected population is sometimes called the wild type. A diploid organism may be homozygous or heterozygous for an allele form. A diallelic genetic polymorphism has two forms. A triallelic genetic polymorphism has three forms. A single nucleotide polymorphism (SNP) is one form of genetic polymorphism. In some aspects, one or more genetic polymorphisms include one or more single nucleotide polymorphisms, indels, small insertions, small deletions, structural variant junctions, variable length tandem repeats, flanking sequences, or combinations thereof. One or more genetic polymorphisms may be located within a coding region and / or a non-coding region. One or more genetic polymorphisms may be located within, around, or near a gene, exon, intron, splice site, untranslated region, or combinations thereof. One or more genetic polymorphisms may span at least a gene, exon, intron, and untranslated region.

[0072] Some embodiments may include a nucleic acid sample or molecule comprising one or more genomic regions, wherein at least one of the one or more genomic regions comprises a genomic region feature comprising one or more simple tandem repeats (STRs), unstable expanding repeats, segmental duplications, single and paired read degenerative mapping scores, GRCh37 patches, or combinations thereof. The one or more STRs may include one or more homopolymers, one or more dinucleotide repeats, one or more trinucleotide repeats, or combinations thereof. The one or more homopolymers may be about 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more bases or base pairs. The dinucleotide repeats and / or trinucleotide repeats may be 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50 or more bases or base pairs. The single and paired read degenerative mapping scores may be based on or derived from the alignability by 100-mers of GEM from ENCODE / CRG (Guigo), the alignability by 75-mers of GEM from ENCODE / CRG (Guigo), the 100 base pair box car average for signal mappability, the max of locus and possible pairs for paired read score, or combinations thereof. The genomic region features may include one or more low mean coverage regions from whole genome sequencing (WGS), zero mean coverage regions from WGS, validated compressions, or combinations thereof.Low average coverage regions from WGS may include regions created from Illumina v3 chemistry, regions below the first percentile of a Poisson distribution based on average coverage, or combinations thereof. Zero average coverage regions from WGS may include regions created from Illumina v3 chemistry. Validated compressions may include regions of high mapped depth, regions with two or more observed haplotypes, regions expected to be missing repeats in a reference, or combinations thereof. Genome region features may include one or more alternative or non-reference sequences. One or more alternative or non-reference sequences may include known structural variant junctions, known insertions, known deletions, alternative haplotypes, or combinations thereof. Genome region features may include one or more gene phasing and reconstructed genes. Examples of phasing and reconstructed genes include, but are not limited to, one or more major histocompatibility complexes, blood types, and amylase gene families. One or more major histocompatibility complexes may include one or more HLA class I, HLA class II, or combinations thereof. One or more HLA class I may include HLA-A, HLA-B, HLA-C, or combinations thereof. One or more HLA class II may include HLA-DP, HLA-DM, HLA-DOA, HLA-DOB, HLA-DQ, HLA-DR, or combinations thereof. Blood type genes may include ABO, RHD, RHCE, or combinations thereof.

[0073] Some embodiments may include a nucleic acid sample or molecule comprising one or more genomic regions, wherein at least one of the one or more genomic regions comprises a genomic region feature related to the GC content of one or more nucleic acid molecules. The GC content may refer to the GC content of a nucleic acid molecule. Alternatively, the GC content may refer to the GC content of one or more nucleic acid molecules and may be referred to as the average GC content. As used herein, the terms "GC content" and "average GC content" may be used interchangeably. The GC content of a genomic region may be a high GC content. Typically, a high GC content refers to a GC content of about 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97% or more, or more. In some embodiments, a high GC content may refer to a GC content of about 70% or more. The GC content of a genomic region may be a low GC content. Typically, a low GC content refers to a GC content of about 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10%, 5%, 2% or less, or less.

[0074] Some embodiments may include a nucleic acid sample or molecule comprising one or more genomic regions, wherein at least one of the one or more genomic regions comprises a genomic region feature related to the complexity of one or more nucleic acid molecules. The complexity of a nucleic acid molecule may refer to the randomness of the nucleotide sequence. Low complexity may refer to the pattern, repetition and / or depletion of nucleotides of one or more chemical species in the sequence.

[0075] Some embodiments may include a nucleic acid sample or molecule comprising one or more genomic regions, wherein at least one of the one or more genomic regions comprises a genomic region feature related to the mappability of one or more nucleic acid molecules. The mappability of a nucleic acid molecule may refer to the uniqueness of the alignment of the nucleic acid molecule to a reference sequence. A nucleic acid molecule with low mappability may have a low alignment to the reference sequence.

[0076] D. Predicting Whether a Candidate Variant is a Somatic Variant The two-model classification method is used to predict somatic variants from an attribute table. For example, the attribute table can be subdivided into two datasets as shown in FIG. 4 and processed using a trained model such as the model described in Example 3. The first dataset can include candidate somatic variants identified by one or more bioinformatics tools. The first model can be applied to remove false positives from this dataset. The second dataset can include the remainder of the candidate variants, including false negatives and true negatives. The second model can be applied to salvage false negatives from this dataset. This method can predict somatic variants with an acceptable accuracy rate despite the absence of corresponding normal samples.

[0077] To enhance the controllability of the threshold, the classification problem of somatic variants is decomposed into two sub-problems: (1) removing false positives during tumor-only calls from each variant caller, and (2) salvaging false-negative candidate variants that do not exist in tumor-only calls. The attribute table is subdivided into two datasets. The first dataset includes candidate variants identified by MuTect and MuTect2 (tumor-only context). The first model is trained to remove false positives from this dataset. The second dataset includes the remainder of the candidate variants. The second model is trained to salvage false negatives from this dataset. The models are trained using Microsoft's LightGBM framework (LGBM). The classification results from both of these models are then combined to create a final set of candidate variants.

[0078] E. Create a report identifying candidate variants One or more reports can be created that include some or all of the predicted candidate variants (diagnostic reports and / or prognostic reports). Based on the predicted candidate variants and / or reports, one or more treatments can be administered to or withheld from a patient. For example, to diagnose or characterize cancer, the predicted candidate variants can be compared to one or more databases of known cancer mutations. Variants associated with responsiveness or non-responsiveness to a particular cancer treatment can be identified and treatment recommendations can be provided. Cancer can be treated based on this recommendation.

[0079] IV. Process for Somatic Variants from Unmatched Biological Samples FIG. 9 includes a flowchart 900 showing an example of a method for somatic variant calling from unmatched biological samples, according to some embodiments. The operations described in flowchart 900 may be performed, for example, by a computer system executing a trained machine learning model that includes a filtering model and a rescue model. Although flowchart 900 can describe the operations as sequential processes, in various embodiments, many of the operations may be performed in parallel or simultaneously. Further, the order of the operations may be rearranged. The operations may have additional steps not described in the figure. Further, embodiments of the method may be performed by hardware, software, firmware, middleware, microcode, a hardware description language, or any combination thereof. When executed in software, firmware, middleware, or microcode, program code or code segments for performing the associated tasks may be stored in a computer-readable medium such as a storage medium.

[0080] In operation 910, the computer system obtains nucleic acid sequence data of a target biological sample. The nucleic acid sequence data can be created by sequencing a plurality of nucleic acid molecules of a tumor sample. In some embodiments, the tumor sample is from a human subject. Sequencing can include whole exome sequencing. In some embodiments, this sequencing can include whole genome sequencing. In some embodiments, this sequencing includes shotgun sequencing. In some embodiments, this sequencing includes sequencing a selected portion of the genome or exome.

[0081] In operation 920, the computer system aligns the nucleic acid sequence data with a reference genome. For example, a FASTQ file corresponding to the nucleic acid sequence data can be aligned with the reference genome to create one or more BAM files.

[0082] In operation 930, the computer system identifies a set of candidate variants in the nucleic acid sequence data based on the aligned nucleic acid sequence data. In some examples, the set of candidate variants includes one or more somatic variants and one or more germline variants. A somatic variant refers to a change in DNA that occurs after conception and does not exist in the germline. A germline variant refers to a genetic change in germ cells (sperm and egg cells) that is incorporated into the DNA of all cells in the body of offspring. In some examples, somatic variants indicate the presence or level of cancer in a subject instead of germline variants.

[0083] An attribute table can be created, where the attribute table can include many features for each candidate variant. In some embodiments, the attribute table includes attributes from sequencing data corresponding to a particular candidate variant. The attribute table can include attributes from a file containing processed sequencing data. In some embodiments, the attribute table includes one or more of the following attributes: (a) pileup attributes from a BCFtools output file; (b) allele frequency data; (c) base quality data; (d) read depth data; (e) an estimate of tumor cell fraction (which may be calculated based on the B allele frequency distribution); (f) predicted germline variants; (g) predicted somatic variants; (h) copy number variation data; (i) population frequency data from one or more databases; (j) data from at least one database selected from the group consisting of Cosmic, GnomAD, Dbsnp, and Mills Indels; (k) data regarding the presence of candidate somatic variants in problematic regions of the genome, and (l) data regarding the presence of candidate somatic variants in homopolymers.

[0084] In operation 940, the computer system processes a set of candidate variants using a trained machine learning model to identify somatic variants without using the nucleic acid sequence data of the corresponding biological sample of interest. In some examples, the trained machine learning model includes a gradient boosting decision tree that facilitates a significant reduction in the false positive rate corresponding to somatic variant calls. In some embodiments, the trained machine learning model includes a two-model classification method. The trained machine learning model may include a filtering model that removes false positives. The trained machine learning model may also include a rescue model that remedies false negatives. In some embodiments, the attribute table includes attributes from sequencing data.

[0085] In operation 950, the computer system outputs a report identifying somatic variants. In some embodiments, the report includes information identifying at least one diagnostic marker and at least one prognostic marker. In some embodiments, the absence of somatic variants, treatment recommendations, recommendations to treat a human subject, and / or recommendations not to treat a human subject. In some embodiments, the recommended treatment is performed on the human subject. Process 900 ends there.

[0086] V. Further Considerations A. Exploration Techniques Some embodiments may include one or more labels. The one or more labels may be attached to one or more capture probes, nucleic acid molecules, beads, primers, or combinations thereof. Examples of labels include, but are not limited to, radioisotopes, fluorophores, chemiluminescent groups, chromophores, lumiphores, enzymes, colloidal particles, and fluorescent microparticles, quantum dots, and detectable labels such as antigens, antibodies, haptens, avidin / streptavidin, biotin, haptens, enzyme cofactors / substrates, one or more members of a quenching system, chromogens, haptens, magnetic particles, substances exhibiting nonlinear optics, semiconductor nanocrystals, metal nanoparticles, enzymes, aptamers, and one or more members of a binding pair.

[0087] Some embodiments may include one or more capture probes, a plurality of capture probes, or one or more sets of capture probes. Typically, a capture probe includes a nucleic acid binding site. The capture probe may further include one or more linkers. The capture probe may further include one or more labels. The one or more linkers may attach the one or more labels to the nucleic acid binding site.

[0088] The capture probe may hybridize with one or more nucleic acid molecules in the sample. The capture probe may hybridize with one or more genomic regions. The capture probe may hybridize with one or more genomic regions within, surrounding, near, or encompassing one or more genes, exons, introns, UTRs, or combinations thereof. The capture probe may hybridize with one or more genomic regions encompassing one or more genes, exons, introns, UTRs, or combinations thereof. The capture probe may hybridize with one or more known indels. The capture probe may hybridize with one or more known structural variants.

[0089] Some embodiments may include one or more than one, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, twenty or more, thirty or more, forty or more, fifty or more, sixty or more, seventy or more, eighty or more, ninety or more, one hundred or more, one hundred and twenty-five or more, one hundred and fifty or more, one hundred and seventy-five or more, two hundred or more, two hundred and fifty or more, three hundred or more, three hundred and fifty or more, four hundred or more, five hundred or more, six hundred or more, seven hundred or more, eight hundred or more, nine hundred or more, or one thousand or more capture probes or capture probe sets. The one or more capture probes or capture probe sets may be different, similar, identical, or combinations thereof.

[0090] One or more capture probes may include a nucleic acid binding site that hybridizes to at least a portion of one or more nucleic acid molecules or variants or derivatives thereof in the sample or a subset of nucleic acid molecules. The capture probe may include a nucleic acid binding site that hybridizes with one or more genomic regions. The capture probe may hybridize with different, similar, and / or identical genomic regions. One or more capture probes may be at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99% or more complementary to one or more nucleic acid molecules or variants or derivatives thereof.

[0091] The capture probe may comprise one or more nucleotides. The capture probe may comprise one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, twenty or more, thirty or more, forty or more, fifty or more, sixty or more, seventy or more, eighty or more, ninety or more, one hundred or more, one hundred and twenty-five or more, one hundred and fifty or more, one hundred and seventy-five or more, two hundred or more, two hundred and fifty or more, three hundred or more, three hundred and fifty or more, four hundred or more, five hundred or more, six hundred or more, seven hundred or more, eight hundred or more, nine hundred or more, or one thousand or more nucleotides. The capture probe may comprise about 100 nucleotides. The capture probe may comprise about 10 to about 500 nucleotides, about 20 to about 450 nucleotides, about 30 to about 400 nucleotides, about 40 to about 350 nucleotides, about 50 to about 300 nucleotides, about 60 to about 250 nucleotides, about 70 to about 200 nucleotides, about 80 to about 150 nucleotides. In some embodiments, the capture probe may comprise from about 80 nucleotides to about 100 nucleotides.

[0092] The plurality of capture probes or capture probe sets may comprise two or more capture probes having the same, similar, and / or different nucleic acid binding site sequences, linkers, and / or labels. For example, two or more capture probes may comprise the same nucleic acid binding site. In another example, two or more capture probes may comprise similar nucleic acid binding sites. In yet another example, two or more capture probes may comprise different nucleic acid binding sites. Two or more capture probes may further comprise one or more linkers. Two or more capture probes may further comprise different linkers. Two or more capture probes may further comprise similar linkers. Two or more capture probes may further comprise the same linker. Two or more capture probes may further comprise one or more labels. Two or more capture probes may further comprise different labels. Two or more capture probes may further comprise similar labels. Two or more capture probes may further comprise the same label.

[0093] B. Assay and Amplification Techniques Some embodiments may include performing one or more assays on a sample containing one or more nucleic acid molecules. Producing two or more subsets of nucleic acid molecules may include performing one or more assays. An assay may be performed on a subset of nucleic acid molecules derived from a sample. An assay may be performed on one or more nucleic acid molecules derived from a sample. An assay may be performed on at least a portion of a subset of nucleic acid molecules. An assay may include one or more techniques, reagents, capture probes, primers, labels, and / or components for detection, quantification, and / or analysis of one or more nucleic acid molecules.

[0094] Assays may include, but are not limited to, sequencing, amplification, hybridization, enrichment, separation, elution, fragmentation, detection, quantification of one or more nucleic acid molecules. An assay may include a method for preparing one or more nucleic acid molecules.

[0095] Some embodiments may include performing one or more amplification reactions on one or more nucleic acid molecules in a sample. The term "amplification" refers to any process that creates at least one copy of a nucleic acid molecule. The terms "amplification product" and "amplified nucleic acid molecule" refer to copies of a nucleic acid molecule and can be used interchangeably. The amplification reaction can include a PCR-based method, a non-PCR-based method, or a combination thereof. Examples of non-PCR-based methods include, but are not limited to, multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, rolling circle amplification, or circle-to-circle amplification. PCR-based methods can include, but are not limited to, PCR, HD-PCR, next-generation PCR (Next Gen PCR), digital RTA, or any combination thereof. Additional PCR methods include, but are not limited to, linear amplification, allele-specific PCR, Alu PCR, assembly PCR, asymmetric PCR, droplet PCR, emulsion PCR, helicase-dependent amplification (HDA), hot-start PCR, inverse PCR, linear-after-the-exponential (LATE)-PCR, long PCR, multiplex PCR, nested PCR, hemi-nested PCR, quantitative PCR, RT-PCR, real-time PCR, single-cell PCR, and touchdown PCR.

[0096] Some embodiments may include performing one or more hybridization reactions on one or more nucleic acid molecules in a sample. The hybridization reaction may include hybridization of one or more capture probes to one or more nucleic acid molecules or a subset of nucleic acid molecules in the sample. The hybridization reaction may include hybridizing a set of one or more capture probes to one or more nucleic acid molecules or a subset of nucleic acid molecules in the sample. The hybridization reaction may include one or more hybridization arrays, multiplex hybridization reactions, hybridization chain reactions, isothermal hybridization reactions, nucleic acid hybridization reactions, or combinations thereof. The one or more hybridization arrays may include hybridization array genotyping, hybridization array proportional sensing, DNA hybridization arrays, macroarrays, microarrays, high-density oligonucleotide arrays, genomic hybridization arrays, comparative hybridization arrays, or combinations thereof. The hybridization reaction may include one or more capture probes, one or more beads, one or more labels, one or more subsets of nucleic acid molecules, one or more nucleic acid samples, one or more reagents, one or more wash buffers, one or more elution buffers, one or more hybridization buffers, one or more hybridization chambers, one or more incubators, one or more separators, or combinations thereof.

[0097] Some embodiments may include performing one or more enrichment reactions on one or more nucleic acid molecules in a sample. The enrichment reaction may include contacting the sample with one or more beads or bead sets. The enrichment reaction may include differential amplification of two or more subsets of nucleic acid molecules based on one or more genomic region features. For example, the enrichment reaction may include differential amplification of two or more subsets of nucleic acid molecules based on GC content. Alternatively, or in addition, the enrichment reaction may include differential amplification of two or more subsets of nucleic acid molecules based on methylation sites. The enrichment reaction may include one or more hybridization reactions. The enrichment reaction may further include separation and / or purification of one or more hybridized nucleic acid molecules, one or more bead-bound nucleic acid molecules, one or more free nucleic acid molecules (e.g., nucleic acid molecules without capture probes, nucleic acid molecules without beads), one or more labeled nucleic acid molecules, one or more unlabeled nucleic acid molecules, one or more amplification products, one or more non-amplified nucleic acid molecules, or combinations thereof. Alternatively, or in addition, the enrichment reaction may include enriching one or more cell types in the sample. One or more cell types may be enriched by flow cytometry.

[0098] One or more enrichment reactions may produce one or more enriched nucleic acid molecules. The enriched nucleic acid molecules may comprise nucleic acid molecules or variants or derivatives thereof. For example, the enriched nucleic acid molecules may comprise one or more hybridized nucleic acid molecules, one or more bead-bound nucleic acid molecules, one or more free nucleic acid molecules (e.g., nucleic acid molecules without capture probes, nucleic acid molecules without beads), one or more labeled nucleic acid molecules, one or more unlabeled nucleic acid molecules, one or more amplification products, one or more non-amplified nucleic acid molecules, or combinations thereof. The enriched nucleic acid molecules may be distinguished from the non-enriched nucleic acids by GC content, molecular size, genomic region, genomic region characteristics, or combinations thereof. The enriched nucleic acid molecules may be derived from one or more assays, supernatants, eluates, or combinations thereof. The enriched nucleic acid molecules may differ from the non-enriched nucleic acid molecules in average size, average GC content, genomic region, or combinations thereof.

[0099] Some embodiments may include performing one or more separation or purification reactions on one or more nucleic acid molecules in a sample. The separation or purification reaction may include contacting the sample with one or more beads or bead sets. The separation or purification reaction may include one or more hybridization reactions, enrichment reactions, amplification reactions, sequencing reactions, or combinations thereof. The separation or purification reaction may include the use of one or more separators. The one or more separators may include a magnetic separator. The separation or purification reaction may include separating bead-bound nucleic acid molecules from nucleic acid molecules without beads. The separation or purification reaction may include separating nucleic acid molecules hybridized with capture probes from nucleic acid molecules without capture probes. The separation or purification reaction may include separating a first subset of nucleic acid molecules from a second subset of nucleic acid molecules, wherein the first subset of nucleic acid molecules differs from the second subset of nucleic acid molecules in average size, average GC content, genomic region, or combinations thereof.

[0100] Some embodiments may include performing one or more elution reactions on one or more nucleic acid molecules in a sample. The elution reaction may include contacting the sample with one or more beads or bead sets. The elution reaction may include separating bead-bound nucleic acid molecules from bead-free nucleic acid molecules. The elution reaction may include separating capture probe-hybridized nucleic acid molecules from capture probe-free nucleic acid molecules. The elution reaction may include separating a first subset of nucleic acid molecules from a second subset of nucleic acid molecules, where the first subset of nucleic acid molecules is different from the second subset of nucleic acid molecules in average size, average GC content, genomic region, or a combination thereof.

[0101] Some embodiments may include one or more fragmentation reactions. The fragmentation reaction may include fragmenting one or more nucleic acid molecules or subsets of nucleic acid molecules in a sample to produce one or more fragmented nucleic acid molecules. The one or more nucleic acid molecules may be fragmented by sonication, needle shearing, spraying, shearing (e.g., acoustic shearing, mechanical shearing, point-sink shearing, etc.), passage through a French press cell, or enzymatic digestion. The enzymatic digestion may be caused by nuclease digestion (e.g., micrococcal nuclease digestion, endonuclease, exonuclease, ribonuclease H (RNAse H) or deoxyribonuclease I (DNase I)). Fragmentation of the one or more nucleic acid molecules may produce fragments having a size of about 100 base pairs to about 2000 base pairs, about 200 base pairs to about 1500 base pairs, about 200 base pairs to about 1000 base pairs, about 200 base pairs to about 500 base pairs, about 500 base pairs to about 1500 base pairs, about 500 base pairs to about 1000 base pairs. The one or more fragmentation reactions may produce fragments having a size of about 50 base pairs to about 1000 base pairs. The one or more fragmentation reactions may produce fragments having a size of about 100 base pairs, 150 base pairs, 200 base pairs, 250 base pairs, 300 base pairs, 350 base pairs, 400 base pairs, 450 base pairs, 500 base pairs, 550 base pairs, 600 base pairs, 650 base pairs, 700 base pairs, 750 base pairs, 800 base pairs, 850 base pairs, 900 base pairs, 950 base pairs, 1000 base pairs, or greater.

[0102] Fragmenting one or more nucleic acid molecules may include mechanically shearing one or more nucleic acid molecules in a sample for a period of time. The fragmentation reaction may occur for at least about 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500 seconds or more.

[0103] Fragmenting one or more nucleic acid molecules may include contacting the nucleic acid sample with one or more beads. Fragmenting one or more nucleic acid molecules may include contacting the nucleic acid sample with a plurality of beads, and the ratio of the volume of the plurality of beads to the volume of the nucleic acid sample is about 0.10, 0.20, 0.30, 0.40, 0.50, 0.60, 0.70, 0.80, 0.90, 1.00, 1.10, 1.20, 1.30, 1.40, 1.50, 1.60, 1.70, 1.80, 1.90, 2.00 or more. Fragmenting one or more nucleic acid molecules may include contacting the nucleic acid sample with a plurality of beads, and the ratio of the volume of the plurality of beads to the volume of the nucleic acid is about 2.00, 1.90, 1.80, 1.70, 1.60, 1.50, 1.40, 1.30, 1.20, 1.10, 1.00, 0.90, 0.80, 0.70, 0.60, 0.50, 0.40, 0.30, 0.20, 0.10, 0.05, 0.04, 0.03, 0.02, 0.01 or less.

[0104] Some embodiments may include performing one or more detection reactions on one or more nucleic acid molecules in a sample. The detection reaction may include one or more sequencing reactions. Alternatively, performing the detection reaction may include optical sensing, electrical sensing, or a combination thereof. Optical sensing may include optical sensing of photoluminescence photon emission, fluorescence photon emission, pyrophosphate photon emission, chemiluminescence photon emission, or a combination thereof. Electrical sensing may include electrical sensing of ion concentration, ion current modulation, nucleotide electric field, nucleotide tunnel current, or a combination thereof.

[0105] Some embodiments may include performing one or more quantification reactions on one or more nucleic acid molecules in a sample. The quantification reaction may include sequencing, PCR, qPCR, digital PCR, or a combination thereof.

[0106] Some embodiments may include one or more samples. Some embodiments may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more samples. The sample may be derived from a subject. Two or more samples may be derived from a single subject. Two or more samples may be derived from 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more different subjects. The subject may be a mammal, reptile, amphibian, bird, and fish. The mammal may be a human, ape, orangutan, monkey, chimpanzee, cow, pig, horse, rodent, bird, reptile, dog, cat, or other animal. The reptile may be a lizard, snake, alligator, turtle, crocodile, and turtle. The amphibian may be a frog, toad, newt, and salamander. Examples of birds include, but are not limited to, ducks, geese, penguins, ostriches, and owls. Examples of fish include, but are not limited to, catfish, eels, sharks, and tuna. Preferably, the subject is a human. The subject may suffer from a disease or condition (e.g., cancer).

[0107] Two or more samples may be collected at more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 locations or more. The time points may occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 hours or more. The time points may occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 days or more. The time points may occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 weeks or more. The time points may occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 months or more. The time points may occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 years or more.

[0108] The sample may be derived from a body fluid, cell, skin, tissue, organ, or a combination thereof. The sample may be blood, plasma, a blood fraction, saliva, sputum, urine, semen, vaginal fluid, cerebrospinal fluid, feces, a cell or tissue biopsy material. The sample may be derived from the adrenal gland, appendix, bladder, brain, ear, esophagus, eye, gallbladder, heart, kidney, large intestine, liver, lung, mouth, muscle, nose, pancreas, parathyroid gland, pineal gland, pituitary gland, skin, small intestine, spleen, stomach, thymus, thyroid gland, trachea, uterus, vermiform appendix, cornea, skin, heart valve, artery, or vein.

[0109] The sample may contain one or more nucleic acid molecules. The nucleic acid molecules may be DNA molecules, RNA molecules (e.g., messenger RNA (mRNA), complementary RNA (cRNA) or microRNA (miRNA)), and DNA / RNA hybrids. Examples of DNA molecules include, but are not limited to, double-stranded DNA, single-stranded DNA, single-stranded DNA hairpins, complementary DNA (cDNA), and genomic DNA. The nucleic acid may be an RNA molecule such as double-stranded RNA, single-stranded RNA, non-coding RNA (ncRNA), RNA hairpins, and messenger RNA (mRNA). Examples of non-coding RNA (ncRNA) include, but are not limited to, small interfering RNA (siRNA), microRNA (miRNA), small nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), transcription initiation RNA (tiRNA), PASR, TASR, aTASR, TSSa-RNA, small nuclear RNA (snRNA), RE-RNA, uaRNA, x-ncRNA, hYRNA, usRNA, snaR, and vtRNA.

[0110] Some embodiments may include one or more containers. Some embodiments may include one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, twenty or more, thirty or more, forty or more, fifty or more, sixty or more, seventy or more, eighty or more, ninety or more, one hundred or more, one hundred twenty-five or more, one hundred fifty or more, one hundred seventy-five or more, two hundred or more, two hundred fifty or more, three hundred or more, three hundred fifty or more, four hundred or more, five hundred or more, six hundred or more, seven hundred or more, eight hundred or more, nine hundred or more, or one thousand or more containers. The one or more containers may be different, similar, identical, or combinations thereof. Examples of containers include, but are not limited to, plates, microplates, PCR plates, wells, microwells, tubes, Eppendorf tubes, vials, arrays, microarrays, and chips.

[0111] Some embodiments may include one or more reagents. Some embodiments may include one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, twenty or more, thirty or more, forty or more, fifty or more, sixty or more, seventy or more, eighty or more, ninety or more, one hundred or more, one hundred and twenty-five or more, one hundred and fifty or more, one hundred and seventy-five or more, two hundred or more, two hundred and fifty or more, three hundred or more, three hundred and fifty or more, four hundred or more, five hundred or more, six hundred or more, seven hundred or more, eight hundred or more, nine hundred or more, or one thousand or more reagents. The one or more reagents may be different, similar, identical, or combinations thereof. The reagents may improve the efficiency of one or more assays. The reagents may enhance the stability of nucleic acid molecules or variants or derivatives thereof. The reagents may include, but are not limited to, enzymes, proteases, nucleases, molecules, polymerases, reverse transcriptases, ligases, and chemicals. Some embodiments may include performing an assay that includes one or more antioxidants. Generally, an antioxidant is a molecule that suppresses the oxidation of another molecule. Examples of antioxidants include, but are not limited to, ascorbic acid (e.g., vitamin C), glutathione, lipoic acid, uric acid, carotenes, α-tocopherol (e.g., vitamin E), ubiquinol (e.g., coenzyme Q), and vitamin A.

[0112] Some embodiments may include one or more buffers or solutions. The one or more buffers or solutions may be different, similar, identical, or combinations thereof. The buffer or solution may improve the efficiency of one or more assays. The buffer or solution may enhance the stability of nucleic acid molecules or variants or derivatives thereof. The buffer or solution may include, but are not limited to, wash buffers, elution buffers, and hybridization buffers.

[0113] Some embodiments may include one or more beads, a plurality of beads, or one or more bead sets. Some embodiments may include one or more bead sets of 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more beads. The one or more beads or bead sets may be different, similar, identical, or combinations thereof. The beads may be magnetic, antibody-coated, protein A-crosslinked, protein G-crosslinked, streptavidin-coated, oligonucleotide-binding, silica-coated, or combinations thereof. Examples of beads include, but are not limited to, Ampure beads, Ampure XP beads, streptavidin beads, agarose beads, magnetic beads, Dynabeads®, MACS® microbeads, antibody-binding beads (e.g., anti-immunoglobulin microbeads), protein A-binding beads, protein G-binding beads, protein A / G-binding beads, protein L-binding beads, oligo dT-binding beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorescent dye microbeads, and BcMag™ carboxy-terminated magnetic beads. In some aspects, the one or more beads include one or more Ampure beads. Alternatively, or in addition, the one or more beads include Ampure XP beads.

[0114] Some embodiments may include one or more primers, a plurality of primers, or one or more primer sets. The primer may further include one or more linkers. The primer may further include one or more labels. The primer may be used in one or more assays. For example, the primer is used in one or more sequencing reactions, amplification reactions, or combinations thereof. Some embodiments may include one or more primers or primer sets of 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more. The primer may include about 100 nucleotides. The primer may include about 10 to about 500 nucleotides, about 20 to about 450 nucleotides, about 30 to about 400 nucleotides, about 40 to about 350 nucleotides, about 50 to about 300 nucleotides, about 60 to about 250 nucleotides, about 70 to about 200 nucleotides, about 80 to about 150 nucleotides. In some aspects, the primer includes from about 80 nucleotides to about 100 nucleotides. One or more primers or primer sets may be different, similar, identical, or combinations thereof.

[0115] The primer may hybridize with at least a part of one or more nucleic acid molecules or variants or derivatives thereof in the sample, or with a subset of nucleic acid molecules. The primer may hybridize with one or more genomic regions. The primer may hybridize with different, similar, and / or identical genomic regions. One or more primers may be at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99% or more complementary to one or more nucleic acid molecules or variants or derivatives thereof.

[0116] The primer may comprise one or more nucleotides. The primer may comprise one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, twenty or more, thirty or more, forty or more, fifty or more, sixty or more, seventy or more, eighty or more, ninety or more, one hundred or more, one hundred and twenty-five or more, one hundred and fifty or more, one hundred and seventy-five or more, two hundred or more, two hundred and fifty or more, three hundred or more, three hundred and fifty or more, four hundred or more, five hundred or more, six hundred or more, seven hundred or more, eight hundred or more, nine hundred or more, or one thousand or more nucleotides. The primer may comprise about 100 nucleotides. The primer may comprise about 10 to about 500 nucleotides, about 20 to about 450 nucleotides, about 30 to about 400 nucleotides, about 40 to about 350 nucleotides, about 50 to about 300 nucleotides, about 60 to about 250 nucleotides, about 70 to about 200 nucleotides, or about 80 to about 150 nucleotides. In some embodiments, the primer comprises from about 80 nucleotides to about 100 nucleotides.

[0117] A plurality of primers or primer sets may include two or more primers having the same, similar, and / or different sequences, linkers, and / or labels. For example, two or more primers may include the same sequence. In another example, two or more primers may include similar sequences. In yet another example, two or more primers may include different sequences. Two or more primers may further include one or more linkers. Two or more primers may further include different linkers. Two or more primers may further include similar linkers. Two or more primers may further include the same linker. Two or more primers may further include one or more labels. Two or more primers may further include different labels. Two or more primers may further include similar labels. Two or more primers may further include the same label.

[0118] Capture probes, primers, labels, and / or beads may include one or more nucleotides. The one or more nucleotides may include RNA, DNA, a mixture of DNA and RNA residues, or modified analogs thereof such as 2'-O-Me or 2'-fluoro (2'-F), locked nucleic acid (LNA), or abasic sites.

[0119] Some embodiments may include one or more labels. Some embodiments may include one or more labels of 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more. The one or more labels may be different, similar, identical, or combinations thereof.

[0120] Examples of labels include, but are not limited to, chemical, biochemical, biological, chromogenic, enzymatic, fluorescent, and luminescent labels well-known in the art. Labels include dyes, photocrosslinkers, cytotoxic compounds, drugs, affinity labels, photoaffinity labels, reactive compounds, antibodies or antibody fragments, biomaterials, nanoparticles, spin labels, fluorophores, metal-containing moieties, radioactive moieties, novel functional groups, groups that interact covalently or noncovalently with other molecules, photocaged moieties, radiation-excitable moieties, ligands, photoisomerizable moieties, biotin, biotin analogs, moieties incorporating heavy atoms, chemically cleavable groups, photocleavable groups, redox-active agents, isotope-labeled moieties, biophysical probes, phosphorescent groups, chemiluminescent groups, high electron density groups, magnetic groups, intercalating groups, chromophores, energy transfer agents, bioactive agents, detectable labels, or combinations thereof.

[0121] The label may be a chemical label. Examples of chemical labels can include, but are not limited to, biotin and radioisotopes (e.g., iodine, carbon, phosphate, hydrogen).

[0122] The methods, kits, and compositions disclosed herein may include a biological label. Biological labels may include, but are not limited to, metabolic labels including bioorthogonal azide-modified amino acids, saccharides, and other compounds.

[0123] The methods, kits, and compositions disclosed herein may include an enzyme label. Enzyme labels can include, but are not limited to, horseradish peroxidase (HRP), alkaline phosphatase (AP), glucose oxidase, and β-galactosidase. The enzyme label may be luciferase.

[0124] The methods, kits, and compositions disclosed herein may include a fluorescent label. The fluorescent label may include an organic dye (e.g., FITC), a biological fluorophore (e.g., green fluorescent protein), or a quantum dot. A non-limiting list of fluorescent labels includes fluorescein isothiocyanate (FITC), DyLight Fluor, fluorescein, rhodamine (tetramethylrhodamine isothiocyanate, TRITC), coumarin, lucifer yellow, and BODIPY. The label may be a fluorophore. Exemplary fluorophores include, but are not limited to, indocarbocyanine (C3), indodicarbocyanine (C5), Cy3, Cy3.5, Cy5, Cy5.5, Cy7, Texas Red, Pacific Blue, Oregon Green 488, Alexa Fluor®-355, Alexa Fluor 488, Alexa Fluor 532, Alexa Fluor 546, Alexa Fluor-555, Alexa Fluor 568, Alexa Fluor 594, Alexa Fluor 647, Alexa Fluor 660, Alexa Fluor 680, JOE, Lissamine, Rhodamine Green, BODIPY, fluorescein isothiocyanate (FITC), carboxyfluorescein (FAM), phycoerythrin, rhodamine, dichlororhodamine (dRhodamine), carboxytetramethylrhodamine (TAMRA), carboxy-X-rhodamine (ROX™), LIZ™, VIC™ NED™ PET™, SYBR, PicoGreen, RiboGreen, and the like. The fluorescent label may be a green fluorescent protein (GFP), a red fluorescent protein (RFP), a yellow fluorescent protein, a phycobiliprotein (e.g., allophycocyanin, phycocyanin, phycoerythrin, and phycoerythrocyanin).

[0125] Some embodiments may include one or more linkers. Some embodiments may include one or more than one, more than one, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, twenty or more, thirty or more, forty or more, fifty or more, sixty or more, seventy or more, eighty or more, ninety or more, one hundred or more, one hundred twenty-five or more, one hundred fifty or more, one hundred seventy-five or more, two hundred or more, two hundred fifty or more, three hundred or more, three hundred fifty or more, four hundred or more, five hundred or more, six hundred or more, seven hundred or more, eight hundred or more, nine hundred or more, or one thousand or more linkers. The one or more linkers may be different, similar, identical, or combinations thereof.

[0126] Suitable linkers include any chemical or biological compound that can attach to a label, primer, and / or capture probe disclosed herein. If the linker attaches to both the label and the primer or capture probe, a suitable linker will be able to sufficiently separate the label from the primer or capture probe. A suitable linker will not significantly interfere with the ability of the primer and / or capture probe to hybridize to a nucleic acid molecule, a portion thereof, or a variant or derivative thereof. A suitable linker will not significantly interfere with the ability to detect the label. The linker may be rigid. The linker may be flexible. The linker may be semi-rigid. The linker may be stable to proteolysis (e.g., resistant to protein cleavage). The linker may be unstable to proteolysis (e.g., sensitive to protein cleavage). The linker may be helical. The linker may be non-helical. The linker may be coiled. The linker may be three-stranded. The linker may include a rotational conformation. The linker may be single-stranded. The linker may be long-chain. The linker may be short-chain. The linker may be at least about 5 residues, at least about 10 residues, at least about 15 residues, at least about 20 residues, at least about 25 residues, at least about 30 residues, at least about 40 residues, or more.

[0127] Examples of linkers include, but are not limited to, hydrazones, disulfides, thioethers, and peptide linkers. The linker may be a peptide linker. The peptide linker may contain a proline residue. The peptide linker may contain arginine, phenylalanine, threonine, glutamine, glutamic acid, or any combination thereof. The linker may be a heterobifunctional crosslinker.

[0128] Some embodiments may include performing one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, eleven or more, twelve or more, thirteen or more, fourteen or more, fifteen or more, twenty or more, twenty-five or more, thirty or more, thirty-five or more, forty or more, forty-five or more, or fifty or more assays on a sample containing one or more nucleic acid molecules. The two or more assays may be different, similar, identical, or combinations thereof. For example, some embodiments include performing two or more sequencing reactions. In another example, some embodiments include performing two or more assays, wherein at least one of the two or more assays includes a sequencing reaction. In yet another example, some embodiments include performing two or more assays, wherein at least two of the two or more assays include a sequencing reaction and a hybridization reaction. The two or more assays may be performed over time, simultaneously, or combinations thereof. For example, two or more sequencing reactions may be performed simultaneously. In another example, some embodiments include performing a hybridization reaction and then a sequencing reaction. In yet another example, some embodiments include performing two or more hybridization reactions simultaneously and then performing two or more sequencing reactions simultaneously. The two or more assays may be performed by one or more devices. For example, two or more amplification reactions may be performed by a PCR device. In another example, two or more sequencing reactions may be performed by two or more sequencers.

[0129] C. Apparatus Some embodiments may include one or more devices. Some embodiments may include one or more assays that include one or more devices. Some embodiments may include the use of one or more devices to perform one or more steps or assays. Some embodiments may include the use of one or more devices in one or more steps or assays. For example, performing a sequencing reaction may include one or more sequencers. In another example, creating a subset of nucleic acid molecules may include the use of one or more magnetic separators. In yet another example, one or more processors may be used in the analysis of one or more nucleic acid samples. Exemplary devices include, but are not limited to, sequencers, thermocyclers, real-time PCR instruments, magnetic separators, transmission devices, hybridization chambers, electrophoresis apparatuses, centrifuges, microscopes, imaging devices, fluorometers, luminometers, plate readers, computers, processors, and bioanalyzers.

[0130] Some embodiments may include one or more sequencers. The one or more sequencers may include one or more HiSeq, MiSeq, HiScan, Genome Analyzer IIx, SOLiD sequencers, Ion Torrent PGM, 454 GS Junior, Pac Bio RS, or combinations thereof. The one or more sequencers may include one or more sequencing platforms. The one or more sequencing platforms may include the GS FLX by 454 Life Technologies / Roche, the Genome Analyzer by Solexa / Illumina, the SOLiD by Applied Biosystems, the CGA Platform by Complete Genomics, the PacBio RS by Pacific Biosciences, or combinations thereof.

[0131] Some embodiments may include one or more thermocyclers. The one or more thermocyclers may be used to amplify one or more nucleic acid molecules. Some embodiments may include one or more real-time PCR instruments. The one or more real-time PCR instruments may include a thermal cycler and a fluorometer. The one or more thermocyclers may be used to amplify and detect one or more nucleic acid molecules.

[0132] Some embodiments may include one or more magnetic separators. The one or more magnetic separators may be used to separate paramagnetic and ferromagnetic particles from a suspension. The one or more magnetic separators may include one or more LifeStepTM Biomagnetic Separators, SPHERO™ FlexiMag Separators, SPHERO™ MicroMag Separators, SPHERO™ HandiMag Separators, SPHERO™ MiniTube Mag Separators, SPHERO™ UltraMag Separators, DynaMag™ Magnets, DynaMag™ 2 Magnets, or combinations thereof.

[0133] Some embodiments may include one or more bioanalyzers. Generally, a bioanalyzer is a capillary electrophoresis apparatus using a microchip capable of analyzing RNA, DNA, and proteins. The one or more bioanalyzers may include the Agilent 2100 Bioanalyzer.

[0134] Some embodiments may include one or more processors. The one or more processors may analyze, collect, store, sort, combine, evaluate, or otherwise process one or more data and / or results of one or more assays, one or more data and / or results based on or derived from one or more assays, one or more outputs of one or more assays, one or more outputs based on or derived from one or more assays, one or more outputs from one or more data and / or results, one or more outputs based on or derived from one or more data and / or results, or combinations thereof. The one or more processors may transmit one or more data, results, or outputs of one or more assays, one or more data, results, or outputs based on or derived from one or more assays, one or more outputs of one or more data or results, one or more outputs based on or derived from one or more data or results, or combinations thereof. The one or more processors may receive and / or store requests from a user. The one or more processors may generate or create one or more data, results, outputs. The one or more processors may generate or create one or more biomedical reports. The one or more processors may transmit one or more biomedical reports. The one or more processors may analyze, collect, store, sort, combine, evaluate, or otherwise process information of one or more databases, one or more data or results, one or more outputs, or combinations thereof. The one or more processors may analyze, collect, store, sort, combine, evaluate, or otherwise process information of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30 or more databases. The one or more processors may transmit one or more requests, data, results, outputs, and / or information to one or more users, processors, computers, computer systems, memory locations, devices, databases, or combinations thereof.One or more processors may receive one or more requests, data, results, outputs, and / or information from one or more users, processors, computers, computer systems, memory locations, devices, databases, or combinations thereof. One or more processors may retrieve one or more requests, data, results, outputs, and / or information from one or more users, processors, computers, computer systems, memory locations, devices, databases, or combinations thereof.

[0135] Some embodiments may include one or more memory locations. The one or more memory locations may store information, data, results, outputs, requests, or combinations thereof. The one or more memory locations may receive information, data, results, outputs, requests, or combinations thereof from one or more users, processors, computers, computer systems, devices, or combinations thereof.

[0136] The methods described herein can be performed with the aid of one or more computers and / or computer systems. The computer or computer system may include an electronic storage location (e.g., a database, a memory) having machine-executable code for performing the methods provided herein, and one or more processors for executing the machine-executable code.

[0137] The code can be configured to be used with a machine having a processor adapted to pre-compile and execute the code, or can be compiled during run-time. The code can be supplied in a programming language that can be selected to be executed in a pre-compiled or compiled manner.

[0138] One or more computers and / or computer systems may analyze, collect, store, sort, combine, evaluate, or otherwise process one or more data and / or results of one or more assays, one or more data and / or results based on or derived from one or more assays, one or more outputs of one or more assays, one or more outputs based on or derived from one or more assays, one or more outputs from one or more data and / or results, one or more outputs based on or derived from one or more data and / or results, or combinations thereof. One or more computers and / or computer systems may transmit one or more data, results, or outputs of one or more assays, one or more data, results, or outputs based on or derived from one or more assays, one or more outputs of one or more data or results, one or more outputs based on or derived from one or more data or results, or combinations thereof. One or more computers and / or computer systems may receive and / or store requests from users. One or more computers and / or computer systems may generate or create one or more data, results, outputs. One or more computers and / or computer systems may generate or create one or more biomedical reports. One or more computers and / or computer systems may transmit one or more biomedical reports. One or more computers and / or computer systems may analyze, collect, store, sort, combine, evaluate, or otherwise process information of one or more databases, one or more data or results, one or more outputs, or combinations thereof. One or more computers and / or computer systems may analyze, collect, store, sort, combine, evaluate, or otherwise process information of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30 or more databases.One or more computers and / or computer systems may transmit one or more requests, data, results, outputs, and / or information to one or more users, processors, computers, computer systems, memory locations, devices, or combinations thereof. One or more computers and / or computer systems may receive one or more requests, data, results, outputs, and / or information from one or more users, processors, computers, computer systems, memory locations, devices, or combinations thereof. One or more computers and / or computer systems may retrieve one or more requests, data, results, outputs, and / or information from one or more users, processors, computers, computer systems, memory locations, devices, databases, or combinations thereof.

[0139] D. Database Some embodiments may include one or more databases. Some embodiments may include at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, or more databases. The databases may include genomic, proteomic, pharmacogenomic, biomedical, and scientific databases. The databases may be publicly available databases. Alternatively, or in addition, the databases may be proprietary databases. The databases may be commercial databases. Examples of databases include, but are not limited to, Cosmic, GnomAD, Dbsnp, Mills Indels, MendelDB, PharmGKB, Varimed, Regulome, curated BreakSeq junctions, Online Mendelian Inheritance in Man (OMIM), Human Gene Mutation Database (HGMD), NCBI db SNP, NCBI RefSeq, GENCODE, GO (Gene Ontology), and Kyoto Encyclopedia of Genes and Genomes (KEGG).

[0140] Some embodiments may include analyzing one or more databases. Some embodiments may include analyzing at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30 or more databases. Analyzing one or more databases may include one or more algorithms, computers, processors, memory locations, devices, or combinations thereof.

[0141] Some embodiments may include identifying one or more nucleic acid regions based on data and / or information from one or more databases. Some embodiments may include identifying one or more sets of nucleic acid regions based on data and / or information from one or more databases. Some embodiments may include identifying one or more nucleic acid regions and / or sets of nucleic acid regions based on data and / or information from at least two or more databases. Some embodiments may include identifying one or more nucleic acid regions and / or sets of nucleic acid regions based on data and / or information from at least three or more databases. Some embodiments may include identifying one or more nucleic acid regions and / or sets of nucleic acid regions based on data and / or information from at least about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30 or more databases.

[0142] Some embodiments may include analyzing one or more results based on data and / or information from one or more databases. Some embodiments may include analyzing one or more sets of results based on data and / or information from one or more databases. Some embodiments may include analyzing one or more totaled results based on data and / or information from one or more databases. Some embodiments may include analyzing one or more results, sets of results, and / or totaled results based on data and / or information from at least about two or more databases. Some embodiments may include analyzing one or more results, sets of results, and / or totaled results based on data and / or information from at least about three or more databases. Some embodiments may include analyzing one or more results, sets of results, and / or totaled results based on data and / or information from at least about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30 or more databases.

[0143] Some embodiments may include comparing one or more results based on data and / or information from one or more databases. Some embodiments may include comparing one or more sets of results based on data and / or information from one or more databases. Some embodiments may include comparing one or more totaled results based on data and / or information from one or more databases. Some embodiments may include comparing one or more results, sets of results, and / or totaled results based on at least about two or more database data and / or information. Some embodiments may include comparing one or more results, sets of results, and / or totaled results based on data and / or information from at least about three or more database data. Some embodiments may include comparing one or more results, sets of results, and / or totaled results based on data and / or information from at least about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30 or more databases.

[0144] Some embodiments may include biomedical databases, genomic databases, biomedical reports, disease reports, case-control analyses, and rare variant discovery analyses based on data and / or information from one or more databases, one or more assays, one or more data or results, one or more outputs based on or derived from one or more assays, one or more outputs based on or derived from one or more data or results, or combinations thereof.

[0145] E. Analysis Some embodiments may include one or more data, one or more data sets, one or more aggregated data, one or more aggregated data sets, one or more results, one or more sets of results, one or more aggregated results, or combinations thereof. The data and / or results may be based on one or more assays, one or more databases, or combinations thereof. Some embodiments may include the analysis of one or more data, one or more data sets, one or more aggregated data, one or more aggregated data sets, one or more results, one or more sets of results, one or more aggregated results, or combinations thereof. Some embodiments may include the processing of one or more data, one or more data sets, one or more aggregated data, one or more aggregated data sets, one or more results, one or more sets of results, one or more aggregated results, or combinations thereof.

[0146] Some embodiments may include at least one analysis and at least one processing of one or more data, one or more data sets, one or more aggregated data, one or more aggregated data sets, one or more results, one or more sets of results, one or more aggregated results, or combinations thereof. Some embodiments may include one or more analyses and one or more processings of one or more data, one or more data sets, one or more aggregated data, one or more aggregated data sets, one or more results, one or more sets of results, one or more aggregated results, or combinations thereof. Some embodiments may include at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 or more different analyses of one or more data, one or more data sets, one or more aggregated data, one or more aggregated data sets, one or more results, one or more sets of results, one or more aggregated results, or combinations thereof. Some embodiments may include at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 or more different processings of one or more data, one or more data sets, one or more aggregated data, one or more aggregated data sets, one or more results, one or more sets of results, one or more aggregated results, or combinations thereof. The one or more analyses and / or the one or more processings may occur simultaneously, over time, or in combination thereof.

[0147] One or more analyses and / or one or more processes may occur at a point in time exceeding 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000. The point in time may occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 hours or more. The point in time may occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 days or more. The point in time may occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 weeks or more. The point in time may occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 months or more. The point in time may occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 years or more.

[0148] Some embodiments may include one or more data. The one or more data may include one or more raw data based on or derived from one or more assays. The one or more data may include one or more raw data based on or derived from one or more databases. The one or more data may include at least partially analyzed data based on or derived from one or more raw data. The one or more data may include at least partially processed data based on or derived from one or more raw data. The one or more data may include fully analyzed data based on or derived from one or more raw data. The one or more data may include fully processed data based on or derived from one or more raw data. The data may include sequencing read data or expression data. The data may include biomedical, scientific, pharmacological, and / or genetic information.

[0149] Some embodiments may include one or more aggregated data. The one or more aggregated data may include two or more data. The one or more aggregated data may include two or more data sets. The one or more aggregated data may include one or more raw data based on or derived from one or more assays. The one or more aggregated data may include one or more raw data based on or derived from one or more databases. The one or more aggregated data may include at least partially analyzed data based on or derived from one or more raw data. The one or more aggregated data may include at least partially processed data based on or derived from one or more raw data. The one or more aggregated data may include fully analyzed data based on or derived from one or more raw data. The one or more aggregated data may include fully processed data based on or derived from one or more raw data. The one or more aggregated data may include sequencing read data or expression data. The one or more aggregated data may include biomedical, scientific, pharmacological, and / or genetic information.

[0150] Some embodiments may include one or more data sets. The one or more data sets may include one or more data. The one or more data sets may include one or more aggregated data. The one or more data sets may include one or more raw data based on or derived from one or more assays. The one or more data sets may include one or more raw data based on or derived from one or more databases. The one or more data sets may include at least partially analyzed data based on or derived from one or more raw data. The one or more data sets may include at least partially processed data based on or derived from one or more raw data. The one or more data sets may include fully analyzed data based on or derived from one or more raw data. The one or more data sets may include fully processed data based on or derived from one or more raw data. The data set may include sequencing read data or expression data. The data set may include biomedical, scientific, pharmacological, and / or genetic information.

[0151] Some embodiments may include one or more aggregated data sets. The one or more aggregated data sets may include two or more data. The one or more aggregated data sets may include two or more aggregated data. The one or more aggregated data sets may include two or more data sets. The one or more aggregated data sets may include one or more raw data based on or derived from one or more assays. The one or more aggregated data sets may include one or more raw data based on or derived from one or more databases. The one or more aggregated data sets may include at least partially analyzed data based on or derived from one or more raw data. The one or more aggregated data sets may include at least partially processed data based on or derived from one or more raw data. The one or more aggregated data sets may include fully analyzed data based on or derived from one or more raw data. The one or more aggregated data sets may include fully processed data based on or derived from one or more raw data. Some embodiments may further include further processing and / or analysis of the aggregated data sets. The one or more aggregated data sets may include sequencing read data or expression data. The one or more aggregated data sets may include biomedical, scientific, pharmacological, and / or genetic information.

[0152] Some embodiments may include one or more results. The one or more results may include one or more data, data sets, aggregated data, and / or aggregated data sets. The one or more results may be based on or derived from one or more data, data sets, aggregated data, and / or aggregated data sets. The one or more results may be generated from one or more assays. The one or more results may be based on or derived from one or more assays. The one or more results may be based on or derived from one or more databases. The one or more results may include at least partially analyzed results based on or derived from one or more data, data sets, aggregated data, and / or aggregated data sets. The one or more results may include at least partially processed results based on or derived from one or more data, data sets, aggregated data, and / or aggregated data sets. The one or more results may include fully analyzed results based on or derived from one or more data, data sets, aggregated data, and / or aggregated data sets. The one or more results may include fully processed results based on or derived from one or more data, data sets, aggregated data, and / or aggregated data sets. The results may include sequencing read data or expression data. The results may include biomedical, scientific, pharmacological, and / or genetic information.

[0153] Some embodiments may include one or more sets of results. The one or more sets of results may include one or more data, data sets, aggregated data, and / or aggregated data sets. The one or more sets of results may be based on or derived from one or more data, data sets, aggregated data, and / or aggregated data sets. The one or more sets of results may be generated from one or more assays. The one or more sets of results may be based on or derived from one or more assays. The one or more sets of results may be based on or derived from one or more databases. The one or more sets of results may include at least partially analyzed set results based on or derived from one or more data, data sets, aggregated data, and / or aggregated data sets. The one or more sets of results may include at least partially processed set results based on or derived from one or more data, data sets, aggregated data, and / or aggregated data sets. The one or more sets of results may include fully analyzed set results based on or derived from one or more data, data sets, aggregated data, and / or aggregated data sets. The one or more sets of results may include fully processed set results based on or derived from one or more data, data sets, aggregated data, and / or aggregated data sets. The set results may include sequencing read data or expression data. The set results may include biomedical, scientific, pharmacological, and / or genetic information.

[0154] Some embodiments may include one or more aggregated results. The aggregated results may include one or more results, paired results, and / or aggregated paired results. The aggregated results may be based on or derived from one or more results, paired results, and / or aggregated paired results. One or more aggregated results may include one or more data, data sets, aggregated data, and / or aggregated data sets. One or more aggregated results may be based on or derived from one or more data, data sets, aggregated data, and / or aggregated data sets. One or more aggregated results may be generated from one or more assays. One or more aggregated results may be based on or derived from one or more assays. One or more aggregated results may be based on or derived from one or more databases. One or more aggregated results may include at least partially analyzed aggregated results based on or derived from one or more data, data sets, aggregated data, and / or aggregated data sets. One or more aggregated results may include at least partially processed aggregated results based on or derived from one or more data, data sets, aggregated data, and / or aggregated data sets. One or more aggregated results may include fully analyzed aggregated results based on or derived from one or more data, data sets, aggregated data, and / or aggregated data sets. One or more aggregated results may include fully processed aggregated results based on or derived from one or more data, data sets, aggregated data, and / or aggregated data sets. The aggregated results may include sequencing read data or expression data. The aggregated results may include biomedical, scientific, pharmacological, and / or genetic information.

[0155] Some embodiments may include aggregated set results. The aggregated set results may include one or more results, set results, and / or aggregated set results. The aggregated set results may be based on or derived from one or more results, set results, and / or aggregated set results. One or more aggregated set results may include one or more data, data sets, aggregated data, and / or aggregated data sets. One or more aggregated set results may be based on or derived from one or more data, data sets, aggregated data, and / or aggregated data sets. One or more aggregated set results may be generated from one or more assays. One or more aggregated set results may be based on or derived from one or more assays. One or more aggregated set results may be based on or derived from one or more databases. One or more aggregated set results may include aggregated set results that have been at least partially analyzed based on or derived from one or more data, data sets, aggregated data, and / or aggregated data sets. One or more aggregated set results may include aggregated set results that have been at least partially processed based on or derived from one or more data, data sets, aggregated data, and / or aggregated data sets. One or more aggregated set results may include fully analyzed aggregated set results based on or derived from one or more data, data sets, aggregated data, and / or aggregated data sets. One or more aggregated set results may include fully processed aggregated set results based on or derived from one or more data, data sets, aggregated data, and / or aggregated data sets. The aggregated set results may include sequencing read data or expression data. The aggregated set results may include biomedical, scientific, pharmacological, and / or genetic information.

[0156] Some embodiments may include one or more outputs, a set of outputs, a combined output, and / or a combined set of outputs. The methods, libraries, kits, and systems described herein may include creating one or more outputs, a set of outputs, a combined output, and / or a combined set of outputs. A set of outputs may include one or more outputs, one or more combined outputs, or a combination thereof. A combined output may include one or more outputs, one or more sets of outputs, one or more combined sets of outputs, and / or a combination thereof. A combined set of outputs may include one or more outputs, one or more sets of outputs, one or more combined outputs, or a combination thereof. One or more outputs, a set of outputs, a combined output, and / or a combined set of outputs may be based on or derived from one or more data, one or more data sets, one or more combined data, one or more combined data sets, one or more results, one or more sets of results, one or more combined results, or a combination thereof. One or more outputs, a set of outputs, a combined output, and / or a combined set of outputs may be based on or derived from one or more databases. One or more outputs, a set of outputs, a combined output, and / or a combined set of outputs may include one or more biomedical reports, biomedical outputs, rare variant outputs, pharmacogenetic outputs, population study outputs, case-control outputs, biomedical databases, genomic databases, disease databases, and net content.

[0157] Some embodiments may include one or more biomedical outputs, one or more sets of biomedical outputs, one or more aggregated biomedical outputs, one or more aggregated sets of biomedical outputs. The methods, libraries, kits, and systems herein may include creating one or more biomedical outputs, one or more sets of biomedical outputs, one or more aggregated biomedical outputs, one or more aggregated sets of biomedical outputs. A set of biomedical outputs may include one or more biomedical outputs, one or more aggregated biomedical outputs, or combinations thereof. An aggregated biomedical output may include one or more biomedical outputs, one or more sets of biomedical outputs, one or more aggregated sets of biomedical outputs, or combinations thereof. An aggregated set of biomedical outputs may include one or more biomedical outputs, one or more sets of biomedical outputs, one or more aggregated biomedical outputs, or combinations thereof. One or more biomedical outputs, one or more sets of biomedical outputs, one or more aggregated biomedical outputs, one or more aggregated sets of biomedical outputs may be based on or derived from one or more data, one or more datasets, one or more aggregated data, one or more aggregated datasets, one or more results, one or more sets of results, one or more aggregated results, one or more outputs, one or more sets of outputs, one or more aggregated outputs, one or more aggregated sets of outputs, or combinations thereof. One or more biomedical outputs may include biomedical information of a subject. The biomedical output of a subject may predict, diagnose, and / or foreknow one or more biomedical characteristics. One or more biomedical characteristics may include symptoms of a disease or condition, genetic risk of a disease or condition, reproductive risk, genetic risk to a fetus, risk of drug adverse reaction, effectiveness of drug therapy, prediction of optimal drug dosage, transplant immune tolerance, or combinations thereof.

[0158] Some embodiments may include one or more biomedical reports. The methods, libraries, kits, and systems herein may include creating one or more biomedical reports. The one or more biomedical reports may be based on or derived from one or more data, one or more data sets, one or more aggregated data, one or more aggregated data sets, one or more results, one or more sets of results, one or more aggregated results, one or more outputs, one or more sets of outputs, one or more aggregated outputs, one or more sets of aggregated outputs, one or more biomedical outputs, one or more sets of biomedical outputs, aggregated biomedical outputs, one or more sets of biomedical outputs, or combinations thereof. The biomedical report may predict, diagnose, and / or prognose one or more biomedical characteristics. The one or more biomedical characteristics may include symptoms of a disease or condition, genetic risk of a disease or condition, reproductive risk, genetic risk to a fetus, risk of drug adverse reaction, effectiveness of drug therapy, prediction of optimal drug dosage, transplant tolerance, or combinations thereof.

[0159] Some embodiments may also include transmitting one or more data, information, results, outputs, reports, or combinations thereof. For example, transmitting data / information based on or derived from one or more assays to another device and / or instrument. In another example, transmitting data, results, outputs, biomedical outputs, biomedical reports, or combinations thereof to another device and / or instrument. Information obtained from an algorithm may also be transmitted to another device and / or instrument. Information based on the analysis of one or more databases may be transmitted to another device and / or instrument. Transmission of data / information may include transmitting the data / information from a first source to a second source. The first and second sources may be at the same approximate location (e.g., in the same room, building, block, campus, etc.). Alternatively, the first and second sources may be at multiple locations (e.g., in multiple cities, states, countries, continents, etc.). Data, results, outputs, biomedical outputs, biomedical reports can be transmitted to a patient and / or a healthcare provider.

[0160] The transmission may be based on the analysis of one or more data, results, information, databases, outputs, reports, or combinations thereof. For example, the transmission of a second report is based on the analysis of a first report. Alternatively, the transmission of a report is based on the analysis of one or more data or results. The transmission may be based on the receipt of one or more requests. For example, the transmission of a report may be based on receiving a request from a user (such as a patient, healthcare provider, individual, etc.).

[0161] The transmission of data / information may include digital transmission or analog transmission. Digital transmission may include the physical transmission of data (digital bitstream) through point-to-point or point-to-multipoint communication lines. Examples of such lines are copper wires, optical fibers, wireless communication lines, and recording media. The data may be written as an electromagnetic signal such as a voltage, radio wave, microwave, or infrared signal.

[0162] Analog transmission may include the transmission of a continuously variable analog signal. A message can be represented either by a series of pulses by line coding (baseband transmission) using digital modulation methods or by a finite set of continuously variable waveforms (band transmission). Band modulation and the corresponding demodulation (also known as detection) can be performed by a modem device. According to the most common definition of a digital signal, both a baseband signal representing a bitstream and a band signal are considered digital transmissions, while in another definition, only the baseband signal is considered digital, and the band transmission of digital data is considered a form of digital-to-analog conversion.

[0163] Some embodiments may include one or more sample identifiers. The sample identifier may include a label, barcode, and other indicators that can be associated with one or more samples and / or subsets of nucleic acid molecules. Some embodiments may include one or more processors, one or more memory locations, one or more computers, one or more monitors, one or more computer softwares, one or more algorithms for associating data, results for samples, outputs, biomedical outputs, and / or biomedical reports.

[0164] Some embodiments may include a processor that correlates the expression level of one or more nucleic acid molecules with the prognosis of disease progression. Some embodiments may include one or more various correlation techniques, including a lookup table, algorithm, multivariate model, and linear or non-linear combinations of expression models or algorithms. The expression level may be converted into one or more likelihood scores to indicate the likelihood that a patient from whom the sample is provided exhibits a specific disease progression. The method and / or algorithm can be provided in a machine-readable format, and the method and / or algorithm can further optionally direct a treatment modality for the patient or patient classification.

[0165] In some cases, the methods and systems described herein are used to generate an output that includes the detection and / or quantification of genomic DNA regions, such as regions containing DNA polymorphisms (e.g., germline variants or somatic variants). In some cases, the detection of one or more genomic regions is based on one or more algorithms, depending on the data input or database source described elsewhere herein. Each of the one or more algorithms can be used to receive, combine, and generate data including the detection of genomic regions (i.e., genetic polymorphisms). In some embodiments, the method and system can include the detection of genomic regions based on one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more algorithms. The algorithms can be machine learning algorithms, computer-executable algorithms, machine-executable algorithms, automated algorithms, and the like.

[0166] The data generated for each nucleic acid sample can be analyzed using feature selection techniques that include filter techniques that evaluate the relevance of features by examining the intrinsic properties of the data, wrapper methods that embed a model hypothesis within a feature subset search, and embedded techniques where the search for the optimal set of features is incorporated into an algorithm or model.

[0167] In some cases, the detection of one or more genomic regions is based on one or more statistical models. Statistical models or filtering techniques useful in the methods of the present invention include (1) parametric methods such as the t-test of two samples, ANOVA analysis, Bayesian framework, and use of the gamma distribution model, (2) non-parametric methods such as the Wilcoxon rank sum test, between-class within-class sum of squares test, rank product method, random permutation method, or TNoM which includes setting a threshold point for the fold change of the difference in expression between two datasets and then detecting the threshold point of each gene that minimizes the number of misclassifications, and (3) multivariate methods such as bivariate methods, correlation-based feature selection method (CFS), minimum redundancy maximum relevance method (MRMR), Markov blanket filter method, Markov model, hidden Markov model (HMM), and uncorrelated shrinkage centroid method. In some cases, the hidden Markov model (HMM) is given internal states, and the internal states are set according to the total copy number of chromosomes in the first or second nucleic acid sample. In one example, for a chromosomal diploid, the internal states of the HMM can be homozygous deletion (locally zero copies), heterozygous deletion (locally one copy), normal (locally two copies), duplication (two or more copies), and reference gap (existing as a state to distinguish a gap from a homozygous deletion). In another example, for a haploid chromosome (e.g., male X or Y), the internal states of the HMM can be homozygous deletion (locally zero copies), normal (locally two copies), duplication (two or more copies), and reference gap (existing as a state to distinguish a gap from a homozygous deletion). For example, for a haploid chromosome, a heterozygous deletion state may not be available. In another example, for trisomy and / or tetrasomy, additional intermediate HMM states may have additional intermediate states, and the intermediate states can account for various CNV possibilities. In another embodiment, the hidden Markov model is used to filter the output by testing the measured insertion size of reads proximal to the limit points of the detected features.

[0168] Other models or algorithms useful in the method of the present invention include sequential search methods, genetic algorithms, distribution estimation algorithms, random forest algorithms, support vector machine weight vector algorithms, logistic regression weight algorithms, and the like. Bioinformatics. 2007 Oct 1; 23(19):2507-17 outlines the advantages of the above algorithms or models used in data analysis. Although not limited to exemplary algorithms, methods for directly handling a large number of variables such as principal component analysis algorithms, partial least squares methods, independent component analysis algorithms, statistical models, and methods based on machine learning techniques, and methods for reducing the number of variables are included. Statistical models include penalized logistic regression, microarray prediction analysis (PAM), methods based on shrunken centroids, support vector machine analysis, and regularized linear discriminant analysis. Machine learning includes fully connected neural networks, convolutional neural networks, 1D convolutional neural networks, 2D convolutional neural networks, gradient boosting decision trees (e.g., XGBoost framework, LightGBM framework), bagging methods, boosting methods, random forest algorithms, and combinations thereof. Cancer Inform. 2008; 6: 77-97 outlines the above techniques used in data analysis. In some embodiments, the trained machine learning model includes a gradient boosting decision tree (including, for example, the LightGBM framework). In some embodiments, the trained machine learning model includes a convolutional neural network (e.g., a one-dimensional convolutional neural network or a two-dimensional convolutional neural network). In some embodiments, the trained machine learning model includes a fully connected neural network.

[0169] Machine learning can include deep learning. Deep learning can be used to increasingly capture the internal structure of large, high-dimensional datasets (e.g., nucleic acid sequencing data). Deep models can often enable the discovery of high-level features, have improved performance compared to conventional models, increased interpretability, and provide further understanding of the structure of biological data.

[0170] A trained machine learning model can include a fully connected neural network. A fully connected neural network can include a series of fully connected layers. Each output dimension can depend on each input dimension. A fully connected neural network can be a feedforward network.

[0171] A trained machine learning model can include a convolutional neural network. A convolutional neural network can rely on local connections and shared weights between units to obtain translation invariant descriptors, followed by feature pooling (subsampling). A basic convolutional neural network architecture can optionally include an arbitrarily selected convolutional and pooling layer, optionally followed by a fully connected layer for supervised prediction. In practice, a convolutional neural network can consist of multiple (e.g., more than 10) convolutional and pooling layers to better model the input space. In some cases, convolutional neural networks require large datasets for successful training. In some cases, convolutional neural networks can use fewer parameters than fully connected neural networks by computing convolutions over small regions of the input space and sharing parameters between regions. A convolutional neural network can be a one-dimensional (1D) convolutional neural network. A convolutional neural network can be a two-dimensional (2D) convolutional neural network. In some embodiments, a convolutional neural network includes three or more dimensions.

[0172] The trained machine learning model can include gradient boosting decision trees. Gradient boosting is a machine learning that can be used for regression and classification problems, and can create an ensemble of weak prediction models, such as decision trees. Gradient boosting decision trees can include, for example, the XGBoost framework or the LightGBM framework.

[0173] The machine learning model can include hyperparameters. Hyperparameters are external configurations for the model, and their values cannot be estimated from the data. Hyperparameters can be tuned, for example, for known prediction model problems. In some cases, hyperparameters are used during the process to help estimate model parameters. In some cases, hyperparameters may be specified by a physician. In some cases, hyperparameters can be set using heuristic methods.

[0174] In some embodiments, the HMM-based detection algorithm can detect large or fairly large CNVs "in segments." In some cases, due to fluctuations in the coverage signal, there may be small detection gaps along the length of the true CNV. In one example, a 1 megabase pair (Mbp) deletion may be detected as a few separate nominal detections with small gaps in between. To mitigate this, a merge operation can be used that identifies pairs of adjacent detections separated by a gap smaller than either of the two flanking detections. The merge operation then measures the median coverage level of the gap. If the median coverage exceeds a predefined threshold, the two detections are merged into a single large detection spanning the two original detections (including the enclosed detection gap). In one example, the true feature spans both detections and the gap is a statistical artifact. Using real sequencing data from samples known to have large CNVs, this merge operation may allow for fairly good fidelity with respect to the true features of the CNV.

[0175] The methods and systems provided herein may further include the use of the feature selection algorithms provided herein. In some embodiments of the invention, feature selection is effected by the use of the LIMMA software package (Smyth, G. K. (2005). Limma: Linear Models for Microarray Data. In: Bioinformatics and Computational Biology Solutions using R and Bioconductor, R. Gentleman, V. Carey, S. Dudoit, R. Irizarry, W. Huber (eds.), Springer, New York, pages 397-420).

[0176] In some embodiments of the present invention, diagonal discriminant analysis, k-nearest neighbor algorithm, support vector machine (SVM) algorithm, linear support vector machine, random forest algorithm, or probabilistic model utilization method or combinations thereof are provided for detecting one or more genomic regions. In some embodiments, identified markers that distinguish samples (e.g., diseased or normal) or distinguish genomic regions (e.g., copy number polymorphism vs. normal) are selected based on the statistical significance of the difference in expression levels between classifications of interest. In some cases, the statistical significance is adjusted by providing another correction for Benjamini Hochberg or false discovery rate (FDR).

[0177] In some cases, the algorithm may be supplemented with a meta-analysis technique such as that described by Fishel and Kaufman et al. 2007 Bioinformatics 23(13): 1599-606. In some cases, the algorithm may be supplemented with a meta-analysis technique such as reproducibility analysis. In some cases, reproducibility analysis selects markers that appear in at least one predictive expression product marker set.

[0178] Statistical evaluation of genomic region detection may provide one or more quantitative values that are one or more of the following indicators: the likelihood that the diagnosis is accurate; the likelihood of a disorder, disease, condition, and the like; the likelihood of a specific disorder, disease, condition; and the likelihood that a specific therapeutic intervention will be successful. Thus, a physician who is likely not trained in genetics or molecular biology need not understand the raw data. Rather, the data are presented directly to the physician in the form of quantitative values that guide patient care. The results can be statistically evaluated using many methods known in the art including, but not limited to: Student's t-test, two-sided t-test, Pearson's rank sum analysis, hidden Markov model analysis, q-q plot analysis, principal component analysis, one-way ANOVA, two-way ANOVA, LIMMA, and the like.

[0179] F. Disease or Condition Some embodiments include predicting, diagnosing, and / or prognosticating the condition or outcome of a disease or condition of a subject based on one or more biomedical outputs. Predicting, diagnosing, and / or prognosticating the condition or outcome of a disease of a subject can include diagnosing the disease or condition, identifying the disease or condition, determining the state of the disease or condition, assessing the risk of the disease or condition, assessing the risk of disease recurrence, assessing the effectiveness of a drug, assessing the risk of drug adverse reactions, predicting an optimal drug dosage, predicting drug resistance, or combinations thereof.

[0180] The samples disclosed herein may be from subjects suffering from cancer. The sample may include malignant tissue, benign tissue, or combinations thereof. The cancer may be recurrent cancer and / or refractory cancer. Examples of cancer include, but are not limited to, sarcoma, cytoma, lymphoma, or leukemia. In some cases, samples containing cancer tissue are obtained, but no matched normal samples are obtained. In some cases, no matched normal samples are available. In some cases, matched normal samples are obtained (e.g., for training and testing the models disclosed herein).

[0181] A sarcoma is a cancer of bone, cartilage, fat, muscle, blood vessels, or other connective or supportive tissue. Sarcomas include, but are not limited to, bone cancer, fibrosarcoma, chondrosarcoma, Ewing sarcoma, malignant angioendothelioma, malignant schwannoma, bilateral vestibular schwannoma, osteosarcoma, soft tissue sarcoma (e.g., alveolar soft part sarcoma, angiosarcoma, cystosarcoma phylloides, dermatofibrosarcoma, desmoid tumor, epitheloid sarcoma, extraskeletal osteosarcoma, fibrosarcoma, hemangiopericytoma, angiosarcoma, Kaposi sarcoma, leiomyosarcoma, liposarcoma, lymphangiosarcoma, lymphoma, malignant fibrous histiocytoma, neurofibrosarcoma, rhabdomyosarcoma, and synovial sarcoma).

[0182] Carcinoma is a cancer that begins in epithelial cells, which cover the body's surface, produce hormones, and make up glands. By way of non-limiting example, carcinomas include breast cancer, pancreatic cancer, lung cancer, colon cancer, colorectal cancer, rectal cancer, kidney cancer, bladder cancer, stomach cancer, prostate cancer, liver cancer, ovarian cancer, brain cancer, vaginal cancer, vulvar cancer, uterine cancer, oral cancer, penile cancer, testicular cancer, esophageal cancer, skin cancer, fallopian tube cancer, head and neck cancer, gastrointestinal stromal cancer, adenocarcinoma, cutaneous melanoma or uveal melanoma, anal area cancer, small intestine cancer, endocrine system cancer, thyroid cancer, parathyroid cancer, adrenal cancer, urethral cancer, renal pelvic cancer, ureter cancer, endometrial cancer, cervical cancer, pituitary cancer, central nervous system (CNS) tumors, primary CNS lymphoma, brainstem glioma, and spinal cord axis tumors. The cancer may be basal cell carcinoma, squamous cell carcinoma, melanoma, non-melanoma cancer, or actinic (solar) keratosis.

[0183] The cancer may be lung cancer. Lung cancer can begin in the airways that enter the lungs (bronchi) from the trachea into the side passages or in the small air sacs of the lungs (alveoli). Lung cancer includes non-small cell lung cancer (NSCLC), small cell lung cancer, and mesothelioma. Examples of NSCLC include squamous cell carcinoma, adenocarcinoma, and large cell carcinoma. Mesothelioma may be a cancerous tumor of the lining of the lungs and chest cavity (pleura) or the lining of the abdomen (peritoneum). Mesothelioma may be due to asbestos exposure. The cancer may be a brain cancer such as glioblastoma.

[0184] The cancer may be a central nervous system (CNS) tumor. The CNS tumor may be classified as a glioma or a non-glioma. The glioma may be a malignant glioma, a high-grade glioma, or a diffuse intrinsic pontine glioma. Examples of gliomas include astrocytomas, oligodendrogliomas (or a mixture of oligodendroglioma elements and astrocytoma elements), and ependymomas. Examples of astrocytomas include, but are not limited to, low-grade astrocytomas, anaplastic astrocytomas, glioblastoma multiforme, pilocytic astrocytomas, pleomorphic xanthoastrocytomas, and subependymal giant cell astrocytomas. Oligodendrogliomas include low-grade oligodendrogliomas (or oligoastrocytomas) and anaplastic oligodendriogliomas. Non-gliomas include meningiomas, pituitary adenomas, primary CNS lymphomas, and medulloblastomas. The cancer may be a meningioma.

[0185] The leukemia may be acute lymphocytic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, or chronic myeloid leukemia. Further types of leukemia include hairy cell leukemia, chronic myelomonocytic leukemia, and juvenile myelomonocytic leukemia.

[0186] Lymphoma is a cancer of lymphocytes and can originate from either B lymphocytes or T lymphocytes. The two main types of lymphoma are Hodgkin lymphoma, which was formerly known as Hodgkin's disease, and non-Hodgkin lymphoma. Hodgkin lymphoma is characterized by the presence of Reed-Sternberg cells. Non-Hodgkin lymphoma encompasses all lymphomas that are not Hodgkin lymphoma. Non-Hodgkin lymphoma can be low-grade lymphoma and intermediate-grade lymphoma. Examples of non-Hodgkin lymphoma include, but are not limited to, diffuse large B-cell lymphoma, follicular lymphoma, mucosa-associated lymphatic tissue lymphoma (MALT), small lymphocytic lymphoma, mantle cell lymphoma, Burkitt lymphoma, mediastinal large B-cell lymphoma, Waldenström's macroglobulinemia, nodular marginal zone B-cell lymphoma (NMZL), splenic marginal zone lymphoma (SMZL), extranodal marginal zone B-cell lymphoma, intravascular large B-cell lymphoma, primary effusion lymphoma, and lymphomatoid granulomatosis.

[0187] Some embodiments may include treating and / or preventing a subject's disease or condition based on one or more biomedical outputs. The one or more biomedical outputs may recommend one or more treatments. The one or more biomedical outputs may propose, select, direct, recommend, or otherwise determine a treatment and / or prevention process for a disease or condition. The one or more biomedical outputs may recommend changing or continuing one or more treatments. Changing one or more treatments may include performing, initiating, decreasing, increasing, and / or ending one or more treatments. The one or more treatments may include anti-cancer therapy, anti-viral therapy, anti-bacterial therapy, anti-fungal therapy, immunosuppressive therapy, or combinations thereof. The one or more treatments may treat, alleviate, or prevent one or more diseases or symptoms.

[0188] Examples of anti-cancer therapies include, but are not limited to, surgery, chemotherapy, radiation therapy, immunotherapy / bio-therapy, and photodynamic therapy. Anti-cancer therapies may include chemotherapy, monoclonal antibodies (e.g., rituximab, trastuzumab), cancer vaccines (e.g., therapeutic vaccines, prophylactic vaccines), gene therapy, or combinations thereof.

[0189] G. Systems, Kits, and Libraries Certain embodiments can be implemented by means of a system, kit, library, or combinations thereof. The methods of the present invention may include one or more systems. A system can be implemented by means of a kit, library, or both. A system may include one or more components for performing any of the methods or steps of some embodiments. For example, a system may include one or more kits, devices, libraries, or combinations thereof. A system may include one or more sequencers, processors, memory locations, computers, computer systems, or combinations thereof. A system may include a transmission device.

[0190] A kit may include various reagents for performing the various operations disclosed herein, including sample processing and / or analysis operations. A kit may include instructions for performing at least a portion of the operations disclosed herein. A kit may include one or more capture probes, one or more beads, one or more labels, one or more linkers, one or more devices, one or more reagents, one or more buffers, one or more samples, one or more databases, or combinations thereof.

[0191] The library may include one or more capture probes. The library may include one or more subsets of nucleic acid molecules. The library may include one or more databases. The library may be created or generated from any of the methods, kits, or systems disclosed herein. A database library may be created from one or more databases. A method of creating one or more libraries may include (a) integrating information from one or more databases to create an integrated dataset, (b) analyzing this integrated dataset, and (c) creating one or more database libraries from this integrated dataset.

[0192] Specific implementations have been illustrated and described from the foregoing, but various improvements may be made thereto, and it should be understood that various improvements are contemplated herein. Embodiments of one aspect may be combined with or modified by another aspect. The present invention is not intended to be limited by the specific examples provided herein. Although the present invention has been described with reference to the foregoing specification, the description of the embodiments of the present invention and the drawings herein are not intended to be construed in a limiting sense. Furthermore, it is to be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative proportions described herein that depend on various conditions and variables. Various improvements in the form and details of the embodiments of the present invention will be apparent to those skilled in the art. Accordingly, the present invention is also considered to encompass all such improvements, modifications, and equivalents.

[0193] VI. Computing Environment FIG. 10 shows an example of a computer system 1000 for executing a part of the embodiments disclosed in this specification. The computer system 1000 may have a distributed architecture, and some of the components (e.g., memory and processors) are part of an end-user device, and some other similar components (e.g., memory and processors) are part of a computer server. The computer system 1000 includes at least a processor 1002, a memory 1004, a storage device 1006, an input / output (I / O) peripheral device 1008, a communication peripheral device 1010, and an interface bus 1012. The interface bus 1012 is configured to communicate, transmit, and transfer data, control, and commands among various components of the computer system 1000. The processor 1002 may include one or more processing units such as a CPU, GPU, TPU, systolic array, or SIMD processor. The memory 1004 and the storage device 1006 include computer-readable storage media such as RAM, ROM, electrically erasable programmable read-only memory (EEPROM), hard drives, CD-ROMs, optical storage devices, magnetic storage devices, e.g., flash (registered trademark) memory, and other tangible recording media. Any of such computer-readable storage media may be configured to store instructions or program code embodying an aspect. The memory 1004 and the storage device 1006 also include a computer-readable signal medium. The computer-readable signal medium includes a propagated data signal having computer-readable program code embodied therein. Such a propagated signal may take any of various forms including, but not limited to, an electromagnetic signal, an optical signal, or a combination thereof. The computer-readable signal medium includes any computer-readable medium that is not a computer-readable storage medium and that can communicate, transmit, or carry a program for use with the computer system 1000.

[0194] Furthermore, the memory 1004 includes an operating system, programs, and applications. The processor 1002 is configured to execute the stored instructions and includes, for example, a logic processing unit, a microprocessor, a digital signal processor, and other processors. The memory 1004 and / or the processor 1002 can be virtualized and provided, for example, within another computing system such as a cloud network or a data center. The I / O peripheral device 1008 includes a user interface such as a keyboard, a screen (e.g., a touch screen), a microphone, a speaker, and other input / output devices, and computing components such as a graphics processing unit, a serial port, a parallel port, a universal serial bus, and other input / output peripheral devices. The I / O peripheral device 1008 is connected to the processor 1002 by any of the ports coupled to the interface bus 1012. The communication peripheral device 1010 is configured to facilitate communication between the computer system 1000 on the communication network and other computing devices and includes, for example, a network interface controller, a modem, wireless and wired interface cards, an antenna, and other communication peripheral devices.

[0195] Although the subject matter has been described in detail with respect to its particular embodiments, those skilled in the art will, upon understanding the foregoing, readily be able to produce alternatives, modifications, and equivalents of such embodiments. Accordingly, this disclosure is presented for purposes of illustration rather than limitation, and it is to be understood that such improvements, modifications, and / or additions to the subject matter are not excluded as would be readily understood by those skilled in the art. On the contrary, the methods and systems described herein may be embodied in other various forms, and moreover, various omissions, substitutions, and changes in the form of the methods and systems described herein may be made without departing from the spirit of this disclosure. The appended claims and their equivalents are intended to cover such forms and improvements as fall within the scope and spirit of this disclosure.

[0196] Unless otherwise specifically stated, throughout this specification, discussions using terms such as "processing", "computing", "calculating", "determining", and "identifying" or the like refer to the operation or process of one or more computers or one or more similar electronic computing devices, such as manipulating or transforming data represented as physical, electronic, or magnetic quantities within the memory, registers, or other information recording devices, transmission devices, or display devices of a computing platform.

[0197] One or more systems discussed in this specification are not limited to any particular hardware architecture or configuration. A computing device can include any suitable arrangement of components that provide results that determine one or more inputs. Suitable computing devices include computing systems that use a general-purpose microprocessor that accesses stored software to program or configure the computing device from a general-purpose computing device to a specialized computing device that executes one or more embodiments of the subject matter. Any suitable programming, script, or other type of language or combination of languages may be used in the software used to program or configure the computing device to execute the teachings contained herein.

[0198] Embodiments of some embodiments may be implemented in the operation of such computing devices. The order of the blocks presented in the above examples can be varied, for example, by rearranging, combining, and / or decomposing the blocks into sub-blocks. Specific blocks or processes can be executed in parallel.

[0199] In particular, conditional syntax used in this specification, such as "can", "could", "might", "may", "e.g.", and the like, is understood within the context in which it is used, unless otherwise specifically stated, and generally a particular example is meant to represent a particular feature, element, and / or step and does not include other examples. Thus, such conditional syntax generally means that a feature, element, and / or step does not require one or more examples at all, or that one or more examples do not necessarily include the logic for determining whether these features, elements, and / or steps are included in or implemented in any particular example, regardless of the presence or absence of author input or instructions.

[0200] The terms "including", "including", "having", and the like mean the same thing, are open-ended and used inclusively, and do not exclude additional elements, features, acts, operations, etc. The term "or" is also used in an inclusive sense (and not in an exclusive sense) such that, for example, when used in connection with a list of elements, the term "or" means one, some, or all of the elements in the list. The use of "adapted to" or "configured to" in this specification means an open and inclusive syntax that does not exclude an apparatus adapted or configured to perform additional tasks or steps. Moreover, the use of "based on" means an open and inclusive syntax in that a process, step, calculation, or other act "based on" one or more recited conditions or values may in fact be based on additional conditions or values beyond those recited. Similarly, the use of "based at least in part on" means an open and inclusive syntax in that a process, step, calculation, or other act "based at least in part on" one or more recited conditions or values may in fact be based on additional conditions or values beyond those recited. The headings, tables, and numbering included in this specification are for ease of explanation only and are not intended to be limiting.

[0201] The various features and processes described above may be used independently of each other or combined in various ways. All possible combinations and sub - combinations are intended to be within the scope of this disclosure. Additionally, a particular method or process block may be omitted in some executions. The methods and processes described herein are also not limited to any order, and the blocks or states related thereto can be executed in other appropriate orders. For example, the described blocks or states may be executed in an order other than the specifically disclosed one, or multiple blocks or states may be combined into a single block or state. The example blocks or states may be executed sequentially, in parallel, or in some other way. Blocks or states may be added to or removed from the disclosed examples. Similarly, the example systems and components described herein may be configured differently from those described. For example, elements may be added to, removed from, or rearranged from the disclosed examples.

Claims

1. Obtaining nucleic acid sequence data of a target biological sample, Aligning the nucleic acid sequence data with a reference genome, Identifying a set of candidate variants in the nucleic acid sequence data based on the aligned nucleic acid sequence data, wherein the set of candidate variants includes one or more somatic variants and one or more germline variants, Identifying the somatic variants by processing the set of candidate variants using a trained machine learning model without using the nucleic acid sequence data of the corresponding biological sample of the target, wherein the corresponding biological sample of the target indicates the absence of a tumor, and Outputting a report identifying the somatic variants, A method comprising the above steps.

2. The method according to claim 1, wherein the biological sample is a tumor sample of the target.

3. The method according to claim 2, wherein the target is a human target.

4. The method according to claim 1, wherein the trained machine learning model includes a gradient boosting decision tree.

5. The method according to claim 1, wherein the trained machine learning model includes two classification models.

6. The method according to claim 1, wherein the trained machine learning model includes a filtering model.

7. The method according to claim 1, wherein the trained machine learning model includes a relief model.

8. The method according to claim 1, wherein the trained machine learning model is trained using training data corresponding to a set of corresponding tumor-normal pairs.

9. The method according to claim 1, wherein the trained machine learning model is trained by tuning one or more hyperparameters via random search.

10. The method according to claim 1, wherein the report identifies at least one biomarker.

11. The method according to claim 1, wherein the report identifies at least one prognostic marker.

12. The method according to claim 1, wherein the report identifies the presence or absence of the one or more somatic variants.

13. The method according to claim 1, wherein the report identifies treatment recommendations.

14. The method according to claim 13, wherein the treatment recommendations include recommendations to treat the human target.

15. The method according to claim 14, further comprising performing the treatment on the human target.

Citation Information

Patent Citations

  • No-contrast somatic cell mutation detecting method and device

    CN109903811A

  • Machine learning system and method for somatic mutation discovery

    US20190189242A1

  • Methods for classifying somatic variations

    WO2018064547A1