Compositions, methods and systems for processing or analyzing multi-species nucleic acid samples

By combining nucleic acid probe pools and computer processing units, the high cost and insufficient regional capture of whole-genome and exome sequencing in existing technologies have been solved, enabling efficient and economical sequencing and analysis of human and non-human genomes.

CN120905355APending Publication Date: 2025-11-07PERSONALIS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511045707.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2018-08-07
Filing Date
2019-05-24
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing whole-genome and exome sequencing methods are expensive and cannot effectively capture non-exome regions of biomedical interest. They also perform poorly in high-CG-content regions and cannot provide adequate and cost-effective sequencing of repetitive elements in the genome.

Method used

Nucleic acid probe pools are used to generate subgroups of nucleic acid molecules from biological samples, including probes targeting human and non-human genomes. Sequence alignment and analysis are performed using hybridization reactions and sequencing technologies, combined with computer processing units, to generate biomedical reports.

Benefits of technology

It enables efficient and cost-effective sequencing of human and non-human nucleic acids, and can simultaneously detect candidate tumor neoantigens, non-human species, and CDR3 sequences, providing more comprehensive genomic analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120905355A_ABST
    Figure CN120905355A_ABST
Patent Text Reader

Abstract

The present application relates to compositions, methods, and systems for processing or analyzing multi-species nucleic acid samples. Provided herein are compositions, methods, and systems for sample processing and / or data analysis. Sample processing may include nucleic acid sample processing and subsequent sequencing. The methods and systems of the present disclosure are useful, for example, for analyzing nucleic acid samples from humans, non-humans, and combinations thereof.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional of application with the application number 201980050741.1, filed on May 24, 2019, and the title “Compositions, Methods, and Systems for Processing or Analyzing Multi-Species Nucleic Acid Samples,” the disclosure of which is incorporated herein by reference in its entirety. CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Patent Application No. 16 / 056,982, filed August 7, 2018, and U.S. Provisional Patent Application No. 62 / 678,475, filed May 31, 2018, the entire contents of which are incorporated herein by reference. BACKGROUND

[0003] Current methods for whole genome and / or exome sequencing can be expensive and fail to capture many biomedically important variants. For example, commercially available exome enrichment kits (e.g., Illumina’s TruSeq Exome Enrichment and Agilent’s SureSelect Exome Enrichment) can fail to target non-exomic regions and exomic regions of biomedical interest. Typically, whole genome and / or exome sequencing performed using standard sequencing methods performs poorly in regions of content with very high CG content (>70%). Furthermore, whole genome and / or exome sequencing also fails to provide adequate and / or cost-effective sequencing of repetitive elements in the genome.

[0004] The methods disclosed herein provide specialized sequencing protocols or techniques to address these issues and extend the analysis to human and non-human genomes in a single sample. SUMMARY

[0005] Disclosed herein is a method for processing a biological sample obtained from a subject, comprising (a) generating a subset of nucleic acid molecules from the biological sample using a pool of nucleic acid probes, wherein the probes comprise (i) a first plurality of nucleic acid probes configured to target elements of a human genome and (ii) a second plurality of nucleic acid probes configured to target elements of one or more non-human genomes; and (b) performing an assay on the subset of nucleic acid molecules to generate sequence information comprising (i) human nucleic acids from the biological sample of the subject and (ii) non-human nucleic acids from the biological sample of the subject. In some cases, the first plurality of nucleic acid probes of (i) is configured to target elements derived from a human genome. In some cases, the subject can be a human. In some cases, the second plurality of nucleic acid probes is configured to target elements in non-human genomic sequences from one or more species selected from the group consisting of viruses, bacteria, bacteriophages, fungi, protists, archaea, amoebas, worms, algae, genetically modified cells, and genetically modified vectors. In an aspect, generating the subset of nucleic acid molecules from the biological sample comprises performing one or more hybridization reactions. In some cases, the method further comprises obtaining the biological sample from the subject. In some cases, the biological sample of the subject can be derived from a tumor biopsy, whole blood, or plasma. In some aspects, the method further comprises aligning sequences of the subset of nucleic acid molecules to one or more reference sequences. In some cases, the one or more reference sequences comprises a plurality of reference sequences. In some cases, the plurality of reference sequences corresponds to two or more different species. The method can further comprise identifying the source of the nucleic acid molecules in the subset based on the alignment. In some cases, the method can further comprise generating an output comprising identifying the source of the nucleic acid molecules in the biological sample. In some cases, the concentration of the second plurality of nucleic acid probes is greater than the concentration of the first plurality of nucleic acid probes in the pool of nucleic acid probes. In an aspect, the relative concentration of the second plurality of nucleic acid probes in the pool of probes is greater than the relative concentration of the first plurality of nucleic acid probes in the pool of nucleic acid probes. In an aspect, the first plurality of nucleic acid probes comprises a human exome capture probe set. In some cases, the first plurality of nucleic acid probes comprises probes configured to target junction sequences resulting from human V(D)J recombination or rearrangement. In some cases, the second plurality of nucleic acid probes comprises one or more probes configured to target human papillomavirus E6 and / or E7 genes. In some cases, the second plurality of nucleic acid probes comprises probes configured to target one or more elements of a bacterial 16S ribosomal RNA gene. In an aspect, the assay of (b) comprises performing sequencing to generate paired-end read sequences of 130 bases to 280 bases in length.In an aspect, the method further comprises generating one or more biomedical reports comprising one or more sets of data selected from the group consisting of: (i) candidate tumor neoantigens, (ii) detected non-human species, (iii) detected CDR3 sequences, and any combination thereof. In an aspect, the detected non-human species is an antigen. In some cases, the CDR3 sequences are generated by V(D)J rearrangement or recombination. In some cases, the CDR3 sequences correspond to an immune response to an antigen. In some cases, the one or more biomedical reports comprise (i)-(iii).

[0006] Disclosed herein is a method for processing a biological sample obtained from a subject, the method comprising (a) generating a subset of nucleic acid molecules from the biological sample, wherein the subset of nucleic acid molecules comprises (i) a first plurality of nucleic acid molecules from the subject and (ii) a second plurality of nucleic acid molecules that are not from the subject, and wherein the abundance of the first plurality of nucleic acid molecules is greater than the abundance of the second plurality of nucleic acid molecules in the biological sample; (b) performing an assay on the subset of nucleic acid molecules to generate sequence information comprising: (i) the first plurality of nucleic acid molecules and (ii) the second plurality of nucleic acid molecules. In some cases, the nucleic acid molecules of (i) are derived from the genome of the subject. In some cases, the subject is a human. In some cases, the second plurality of nucleic acid molecules that are not from the subject comprises one or more members selected from the group consisting of: a virus, a bacterium, a bacteriophage, a fungus, a protist, an archaeon, an amoeba, a worm, an alga, a genetically modified cell, and a genetically modified vector. In an aspect, generating the subset of nucleic acid molecules from the biological sample comprises performing one or more hybridization reactions. In an aspect, the method further comprises obtaining the biological sample from the subject. In some cases, the biological sample from the subject is derived from a tumor biopsy, whole blood, or plasma. In an aspect, the method further comprises aligning the sequences of the subset of nucleic acid molecules to one or more reference sequences. In an aspect, the one or more reference sequences comprises a plurality of reference sequences. In an aspect, the plurality of reference sequences corresponds to two or more different species. The method can further comprise identifying the origin of the nucleic acid molecules in the subset based on the alignment. In an aspect, the method further comprises generating an output comprising the identified origin of the nucleic acid molecules in the biological sample. In some cases, the abundance of the first plurality of nucleic acid molecules is greater than the abundance of the first plurality of nucleic acid molecules in the biological sample. In some cases, the relative abundance of the second plurality of nucleic acid molecules in the subset is greater than the relative abundance of the first plurality of nucleic acid molecules in the subset.

[0007] Disclosed herein is a composition comprising a pool of probes configured to hybridize to (i) one or more human sequences from a subject and (ii) one or more non-human sequences from a subject. In an aspect, the pool of probes is a plurality of capture probes. In an aspect, the pool of probes is a plurality of amplification probes.

[0008] Disclosed herein is a system for processing a biological sample of a subject, comprising: a processing unit comprising one or more computer processors individually or collectively programmed to perform assays on a subset of nucleic acid molecules to produce sequence information comprising i) human nucleic acids from a biological sample of a subject and ii) non-human nucleic acids from a biological sample of a subject, the subset of nucleic acid molecules produced from the biological sample using a pool of nucleic acid probes, wherein the probes comprise (i) a first plurality of nucleic acid probes configured to target elements of the human genome and (ii) a second plurality of nucleic acid probes configured to target elements of one or more non-human genomes; and a computer memory configured to store the sequence information. In some embodiments, the one or more computer processors are programmed to generate an alignment of the sequences to one or more reference sequences. In some embodiments, the one or more reference sequences comprise a plurality of reference sequences, and wherein the plurality of reference sequences correspond to two or more different species. In some embodiments, the one or more computer processors are programmed to identify the origin of the nucleic acid molecules in the subset based on the alignment. In some embodiments, the one or more computer processors are programmed to produce one or more biomedical reports comprising information selected from the group consisting of: (i) candidate tumor neoantigens, (ii) detected non-human species, (iii) detected complementarity determining region 3 (CDR3) sequences, and any combination thereof.

[0009] Disclosed herein is a system for processing a biological sample of a subject, comprising: a processing unit comprising one or more computer processors individually or collectively programmed to perform an assay on a subset of nucleic acid molecules to produce sequence information comprising sequences of: i) a first plurality of nucleic acid molecules and (ii) a second plurality of nucleic acid molecules, the subset of nucleic acid molecules being produced from the biological sample, wherein the subset of nucleic acid molecules comprises (i) a first plurality of nucleic acid molecules from the subject and (ii) a second plurality of nucleic acid molecules not from the subject, wherein the abundance of the first plurality of nucleic acid molecules is greater than the abundance of the second plurality of nucleic acid molecules in the biological sample; and a computer memory configured to store the sequence information. In some embodiments, the one or more computer processors are programmed to generate an alignment of the sequences to one or more reference sequences. In some embodiments, the one or more reference sequences comprise a plurality of reference sequences, and wherein the plurality of reference sequences correspond to two or more different species. In some embodiments, the one or more computer processors are programmed to identify the origin of the nucleic acid molecules in the subset based on the alignment. In some embodiments, the one or more computer processors are programmed to produce one or more biomedical reports containing information selected from the group consisting of: (i) candidate tumor neoantigens, (ii) detected non-human species, (iii) detected complementarity determining region 3 (CDR3) sequences, and any combination thereof.

[0010] Another aspect of the disclosure provides a non-transitory computer readable medium comprising machine executable code that, when executed by one or more computer processors, implements any of the methods described above or elsewhere herein.

[0011] Another aspect of the disclosure provides a system comprising one or more computer processors and a computer memory coupled with the same. The computer memory comprises machine executable code that, when executed by the one or more computer processors, implements any of the methods described above or elsewhere herein.

[0012] Other aspects and advantages of the disclosure will become apparent to those skilled in the art from the following detailed description, taken in conjunction with the accompanying drawings, illustrating the principles of the disclosure by way of representative embodiments. As will be realized, the disclosure is capable of other and different embodiments and its several details are capable of modifications in various apparent respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature and not restrictive. INCORPORATION BY REFERENCE

[0013] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. BRIEF DESCRIPTION OF DRAWINGS

[0014] The novel features of the application are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present application will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the application are utilized, and the accompanying drawings (also “figure” and “figures” herein), of which: FIG. 1 Methods of generating subsets of nucleic acid molecules using a first plurality of nucleic acid probes and a second plurality of nucleic acid probes are schematically illustrated; FIG. 2 Methods of generating subsets of nucleic acid molecules using a first plurality of nucleic acid probes and a second plurality of nucleic acid probes are schematically illustrated. The first plurality of nucleic acid probes can bind to non-human genomes; FIG. 3 Methods of generating subsets of nucleic acid molecules using a first plurality of nucleic acid probes and a second plurality of nucleic acid probes are schematically illustrated. The first plurality of nucleic acid probes can bind to fragmented transcriptomes; FIG. 4 Methods of generating subsets of nucleic acid molecules using a plurality of nucleic acid probes are schematically illustrated. The plurality of nucleic acid probes can include a set of probes targeting the exome of a subject and a set of probes targeting non-subject nucleic acid sequences; FIG. 5 Computer systems programmed or otherwise configured to implement the methods of the present disclosure are illustrated; FIG. 6A A schematic of a workflow is illustrated. Prep 1 and Prep 2 refer to subsets of nucleic acids. Assay 1, Analysis 1, and Output refer to any assays, analyses, and outputs described herein. FIG. 6B A schematic of a workflow is illustrated. Prep 1 and Prep 2 refer to subsets of nucleic acids. Assay 1, Assay 2, Analysis 1, and Output refer to any assays, analyses, and outputs described herein. FIG. 6C A schematic of a workflow is illustrated. Prep 1 and Prep 2 refer to subsets of nucleic acids. Assay 1, Assay 2, Analysis 1, Analysis 2, and Output refer to any assays, analyses, and outputs described herein. FIG. 6DA schematic showing a workflow comprising (1) separating a nucleic acid sample into multiple subgroups that are treated with multiple protocols. These protocols can involve enrichment for different genomic regions or non-genomic regions, and include one or more different amplification operations to prepare libraries of nucleic acid molecules for assays. Certain of these libraries can be pooled (2) for assays. Results from some assays can be pooled (3) for subsequent analysis. Variant calling or other assessment of sequence or genetic state can be further pooled (4) to produce a combined assessment at the east locus that the assays are directed at. Protocols 1-4 refer to any of the methods described herein. Assay 1, Assay 2, Assay 3, Analysis 1, Analysis 2, and Output refer to any of the assays, analyses, and outputs described herein; FIG. 7 An example of an assay workflow described herein is depicted; FIG. 8 A schematic showing a workflow of the present disclosure; FIG. 9 A schematic showing a workflow of the present disclosure; FIG. 10 An effect of shear time on fragment size is shown; FIG. 11 An effect of bead ratio on fragment size is shown; FIG. 12 An effect of shear time on fragment size is shown; FIG. 13 A schematic depicting a nucleic acid library construction workflow; FIG. 14 A method for developing multi-threaded assays for a variety of biomedical applications is shown; FIG. 15 An example of an assay workflow comprising multiple subgroups of DNA enriched for different genomic regions is depicted, the subgroups undergoing some independent processing operations before being pooled for a sequencing assay. Read segments from two or more subgroups are combined a) in a sequencing device or b) subsequently in a computer (e.g., using one or more algorithms) to produce a single test result for the region addressed by the two or more subgroups in combination, and to produce a pool of data that can be used for one or more biomedical reports. Pullouts can include human target sequences and non-human target sequences; FIG. 16Examples of assay workflows comprising multiple subgroups of DNA enriched for different genomic regions are depicted, which undergo some independent processing operations before being independently sequenced and analyzed for variants. Variants from two subgroups can be combined to produce results for regions addressed by the combination of the two or more subgroups, and to produce a pool of data that can be used in one or more biomedical reports. Supplementary extracts can include human target sequences and non-human target sequences; FIG. 17 Examples of assay workflows comprising multiple subgroups of DNA enriched for different genomic regions are depicted, which undergo some independent processing operations before being independently sequenced and analyzed for variants. Variants from two subgroups can be combined to produce results for regions addressed by the combination of the two or more subgroups, and to produce a pool of data that can be used in one or more biomedical reports. Supplementary extracts can include human target sequences and non-human target sequences; FIG. 18 Multi-threaded assays are depicted, which include producing two subgroups of DNA by size selection, and further dividing into two subgroups of DNA enriched for different genomic regions based on GC content. Longer molecules can be sequenced using techniques suitable for longer molecules. The two shorter molecule subgroups can be further prepared and amplified according to protocols suitable for T m subgroups are then combined for sequencing on a high-throughput short read sequencer, HiSeq. Raw data from sequencing can be combined and analyzed (e.g., by one or more software programs or algorithms) to produce a single best result for all regions addressed by the subgroups, and to produce a pool of data that can be used in one or more biomedical reports. Supplementary extracts can include human target sequences and non-human target sequences. DETAILED DESCRIPTION

[0015] While various embodiments of the application have been shown and described in the herein, it will be clear to those of ordinary skill in the art that many changes, modifications, and alternatives can be made thereto without departing from the application. It should be understood that various alternatives to the embodiments of the application described herein can be employed.

[0016] Humans often carry large numbers of microbial species, in some cases over 1000. These microbial species have been documented by the Human Microbiome Project and other studies, and can include many viruses, bacteria, bacteriophages, fungi, protists, archaea, and some amoebae, helminths, and algae. These non-human species can be beneficial, but they can also cause or modulate disease, including cancer. Their effects can be direct (e.g. causing human cells to mutate leading to cancer) or indirect (e.g. stimulating the immune system, which affects its ability to fight disease). Microbial species in an individual can be identified and characterized by sequencing nucleic acids extracted from a sample from a person. Many of these microbes live at the interface of the human body and its surrounding environment (e.g. skin, saliva, nasal passages, gut, intestines, genitalia). Samples can be taken from these interfaces by swabbing, biopsy, taking stool or urine, taking saliva, or similar methods. A variety of methods have been developed to sequence analyze microbial nucleic acids in these samples. These include both targeted PCR amplification (e.g. of a sub-portion of the 16S ribosomal RNA gene), and non-targeted (metagenomic) methods using deep sequencing. Many of these methods involve sample types and / or enrichment methods that attempt to minimize the amount of human DNA, optimizing sensitivity to the microbial targets of interest. The human genome (approximately 3 billion bases) is about 1,000 times larger than a typical bacterial genome (less than one million bases up to several million bases), and thousands of times larger than most viral genome sizes (typically thousands up to tens of thousands of bases). Thus, the human DNA content can occupy a large portion of the detection capacity if no sampling methods or detection techniques are used to avoid or reduce it.

[0017] The progression of many diseases is also affected by human cell genetics. This can include inherited genetic variants, somatic variants in cancer, VDJ recombination in immune cells, differential gene expression in different cell types, and other characteristics. These can also be determined by sequencing nucleic acids, typically from blood (DNA or cell-free DNA most commonly from PBMCs) or from nucleic acids from diseased tissue (e.g. tumor biopsies). A variety of methods have been developed to sequence analyze nucleic acids in these samples, including amplicon panels (typically for up to several hundred genes), hybrid capture (typically for large numbers of genes, including exome), and non-targeted methods (whole genome sequencing). Many of these methods involve sample types and / or enrichment methods optimized for their intended human nucleic acid targets.

[0018] Sample types for microbial analysis can differ from sample types for human genetic analysis. For example, while a stool sample can be useful for analyzing the gut microbiome, a stool sample can be a poor sample for detecting inherited human genetic diseases (such as cystic fibrosis). White blood cells (PBMCs) in blood can be sequenced to look for causes of inherited genetic diseases, but are generally not a good choice for bacterial analysis because the immune system largely excludes bacteria from the blood. Assay methods for human and microbial analysis also differ. For example, either can use PCR, but because PCR produces very narrow focused results (i.e., amplicons each typically represent a small fraction of the genome of the target species), human and microbial targets are not generally combined in a single PCR-based assay. Non-targeted assays are available for human genetics (i.e., whole human genome sequencing) and microbial metagenomics, but because of the large difference in genome size described above, it is generally best to perform the assays separately for each, and to optimize reagents and assays for each. For the same reason, hybrid capture assay technologies (e.g., exome) have been developed for specific (usually mammalian) species (e.g., human exons or bovine exons).

[0019] Cancer can be a special case. Many tumors use checkpoint genes and other methods to partially or completely exclude the immune system. Thus, microbes that manage to penetrate into a tumor can be able to survive and even multiply there. Complete live bacteria (e.g., Fusobacterium) can be cancer cells and include cycles of cell division and live entirely within them by translocation. Bacteria can cause cancer (e.g., Helicobacter pylori is a cause of stomach cancer), they can affect the progression of cancer and respond to cancer therapeutics (e.g., immune checkpoint inhibitors). Viruses can also enter cells and in some cases be a cause of cancer and / or can integrate their genomes into human chromosomes. Thus, a tumor biopsy can contain human and microbiome species, and their nucleic acids. When these cells die, their multi-species nucleic acids can drop into their surroundings and eventually be detected as cell-free nucleic acids in plasma.

[0020] Sample amounts for tumor biopsies are often very limited, making it more difficult to perform multiple different assays for both human genetic targets and microbial genetic targets in the same sample. When sample amounts are limited, an integrated assay that can provide sequence data from both human and microbiome species at the same time can be advantageous. The amount of cell-free DNA and RNA in plasma is also often very limited. An integrated assay that can provide sequence data from both human and microbiome species at the same time from a single small plasma sample can also be advantageous.

[0021] Disclosed herein is an assay that supports the parallel detection of extensive human genetic data as well as microbiome data from the same sample at the same time. A single assay can be performed that does not require any more sample than what would be required for an equivalent human-only assay. The assay can use a human exome capture kit (e.g., Agilent Clinical Research Exome v2). The kit or composition of probe pools can use hybridization probes that are complementary to the human sequences that they target. By using over 50,000 capture probes, methods, kits, or compositions can target the exons of most predetermined human genes. For example, we add a set of additional capture probes that have been designed to target non-human sequences before the nucleic acids in a cancer sample are subjected to a hybridization reaction with the probes of the kit. The non-human sequences can be from viral, bacterial, fungal, or archaeal genomes, i.e., from the microbiome of a human. Once the human exome probes are combined with the probes for non-human microbiome, the probe mixture can be used for a hybridization-based capture reaction with nucleic acids extracted from a patient sample FIG. 2 ). Afterwards, the captured nucleic acids can be sequenced. In our lab, this sequencing is performed using an Illumina NovaSeq-6000 DNA sequencing instrument.

[0022] By aligning to human and microbiome species reference sequences, the mixed (human and non-human) DNA sequences that result from this process can be isolated. Isolating sequences by alignment is possible because the human genome has diverged significantly from microbial genomes over the course of evolution.

[0023] Because PCR is difficult to target multiple genes at the same time, using capture probes can be advantageous over polymerase chain reaction (PCR)-based assays to enrich and target sequences of interest. To target multiple genes, PCR can require a large number of primers, e.g., as many as 100,000 primers to amplify and target 50,000 sequences, and requires enzymatic manipulations to produce nucleic acid molecules for identification. It can also be difficult to optimize the ratios of PCR primers for multiple genes, and can result in biased amplification results that can not represent the relative amounts of target nucleic acids in a nucleic acid sample.

[0024] As used in the specification and claims, the singular forms “a,” “an” and “the” include plural referents unless the context clearly dictates otherwise. For example, the term “a chimeric transmembrane receptor polypeptide” includes multiple chimeric transmembrane receptor polypeptides.

[0025] The term "about" or "approximately" means within an acceptable range of error for a particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, "about" can mean within 1 or more than 1 standard deviation, per the practice in the art. Alternatively, "about" can mean ranges approximately 20%, approximately 10%, approximately 5%, or approximately 1% of a given value. Alternatively, e.g., in relation to a biological system or process, the term can mean within an order of magnitude, preferably within 5-fold and more preferably within 2-fold of a value. Where particular values described in the application and claims depend on measurements, unless otherwise stated, the term "about" should be assumed to mean within an acceptable range of error for a particular value.

[0026] As used herein, "cell" generally refers to a biological cell. A cell can be a basic structural, functional, and / or biological unit of a living organism. A cell can be derived from any organism having one or more cells. Some non-limiting examples include: a prokaryotic cell, a eukaryotic cell, a bacterial cell, an archaeal cell, a cell of a single-celled eukaryotic organism, a protozoan cell, a cell from a plant (e.g., a cell from a plant crop, a fruit, a vegetable, a grain, a soybean, a corn, a maize, a wheat, a seed, a tomato, a rice, a cassava, a sugarcane, a pumpkin, a hay, a potato, a cotton, a hemp, a tobacco, a flowering plant, a conifer, a gymnosperm, a fern, a clubmoss, a hornwort, a liverwort, a moss), an algal cell (e.g., a Botryococcus braunii, a Chlamydomonas reinhardtii, a Nannochloropsis gaditana, a Chlorella pyrenoidosa, a Sargassum patens C. Agardh, etc.), a seaweed, a fungal cell (e.g., a yeast cell, a cell from a mushroom), an animal cell, a cell from an invertebrate (e.g., a fruit fly, a cnidarian, an echinoderm, a nematode, etc.), a cell from a vertebrate (e.g., a fish, an amphibian, a reptile, a bird, a mammal), a cell from a mammal (e.g., a pig, a cow, a goat, a sheep, a rodent, a rat, a mouse, a non-human primate, a human, etc.), and the like. Sometimes a cell is not from a natural organism (e.g., a cell can be artificially synthesized, sometimes referred to as a synthetic cell).

[0027] As used herein, the term "nucleotide" generally refers to a combination of a base-sugar-phosphate. A nucleotide can comprise a synthetic nucleotide. A nucleotide can comprise a synthetic nucleotide analog. Nucleotides can be monomeric units of nucleic acid sequences (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide can include ribonucleoside triphosphates adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP), and deoxyribonucleoside triphosphates, such as dATP, dCTP, dITP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives can include, for example, [aS]dATP, 7-deaza-dGTP, and 7-deaza-dATP, as well as nucleotide derivatives that confer nuclease resistance on nucleic acid molecules comprising them. The term nucleotide as used herein can refer to a dideoxyribonucleoside triphosphate (ddNTP) and derivatives thereof. Illustrative examples of dideoxyribonucleoside triphosphates can include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. Nucleotides can be unlabeled or detectably labeled. Labeling can also be done with quantum dots. Detectable labels can include, for example, radioisotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzymatic labels. Fluorescent labels for nucleotides can include, but are not limited to, fluorescein, 5-carboxyfluorescein (FAM), 2'7'-dimethoxy-4'5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N',N'-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-(4'dimethylaminophenylazo)benzoic acid (DABCYL), Cascade Blue, Oregon Green, Texas Red, cyanine, and 5-(2'-aminoethyl)aminonaphthalene-l-sulfonic acid (EDANS).Specific examples of fluorescently labeled nucleotides can include [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP available from Perkin Elmer of Foster City, California; FluoroLink Deoxy Nucleotides, FluoroLink Cy3-dCTP, FluoroLink Cy5-dCTP, FluoroLink Fluor X-dCTP, FluoroLink Cy3-dUTP, and FluoroLink Cy5-dUTP available from Amersham of Arlington Heights, Illinois; Fluorescein 15-dATP, Fluorescein-12-dUTP, Tetramethylrhodamine 6-dUTP, IR770-9-dATP, Fluorescein-12-ddUTP, Fluorescein-12-UTP, and Fluorescein-15-2'-dATP available from Boehringer Mannheim of Indianapolis, Indiana; and Chromophore-labeled Nucleotides, BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-dUTP, Cascade Blue-7-dUTP, Fluorescein-12-UTP, Fluorescein-12-dUTP, Oregon Green 488-5-dUTP, Rhodamine Green-5-UTP, Rhodamine Green-5-dUTP, Tetramethylrhodamine-6-UTP, Tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP available from Molecular Probes of Eugene, Oregon. Nucleotides can also be labeled or tagged by chemical modification. A single nucleotide that is chemically modified can be biotin-dNTP. Some non-limiting examples of biotinylated dNTPs can include biotin-dATP (e.g., biotin-N6-ddATP, biotin-l4-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP).

[0028] As used herein, the term“genome” is generally used to refer to a portion of a subject’s genome or to the entire genome of a subject. For example, a genome can refer to a subject’s genetic sequence. A genome can refer to the entire genomic sequence of a subject.

[0029] The terms“polynucleotide,”“oligonucleotide,” and“nucleic acid” are used interchangeably and refer to a polymeric form of either deoxyribonucleotides or ribonucleotides in either single- or double-stranded form, or a polymeric form of either deoxyribonucleotides or ribonucleotides, which can contain natural, modified, or non-natural nucleotides, and which can be inter or intra stranded. A polynucleotide can be exogenous or endogenous to a cell. A polynucleotide can exist in a cell-free environment. A polynucleotide can be a gene or a fragment thereof. A polynucleotide can be DNA. A polynucleotide can be RNA. A polynucleotide can have any three-dimensional structure and can perform any function. A polynucleotide can include one or more analogs of a nucleotide (e.g., altered backbone, sugar or nucleobase). If present, modifications to the nucleotide structure can be made by either pre- or post-assembly of the polymer. Some non-limiting examples of analogs include: 5-bromouracil, peptide nucleic acids, xeno nucleic acids, morpholinos, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein attached to a sugar), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouracil, dihydrouracil, qosine, and wyosine. Non-limiting examples of polynucleotides include coding or non-coding regions of a gene or gene fragment, loci defined

[0030] As used herein, the term "gene" refers to a nucleic acid (e.g., DNA, e.g., genomic DNA and cDNA) involved in encoding an RNA transcript and its corresponding nucleotide sequence. As used herein, the term includes intervening non-coding regions as well as regulatory regions, and can include 5' and 3' ends. In certain uses, the term encompasses transcribed sequences, including 5' and 3' untranslated regions (5'-UTR and 3'-UTR), exons, and introns. In certain genes, the transcribed region will include an "open reading frame" that encodes a polypeptide. In certain uses of the term, a "gene" includes only the coding sequences (e.g., "open reading frame" or "coding region") required to encode a polypeptide. In some cases, a gene does not encode a polypeptide, e.g., ribosomal RNA genes (rRNA) and transfer RNA (tRNA) genes. In some cases, the term "gene" includes not only transcribed sequences, but also non-transcribed regions (including upstream and downstream regulatory regions), enhancers, and promoters. A gene can refer to an "endogenous gene" or native gene in its natural location in the genome of an organism. A gene can refer to an "exogenous gene" or non-native gene. A non-native gene can refer to a gene that is not normally present in the host organism, but has been introduced by gene transfer. A non-native gene can also refer to a gene that is not in its natural location in the genome of an organism, such as a genetically modified organism. A non-native gene can also refer to a naturally occurring nucleic acid or polypeptide sequence that includes a mutation, insertion, and / or deletion (e.g., a non-native sequence).

[0031] As used herein, the term "percent (%) identity" refers to the percentage of amino acid or nucleic acid residues in a candidate sequence that are identical with a reference sequence, after aligning the sequences and introducing gaps, if necessary, to achieve the maximum percent identity (e.g., gaps can be introduced in one or both of a candidate sequence and a reference sequence to achieve the best alignment, and non-identical sequences can be disregarded for comparison purposes). To determine percent identity, the alignment can be achieved in various ways, e.g., using computer software, such as the BLAST, ALIGN or Megalign (DNASTAR) software. The percent identity of two sequences can be determined by using BLAST to align the test and comparison sequences, determining the number of amino acids or nucleotides in the test sequence that are identical with the comparison sequence at the same position, and dividing the number of identical amino acids or nucleotides by the number of amino acids or nucleotides in the comparison sequence.

[0032] As used herein, the term“subject” generally refers to any animal, such as a mammal or a marsupial. The subject can be a patient. The subject can be symptomatic or asymptomatic for a disease or condition. The subject can be a primate (e.g., a human), a non-human primate (e.g., a rhesus or other type of macaque monkey), a dog, a cat, a mouse, a pig, a horse, a donkey, a cow, a sheep, a rat, and a poultry. Mammals include, but are not limited to, murines, simians, humans, farm animals, sport animals, and pets. Also included are tissues, cells, and their progeny of biological entities obtained in vivo or cultured in vitro. A host is an organism that can harbor a non-host organism. The subject can be asymptomatic for a disease. Alternatively, the subject is not asymptomatic for a disease.

[0033] As used herein, the term“treatment” refers to an approach for obtaining beneficial or desired results, including but not limited to therapeutic benefit and / or prophylactic benefit. For example, treatment can include administration of a system or population of cells disclosed herein. Therapeutic benefit refers to any therapeutically relevant improvement or effect on a disease, condition, or symptom(s) in which one or more treatments. For prophylactic benefit, a composition can be administered to a subject at risk of developing a particular disease, condition, or symptom, or to a subject reporting one or more physiological symptoms of a disease, even though the disease, condition, or symptom can not yet be present.

[0034] In some cases, the present disclosure also provides compositions and methods for processing and analyzing a biological sample. In some cases, a biological sample from a subject can comprise nucleic acid from the subject and nucleic acid molecules that are not from the subject. In some cases, the biological sample can comprise nucleic acid from human and non-human genomes. In some cases, the non-human genomes can be from viruses, bacteria, bacteriophages, fungi, protists, archaea, amoebas, worms, algae, or combinations thereof. In some cases, the source of the non-human genomes can be beneficial to the human host. In some cases, the source of the non-human genomes can be involved in modulating a disease, such as cancer. In some cases, the source of the non-human organisms comprising the non-human genomes can manifest directly by causing mutations in human cells, thereby causing cancer in the subject. In some cases, the source of the non-human genomes can manifest indirectly, for example, by stimulating the immune system of the human host to affect its ability to resist disease. In some cases, the present disclosure also provides a method comprising identifying the presence of non-human genomes and human genomes in a sample.

[0035] The sample can be from skin, saliva, nasal passages, intestinal tract, intestines, genitalia, or combinations thereof. In some cases, the sample can be obtained by swabbing, biopsy, collecting stool, collecting urine, collecting saliva, and the like.

[0036] The human genome (about 3 billion bases) is 1000 times larger than a non-human genome (e.g., a bacterial genome), and thousands of times larger than most viral genomes. In some cases, the method can include a sampling method to enrich for non-human genomes in a mixed sample.

[0037] The present disclosure also provides a method that includes performing a sequence (or sequencing) analysis of microbial nucleic acids from a sample. The sequencing analysis can include PCR amplification (e.g., of a segment of the 16S ribosomal RNA gene), as well as non-targeted metagenomic methods using deep sequencing. The method can include selection of a sample type and / or enrichment method, which can seek to minimize the amount of human genome, thereby optimizing sensitivity to non-human genomes.

[0038] In an aspect, the first plurality of nucleic acid molecules can be from a subject. The subject can be a human, and thus the present disclosure also provides human nucleic acid molecules. The present disclosure also provides use of a first subset of a plurality of nucleic acid probes or products that can be from a human genome. The present disclosure also provides a human genome. In some cases, the human genome can include, for example, inherited genetic variants, somatic variants, VDJ recombination in immune cells, differential gene expression in different cell types, and other characteristics. At least some of the foregoing can contribute to or be associated with a disease in the subject (e.g., somatic variants in cancer). The genome can comprise genes, exons, UTRs, regulatory regions, splice sites, recombination genes, alternative sequences, reassembled genes, gene phasing, exogenous sequences, and the like.

[0039] In some cases, nucleic acids in a biological sample can be analyzed by sequencing. In some cases, nucleic acids, typically from blood or diseased tissue, can be sequenced. In some cases, a blood sample can comprise peripheral blood mononuclear cells, cell-free DNA, or a combination thereof. In some cases, the diseased tissue can comprise a cancer. In some cases, sequence analysis of nucleic acids from a blood sample or a diseased tissue sample can include amplicon panel generation, hybrid capture, and non-targeted methods, such as whole genome sequencing. In some cases, the method can involve selection of a sample type, a human genome enrichment method, and combinations thereof.

[0040] In some cases, a biological sample can comprise subject nucleic acids, non-subject nucleic acids, and combinations thereof. In some cases, a biological sample can comprise host nucleic acids, non-host nucleic acids, and combinations thereof. In some cases, a biological sample can comprise a genome encoding a receptor. In some cases, the receptor can be from an immune cell. In some cases, the receptor from an immune cell can be a T cell receptor (TCR), a B cell receptor (BCR), a chimeric antigen receptor (CAR), and the like.

[0041] In some cases, the methods provided herein include assaying nucleic acid molecules to generate sequence information. The assay can include sequencing nucleic acids comprising VDJ rearrangements or VDJ recombination. In some cases, VDJ rearrangements or recombination can refer to cell receptors. For example, somatic hypermutation can generate B cell receptor (BCR) sequences encoding antibodies to achieve significant antigenic diversity. In some cases, BCRs can be sequenced. The methods provided herein can include sequencing BCRs to elucidate how antibodies develop. For example, the methods can include sequence analysis to annotate each base as coming from a particular gene in the V, D, or J genes, or from N addition (also known as non-templated insertion). In some cases, VDJ recombination can generate CDR3 sequences. In some cases, CDR3 sequences can generate polypeptides that bind to antigens, e.g., tumor antigens. In some cases, CDR3 sequences can generate polypeptides that bind to antigens, e.g., neoantigens.

[0042] In some cases, the methods provided herein can comprise sequencing a nucleic acid encoding a candidate tumor neoantigen or associated with a candidate tumor neoantigen. In some cases, the methods provided herein can comprise sequencing a detected CDR3 sequence. In some aspects, a variety of methods can be used to identify a subject TCR. In some cases, a TCR can be identified using whole-exomic sequencing. For example, a TCR can target a neoantigen or neoepitope identified by whole-exomic sequencing of a target cell. Alternatively, a TCR can be identified from an autologous, allogeneic, or xenogeneic repertoire. In some cases, a gene that can comprise a mutation that produces a neoantigen or neoepitope can be ABL1, ACOl 1997, ACVR2A, AFP, AKT1, ALK, ALPPL2, ANAPC1, APC, ARID1A, AR, AR-v7, ASCL2, β2M, BRAF, BTK, C15ORF40, CDH1, CLDN6, CNOT1, CT45A5, CTAG1B, DCT, DKK4, EEF1B2, EEF1DP3, EGFR, EIF2B3, env, EPHB2, ERBB3, ESR1, ESRP1, FAM11 IB, FGFR3, FRG1B, GAGE1, GAGE 10, GATA3, GBP3, HER2, IDH1, JAK1, KIT, KRAS, LMAN1, MABEB 16, MAGEA1, MAGEA10, MAGEA4, MAGEA8, MAGEB 17, MAGEB4, MAGEC1, MEK, MLANA, MLL2, MMP13, MSH3, MSH6, MYC, NDUFC2, NRAS, NY-ESO, PAGE2, PAGE5, PDGFRa, PIK3CA, PMEL, pol protein, POLE, PTEN, RAC1, RBM27, RNF43, RPL22, RUNX1, SEC31A, SEC63, SF3B 1, SLC35F5, SLC45A2, SMAP1, SMAP1, SPOP, TFAM, TGFBR2, THAP5, TP53, TTK, TYR, UBR5, VHL, XPOT.

[0043] The present disclosure also provides methods comprising obtaining or providing a subset of nucleic acid samples or nucleic acid molecules comprising one or more genomes. The methods disclosed herein can analyze a subset of nucleic acids from nucleic acid molecules produced from a biological sample. The subset can comprise nucleic acid molecules from a subject and nucleic acid molecules that are not from the subject. For example, human and microbial sequences can be enriched using target capture and sequencing. The one or more genomes can comprise one or more genomic features.

[0044] A genomic feature can include an entire genome or a portion thereof. A genomic feature can include an entire exome or a portion thereof. A genomic feature can include one or more sets of genes. A genomic feature can include one or more genes. A genomic feature can include one or more sets of regulatory elements. A genomic feature can include one or more regulatory elements. A genomic feature can include one or more sets of polymorphisms. A genomic feature can include one or more polymorphisms. In some cases, a polymorphism refers to a mutation in a genotype. A polymorphism can include a change of one or more bases, an insertion of one or more bases, a duplication or deletion of one or more bases. A genomic feature can include copy number variants (CNVs), transversions, other rearrangements, and other forms of genetic variation. In some cases, one or more features of a subset of nucleic acid samples can be a polymorphism marker, including a restriction fragment length polymorphism, a variable number of tandem repeat (VNTR), a hypervariable region, a minisatellite sequence, a dinucleotide repeat, a trinucleotide repeat, a tetranucleotide repeat, a simple sequence repeat, and an insertion element (e.g., Alu). In some cases, a difference between a first subset of nucleic acid molecules and a second subset of nucleic acid molecules can be a polymorphism marker, including a restriction fragment length polymorphism, a variable number of tandem repeat (VNTR), a hypervariable region, a minisatellite sequence, a dinucleotide repeat, a trinucleotide repeat, a tetranucleotide repeat, a simple sequence repeat, and an insertion element (e.g., Alu). The most frequently occurring allele form in a selected population is sometimes referred to as the wild-type form. A diploid organism can be homozygous or heterozygous for an allele form. A biallelic polymorphism has two forms. A triallelic polymorphism has three forms. A polymorphism can include a single nucleotide polymorphism (SNP). In some aspects of the disclosure, one or more polymorphisms include one or more single nucleotide variations, inDels, small insertions, small deletions, structural variant junctions, variable length tandem repeats, flanking sequences, or combinations thereof. One or more polymorphisms can be located within, around, or near a coding and / or non-coding region. One or more polymorphisms can span at least a portion of a gene, exon, intron, untranslated region. In some cases, a genomic feature can be related to one or more GC content, complexity, and / or mappability of nucleic acid molecules. A genomic feature can include one or more simple tandem repeats (STRs), unstable expanded repeat sequences, segmental duplications, denatured mapping scores of single and paired reads, GRCh38 or GRCh37 patches, or combinations thereof. A genomic feature can include one or more low average coverage regions from whole genome sequencing (WGS), zero average coverage regions from WGS, verified compaction, or combinations thereof. A genomic feature can include one or more alternative or non-reference sequences.

[0045] Genomic features can include one or more genetic phenotyping and reassembled genes. Examples of phenotyping and reassembling genes include, but are not limited to, one or more major histocompatibility complex, blood group, and amylase gene families. In some cases, genetic phenotyping and / or reassembled genes can include genes related to blood group. Blood group genes can include ABO, RHD, RHCE, or combinations thereof.

[0046] In some cases, genomic features can include reassembled genes. Reassembled genes can include genes involved in immune response. Genes involved in immune response can include genes related to major histocompatibility complex, immune receptors, and cellular functions. One or more major histocompatibility complex can include one or more HLA class I, HLA class II, or combinations thereof. HLA class I can be any of HLA-A, HLA-B, HLA-C, or combinations thereof. HLA class II can be any of HLA-DP, HLA-DM, HLA-DOA, HLA-DOB, HLA-DQ, HLA-DR, or combinations thereof. Genes involved in immune response can be RAG1, RAG2, and combinations thereof. In some cases, reassembled genes can include genes involved in VDJ recombination. For example, to establish diversity in B cell and T cell receptors (BCR and TCR), genes can be generated by recombining pre-existing gene segments. In some cases, different combinations of a limited set of gene segments can yield a repertoire capable of identifying an unlimited number of foreign or non-human genomes. VDJ recombination can include cleaving DNA containing recombination signal sequences (RSS). In some cases, fragmented sequences can be reassembled using cellular repair mechanisms. The present disclosure also provides methods including sequencing fragmented portions of genomes from VDJ recombination. For example, the present disclosure provides methods including sequencing at VDJ recombination sites.

[0047] In some aspects of the present disclosure, one or more genomic features can not be mutually exclusive. For example, a genomic feature comprising the entire genome or a portion thereof can overlap with another genomic feature (e.g., the entire exome or a portion thereof, one or more genes, one or more regulatory elements, etc.). Alternatively, or additionally, one or more genomic features can be mutually exclusive. For example, a genome comprising the non-coding portion of the entire genome can not overlap with a genomic feature (e.g., the coding portion of the exome or a portion thereof or a gene). Alternatively, or additionally, one or more genomic features are partially exclusive or partially inclusive. For example, a genome comprising the entire exome or a portion thereof can overlap with a genome portion comprising the exon portion of a gene. However, a genome comprising the entire exome or a portion thereof can not overlap with a genome comprising the intron portion of a gene. Thus, a genomic feature comprising a gene or a portion thereof can be partially exclusive and / or partially inclusive of a genomic feature comprising the entire exome or a portion thereof. In some cases, a genetic feature can be associated with a species, such that more than one species can be distinguished. In some cases, a genetic feature can be associated with a species, such that a human genetic feature and a bacterial genetic feature can be distinguished. In some cases, a first subset of nucleic acid molecules is specific to a first species and a second subset of nucleic acid molecules is specific to a second species.

[0048] A biological sample can comprise a human genome, a non-human genome, or a combination thereof. In some cases, a biological sample can be processed. In some cases, a biological sample can be enriched for nucleic acid sequences of a subject (e.g., a human), or nucleic acid sequences that are not from a subject (e.g., a non-human) can be enriched, while both human and non-human genomes in the sample are detected. A biological sample can comprise cell-free nucleic acid molecules, such as cell-free DNA (cfDNA) or cell-free RNA. Cell-free nucleic acid molecules can be circulating tumor nucleic acid molecules (e.g., circulating tumor DNA). Cell-free DNA can comprise mutations that can be indicative of, associated with, or linked to a disease, such as a cancer.

[0049] A biological sample can have nucleic acid molecules comprising engineered sequences. For example, nucleic acids in a biological sample can comprise exogenous or replacement sequences (e.g., tags), exogenous receptors (e.g., chimeric antigen receptor (CAR) receptors), plasmid sequences, and neoantigen-specific sequences, to name a few. In some cases, engineered sequences can be used as diagnostic markers. In some cases, engineered sequences can be utilized to determine whether a therapeutic agent is transported to a target, such as a tumor target. Replacement sequences can be exogenous or endogenous. In some cases, replacement sequences can include or be from plasmid sequences. Plasmid sequences can be DNA or RNA. In some cases, plasmid sequences can also be DNA minicircle sequences or dogbone sequences.

[0050] The present disclosure also provides methods comprising obtaining or providing a subset of nucleic acid samples or nucleic acid molecules comprising one or more transcriptomes. The nucleic acid samples can comprise mRNA. The mRNA can be from different tissues of a subject. The amount of mRNA in the sample can be used to analyze the expression level of mRNA or protein in the tissue and specific kinds of the subject. For example, the amount of mRNA in the sample can be related to a particular characteristic or disease (e.g., cancer) in the subject. Additionally, the amount of mRNA in the sample can be indicative of a change or relative difference in expression in a particular tissue type in the subject. In some cases, the mRNA can be processed into cDNA. For example, the mRNA can be reverse transcribed using reverse transcriptase to synthesize cDNA molecules. The cDNA molecules can be captured, isolated, enriched, amplified, sequenced, or other reactions can be performed on the nucleic acids as described elsewhere herein.

[0051] Nucleic acids from a biological sample comprising sequences from a subject and non-subject can be enriched and sequenced in a single pool and separated by computer. For example, a mixed human genome and non-human genome can be separated by alignment to human and non-human species reference sequences. Separating sequences by alignment is possible because the human genome has diverged from microbial genomes over the course of evolution. Genomes, for example DNA from non-human genomes (e.g., microbial species), generally do not align to the human genome and vice versa. Non-human genomes, for example microbial genomes, can have segments that are identical or similar to other non-human genomes. More than one non-human genome in a mixture of human and non-human genomes can require additional alignment to identify the non-human species.

[0052] In some cases, the method can comprise producing a subset of nucleic acid molecules from the biological sample. In some cases, the method can comprise producing a subset of nucleic acid molecules from the biological sample by performing one or more hybridization reactions. The hybridization reactions can comprise enrichment. In some cases, enrichment can be performed. Enrichment can be performed by various methods. In some cases, enrichment can be performed by hybrid capture, array capture, bead capture, and the like. In some cases, hybrid capture can be in solution or on a solid support, such as on an array. In some cases, enrichment can be performed by molecular inversion probes (MIPs). Enrichment can be performed by amplification, such as using PCR. In some cases, the method can comprise producing a subset of nucleic acids by amplifying the human genome or non-human genome for the purpose of enrichment. In some cases, amplification can comprise one of the following: polymerase chain reaction (PCR)-based techniques (e.g., solid-phase PCR, RT-PCR, qPCR, multiplex PCR, touchdown PCR, nano-PCR, nested PCR, hot-start PCR, and the like), helicase-dependent amplification (HDA), loop-mediated isothermal amplification (LAMP), self-sustained sequence replication (3SR), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), rolling circle amplification (RCA), ligase chain reaction (LCR), and any other suitable amplification technique.

[0053] The nucleic acid probe pool can comprise a human exome capture kit, such as Agilent Clinical Research Exome v2 (Agilent Technologies, Santa Clara, CA) FIG. 1 The nucleic acid probe pool can comprise hybridization probes complementary to the human sequences they target. The nucleic acid probe pool can comprise about 50,000 capture probes to target the genomes of most predetermined human genes, and the like. The capture probes can be designed to target genes, exons, UTRs, regulatory regions, splice sites, recombined genes, alternative sequences, and other genomic contents. In some cases, the method can comprise a nucleic acid probe pool designed to target non-human sequences. In some cases, the nucleic acid probe pool can be specific to non-human genomes, such as from viruses, bacteria, fungi, or archaea (i.e., from the human microbiome). In some cases, the nucleic acid probe pool can be combined with a second nucleic acid probe pool specific to a second species as compared to the first nucleic acid probe pool FIG. 2 The nucleic acid probe pool can be configured to bind human sequences and non-human sequences, and the nucleic acid probe pool can bind sequences from a fragmented transcriptome FIG. 3 In some cases, the nucleic acid probe pool can bind human sequences, and the nucleic acid probe pool can bind non-human sequences FIG. 4). In some cases, a pool of nucleic acid probes can be used for a hybridization-based capture reaction with nucleic acids extracted from a patient sample. In some cases, the method can include sequencing the pool of nucleic acids that have been captured. The captured or enriched nucleic acids can be human sequences, non-human sequences, or human and non-human sequences, and sequencing can include an Illumina NovaSeq-6000 DNA sequencer. In some cases, capture probes targeting microbial species can be designed to target regions of one or more species-distinct microbial species sequences, or directly adjacent to one or more species-distinct regions, such as non-human or human. In some cases, the method includes targeting different regions between non-human sequences, such as microbial sequences, enabling capture of nucleic acids from a large number of potential non-human microbiome species using a small number of capture probes. By using capture probes that can contain a region of consensus sequence but are adjacent to a variable region, many non-human sequences that can be captured can span both.

[0054] In some cases, sequences from non-human genomes can then be assigned to their source species by one or more species-distinct regions. For example, the 16S ribosomal RNA gene, which is present in almost all bacteria, has approximately 9 regions in which the sequence varies by species, interspersed with regions that share sequence. In some cases, capture molecules that extend from these shared regions into the variable regions can then be assigned to their source species based on the sequence from the variable region portion. Fungal nucleic acid sequences can be similarly assessed by using the partially conserved D2 region of the large subunit ribosomal RNA gene of fungal genomes. Exomes of genomes can be analyzed. Intronic regions of genomes can be analyzed. Exomes primarily target the coding regions of the human genome and can comprise less than 2% of the total human genome. By excluding most of the intronic and intergenic portions of the human genome, the number of human sequences can be reduced by about 98%. Exons can be added to include non-coding content. Sequencing of exomes, for example in cancer, can allow for deep sequencing, improving detection of somatic variants with low allele frequencies, and can also improve detection of non-human sequences that are co-captured from a sample.

[0055] The present disclosure also provides compositions and methods for processing a biological sample. A biological sample can be obtained from a subject, such as an adult or a child. In some cases, a method for processing a biological sample can comprise (a) generating a subset of nucleic acid molecules from a biological sample using a pool of nucleic acid probes, wherein the probes comprise (i) a first plurality of nucleic acid probes configured to target elements of a human genome; and (ii) a second plurality of nucleic acid probes configured to target elements of one or more non-human genomes; (b) performing an assay on the subset of nucleic acid molecules to generate sequence information comprising (i) human nucleic acids from the biological sample of the subject and (ii) non-human nucleic acids from the biological sample of the subject.

[0056] The disclosed methods may include detecting, monitoring, quantifying, or evaluating one or more non-human nucleic acid molecules or one or more diseases or conditions caused by one or more non-human genomes or non-host genomes. In some cases, the capture probe may target different genera. In some cases, the capture probe may target different species. In some aspects, the capture probe may target different orders of more than one organism. In some cases, the capture probe may target the plant kingdom, animal kingdom, fungi, protists, eubacteria, and / or archaea. In some cases, the capture probe may target viruses, bacteria, bacterial phages, fungi, protists, archaea, amoebas, worms, algae, genetically modified cells, genetically modified vectors, and combinations thereof. In some cases, the non-human sequence may be bacterial. For example, bacterial sequences can originate from acidiobacteria, actiniobacteria, aquagenic bacteria, armatimonadetes, Bacteroidetes, caldiserica, chlamydiae, chlorobacteria, chloroflexi, chrysiogenetes, cyanobacteria, deferribacteres, deinococcus-thermus, dictyoglomi, and elusimicrobia. Phylum: Fibrobacteres, Firmicutes, Fusobacteria, Gemmatimonadetes, Lentisphaerae, Nitrospirae, Planctomycetes, Proteobacteria, Spirochaetaes, Synergistetes, Tenericutes, Thermodesulfobacteria, Thermomicrobia, Thermotogae, and / or Verrucomicrobia. In some cases, non-human sequences may originate from, but are not limited to, *Bordetella* genus. Bordetella ), genus *Borrelia* Borrelia Brucella ( ) Brucella ), Campylobacter spp.Campylobacter Chlamydia ( ) Chlamydia Chlamydia genus ( Chlamydophila Clostridium ( Clostridium Corynebacterium spp. Corynebacterium ), Enterococcus spp. Enterococcus ), Escherichia coli ( Escherichia ), Francisella genus ( Francisella Haemophilus spp. Haemophilus ), Helicobacter spp. Helicobacter Legionella ( ) Legionella Leptospira ( ) Leptospira Listeria ( ) Listeria ), Mycobacterium ( Mycobacterium Mycoplasma genus Mycoplasma ), Neisseria ( Neisseria ), Pseudomonas spp. Pseudomonas ), Rickettsia spp. ( Rickettsia Salmonella ( Salmonella ), Shigella spp. Shigella Staphylococcus spp. Staphylococcus Streptococcus ( Streptococc us), Treponema genus (Tre ponema ), Vibrio genus ( Vibrio ) or Yersinia spp. ( Yersinia Other pathogens include, but are not limited to, Mycobacterium tuberculosis (…). Mycobacterium tuberculosisStreptococcus, Pseudomonas, Shigella, Campylobacter, and Salmonella. In some cases, the capture probe can target a fungus, such as blastocladiomycota, chytridiomycota, Glomeromycota, Microsporidia, Neocallimastigomycota, Deuteromycota, Ascomycota, Pezizomycotina, Saccharomycotina, Taphrinomycotina, Basidiomycota, Agaricomycotina, Pucciniomycotina, Ustilaginomycotina, Entomophthoromycotina, Kickxellomycotina, Mucoromycotina, Zoopagomycotina, and the like. In some cases, the non-human or non-host can be from a cow, a horse, a fish, a donkey, a rabbit, a rat, a mouse, a hamster, a dog, a cat, a pig, a snake, a sheep, a goat, and the like.

[0057] The disease or condition caused by or associated with one or more non-human genomes can include tuberculosis, pneumonia, foodborne illness, tetanus, typhoid fever, diphtheria, syphilis, leprosy, bacterial vaginosis, bacterial meningitis, bacterial pneumonia, urinary tract infection, bacterial gastroenteritis, bacterial skin infection, or any combination thereof. Examples of bacterial skin infections include, but are not limited to, impetigo, which can be caused by Staphylococcus aureus Staphylococcus aureus ) or Streptococcus pyogenes Streptococcus pyogenes ) and cellulitis, which can be caused by normal skin flora or exogenous bacteria.

[0058] The non-subject nucleic acid sequence can be derived from a fungus, such as Candida Candida ), Aspergillus Aspergillus ), Cryptococcus Cryptococcus ), Histoplasma Histoplasma ), Pneumocystis Pneumocystis ), and Stachybotrys Stachybotrys ). Examples of diseases or conditions caused by fungi include, but are not limited to, tinea cruris, yeast infection, tinea versicolor, and athlete's foot.

[0059] In some cases, the non-subject or non-host nucleic acid sequence can be from a protist. Protists can include protists, protophytes, molds, and combinations thereof. The protist can be a primary chromist organism. The primary chromist organism can be Rhodophyta or Glaucophyta. The protist can be a Sar or Harosa. The SAR can be a clade that includes stramenopiles, alveolates, and Rhizaria (SAR). Additionally, the clade SAR can include stramenopiles, Alveolata, Apicomplexa, Ciliophora, Dinoflagellata, Rhizaria, Cercozoa, Foraminifera, Radiolaria, and combinations thereof. In some cases, the protist can be Excavata. The Excavata can be Euglenozoa, Percolozoa, Metamonada, and combinations thereof. In some cases, the non-host can be Amoebozoa, Hacrobia, Apusozoa, Opisthokonta, and / or Choanozoa.

[0060] The non-subject nucleic acid sequence can be derived from a virus. Examples of viruses include, but are not limited to, adenovirus, coxsackievirus, Epstein-Barr virus, hepatitis virus (e.g., hepatitis A, B, and C), herpes simplex virus (type 1 and 2), cytomegalovirus, herpesvirus, HIV, influenza virus, measles virus, mumps virus, papillomavirus, parainfluenza virus, poliovirus, respiratory syncytial virus, rubella virus, and varicella-zoster virus. Examples of diseases or conditions caused by viruses include, but are not limited to, the common cold, influenza, hepatitis, AIDS, chickenpox, rubella, mumps, measles, warts, and poliomyelitis.

[0061] The non-subject nucleic acid can be derived from a protozoan, such as an Acanthamoeba (e.g., Acanthamoeba astronyxis (A. astronyxis) Acanthamoeba (A. castellanii) A. astronyxis (A. keriothrix) A. castellanii (A. polyphaga) A. (A. qudrafolia) culbertsoni (A. hartmannella) A. hatchetti (A. limax) A. polyphaga (A. spindola) A. rhysodes (A. virida) A. healyi ,A. divionensis ), Balantidium (e.g., Balantidium coli), Brachiola ), Cryptosporidium (e.g., Cryptosporidium parvum), B. connori ), Cryptosporidium (e.g., Cryptosporidium parvum), B. vesicularum ), Cryptosporidium (e.g., Cryptosporidium parvum), Cryptosporidium ), Cryptosporidium (e.g., Cryptosporidium parvum), C. parvum ), Cryptosporidium (e.g., Cryptosporidium parvum), Cyclospora ), Cryptosporidium (e.g., Cryptosporidium parvum), C. cayetanensis ), Cryptosporidium (e.g., Cryptosporidium parvum), Encephalitozoon ), Cryptosporidium (e.g., Cryptosporidium parvum), E. cuniculi ), Cryptosporidium (e.g., Cryptosporidium parvum), E. hellem ), Cryptosporidium (e.g., Cryptosporidium parvum), E. intestinalis ), Cryptosporidium (e.g., Cryptosporidium parvum), Entamoeba ), Cryptosporidium (e.g., Cryptosporidium parvum), E. histolytica ), Cryptosporidium (e.g., Cryptosporidium parvum), Enterocytozoon ), Cryptosporidium (e.g., Cryptosporidium parvum), E. bieneusi ), Cryptosporidium (e.g., Cryptosporidium parvum), Giardia ), Cryptosporidium (e.g., Cryptosporidium parvum), G. lamblia ), Cryptosporidium (e.g., Cryptosporidium parvum), Isospora ), Cryptosporidium (e.g., Cryptosporidium parvum), I. belli ), Cryptosporidium (e.g., Cryptosporidium parvum), Microsporidium ), Cryptosporidium (e.g., Cryptosporidium parvum), M. africanum ), Cryptosporidium (e.g., Cryptosporidium parvum), M. ceylonensis ), Cryptosporidium (e.g., Cryptosporidium parvum), Naegleria ), Cryptosporidium (e.g., Cryptosporidium parvum), N. fowleri ), Cryptosporidium (e.g., Cryptosporidium parvum), Nosema ), Cryptosporidium (e.g., Cryptosporidium parvum), N. algerae ), Cryptosporidium (e.g., Cryptosporidium parvum), N. ocularum ), Cryptosporidium (e.g., Cryptosporidium parvum), Pleistophora ), Cryptosporidium (e.g., Cryptosporidium parvum), Trachipleistophora ), Cryptosporidium (e.g., Cryptosporidium parvum), T. anthropophthera ), Cryptosporidium (e.g., Cryptosporidium parvum), T. hominis ), Cryptosporidium (e.g., Cryptosporidium parvum), Vittaforma ), Cryptosporidium (e.g., Cryptosporidium parvum), V. corneae ), Cryptosporidium (e.g., Cryptosporidium parvum).

[0062] Nucleic acids can be extracted and / or isolated from a biological sample of a subject, for example, by performing a separation of a cell fraction. In variations, sample processing can thus include one or more of: lysing the sample, disrupting membranes in cells of the sample, isolating unwanted elements (e.g., RNA, proteins) from the sample, purifying nucleic acids (e.g., DNA) in the sample to produce a nucleic acid sample comprising non-human microbiome nucleic acid content and human genomic nucleic acid content of the sample, amplifying nucleic acids from the nucleic acid sample, further purifying the amplified nucleic acids of the nucleic acid sample, sequencing the amplified nucleic acids of the nucleic acid sample, and any combination thereof. In variations, lysing the sample and / or disrupting membranes in cells of the sample can include physical methods of cell lysis / membrane disruption (e.g., bead beating, nitrogen decompression, homogenization, sonication) that omit certain reagents that can introduce bias in representation of certain microbial species at the time of sequencing. Additionally or alternatively, lysing or disruption can involve chemical methods (e.g., use of detergents, use of solvents, use of surfactants, etc.).

[0063] In variations, isolating unwanted elements from the sample can include removing RNA using RNase and / or removing proteins using proteases. In variations, purifying nucleic acids in the sample to produce a nucleic acid sample can include one or more of: precipitating nucleic acids from the biological sample (e.g., using alcohol-based precipitation methods); liquid-liquid based purification techniques (e.g., phenol-chloroform extraction); chromatography-based purification techniques (e.g., column adsorption); purification techniques involving use of particles (e.g., magnetic beads, buoyant beads, beads with a size distribution, ultrasound-responsive beads, etc.) configured to bind nucleic acids and configured to release nucleic acids bound by the binding moiety in the presence of an elution environment (e.g., with an elution solution, providing a change in pH, providing a change in temperature, etc.), and any other suitable purification techniques.

[0064] Nucleic acids can be extracted and / or isolated from a biological sample so that the extraction and isolation and / or isolation can be performed in an environment that is sterilized of any contaminating substances (e.g., substances that can affect nucleic acids in the sample or can contribute to contaminating nucleic acids) (e.g., a sterile laboratory fume hood, a sterile room), can control environmental temperature, control oxygen content, control carbon dioxide content, and / or control light exposure (e.g., exposure to ultraviolet light). Extraction can include lysis to disrupt cell membranes and facilitate release of nucleic acids from cells in the biological sample. In one non-limiting example, lysis can include a bead mill apparatus (e.g., a tissue lyser) configured to be used with beads that are mixed with the sample and used to agitate the biological contents of the sample. In some cases, processing of the biological sample can include a combination of one or more of: a lysis reagent (e.g., a protease), a heating module, and any other suitable lysis device.

[0065] To isolate nucleic acids from a lysed sample, the non-nucleic acid content of the sample is separated from the nucleic acid content of the sample. Purification modules of the sample processing method can include force-based separations, size-based separations, binding moiety-based separations (e.g., with magnetic binding moieties, with buoyant binding moieties, etc.), and / or any other suitable form of separation. For example, purification operations of the method can include one or more of a centrifuge to facilitate supernatant extraction, a filter (e.g., a filter plate), a fluidic transport module configured to bind lysed sample to a moiety that binds nucleic acid content and / or sample waste, a wash reagent delivery system, an elution reagent delivery system, and any other suitable equipment for purifying nucleic acid content in a sample.

[0066] A subset of nucleic acid molecules can be subjected to an assay to generate sequence information. Assays to generate sequence information can induce a sequencing reaction. In some cases, sequencing can be directed to RNA. For example, sequencing can be directed to RNA transcription. RNA sequencing can include any of chromatin isolation by RNA purification (ChlRP-Seq), global run-on sequencing (GRO-Seq), ribosome profiling sequencing (Ribo-Seq) / ARTseq™, RNA immunoprecipitation sequencing (RIP-Seq), high-throughput sequencing of CLIP cDNA libraries (HITS-CLIP), cross-linking and immunoprecipitation sequencing (CLIP-Seq), photoactivatable ribonucleoside-enhanced cross-linking and immunoprecipitation (PAR-CLIP), single-nucleotide resolution CLIP (iCLIP), sequencing of natural elongation transcripts (NET-Seq), targeted purification of polyribosome mRNA (TRAP-Seq), cross-linking, ligation, and sequencing of hybrids (CLASH-Seq), parallel analysis of RNA ends sequencing (PARE-Seq), genome-wide mapping of uncapped transcripts (GMUCT), transcript isoform sequencing (TIF-Seq), paired-end analysis of TSS (PEAT), and any combination thereof. In some cases, sequencing can comprise RNA structure. Sequencing of RNA structure can include any of selective 2'-hydroxyl acylation analyzed by primer extension sequencing (SHAPE-Seq), parallel analysis of RNA structure (PARS-Seq), fragmentation sequencing (FRAG-Seq), CXXC affinity purification sequencing (CAP-Seq), calf intestinal alkaline phosphatase-tobacco acid pyrophosphatase sequencing (CIP-TAP), inosine chemical erasure sequencing (ICE), m6A-specific methylated RNA immunoprecipitation sequencing (MeRIP-Seq), and any combination thereof. In some cases, sequencing can include low-level RNA detection. Low-level RNA detection can include digital RNA sequencing, single-cell full-transcript amplification (Quartz-Seq), primer-based RNA sequencing by design (DP-Seq), conversion mechanism of the 5' end of RNA templates version 2 (Smart-Seq2), unique molecular identifiers (UMI), cell expression by linear amplification sequencing (CEL-Seq), single-cell labeled reverse transcription sequencing (STRT-Seq), and any combination thereof. In some cases, sequencing can be directed to DNA. DNA sequencing can include low-level DNA detection.DNA sequencing, including low-level DNA detection, can include at least one of single molecule molecular inversion probes (smMIP), multiple displacement amplification (MDA), multiple annealing and looping-based amplification cycles (MALBAC), oligonucleotide selection sequencing (OS-Seq), duplex sequencing (Duplex-Seq), and any combination thereof. In some aspects, sequencing can include DNA methylation. DNA methylation can include at least one of bisulfite sequencing (BS-Seq), post-bisulfite adaptor tagging (PBAT), tagmentation-based whole-genome bisulfite sequencing (T-WGBS), oxidative bisulfite sequencing (oxBS-Seq), Tet-assisted bisulfite sequencing (TAB-Seq), methylation DNA immunoprecipitation sequencing (MeDIP-Seq), MethylCap sequencing, methyl binding domain capture (MBDCap) sequencing, reduced representation bisulfite sequencing (RRBS-Seq), and combinations thereof. In some cases, sequencing can include DNA-protein interaction. For example, sequencing including DNA-protein interaction can include: DNase1 hypersensitive site sequencing (DNase-Seq), MNase-assisted nucleosome isolation sequencing (MAINE-Seq), chromatin immunoprecipitation sequencing (ChIP-Seq), formaldehyde-assisted regulation element isolation (FAIRE-Seq), assay for transposase accessible chromatin sequencing (ATAC-Seq), chromatin interaction analysis with paired-end tags sequencing (ChlA-PET), chromatin conformation capture (Hi-C / 3C-Seq), circular chromatin conformation capture (4-C or 4C-Seq), chromatin conformation capture carbon copy (5-C), and combinations thereof. In some cases, sequencing can include rearrangement. Sequencing of sequence rearrangement can include at least one of retrotransposon capture sequencing (RC-Seq), transposon sequencing (Tn-Seq) or insertion sequencing (INSeq), translocation capture sequencing (TC-Seq), and combinations thereof.

[0067] Sequencing analysis can include PCR amplification, such as of a segment of the 16S ribosomal RNA gene, and a non-targeted metagenomic approach using deep sequencing. In some embodiments, the method can include a method for next-generation amplification and sequencing, including: simultaneously amplifying the entire 16S region of each of a set of microorganisms, fragmenting amplicon of the entire 16S region of each of the set of microorganisms to produce a set of amplicon fragments, and generating an analysis based on the set of amplicon fragments, wherein the analysis includes at least one of microorganism population characteristics, microorganism species identification, and identification of target microorganism sequences. In some cases, whole exome sequencing can be utilized.

[0068] In some cases, the method can include performing the alignment at a genetic level of the non-human genome. For example, the alignment can include aligning the 16S sequence relative to an 18S sequence, relative to an ITS sequence, and / or the like. Thus, the output can be used to identify features of interest that can be used to characterize the microbiome of the biological sample, where the features can be non-human (e.g., presence of a genus of bacteria), genetic-based (e.g., based on the performance of a particular genetic region and / or sequence), and / or based on any other suitable level.

[0069] In variations, alignment and mapping to a reference non-human genome (e.g., a bacterial genome) (e.g., provided by the National Center for Biotechnology Information) can be performed using an alignment algorithm that includes one or more of the following: a Needleman-Wunsch algorithm that performs a global alignment of two reads (e.g., a sequencing read and a reference read) based on a scoring of the global alignment (e.g., in terms of insertions, deletions, matches, mismatches) with a termination condition; a Smith-waterman algorithm that performs a local alignment of two reads (e.g., a sequencing read and a reference read) and scores the local alignment (e.g., in terms of insertions, deletions, matches, mismatches); a Basic Local Alignment Search Tool (BLAST) that identifies regions of local similarity between sequences (e.g., a sequencing read and a reference read); an FPGA-accelerated alignment tool; a BWT index using the BWA tool; a BWT index using the SOAP tool; a BWT index using the Bowtie tool; a sequence search and alignment by hashing algorithm (SSAHA2) that uses word hashing and dynamic programming to map nucleic acid sequencing reads to a genomic reference sequence; and any other suitable alignment algorithm. Mapping of unidentifiable sequences can also include mapping to a reference viral genome and / or a fungal genome in order to further identify viral and / or fungal components of an individual’s microbiome. For example, PCR can be performed with multiple markers (e.g., a first marker, a second marker, a third marker, an Nth marker) in parallel or in series and associated with one or more of a bacterial marker, a fungal marker, and a eukaryotic marker. Further, overlapping reads (e.g., reads resulting from paired-end sequencing) can be assembled based on the output of the alignment algorithm, or aligned sequence reads can be merged with a reference sequence (e.g., using a hidden Markov model banding technique, using a Durbin-Holmes technique). However, alignment and mapping can implement any other suitable algorithm or technique. In some cases, sequence reads can be encoded to facilitate alignment and mapping operations. In one example, each base of a sequence can be encoded as a byte according to the permutation 0000TGCA, where the least significant bit is 1 if the base is sequenced as possibly containing base A (e.g., A is represented as 00000001); the next significant bit is 1 if the base is sequenced as possibly containing base C (e.g., C is represented as 00000010); the next significant bit is 1 if the base is sequenced as possibly containing base G (e.g., G is represented as 00000100); and the next significant bit is 1 if the base is sequenced as possibly containing base T (e.g., T is represented as 00001000). In this example, the four most significant bits are set to zero.However, alternative variations of the example can encode the bases in any other suitable manner. Furthermore, the predetermined primer sequences used during amplification can be used to trim the sequence reads to omit the primer sequences to improve efficiency of alignment and mapping.

[0070] A subset of nucleic acid molecules can comprise one or more genomes disclosed herein. A subset of nucleic acid molecules can comprise 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 20 or more, 25 or more, 30 or more, 35 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, or 100 or more genomes. The one or more genomes can be the same, similar, different, or a combination thereof. In some cases, there are two subsets of nucleic acid molecules, FIG. 6A .

[0071] A subset of nucleic acid molecules can comprise one or more genome features disclosed herein. A subset of nucleic acid molecules can comprise 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 20 or more, 25 or more, 30 or more, 35 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, or 100 or more genome features. The one or more genome features can be the same, similar, different, or a combination thereof.

[0072] A subset of nucleic acid molecules can comprise nucleic acid molecules of different sizes. The length of a nucleic acid molecule in a subset of nucleic acid molecules can be referred to as the size of the nucleic acid molecule. The average length of nucleic acid molecules in a subset of nucleic acid molecules can be referred to as the average size of the nucleic acid molecules. As used herein, the terms "size of a nucleic acid molecule," "average size of a nucleic acid molecule," "molecular size," and "average molecular size" can be used interchangeably. The size of a nucleic acid molecule can be used to distinguish between two or more subsets of nucleic acid molecules. The difference in the average size of nucleic acid molecules in a subset of nucleic acid molecules and the average size of nucleic acid molecules in another subset of nucleic acid molecules can be used to distinguish between two subsets of nucleic acid molecules. The average size of nucleic acid molecules in one subset of nucleic acid molecules can be greater than the average size of nucleic acid molecules in at least one other subset of nucleic acid molecules. The average size of nucleic acid molecules in one subset of nucleic acid molecules can be less than the average size of nucleic acid molecules in at least one other subset of nucleic acid molecules. The difference in the average molecular size between two or more subsets of nucleic acid molecules can be at least about 50; 75; 100; 125; 150; 175; 200; 225; 250; 275; 300; 350; 400; 450; 500; 550; 600; 650; 700; 750; 800; 850; 900; 950; 1,000; 1,100; 1,200; 1,300; 1,400; 1,500; 1,600; 1,700; 1,800; 1,900; 2,000; 3,000; 4,000; 5,000; 6,000; 7,000; 8,000; 9,000; 10,000; 15,000; 20,000; 30,000; 40,000; 50,000; 60,000; 70,000; 80,000; 90,000; 100,000; or more bases or base pairs. In some aspects of the disclosure, the difference in the average molecular size between two or more subsets of nucleic acid molecules is at least about 200 bases or base pairs. Alternatively, the difference in the average molecular size between two or more subsets of nucleic acid molecules is at least about 300 bases or base pairs.

[0073] Nucleic acid molecular subgroups can contain nucleic acid molecules of different sequencing sizes. The length of nucleic acid molecules in a nucleic acid molecular subgroup to be sequenced can be referred to as the sequencing size of the nucleic acid molecules. The average length of nucleic acid molecules in a nucleic acid molecular subgroup can be referred to as the average sequencing size of the nucleic acid molecules. As used herein, the terms "sequencing size of nucleic acid molecules," "average sequencing size of nucleic acid molecules," "molecular sequencing size," and "average molecular sequencing size" are used interchangeably. The average molecular sequencing size of one or more nucleic acid molecular subgroups can be at least about 50; 75; 100; 125; 150; 175; 200; 225; 250; 275; 300; 350; 400; 450; 500; 550; 600; 650; 700; 750; 800; 850; 900; 950; 1,000; 1100; 1200; 1300; 1400; 1500; 160 0; 1700; 1800; 1900; 2,000; 3,000; 4,000; 5,000; 6,000; 7,000; 8,000; 9,000; 10,000; 15,000; 20,000; 30,000; 40,000; 50,000; 60,000; 70,000; 80,000; 90,000; 100,000 or more bases or base pairs. The sequencing size of a nucleic acid molecule can be used to distinguish two or more nucleic acid molecule subgroups. The difference between the average sequencing size of nucleic acid molecules in one nucleic acid molecule subgroup and the average sequencing size of nucleic acid molecules in another nucleic acid molecule subgroup can be used to distinguish two nucleic acid molecule subgroups. The average sequencing size of nucleic acid molecules in one nucleic acid molecule subgroup can be greater than the average sequencing size of nucleic acid molecules in at least one other nucleic acid molecule subgroup. The average sequencing size of nucleic acid molecules in one nucleic acid molecular subgroup can be smaller than the average sequencing size of nucleic acid molecules in at least one other nucleic acid molecular subgroup. The difference in average molecular sequencing size between two or more nucleic acid molecular subgroups can be at least about 50; 75; 100; 125; 150; 175; 200; 225; 250; 275; 300; 350; 400; 450; 500; 550; 600; 650; 700; 750; 800; 850; 900; 950; 1,000; 1100; 1200; 1300; 1400; 1500; 1600; 1700; 1800; 1900; 2,000; 3,000; 4,000; 5,000; 6,000; 7,000; 8,000; 9,000; 10,000; 15,000; 20,000; 30,000; 40,000; 50,000; 60,000; 70,000; 80,000; 90,000; 100,000 or more bases or base pairs.In some aspects of the disclosure, the difference in average molecular sequencing size between two or more subsets of nucleic acid molecules is at least about 200 bases or base pairs. Alternatively, the difference in average molecular sequencing size between two or more subsets of nucleic acid molecules is at least about 300 bases or base pairs.

[0074] The methods disclosed herein can include one or more capture probes, a plurality of capture probes, or one or more sets of capture probes. Typically, a capture probe comprises a nucleic acid binding site. A capture probe can hybridize to a captured nucleic acid. A capture probe can comprise a nucleic acid sequence that is complementary to a captured nucleic acid. In some cases, a capture probe can comprise a nucleic acid sequence that is fully complementary to a portion of a captured nucleic acid. For example, each nucleic acid in a capture probe can be complementary to a base in a captured nucleic acid. A capture probe can be longer than a captured nucleic acid. For example, each base in a capture probe can be complementary to a base in a captured nucleic acid, but not all of the bases in the capture probe can be complementary to bases in the captured nucleic acid. A capture probe can be shorter than a captured nucleic acid. For example, each base of a capture probe can be complementary to a base of a captured nucleic acid, but not all of the bases in the captured nucleic acid can be complementary to the capture probe.

[0075] A capture probe can perform a capture reaction in solution. A capture probe can be in solution and can capture a nucleic acid in solution. As described elsewhere herein, a captured nucleic acid can be subsequently isolated and / or eluted. A capture probe can capture a nucleic acid in solution, and then the capture probe can subsequently be attached to a support, such as a solid support (e.g., an array or a bead). In some cases, a support can be formed of a semi-solid material (e.g., a gel).

[0076] Attachment to a support can be non-covalent attachment. For example, a nucleic acid can be captured by a biotinylated probe that subsequently binds to an avidin / streptavidin bead, thereby attaching the captured nucleic acid complex to the avidin / streptavidin bead. Other binding pairs can be used to attach a capture probe to a surface. A capture probe can be covalently attached to a support. For example, a support can have a chemically reactive linker that can react with a capture probe, such that the capture probe is covalently linked to the support.

[0077] A capture probe can be attached to a support and perform a capture reaction. For example, a capture probe can be coupled to a support and subsequently capture a nucleic acid molecule. A support can be a solid or semi-solid (e.g., gel) material. Examples of supports include, but are not limited to, beads, slides, and chips. A support can be, for example, glass, silica, silicon, plastic (e.g., polystyrene), agar, or agarose.

[0078] The capture probes can also comprise one or more linkers. The capture probes can also comprise one or more labels. The one or more linkers can attach the one or more labels to the nucleic acid binding site. In some cases, the capture probes can be designed to hybridize to a shared region of 16S gene sequences. Capture probes that can be designed to target shared regions of 16S gene sequences can be used to capture nucleic acid molecules from a variety of species, even species that have not yet been identified and characterized. In some cases, the methods can comprise a first plurality of nucleic acid probes configured to target elements of human genomic sequences. In some cases, the methods can comprise a second plurality of nucleic acid probes configured to target elements of genomic sequences of non-human species.

[0079] The methods disclosed herein can comprise 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, 1000 or more, 5000 or more, 10,000 or more, 20,000 or more, 30,000 or more, 40,000 or more, 50,000 or more, 60,000 or more, 70,000 or more, 80,000 or more, 90,000 or more, 100,0000 or capture probes or capture probe sets. In some cases, the methods can comprise about 50,000 capture probes. The one or more capture probes or capture probe sets can be different, similar, identical, or a combination thereof. The relative concentration of one or more capture probes or capture probe sets can vary as compared to other capture probes. For example, the concentration of certain capture probes can be higher to capture nucleic acids that can be difficult to capture (e.g., sequences with high GC content, sequences that arise due to recombination of sequences, sequences with a high number of mutations, sequences with a high mutation rate), thereby increasing the likelihood of capturing particular nucleic acids.

[0080] The one or more capture probes can comprise a nucleic acid binding site that hybridizes to at least a portion of one or more nucleic acid molecules or variants thereof or derivatives thereof in the nucleic acid molecule sample or subset of nucleic acid molecules. The capture probes can comprise a nucleic acid binding site that hybridizes to one or more genomes. The capture probes can hybridize to different, similar, and / or identical genomes. The one or more capture probes can be at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99%, or more complementary to one or more nucleic acid molecules or variants thereof or derivatives thereof.

[0081] The capture probes can comprise one or more nucleotides. The capture probes can comprise 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more nucleotides. The capture probes can comprise about 100 nucleotides. The capture probes can comprise about 10 to about 500 nucleotides, about 20 to about 450 nucleotides, about 30 to about 400 nucleotides, about 40 to about 350 nucleotides, about 50 to about 300 nucleotides, about 60 to about 250 nucleotides, about 70 to about 200 nucleotides, or about 80 to about 150 nucleotides. In some aspects of the disclosure, the capture probes comprise about 80 nucleotides to about 100 nucleotides.

[0082] The plurality of capture probes or capture probe sets can comprise two or more capture probes having identical, similar, and / or different nucleic acid binding site sequences, linkers, and / or labels. For example, the two or more capture probes comprise identical nucleic acid binding sites. In another example, the two or more capture probes comprise similar nucleic acid binding sites. In another example, the two or more capture probes comprise different nucleic acid binding sites. The two or more capture probes can also comprise one or more linkers. The two or more capture probes can also comprise different linkers. The two or more capture probes can also comprise similar linkers. The two or more capture probes can also comprise identical linkers. The two or more capture probes can also comprise one or more labels. The two or more capture probes can also comprise different labels. The two or more capture probes can also comprise similar labels. The two or more capture probes can also comprise identical labels.

[0083] An assay can include, but is not limited to, sequencing, amplification, hybridization, enrichment, isolation, elution, fragmentation, detection, quantification of one or more nucleic acid molecules. An assay can include methods for preparing one or more nucleic acid molecules. An assay can include conventional assays, long read, high GC content, and hybridization assays, FIG. 9 Any number of assays can be performed. The number of assays can be 1, 2, 3, 4, 5, 6, 7, 8, 9, or up to 10 assays. For example, FIG. 6B FIG. 6B An illustration showing the use of two assays is shown. Similarly, any number of analyses can be performed on data from one or more assays. The number of analyses can be 1, 2, 3, 4, 5, 6, 7, 8, 9, or up to 10 analyses on data from one or more assays. For example, Figure 6C Two different analyses being performed are shown. In some cases, the analyses are bioinformatics analyses, Figure 8 Any number of protocols can be used for an assay. For example, 1, 2, 3, 4, 5, 6, 7, 8, 9, or up to 10 protocols can be used. Figure 6D An illustration showing the use of 4 protocols is provided.

[0084] The methods disclosed herein can comprise one or more sequencing reactions on one or more nucleic acid molecules in the sample. The methods disclosed herein can comprise 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 15 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 200 or more, 300 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more sequencing reactions on one or more nucleic acid molecules in the sample. The sequencing reactions can be performed simultaneously, sequentially, or in a combination thereof. The sequencing reactions can comprise whole genome sequencing or exome sequencing. The sequencing reactions can comprise Maxim-Gilbert, chain termination, or high-throughput systems. Alternatively or additionally, the sequencing reactions can comprise Helioscope™ single molecule sequencing, nanopore DNA sequencing, large-scale parallel signature sequencing (MPSS) by Lynx Therapeutics, 454 pyrosequencing, single molecule real-time (RNAP) sequencing, Illumina (Solexa) sequencing, SOLiD sequencing, Ion Torrent™, Ion semiconductor sequencing, single molecule SMRT (TM) sequencing, Polony sequencing, DNA nanoball sequencing, VisiGen Biotechnologies method, or a combination thereof. Alternatively or additionally, the sequencing reactions can comprise one or more sequencing platforms including, but not limited to, Genome Analyzer IIx, HiSeq, and MiSeq provided by Illumina, single molecule real-time (SMRT™) technology (e.g., PacBio RS system provided by Pacific Biosciences (California) and Solexa Sequencer), true single molecule sequencing (tSMS™) technology (e.g., HeliScope™ Sequencer provided by Helicos Inc. (Cambridge, Massachusetts)). The sequencing reactions can also comprise an electron microscope or a chemically sensitive field effect transistor (chemFET) array. In some aspects of the disclosure, the sequencing reactions comprise capillary sequencing, next generation sequencing, Sanger sequencing, sequencing by synthesis, sequencing by ligation, sequencing by hybridization, single molecule sequencing, or a combination thereof. Sequencing by synthesis can comprise reversible terminator sequencing, progressive single molecule sequencing, sequential flow sequencing, or a combination thereof. Sequential flow sequencing can comprise pyrosequencing, pH-mediated sequencing, semiconductor sequencing, or a combination thereof.

[0085] The methods disclosed herein can comprise performing at least one long read sequencing reaction and at least one short read sequencing reaction. Figure 18 Examples of methods comprising long read and short read sequencing are shown in the Examples. Long read sequencing reactions and / or short read sequencing reactions can be performed on at least a portion of a subset of nucleic acid molecules. Long read sequencing reactions and / or short read sequencing reactions can be performed on at least a portion of two or more subsets of nucleic acid molecules. Long read sequencing reactions and short read sequencing reactions can be performed on at least a portion of one or more subsets of nucleic acid molecules.

[0086] Sequencing of one or more nucleic acid molecules or subsets thereof can comprise at least about 5; 10; 15; 20; 25; 30; 35; 40; 45; 50; 60; 70; 80; 90; 100; 200; 300; 400; 500; 600; 700; 800; 900; 1,000; 1500; 2,000; 2500; 3,000; 3500; 4,000; 4500; 5,000; 5500; 6,000; 6500; 7,000; 7500; 8,000; 8500; 9,000; 10,000; 25,000; 50,000; 75,000; 100,000; 250,000; 500,000; 750,000; 10,000,000; 25,000,000; 50,000,000; 100,000,000; 250,000,000; 500,000,000; 750,000,000; 1,000,000,000 or more sequencing reads.

[0087] The sequencing reaction can comprise sequencing at least about 50; 60; 70; 80; 90; 100; 110; 120; 130; 140; 150; 160; 170; 180; 190; 200; 210; 220; 230; 240; 250; 260; 270; 280; 290; 300; 325; 350; 375; 400; 425; 450; 475; 500; 600; 700; 800; 900; 1,000; 1,500; 2,000; 2,500; 3,000; 3,500; 4,000; 4,500; 5,000; 5,500; 6,000; 6,500; 7,000; 7,500; 8,000; 8,500; 9,000; 10,000; 20,000; 30,000; 40,000; 50,000; 60,000; 70,000; 80,000; 90,000; 100,000 or more bases or base pairs of one or more nucleic acid molecules. The sequencing reaction can comprise sequencing at least about 50; 60; 70; 80; 90; 100; 110; 120; 130; 140; 150; 160; 170; 180; 190; 200; 210; 220; 230; 240; 250; 260; 270; 280; 290; 300; 325; 350; 375; 400; 425; 450; 475; 500; 600; 700; 800; 900; 1,000; 1,500; 2,000; 2,500; 3,000; 3,500; 4,000; 4,500; 5,000; 5,500; 6,000; 6,500; 7,000; 7,500; 8,000; 8,500; 9,000; 10,000; 20,000; 30,000; 40,000; 50,000; 60,000; 70,000; 80,000; 90,000; 100,000 or more contiguous bases or base pairs of one or more nucleic acid molecules.

[0088] The sequencing technology used in the methods of the present disclosure can produce at least 100 reads per run, at least 200 reads per run, at least 300 reads per run, at least 400 reads per run, at least 500 reads per run, at least 600 reads per run, at least 700 reads per run, at least 800 reads per run, at least 900 reads per run, at least 1000 reads per run, at least 5,000 reads per run, at least 10,000 reads per run, at least 50,000 reads per run, at least 100,000 reads per run, at least 500,000 reads per run, or at least 1,000,000 reads per run. Alternatively, the sequencing technology used in the methods of the present disclosure can produce at least 1,500,000 reads per run, at least 2,000,000 reads per run, at least 2,500,000 reads per run, at least 3,000,000 reads per run, at least 3,500,000 reads per run, at least 4,000,000 reads per run, at least 4,500,000 reads per run, or at least 5,000,000 reads per run.

[0089] The sequencing technology used in the methods of the present disclosure can produce at least about 30 base pairs, at least about 40 base pairs, at least about 50 base pairs, at least about 60 base pairs, at least about 70 base pairs, at least about 80 base pairs, at least about 90 base pairs, at least about 100 base pairs, at least about 110, at least about 120 base pairs, at least about 150 base pairs, at least about 200 base pairs, at least about 250 base pairs, at least about 300 base pairs, at least about 350 base pairs, at least about 400 base pairs, at least about 450 base pairs, at least about 500 base pairs, at least about 550 base pairs, about 600 base pairs, at least about 700 base pairs, at least about 800 base pairs, at least about 900 base pairs, or at least about 1,000 base pairs per read. Alternatively, the sequencing technology used in the methods of the present disclosure can produce long sequencing reads. In some cases, the sequencing technology used in the methods of the present disclosure can produce at least about 1200 base pairs / read, at least about 1500 base pairs / read, at least about 1800 base pairs / read, at least about 2000 base pairs / read, at least about 2500 base pairs / read, at least about 3,000 base pairs / read, at least about 3500 base pairs / read, at least about 4,000 base pairs / read, at least about 4,500 base pairs / read, at least about 5,000 base pairs / read, at least about 6,000 base pairs / read, at least about 7,000 base pairs / read, at least about 8,000 base pairs / read, at least about 9,000 base pairs / read, at least about 10,000 base pairs / read, 20,000 base pairs / read, 30,000 base pairs / read, 40,000 base pairs / read, 50,000 base pairs / read, 60,000 base pairs / read, 70,000 base pairs / read, 80,000 base pairs / read, 90,000 base pairs / read, or 100,000 base pairs / read.

[0090] High-throughput sequencing systems can allow for detection of sequenced nucleotides as soon as, or substantially as soon as, the sequenced nucleotides are incorporated into growing chains, i.e., real-time or substantially real-time sequence detection. In some cases, high-throughput sequencing produces at least 1,000, at least 5,000, at least 10,000, at least 20,000, at least 30,000, at least 40,000, at least 50,000, at least 100,000, or at least 500,000 sequence reads per hour; wherein each read is at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 120, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, or at least 500 bases per read. Sequencing can be performed using nucleic acids described herein, e.g., genomic DNA, cDNA derived from RNA transcripts, or RNA, as templates.

[0091] The methods disclosed herein can comprise performing one or more amplification reactions on one or more nucleic acid molecules in a sample. The term “amplification” refers to any process that produces at least one copy of a nucleic acid molecule. The terms “amplicon” and “amplified nucleic acid molecule” refer to a copy of a nucleic acid molecule and can be used interchangeably. An amplification reaction can comprise a PCR-based method, a non-PCR-based method, or a combination thereof. Examples of non-PCR-based methods include, but are not limited to, multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, rolling circle amplification, or circle-to-circle amplification. PCR-based methods can include, but are not limited to, PCR, HD-PCR, next-generation PCR, digital RTA, or any combination thereof. Other PCR methods include, but are not limited to, linear amplification, allele-specific PCR, Alu PCR, assembly PCR, asymmetric PCR, droplet PCR, emulsion PCR, helicase-dependent amplification HDA, hot-start PCR, inverse PCR, linear-after-the-exponential (LATE)-PCR, long PCR, multiplex PCR, nested PCR, semi-nested PCR, quantitative PCR, RT-PCR, real-time PCR, single-cell PCR, and touchdown PCR.

[0092] The methods disclosed herein can comprise one or more hybridization reactions on one or more nucleic acid molecules in a sample. The hybridization reactions can comprise hybridization of one or more capture probes to one or more nucleic acid molecules in a sample or subset of nucleic acid molecules. The hybridization reactions can comprise hybridizing one or more capture probe sets to one or more nucleic acid molecules in a sample or subset of nucleic acid molecules. The hybridization reactions can comprise one or more hybridization arrays, multiplex hybridization reactions, hybridization chain reactions, isothermal hybridization reactions, nucleic acid hybridization reactions, or combinations thereof. The one or more hybridization arrays can comprise hybridization array genotyping, hybridization array ratio sensing, DNA hybridization arrays, macroarrays, microarrays, high-density oligonucleotide arrays, genomic hybridization arrays, comparative hybridization arrays, or combinations thereof. The hybridization reactions can comprise one or more capture probes, one or more beads, one or more labels, one or more subsets of nucleic acid molecules, one or more nucleic acid samples, one or more reagents, one or more wash buffers, one or more elution buffers, one or more hybridization buffers, one or more hybridization chambers, one or more incubators, one or more separators, or combinations thereof.

[0093] The methods disclosed herein can comprise one or more enrichment reactions on one or more nucleic acid molecules in a sample. The enrichment reactions can comprise contacting the sample with one or more beads or bead sets. The enrichment reactions can comprise differential amplification of two or more subsets of nucleic acid molecules based on one or more genomic features. For example, the enrichment reactions comprise differential amplification of two or more subsets of nucleic acid molecules based on GC content. Alternatively or additionally, the enrichment reactions comprise differential amplification of two or more subsets of nucleic acid molecules based on methylation status. The enrichment reactions can comprise one or more hybridization reactions. The enrichment reactions can further comprise isolation and / or purification of one or more hybridized nucleic acid molecules, one or more bead-bound nucleic acid molecules, one or more free nucleic acid molecules (e.g., nucleic acid molecules without capture probes, nucleic acid molecules without beads), one or more labeled nucleic acid molecules, one or more unlabeled nucleic acid molecules, one or more amplicons, one or more unamplified nucleic acid molecules, or combinations thereof. Alternatively or additionally, the enrichment reactions can comprise enriching one or more cell types in the sample. The one or more cell types can be enriched by flow cytometry.

[0094] One or more enrichment reactions can produce one or more enriched nucleic acid molecules. Enriched nucleic acid molecules can comprise nucleic acid molecules or variants thereof or derivatives thereof. For example, enriched nucleic acid molecules include one or more hybridized nucleic acid molecules, one or more bead-bound nucleic acid molecules, one or more free nucleic acid molecules (e.g., nucleic acid molecules without capture probes, nucleic acid molecules without beads), one or more labeled nucleic acid molecules, one or more unlabeled nucleic acid molecules, one or more amplicons, one or more unamplified nucleic acid molecules, or combinations thereof. Enriched nucleic acid molecules can be distinguished from non-enriched nucleic acid molecules by GC content, molecular size, genome, genomic signature, or combinations thereof. Enriched nucleic acid molecules can be derived from one or more assays, supernatants, eluates, or combinations thereof. Enriched nucleic acid molecules can differ from non-enriched nucleic acid molecules in average size, average GC content, genome, or combinations thereof. In some cases, enrichment can include multiple subsets of DNA enriched for different genomic regions that undergo independent processing operations prior to being combined for sequencing assays, Figure 15 In some cases, enrichment can include multiple subsets of DNA enriched for different genomic regions that undergo independent processing operations prior to being independently sequenced and analyzed, Figure 16 .

[0095] Methods disclosed herein can include one or more separation or purification reactions on one or more nucleic acid molecules in a sample. Separation or purification reactions can include contacting a sample with one or more beads or bead sets. Separation or purification reactions can include one or more hybridization reactions, enrichment reactions, amplification reactions, sequencing reactions, or combinations thereof. Separation or purification reactions can include use of one or more separators. One or more separators can include a magnetic separator. Separation or purification reactions can include separating bead-bound nucleic acid molecules from nucleic acid molecules without beads. Separation or purification reactions can include separating capture probe-hybridized nucleic acid molecules from nucleic acid molecules without capture probes. Separation or purification reactions can include separating a first subset of nucleic acid molecules from a second subset of nucleic acid molecules, where the first subset of nucleic acid molecules differs from the second subset of nucleic acid molecules in average size, average GC content, genome, or combinations thereof.

[0096] Methods disclosed herein can include one or more elution reactions on one or more nucleic acid molecules in a sample. Elution reactions can include contacting a sample with one or more beads or bead sets. Elution reactions can include separating bead-bound nucleic acid molecules from nucleic acid molecules without beads. Elution reactions can include separating capture probe-hybridized nucleic acid molecules from nucleic acid molecules without capture probes. Elution reactions can include separating a first subset of nucleic acid molecules from a second subset of nucleic acid molecules, where the first subset of nucleic acid molecules differs from the second subset of nucleic acid molecules in average size, average GC content, genome, or combinations thereof.

[0097] The methods disclosed herein can comprise one or more fragmentation reactions. A fragmentation reaction can comprise fragmenting one or more nucleic acid molecules in a sample or subset of nucleic acid molecules to produce one or more fragmented nucleic acid molecules. The one or more nucleic acid molecules can be fragmented by sonication, needle shearing, nebulization, shearing (e.g., acoustic shearing, mechanical shearing, point-sink shearing), passage through a French pressure cell, or enzymatic digestion. The enzymatic digestion can be performed by nuclease digestion (e.g., micrococcal nuclease digestion, endonuclease, exonuclease, RNase H, or DNase I). Fragmentation of the one or more nucleic acid molecules can produce fragment sizes of about 100 base pairs to about 2000 base pairs, about 200 base pairs to about 1500 base pairs, about 200 base pairs to about 1000 base pairs, about 200 base pairs to about 500 base pairs, about 500 base pairs to about 1500 base pairs, and about 500 base pairs to about 1000 base pairs. The one or more fragmentation reactions can produce fragment sizes of about 50 base pairs to about 1000 base pairs. The one or more fragmentation reactions can produce fragment sizes of about 100 base pairs, 150 base pairs, 200 base pairs, 250 base pairs, 300 base pairs, 350 base pairs, 400 base pairs, 450 base pairs, 500 base pairs, 550 base pairs, 600 base pairs, 650 base pairs, 700 base pairs, 750 base pairs, 800 base pairs, 850 base pairs, 900 base pairs, 950 base pairs, 1000 base pairs, or greater.

[0098] Fragmenting the one or more nucleic acid molecules can comprise mechanically shearing the one or more nucleic acid molecules in the sample over a period of time. The fragmentation reaction can occur for at least about 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, or more seconds.

[0099] Fragmenting the one or more nucleic acid molecules can comprise contacting the nucleic acid sample with one or more beads. Fragmenting the one or more nucleic acid molecules can comprise contacting the nucleic acid sample with a plurality of beads, wherein the ratio of the volume of the plurality of beads to the volume of the nucleic acid sample is about 0.10, 0.20, 0.30, 0.40, 0.50, 0.60, 0.70, 0.80, 0.90, 1.00, 1.10, 1.20, 1.30, 1.40, 1.50, 1.60, 1.70, 1.80, 1.90, 2.00, or more. Fragmenting the one or more nucleic acid molecules can comprise contacting the nucleic acid sample with a plurality of beads, wherein the ratio of the volume of the plurality of beads to the volume of the nucleic acid is about 2.00, 1.90, 1.80, 1.70, 1.60, 1.50, 1.40, 1.30, 1.20, 1.10, 1.00, 0.90, 0.80, 0.70, 0.60, 0.50, 0.40, 0.30, 0.20, 0.10, 0.05, 0.04, 0.03, 0.02, 0.01, or less.

[0100] The methods disclosed herein can comprise performing one or more detection reactions on the one or more nucleic acid molecules in the sample. The detection reactions can comprise one or more sequencing reactions. Alternatively, performing the detection reactions comprises optical sensing, electrical sensing, or a combination thereof. The optical sensing can comprise optical sensing of photoluminescent photon emission, fluorescent photon emission, pyrophosphate photon emission, chemiluminescent photon emission, or a combination thereof. The electrical sensing can comprise electrical sensing of ion concentration, ion current modulation, nucleotide electric field, nucleotide tunneling current, or a combination thereof.

[0101] The methods disclosed herein can comprise performing one or more quantification reactions on the one or more nucleic acid molecules in the sample. The quantification reactions can comprise sequencing, PCR, qPCR, digital PCR, or a combination thereof.

[0102] The methods disclosed herein can comprise one or more samples. The methods disclosed herein can comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, or more samples. The sample can be derived from a subject. Two or more samples can be derived from a single subject. Two or more samples can be derived from 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, or more different subjects. The subject can be a mammal, a reptile, an amphibian, an avian, and a fish. The mammal can be a human, an ape, a chimpanzee, a monkey, a cow, a pig, a horse, a rodent, a dog, a cat, or other animal. The reptile can be a lizard, a snake, an alligator, a turtle, a crocodile, and a terrapin. The amphibian can be a toad, a frog, a newt, and a salamander. Examples of avian include, but are not limited to, a duck, a goose, a penguin, an ostrich, and an owl. Examples of fish include, but are not limited to, a catfish, an eel, a shark, and a swordfish. The subject can be a human. The subject can have a disease or a condition.

[0103] Two or more samples can be collected at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or time points. The time points can occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, or more hours. The time points can occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, or more days. The time points can occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, or more weeks. The time points can occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, or more months. The time points can occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, or more years.

[0104] In some cases, the method can comprise obtaining a biological sample from a subject. The subject can be a human or a non-human. The subject can be an adult or a child. In some cases, the adult subject can be 18 years of age or older. In some cases, the biological sample of the subject can be derived from a tumor biopsy, whole blood, or plasma. In some cases, the biological sample can be from a bodily fluid, a cell, skin, tissue, an organ, or a combination thereof. The sample can be blood, plasma, a blood fraction, saliva, sputum, urine, semen, transvaginal fluid, cerebrospinal fluid, fecal matter, a cell, or a tissue biopsy. The sample can be from the adrenal gland, appendix, bladder, brain, ear, esophagus, eye, gall bladder, heart, kidney, large intestine, liver, lung, mouth, muscle, nose, pancreas, parathyroid gland, pineal gland, pituitary gland, skin, small intestine, spleen, stomach, thymus, thyroid, trachea, uterus, vermiform appendix, cornea, skin, heart valve, artery, or vein.

[0105] A sample can comprise one or more nucleic acid molecules. A nucleic acid molecule can be a DNA molecule, an RNA molecule (e.g., mRNA, cRNA, or miRNA), and a DNA / RNA hybrid. Examples of DNA molecules include, but are not limited to, double-stranded DNA, single-stranded DNA, single-stranded DNA hairpins, cDNA, genomic DNA. A nucleic acid can be an RNA molecule, such as double-stranded RNA, single-stranded RNA, ncRNA, RNA hairpins, and mRNA. Examples of ncRNA include, but are not limited to, siRNA, miRNA, snoRNA, piRNA, tiRNA, PASR, TASR, aTASR, TSSa-RNA, snRNA, RE-RNA, uaRNA, x-ncRNA, hYRNA, usRNA, snaR, and vtRNA.

[0106] A method disclosed herein can comprise one or more containers. A method disclosed herein can comprise 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more containers. One or more containers can be different, similar, identical, or a combination thereof. Examples of containers include, but are not limited to, plates, microplates, PCR plates, wells, microwells, tubes, Eppendorf tubes, vials, arrays, microarrays, and chips.

[0107] The methods disclosed herein can comprise one or more reagents. The methods disclosed herein can comprise 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more reagents. The one or more reagents can be different, similar, identical, or a combination thereof. The reagents can improve the efficiency of one or more assays. The reagents can improve the stability of the nucleic acid molecules, or variants thereof, or derivatives thereof. The reagents can include, but are not limited to, enzymes, proteases, nucleases, molecules, polymerases, reverse transcriptases, ligases, and compounds. The methods disclosed herein can comprise performing an assay comprising one or more antioxidants. Generally, an antioxidant is a molecule that inhibits the oxidation of another molecule. Examples of antioxidants include, but are not limited to, ascorbic acid (e.g., Vitamin C), glutathione, lipoic acid, uric acid, carotenes, alpha-tocopherol (e.g., Vitamin E), ubiquinol (e.g., Coenzyme Q), and Vitamin A.

[0108] The methods disclosed herein can comprise one or more buffers or solutions. The methods disclosed herein can comprise 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more buffers or solutions. The one or more buffers or solutions can be different, similar, identical, or a combination thereof. The buffers or solutions can improve the efficiency of one or more assays. The buffers or solutions can improve the stability of the nucleic acid molecules, or variants thereof, or derivatives thereof. The buffers or solutions can include, but are not limited to, wash buffers, elution buffers, and hybridization buffers.

[0109] The methods disclosed herein can include one or more beads, a plurality of beads, or one or more bead sets. The methods disclosed herein can include 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more one or more beads or bead sets. The one or more beads or bead sets can be different, similar, identical, or a combination thereof. The beads can be magnetic, antibody-coated, protein A-coupled, protein G-coupled, streptavidin-coated, oligonucleotide-conjugated, silica-coated, or a combination thereof. Examples of beads include, but are not limited to, Ampure beads, AMPure XP beads, streptavidin beads, agarose beads, magnetic beads, Dynabeads®, MACS® microbeads, antibody-conjugated beads (e.g., anti-immunoglobulin microbeads), protein A-conjugated beads, protein G-conjugated beads, protein A / G-conjugated beads, protein L-conjugated beads, oligo-dT-conjugated beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorescent dye microbeads, and BcMag™ carboxy-terminated magnetic beads. In some aspects of the disclosure, the one or more beads include one or more Ampure beads. Alternatively or additionally, the one or more beads include AMPure XP beads.

[0110] The methods disclosed herein can comprise one or more primers, a plurality of primers, or one or more primer sets. The primers can further comprise one or more linkers. The primers can further comprise one or more labels. The primers can be used in one or more assays. For example, the primers are used in one or more sequencing reactions, amplification reactions, or a combination thereof. The methods disclosed herein can comprise 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more one or more primers or primer sets. The primers can comprise about 100 nucleotides. The primers can comprise about 10 to about 500 nucleotides, about 20 to about 450 nucleotides, about 30 to about 400 nucleotides, about 40 to about 350 nucleotides, about 50 to about 300 nucleotides, about 60 to about 250 nucleotides, about 70 to about 200 nucleotides, or about 80 to about 150 nucleotides. In some aspects of the disclosure, the primers comprise about 80 nucleotides to about 100 nucleotides. The one or more primers or primer sets can be different, similar, identical, or a combination thereof.

[0111] The primers can hybridize to at least a portion of one or more nucleic acid molecules or variants or derivatives thereof in a sample or subset of nucleic acid molecules. The primers can hybridize to one or more genomes. The primers can hybridize to different, similar, and / or identical genomes. The one or more primers can be complementary to the one or more nucleic acid molecules or variants or derivatives thereof by at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99%, or more.

[0112] A primer can comprise one or more nucleotides. A primer can comprise 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more nucleotides. A primer can comprise about 100 nucleotides. A primer can comprise about 10 to about 500 nucleotides, about 20 to about 450 nucleotides, about 30 to about 400 nucleotides, about 40 to about 350 nucleotides, about 50 to about 300 nucleotides, about 60 to about 250 nucleotides, about 70 to about 200 nucleotides, or about 80 to about 150 nucleotides. In some aspects of the disclosure, a primer comprises about 80 nucleotides to about 100 nucleotides.

[0113] A plurality of primers or primer sets can comprise two or more primers having the same, similar, and / or different sequences, linkers, and / or labels. For example, two or more primers comprise the same sequence. In another example, two or more primers comprise similar sequences. In yet another example, two or more primers comprise different sequences. Two or more primers can also comprise one or more linkers. Two or more primers can also comprise different linkers. Two or more primers can also comprise similar linkers. Two or more primers can also comprise the same linkers. Two or more primers can also comprise one or more labels. Two or more primers can also comprise different labels. Two or more primers can also comprise similar labels. Two or more primers can also comprise the same labels.

[0114] In some cases, universal primers can be used, and the universal primers can comprise one or more of the following: 8F primer, 27F primer, CC[F] primer, 357F primer, 515F primer, 533F primer, 16S.1100.F16 primer, 1237F primer, 519R primer, CD[R] primer, 907R primer, 1391R primer, 1492R(I) primer, 1492R(s) primer, U1492R primer, 928F primer, 336R primer, 1100F primer, 1100R primer, 337F primer, 785F primer, 805R primer, 518R primer, and any other suitable universal primer. Alternatively, for samples for which specific primers can be appropriate, specific primers can be used for amplification. In examples, specific primers can include CYA106 primer (for cyanobacteria), CYA359F primer (for cyanobacteria), 895F primer (for bacteria excluding plastids and cyanobacteria), CYA781R primer (for cyanobacteria), 902R primer (for bacteria excluding plastids and cyanobacteria), 904R primer (for bacteria excluding plastids and cyanobacteria), 1100R primer (for bacteria), 1185mR primer (for bacteria excluding plastids and cyanobacteria), 1185aR primer (for lichen-associated Rhizobiales), 1381R primer (for bacteria excluding Asterochloris sp. plastids), or any other suitable specific primer.

[0115] The capture probes, primers, labels, and / or beads can comprise one or more nucleotides. The one or more nucleotides can comprise RNA, DNA, a mixture of DNA and RNA residues, or modified analogs thereof, such as 2'-OMe or 2'-fluoro (2'-F), locked nucleic acid (LNA), or abasic sites.

[0116] The methods disclosed herein can comprise one or more labels. The methods disclosed herein can comprise 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more of the one or more labels. The one or more labels can be different, similar, identical, or a combination thereof.

[0117] Examples of labeling include, but are not limited to, chemical labeling, biochemical labeling, biological labeling, colorimetric labeling, enzyme-catalyzed labeling, fluorescent labeling, and luminescent labeling. Labels include dyes, photocrosslinking agents, cytotoxic compounds, drugs, affinity labels, photoaffinity labels, reactive compounds, antibodies or antibody fragments, biomaterials, nanoparticles, spin labels, fluorophores, metal-containing moieties, radioactive moieties, novel functional groups, groups that interact covalently or non-covalently with other molecules, photocaged moieties, photochemically excitable moieties, ligands, photoisomerizable moieties, biotin, biotin analogs, heavy atom-binding moieties, chemically cleavable groups, photocleavable groups, redox activators, isotope-labeled moieties, biophysical probes, phosphorescent groups, chemiluminescent groups, electron-dense groups, magnetic groups, intercalation groups, chromophores, energy transfer agents, bioactive agents, detectable labels, or combinations thereof.

[0118] The label can be a chemical label. Examples of chemical labels may include, but are not limited to, biotin and radioactive isotopes (e.g., iodine, carbon, phosphate, hydrogen).

[0119] The methods, kits, and compositions disclosed herein may include biomarkers. Biomarkers may include metabolic markers, including but not limited to bio-orthogonalized azids-modified amino acids, sugars, and other compounds.

[0120] The methods, kits, and compositions disclosed herein may include enzyme labeling. Enzyme labeling may include, but is not limited to, horseradish peroxidase (HRP), alkaline phosphatase (AP), glucose oxidase, and β-galactosidase. Enzyme labeling may also include luciferase.

[0121] The methods, kits, and compositions disclosed herein may include fluorescent labels. Fluorescent labels may be organic dyes (e.g., FITC), biofluoresces (e.g., green fluorescent protein), or quantum dots. A non-limiting list of fluorescent labels includes fluorescein isothiocyanate (FITC), DyLight Fluors, fluorescein, rhodamine (tetramethylrhodamine isothiocyanate, TRITC), coumarin, Lucifer Yellow, and BODIPY. The label may be a fluorophore. Examples of fluorophores include, but are not limited to: indolecarbazine (C3), indoledicarbazine (C5), Cy3, Cy3.5, Cy5, Cy5.5, Cy7, Texas Red, Pacific Blue, Oregon Green 488, and Alexa Fluor. ®-355、Alexa Fluor 488, Alexa Fluor 532, Alexa Fluor 546, Alexa Fluor-555, Alexa Fluor 568, Alexa Fluor 594, Alexa Fluor 647, Alexa Fluor 660, Alexa Fluor 680, JOE, Lissamine, Rhodamine Green, BODIPY, fluorescein isothiocyanate (FITC), carboxyfluorescein (FAM), phycoerythrin, rhodamine, dichlororhodamine (dRhodamine), carboxytetramethylrhodamine (TAMRA), carboxy-X-rhodamine (ROX™), LIZ™, VIC™, NED™, PET™, SYBR, PicoGreen, RiboGreen, and the like. The fluorescent label can be green fluorescent protein (GFP), red fluorescent protein (RFP), yellow fluorescent protein, a phycobiliprotein (e.g., allophycocyanin, phycocyanin, phycoerythrin, and phycoerythrocyanin).

[0122] The methods disclosed herein can comprise one or more linkers. The methods disclosed herein can comprise 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more linkers. The one or more linkers can be different, similar, identical, or a combination thereof.

[0123] Suitable linkers include any compound or biological compound that is capable of connecting to a label, primer, and / or capture probe disclosed herein. If the linker is connected to both a label and a primer or capture probe, the suitable linker can be capable of substantially separating the label and primer or capture probe. The suitable linker can not substantially interfere with the ability of the primer and / or capture probe to hybridize to a nucleic acid molecule, a portion thereof, or a variant or derivative thereof. The suitable linker can not substantially interfere with the ability of the label to be detected. The linker can be rigid. The linker can be flexible. The linker can be semi-rigid. The linker can be proteolytically stable (e.g., resistant to proteolytic cleavage). The linker can be proteolytically unstable (e.g., susceptible to proteolytic cleavage). The linker can be helical. The linker can be non-helical. The linker can be coiled. The linker can be beta-stranded. The linker can comprise a turn conformation. The linker can be single-stranded. The linker can be long-stranded. The linker can be short-stranded. The linker can comprise at least about 5 residues, at least about 10 residues, at least about 15 residues, at least about 20 residues, at least about 25 residues, at least about 30 residues, or at least about 40 residues or more.

[0124] Examples of linkers include, but are not limited to, hydrazone, disulfide, thioether, and peptide linkers. The linker can be a peptide linker. The peptide linker can comprise a proline residue. The peptide linker can comprise arginine, phenylalanine, threonine, glutamine, glutamic acid, or any combination thereof. The linker can be a heterobifunctional crosslinker.

[0125] The methods disclosed herein can include 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 20 or more, 25 or more, 30 or more, 35 or more, 40 or more, 45 or more, or 50 or more assays performed on a sample comprising one or more nucleic acid molecules. Two or more assays can be different, similar, identical, or a combination thereof. For example, the methods disclosed herein include performing two or more sequencing reactions. In another example, the methods disclosed herein include performing two or more assays, wherein at least one of the two or more assays comprises a sequencing reaction. In yet another example, the methods disclosed herein include performing two or more assays, wherein at least two of the two or more assays comprise a sequencing reaction and a hybridization reaction. The two or more assays can be performed sequentially, simultaneously, or a combination thereof. For example, two or more sequencing reactions can be performed simultaneously. In another example, the methods disclosed herein include performing a hybridization reaction followed by a sequencing reaction. In yet another example, the methods disclosed herein include performing two or more hybridization reactions simultaneously, followed by performing two or more sequencing reactions simultaneously. The two or more assays can be performed by one or more devices. For example, two or more amplification reactions can be performed by a PCR machine. In another example, two or more sequencing reactions can be performed by two or more sequencers.

[0126] The methods disclosed herein can include one or more devices. The methods disclosed herein can include one or more assays comprising one or more devices. The methods disclosed herein can include using one or more devices to perform one or more operations or assays. The methods disclosed herein can include using one or more devices in one or more operations or assays. For example, performing a sequencing reaction can include one or more sequencers. In another example, generating a subset of nucleic acid molecules can include using one or more magnetic separators. In yet another example, one or more processors can be used in the analysis of one or more nucleic acid samples. Examples of devices include, but are not limited to, sequencers, thermal cyclers, real-time PCR instruments, magnetic separators, transport devices, hybridization chambers, electrophoresis instruments, centrifuges, microscopes, imagers, fluorometers, luminometers, plate readers, computers, processors, and bioanalyzers.

[0127] The methods disclosed herein can include one or more sequencers. The one or more sequencers can include one or more HiSeq, MiSeq, HiScan, Genome Analyzer IIx, SOLiD sequencers, Ion Torrent PGM, 454 GS Junior, Pac Bio RS, or a combination thereof. The one or more sequencers can include one or more sequencing platforms. The one or more sequencing platforms can include GS FLX by 454 Life Technologies / Roche, Genome Analyzer by Solexa / Illumina, SOLiD by Applied Biosystems, CGA Platform by Complete Genomics, PacBio RS by Pacific Biosciences, or a combination thereof.

[0128] The methods disclosed herein can include one or more thermal cyclers. The one or more thermal cyclers can be used to amplify one or more nucleic acid molecules. The methods disclosed herein can include one or more real-time PCR instruments. The one or more real-time PCR instruments can include a thermal cycler and a fluorometer. The one or more thermal cyclers can be used to amplify and detect one or more nucleic acid molecules.

[0129] The methods disclosed herein can include one or more magnetic separators. The one or more magnetic separators can be used to separate paramagnetic and ferromagnetic particles from a suspension. The one or more magnetic separators can include one or more LifeStep™ BioMag Separator, SPHERO™ FlexiMag Separator, SPHERO™ MicroMag Separator, SPHERO™ HandiMag Separator, SPHERO™ MiniTube Mag Separator, SPHERO™ UltraMag Separator, DynaMag™ Magnet, DynaMag™-2 Magnet, or a combination thereof.

[0130] The methods disclosed herein can include one or more bioanalyzers. In some cases, a bioanalyzer is a chip-based capillary electrophoresis instrument that can analyze RNA, DNA, and proteins. The one or more bioanalyzers can include an Agilent 2100 Bioanalyzer.

[0131] The methods disclosed herein can comprise one or more processors. The one or more processors can analyze, compile, store, sort, combine, evaluate, or otherwise process one or more data and / or results from one or more assays, one or more data and / or results based on or derived from one or more assays, one or more outputs from one or more assays, one or more outputs based on or derived from one or more assays, one or more outputs from one or more data and / or results, one or more outputs based on or derived from one or more data and / or results, or a combination thereof. In some cases, the methods disclosed herein can comprise combining data for analysis, such as Figure 17The one or more processors can send, be based on or derived from, one or more data, results, or outputs from one or more assays; one or more outputs from one or more data or results; be based on or derived from one or more outputs from one or more data or results, or a combination thereof. The one or more processors can receive and / or store a request from a user. The one or more processors can generate or produce one or more data, results, outputs. The one or more processors can generate or produce one or more biomedical reports. The one or more processors can send one or more biomedical reports. The one or more processors can analyze, compile, store, sort, combine, evaluate, or otherwise process information from one or more databases, one or more data or results, one or more outputs, or a combination thereof. The one or more processors can analyze, compile, store, sort, combine, evaluate, or otherwise process information from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, or more databases. The one or more processors can send one or more requests, data, results, outputs, and / or information to one or more users, processors, computers, computer systems, memory locations, devices, databases, or a combination thereof. The one or more processors can receive one or more requests, data, results, outputs, and / or information from one or more users, processors, computers, computer systems, memory locations, devices, databases, or a combination thereof. The one or more processors can retrieve one or more requests, data, results, outputs, and / or information from one or more users, processors, computers, computer systems, memory locations, devices, databases, or a combination thereof. The present disclosure also provides a method that can be used for a variety of biomedical applications. In some cases, variants, genes, reconstituted genes, exons, UTRs, regulatory regions, splice sites, alternative sequences, and other contents of interest of a human or non-human genome are combined from a plurality of databases to generate a set of aggregate contents suitable for a plurality of biomedical reports. This content can then be sorted based on local or global genomic context, nucleotide content, sequencing performance, and interpretation requirements, and then grouped into subgroups for specific protocols, assay development, etc. In some cases, variants, genes, exons, UTRs, regulatory regions, splice sites, alternative sequences, and other contents of interest are combined from a plurality of databases to generate a set of aggregate contents suitable for a plurality of biomedical reports. This content can then be sorted based on local or global genomic context, nucleotide content, sequencing performance, and interpretation requirements, and then subsequently grouped into subgroups for specific protocols and assay development, Figure 14In some cases, the protocol and / or assay can include a supplemental extract. The supplemental extract can comprise human target sequences, non-human target sequences, and combinations thereof. The supplemental extract can comprise nucleic acid molecules from the subject (e.g., derived from cells from tissue from the subject) and nucleic acid molecules that are not from the subject (e.g., from a microbe (commensal or parasitic), a pathogen, or a graft). Figures 15-17 Examples of assay workflows including supplemental extracts for one or more of a plurality of DNA subsets enriched for different genomic regions are provided.

[0132] The methods disclosed herein can include one or more memory locations. The one or more memory locations can store information, data, results, outputs, requests, or combinations thereof. The one or more memory locations can receive information, data, results, outputs, requests, or combinations thereof from one or more users, processors, computers, computer systems, devices, or combinations thereof.

[0133] The methods described herein can be implemented with the aid of one or more computers and / or computer systems. The computer or computer system can include electronic storage locations (e.g., databases, memory) having machine executable code for performing the methods provided in the present disclosure, and one or more processors for executing the machine executable code.

[0134] The methods disclosed herein can include treating and / or preventing a disease or condition in a subject based on one or more biomedical outputs. The one or more biomedical outputs can recommend one or more therapies. The one or more biomedical outputs can suggest, select, specify, recommend, or otherwise determine a treatment and / or prevention procedure for a disease and condition. The one or more biomedical outputs can suggest modifying or continuing one or more therapies. Modifying one or more therapies can include administering, initiating, reducing, increasing, and / or terminating one or more therapies. The one or more therapies include an anti-cancer, anti-viral, anti-bacterial, anti-fungal, immunosuppressive therapy, or combinations thereof. The one or more therapies can treat, ameliorate, or prevent one or more diseases or indications.

[0135] Examples of anti-cancer therapies include, but are not limited to, surgery, chemotherapy, radiation therapy, immunotherapy / biotherapy, photodynamic therapy. Anti-cancer therapies can include chemotherapy, monoclonal antibodies (e.g., rituximab, trastuzumab), cancer vaccines (e.g., therapeutic vaccines, prophylactic vaccines), gene therapy, or combinations thereof.

[0136] The one or more therapies can include an antimicrobial agent. Generally, an antimicrobial agent refers to a substance that kills or inhibits the growth of microorganisms, such as bacteria, fungi, viruses, or protozoa. Antimicrobial drugs kill microorganisms (microbicidal), or prevent the growth of microorganisms (microbistatic). There are two main classes of antimicrobial drugs, which are obtained from natural sources (e.g., antibiotics, protein synthesis inhibitors (e.g., aminoglycosides, macrolides, tetracyclines, chloramphenicols, polypeptides)) and synthetic agents (e.g., sulfonamides, sulfamethoxazole, quinolones). In some cases, the antimicrobial drug is an antibiotic, an antiviral, an antifungal, an antimalarial, an antitubercular, an antileprotic, or an antiprotozoal drug.

[0137] Antibiotics are generally used to treat bacterial infections. Antibiotics can be divided into two categories: bactericidal antibiotics and bacteriostatic antibiotics. Generally, bactericidal agents can directly kill bacteria, while bacteriostatic agents can prevent bacteria from dividing. Antibiotics can be derived from living organisms, or can include synthetic antibacterial agents, such as sulfonamides. Antibiotics can include aminoglycosides, such as amikacin, gentamicin, kanamycin, neomycin, netilmicin, tobramycin, and paromomycin. Alternatively, the antibiotic can be an ansamycin (e.g., geldanamycin, herbimycin), a cabacephem (e.g., flomoxef), a carbapenem (e.g., ertapenem, doripenem, imipenem, cilastatin, meropenem), a glycopeptide (e.g., teicoplanin, vancomycin, telavancin), a lincosamide (e.g., clindamycin, lincomycin, daptomycin), a macrolide (e.g., azithromycin, clarithromycin, dirithromycin, erythromycin, roxithromycin, troleandomycin, telithromycin, spectinomycin, spiramycin), a nitrofurane (e.g., furazidone, furaltadone), and a polypeptide (e.g., bacitracin, colistin, polymyxin B).

[0138] In some cases, the antibiotic therapy includes a cephalosporin, such as cefadroxil, cefazolin, cefalotin, cefalexin, cefaclor, cefamandole, cefoxitin, cefprozil, cefuroxime, cefixime, cefdinir, cefditoren, cefoperazone, cefotaxime, cefpodoxime, ceftazidime, ceftibuten, ceftizoxime, ceftriaxone, cefepime, ceftaroline fosamil, and ceftobiprole.

[0139] The antibiotic therapy can also include a penicillin. Examples of penicillins include amoxicillin, ampicillin, azlocillin, carbenicillin, cloxacillin, dicloxacillin, flucloxacillin, mezlocillin, methicillin, nafcillin, oxacillin, penicillin g, penicillin v, piperacillin, temocillin, and ticarcillin.

[0140] Alternatively, quinolines can be used to treat bacterial infections. Examples of quinolines include ciprofloxacin, enoxacin, gatifloxacin, levofloxacin, lomefloxacin, moxifloxacin, nalidixic acid, norfloxacin, ofloxacin, trovafloxacin, grepafloxacin, sparfloxacin, and temafloxacin.

[0141] In some cases, antibiotic therapy includes a combination of two or more therapies. For example, amoxicillin and clavulanate, ampicillin and sulbactam, piperacillin and tazobactam, or ticarcillin and clavulanate can be used to treat bacterial infections.

[0142] Sulfonamides can also be used to treat bacterial infections. Examples of sulfonamides include, but are not limited to, sulfisoxazole, sulfonamidochrysoidine, sulfacetamide, sulfadiazine, silver sulfadiazine, sulfisomidine, sulfamethoxazole, sulfamethoxazole, sulfasalazine, sulfisoxazole, trimethoprim, trimethoprim-sulfamethoxazole combination (co-trimoxazole) (tmp-smx).

[0143] Tetracyclines are another example of antibiotics. Tetracyclines can inhibit the binding of aminoacyl-tRNA to mRNA ribosome complexes by binding to the 30S ribosomal subunit in the mRNA translation complex. Tetracyclines include demeclocycline, doxycycline, minocycline, oxytetracycline, and tetracycline. Other antibiotics that can be used to treat bacterial infections include arsphenamine, chloramphenicol, fosfomycin, fusidic acid, linezolid, metronidazole, mupirocin, platensimycin, quinupristin / dalfopristin, rifaximin, thiamphenicol, tigecycline, tinidazole, clofazimine, dapsone, capreomycin, cycloserine, ethambutol, ethionamide, isoniazid, pyrazinamide, rifampin, rifamycin, rifabutin, rifapentine, and streptomycin.

[0144] Antiviral therapies are a class of drugs specifically used to treat viral infections. Like antibiotics, specific antiviral drugs are also used for specific viruses. They are relatively harmless to the host and thus can be used to treat infections. Antiviral therapies can inhibit various stages of the viral life cycle. For example, antiviral therapies can inhibit the attachment of a virus to a cell receptor. Such antiviral therapies can include agents that mimic viral associated proteins (VAPs) and bind to cell receptors. Other antiviral therapies can inhibit viral entry, viral uncoating (e.g., amantadine, rimantadine, pleconaril), viral synthesis, viral integration, viral transcription, or viral translation (e.g., favipiravir). In some cases, antiviral therapies are morpholino antisense therapies. Antiviral therapies should be distinguished from virucidal agents, which can effectively inactivate viral particles in vitro.

[0145] Many available antiviral drugs are designed to treat infections of retroviruses (primarily HIV). Antiretroviral drugs can include classes of protease inhibitors, reverse transcriptase inhibitors, and integrase inhibitors. Drugs to treat HIV can include protease inhibitors (e.g., Fortovase, saquinavir, Kaletra, lopinavir, fosamprenavir, fosamprenavir, Norvir, ritonavir, Reyataz, duranavir, Selzentry, Cobic, Atripla), integrase inhibitors (e.g., raltegravir), transcriptase inhibitors (e.g., Epivir, Epzicom, Agenerase, amprenavir, Aptivus, tipranavir, Intelence™, Etrivirine, Edurant, Sustiva, Rescriptor, Intelence, Viread, Epzicom), reverse transcriptase inhibitors (e.g., Delavirdine, efavirenz, Epzicom, havid, nevirapine, zidovudine, AZT, stuvadine, Truvada, Combivir), fusion inhibitors (e.g., Fuzeon, enfuvirtide), chemokine co-receptor antagonists (e.g., Selzentry, emtriva, Emtricitabine, Epzicom, or Trizivir). Alternatively, antiretroviral therapy can be a combination therapy such as Atripla (e.g., Efavirenz, Emtricitabine, and Tenofovir disoproxil fumarate) and Complera (Emtricitabine, Rilpivirine, and Tenofovir disoproxil fumarate). Herpes viruses that can cause cold sores and genital herpes are typically treated with the nucleoside analogue acyclovir. Viral hepatitis (A-E) is caused by five unrelated hepatotropic viruses and is also typically treated with antiviral drugs depending on the type of infection. Influenza A and B viruses are important targets for developing new treatments for influenza to overcome resistance to existing neuraminidase inhibitors such as oseltamivir.

[0146] In some cases, the antiviral therapy can comprise a reverse transcriptase inhibitor. The reverse transcriptase inhibitor can be a nucleoside reverse transcriptase inhibitor or a non-nucleoside reverse transcriptase inhibitor. Nucleoside reverse transcriptase inhibitors can include, but are not limited to, Combivir, emtriva, Epzicom, havid, Trizivir, Epzicom, Truvada, Combivir ec, Epzicom, Viread, Epzicom, and Atripla. Non-nucleoside reverse transcriptase inhibitors can include edurant, Intelence, rescriptor, sustiva, and Viread (immediate release or extended release).

[0147] Protease inhibitors are another example of antiviral drugs, which can include, but are not limited to, agenerase, aptivus, crixivan, fortovase, invirase, kaletra, lexiva, norvir, prevonar, viracept, and viramidine. Alternatively, the antiviral therapy can comprise a fusion inhibitor (e.g., Enfuviride) or an entry inhibitor (e.g., Maraviroc).

[0148] Other examples of antiviral drugs include abacavir, acyclovir, adefovir, amantadine, amprenavir, ampligen, arbidol, atazanavir, atripla, boceprevir, cidofovir, combivir, darunavir, delavirdine, didanosine, docosanol, edoxudine, efavirenz, emtricitabine, enfuvirtide, entecavir, famciclovir, fomivirsen, fosamprenavir, foscarnet, phosphonate, fusion inhibitor, ganciclovir, ibacitabine, imunovir, idoxuridine, imiquimod, indinavir, inosine, integrase inhibitor, interferons (e.g., types I, II, III interferons), lamivudine, lopinavir, loviride, maraviroc, moroxydine, methisazone, nelfmavir, nevirapine, nexavir, nucleoside analogues, oseltamivir, peginterferon alfa-2a, penciclovir, peramivir, pleconaril, podophyllotoxin, protease inhibitor, raltegravir, reverse transcriptase inhibitor, ribavirin, rimantadine, ritonavir, pyramidine, saquinavir, stavudine, tea tree oil, tenofovir, tenofovir disoproxil, tipranavir, trifluridine, trizivir, tromantadine, truvada, valaciclovir, valganciclovir, vicriviroc, vidarabine, viramidine, zalcitabine, zanamivir, and zidovudine.

[0149] Antifungal drugs are drugs that can be used to treat fungal infections, such as athlete's foot, ringworm, candidiasis (thrush), severe systemic infections (e.g., cryptococcal meningitis), and the like. Antifungals work by exploiting differences between mammalian and fungal cells to kill the fungal organism. Unlike bacteria, fungi and humans are eukaryotes. Thus, fungal and human cells are similar at the molecular level, which makes it more difficult to find targets for antifungal drugs to attack, which are not present in the infected organism.

[0150] Antiparasitic drugs are a class of drugs used to treat parasites such as nematodes, cestodes, flukes, infectious protozoa, and amoebae. Like antifungals, they must kill the infectious pest without severely damaging the host.

[0151] Methods of the disclosure can be implemented by a system, a kit, a library, or a combination thereof. Methods of the disclosure can include one or more systems. Systems of the disclosure can be implemented by a kit, a library, or both. A system can include one or more components that perform any of the methods or operations of any of the methods disclosed herein. For example, a system can include one or more kits, devices, libraries, or a combination thereof. A system can include one or more sequencers, processors, memory locations, computers, computer systems, or a combination thereof. A system can include a transmission device.

[0152] A kit can include various reagents for performing various operations disclosed herein, including sample processing and / or analysis operations. A kit can include instructions for performing at least some of the operations disclosed herein. A kit can include one or more capture probes, one or more beads, one or more labels, one or more linkers, one or more devices, one or more reagents, one or more buffers, one or more samples, one or more databases, or a combination thereof.

[0153] A library can include one or more capture probes. A library can include one or more subsets of nucleic acid molecules. A library can include one or more databases. A library can be generated or produced by any of the methods, kits, or systems disclosed herein. A database library can be generated from one or more databases. Methods for generating one or more libraries can include: (a) aggregating information from one or more databases to generate an aggregated dataset; (b) analyzing the aggregated dataset; (c) generating one or more database libraries from the aggregated dataset. Figure 13 Examples of library construction workflows are provided. In some cases, libraries can be merged, Figure 7 .

[0154] Computer system The present disclosure provides a computer system programmed to implement the methods of the present disclosure. Figure 5 A computer system 501 is shown that is programmed or otherwise configured to map and / or align sequence reads to identify a source of a nucleic acid molecule (e.g., human or non-human, host or non-host), identify one or more features (e.g., genetic variants), or any combination thereof. The computer system 501 can modulate various aspects of processing sequencing information as provided in the present disclosure, e.g., aligning sequence reads to one or more reference sequences to identify a source of nucleic acid sequences in a biological sample. The computer system 501 can be an electronic device of a user or a computer system located remotely with respect to the electronic device. The electronic device can be a mobile electronic device.

[0155] The computer system 501 includes a central processing unit (CPU, also "processor" and "computer processor" herein) 505, which can be a single core or multi core processor, or a plurality of processors for parallel processing. The computer system 501 also includes a memory, or a memory location 510 (e.g., random access memory, read only memory, flash memory), storage 515 (e.g., hard disk), a communication interface 520 (e.g., network adapter) for communicating with one or more other systems, and various peripheral devices 525, such as a cache, other memory, data storage and / or electronic display adapters. The memory 510, storage 515, interface 520 and peripheral devices 525 are in communication with the CPU 505 through a communication bus, such as a motherboard. The storage 515 can be a data storage unit (or data repository) for storing data. The computer system 501 can be operatively coupled to a computer network ("network") 530 by means of the communication interface 520. The network 530 can be the Internet, an internet and / or an extranet, or an intranet and / or extranet that in turn can use a private communication network that is communicatively coupled to the Internet. In some cases, the network 530 is a telecommunication and / or data network. The network 530 can include one or more computer servers that enable distributed computing, such as cloud computing. In some cases, the network 530 can implement a peer-to-peer network, by means of the computer system 501, that enables devices coupled to the computer system 501 to behave as a client or a server.

[0156] The CPU 505 can execute a sequence of machine-readable instructions, which can be embodied in a program or software. The instructions can be stored in a memory location, such as the memory 510. The instructions can be directed to the CPU 505, which can subsequently program or otherwise configure the CPU 505 to implement methods of the present disclosure. Examples of operations performed by the CPU 505 can include fetch, decode, execute, and writeback.

[0157] The CPU 505 can be a part of a circuit, such as an integrated circuit. One or more other components of the system 501 can be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).

[0158] The storage 515 can store files, such as drivers, libraries and saved programs. The storage 515 can store user data, e.g., user preferences and user programs. The computer system 501 in some cases can include one or more additional data storage units that are external to the computer system 501, such as located on a remote server that is in communication with the computer system 501 through an intranet or the Internet.

[0159] The computer system 501 can communicate over the network 530 with one or more remote computer systems. For instance, the computer system 501 can communicate with a remote computer system of a user (e.g., a healthcare provider). Examples of remote computer systems include personal computers (e.g., portable PC), slate or tablet computers (e.g., Apple® iPad, Samsung® Galaxy Tab), telephones, smart phones (e.g., Apple® iPhone, Android-enabled device, Blackberry®), or personal digital assistants. The user can access the computer system 501 via the network 1130.

[0160] The methods described herein can be implemented by way of machine (e.g., computer processor) executable code stored on an electronic storage location of the computer system 501, such as, for example, the memory 510 or electronic storage 515. The machine executable or machine readable code can be provided in the form of software. During use, the code can be executed by the processor 505. In some cases, the code can be retrieved from the storage 515 and stored on the memory 510 so that the processor 505 can access the code for processing during use. In some cases, the electronic storage 515 can be eliminated, and the machine executable instructions stored in the memory 510.

[0161] The code can be pre-compiled and configured for use with a machine having a processor adapted to execute the code, or can be compiled for use at runtime. The code can be provided in an object code format that is executable by a processor of the computer system 501, or in a source code format that can be compiled by a processor of the computer system 501.

[0162] Various aspects of the systems and methods (e.g., computer system 501) provided herein can be embodied in programming. Various aspects of the technology can be thought of as "products" or "articles of manufacture" typically in the form of machine (or processor) executable code and / or associated data that is carried on or embodied in a type of machine readable medium. Machine-executable code can be stored on an electronic storage unit, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. "Storage" type media can include any or all of the tangible memory of the computers, processors or the like, or associated modules thereof, such as various semiconductor memories, tape, floppy disks, and the like. All or portions of the software can at times be communicated through an Internet or various other telecommunication networks. Such communications, for example, can enable loading of the software from one computer or processor into another computer or processor, for example, from a management server or host computer into the computer platform of an application server. Thus, another type of media that can bear the software elements includes optical, electrical and electromagnetic waves, such as used across physical interfaces, through wired and optical landline networks and over various air-links. The physical elements that carry such waves, such as wired or wireless links, optical interfaces, etc., also can be considered as media bearing software. As used in this paper, unless restricted to non-transitory, tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.

[0163] Accordingly, a machine readable medium, such as a computer-readable medium, can take many forms, including but not limited to, a tangible storage medium, a carrier wave medium or physical transmission medium. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer(s) or the like, such as can be used to implement the databases, etc. shown in the drawings. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optic cables, including the wires that comprise a bus within a computer system. Carrier-wave transmission media can take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media therefore include for example: a floppy disk, a flexible disk, a hard disk, magnetic tape, any other magnetic medium, a CD-ROM, DVD or DVD-ROM, any other optical medium, punch cards, any other physical storage medium with patterns of holes, a RAM, a ROM, a PROM and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave transporting data or instructions, cables or links transporting such a carrier wave, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer readable media can be involved in carrying one or more sequences of one or more instructions to a processor for execution.

[0164] The computer system 501 can include or be in communication with an electronic display 535, which can comprise a user interface (UI) 540 for providing, for example, one or more biomedical reports including one or more sets of data selected from: i) candidate tumor neoantigens, (ii) detected non-human species, (iii) detected CDR3 sequences, and any combination thereof. Examples of UIs include, without limitation, a graphical user interface (GUI) and a web-based user interface.

[0165] The methods and systems of the present disclosure can be implemented by one or more algorithms. The algorithms can be implemented in software when executed by the central processing unit 505. The algorithms may, for example, map and / or align sequence reads, call variants, annotate sequence information, or any combination thereof.

[0166] EMBODIMENTS Embodiment 1. Preparation of genomic DNA A subset of nucleic acid molecules was prepared from a sample comprising genomic DNA using the following procedure: 1. Shear the sample comprising genomic DNA with the M220 for 15-35 seconds.

[0167] 2. Purify the fragmented gDNA with SPRI beads (1 to 1 ratio of SPRI beads to DNA sample by volume) and elute the DNA into 100 μΐ^of elution buffer (EB).

[0168] 3. Add 50 μΐ^of SPRI beads to 100 μΐ^of DNA.

[0169] 4. Transfer the supernatant to a new tube.

[0170] 5. Elute the DNA from the remaining DNA-bound beads. This eluted DNA is referred to as long inserts.

[0171] 6. Add 10 μΐ^of SPRI beads to the supernatant of Step 4.

[0172] 7. Transfer the supernatant of Step 6 to a new tube.

[0173] 8. Elute the DNA from the remaining DNA-bound beads of Step 6. This eluted DNA is referred to as medium inserts.

[0174] 9. Add 20 μΐ^of SPRI beads to the supernatant of Step 7.

[0175] 10. Transfer the supernatant of Step 9 to a new tube.

[0176] 11. Elute the DNA from the remaining DNA-bound beads of Step 9. This eluted DNA is referred to as short inserts.

[0177] Example 2. Obtaining a biological sample A subject with an evaluable metastatic cancer is subjected to tumor resection. Lymphocytes, tumor infiltrating lymphocytes (TILs), from the tumor are grown and expanded. Multiple single fragments or multiple single cultures of TILs are grown. The single cultures are amplified individually and when sufficient TIL yield (about 10 8 cells) is amplified from each culture, the TILs are cryopreserved and aliquots are taken for immunological testing. Aliquots of the original tumor are subjected to exome and transcriptome sequencing to identify mutations that are uniquely present in the tumor as compared to normal cells. Sequencing also identifies the presence and identity of non-human genomes, including microorganisms.

[0178] Example 3. Extracting genomic material from a biological sample Genomic DNA (gDNA) and total RNA were purified from various tumor and matched normal blood apheresis samples using the QIAGEN AllPrep DNA / RNA kit (cat# 80204) according to the manufacturer’s recommendations. Tumor samples were formalin-fixed, paraffin-embedded (FFPE) and gDNA was extracted using the Covaris truXTRACTM FFPE DNA kit following the manufacturer’s instructions.

[0179] Example 4. Sequencing analysis of biological samples Whole exome library construction and approximately 20,000 coding gene exome capture was performed using Agilent Technologies SureSelect XT Target Enrichment System (Cat# 5190-8646) for the coupled Human All Exon V6 RNA bait (Cat# 5190-8863) (Agilent Technologies, Santa Clara, CA, USA) and bacterial RNA bait. Whole exome sequencing (WES) libraries were sequenced on a NextSeq 500 benchtop sequencer (Illumina, San Diego, CA, USA) subsequently. Libraries were prepared using 3 μg gDNA from fresh tumor tissue samples and 200 ng gDNA from FFPE tumor samples following the manufacturer's protocol. Double-end sequencing was completed using Illumina High Output Flow Cell Kit (300 cycles) (Cat# FC-404-2004). Tu-1, Tu-2A, and Tu-2B samples were initially run on v1 of the reagent / flow cell kit, with subsequent runs of the same library preparation using v2 of the reagent / flow cell kit. Tumor samples were run on v2 reagent / flow cell kit. The average sequencing depth and percentage of tumor (tumor purity) in each sample was determined as estimated by the tumor deconvolution bioinformatics program Allele-Specific Copy Number Analysis (ASCAT) 1. RNA-seq libraries were prepared using 2 μg total RNA and Illumina TruSeq RNA strand library preparation kit following the manufacturer's protocol. RNA-seq libraries were double-end sequenced on a NextSeq 500 benchtop sequencer (Illumina, San Diego, CA, USA). WES was aligned, processed, and variant called using novoalign MPI from novocraft (http: / / www.novocraft.com / ) to align to human genome build hg19. MarkDuplicates tool from Picard was used to mark duplicates. In / del realignment and base recalibration was performed according to the GATK best practices workflow (https: / / www.broadinstitute.org / gatk / ). After cleaning the data, the following criteria were used to create pileup files using samtools (http: / / samtools.sourceforge.net) and call somatic variants using Varscan2 (http: / / varscan.sourceforge.net): tumor and normal read counts of 10 or more, variant allele frequency of 10% or more, tumor variant reads of 4 or more.The variants were then annotated using Annovar (http: / / annovar.openbioinformatics.org). For RNA-seq, STAR (https: / / github.com / alexdobin / STAR) two pass method was used to align to the human genome build hg19. The MarkDuplicates tool from Picard was used to mark duplicates. The GATK SplitNTrim tool was used to split and trim reads. Afterwards, the GATK toolbox was used to perform In / del realignment and base recalibration. The final recalibrated bam files and samtools mpileup were used to create pileup files. Finally, Varscan2 was used to call variants.

[0180] Example 5: Sequence alignment of non-human genomes Genomic information extracted from the sequencing analysis was aligned to the ribosomal RNA gene 16S (genomic reference). Reads that matched exactly to the 16S gene were identified. Capture probes designed to hybridize to the shared regions of the 16S gene can be used to capture nucleic acid molecules from a variety of species, even species that have not yet been identified and characterized. Based on the sequence from the variable region portion, capture molecules that extend their sequence from these shared regions into the variable region were assigned to their source species.

[0181] Example 6. Shearing time and fragment size Genomic DNA (gDNA) was sheared by varying the shearing time of the Covaris settings. The gDNA fragments produced by the various shearing times were then analyzed. The results are shown in Figure 10 and Table 1.

[0182] Table 1. Shearing time and average fragment size

[0183] Example 7. Bead ratio and fragment size The ratio of bead volume to nucleic acid sample volume was varied and the effect of these ratios on the average fragment size was analyzed. As shown in Figure 11 the volume ratio of bead volume to nucleic acid sample volume was varied from 0.8 (line 1), 0.7 (line 2), 0.6 (line 3), 0.5 (line 4), and 0.4 (line 5), resulting in a change in the average size of the DNA fragments. Generally, it appears that the lower the ratio, the larger the average fragment size.

[0184] Example 8. Ligation reaction and fragment size Two different combinations of shearing times and three different ligation reactions were performed on nucleic acid samples. Sample 1 was sheared for 25 seconds and ligation reactions were performed on long insert DNA. Sample 2 was sheared for 32 seconds and ligation reactions were performed on long insert DNA. Sample 3 was sheared for 25 seconds and ligation reactions were performed on medium insert DNA. Sample 4 was sheared for 32 seconds and ligation reactions were performed on medium insert DNA. Sample 5 was sheared for 25 seconds and ligation reactions were performed on short insert DNA. Sample 6 was sheared for 32 seconds and ligation reactions were performed on short insert DNA. Figure 12 The average fragment size of the six reactions is shown.

[0185] Example 9. Capture of nucleic acids of interest A nucleic acid sample is obtained from a subject. This sample can be split into two separate samples and each separate sample can be subjected to different probe pools and conditions. A first pool of biotinylated capture probes is generated by combining an Agilent Clinical Research Exome Kit (based on the exome in the GRCh37 reference genome) with additional probes of interest including exome regions corresponding to the GRCh38 reference sequence, HLA specific probes, T cell receptor and B cell receptor recombination specific probes (i.e. regions corresponding to V(D)J regions), microsatellite instability regions, and tumor virus sequences. Additional probes are titrated into the pool to adjust the relative capture rate of nucleic acids. This can be beneficial for downstream sequencing reactions to increase the depth or sensitivity of sequencing of a set of captured nucleic acids. A second pool of probes (e.g. a “booster set” of probes targeting specific sequences or subsets of sequences) can capture nucleic acids with difficult to capture sequences. The probes of the booster set can include exomes of the GRCh37 and GRCh38 reference sequences with high GC content, probes relevant to cancer therapy, additional T cell receptor and B cell recombination specific probes. The nucleic acid molecules captured by the first pool and the nucleic acid molecules captured by the second pool are combined to generate a combined pool of capture probes and captured nucleic acids. The combined pool can be created by combining the two pools of captured molecules in different proportions to adjust the relative amount of captured nucleic acids corresponding to each pool. This can be beneficial for downstream sequencing reactions to increase the depth or sensitivity of a set of captured nucleic acids. The hybridized capture probes and captured nucleic acids are incubated with magnetic streptavidin beads. A magnetic separator is placed outside of the tube containing the sample and the magnetic streptavidin beads are allowed to settle. The liquid is poured off and additional buffer is added to wash the beads and remove any unbound nucleic acids. An amplification reaction is performed on the captured nucleic acids to attach adapters for further downstream sequencing reactions.

[0186] Example 10. Identification of microbiome and tumor viruses A nucleic acid sample is obtained from a subject. This sample can be split into two separate samples and each separate sample can be subjected to different probe pools and conditions. A first pool of biotinylated capture probes is generated by combining the Agilent Clinical Research Exome Kit (based on the exome in the GRCh37 reference genome) with additional probes of interest including probes homologous to the 16S rRNA region of bacterial, fungal, archaeal, and protist genes including genes of Helicobacter pylori Helicobacter pylori ) and Clostridium Fusobacterium ), probes of viral genes including human papillomavirus, hepatitis B virus, hepatitis C virus, and probes of genes associated with pathogenicity. Additional probes are titrated into the pool to adjust the relative capture rate of nucleic acids. A second pool of probes (e.g., a "booster set" of probes targeting specific sequences or subsets of sequences) can capture nucleic acids with difficult to capture sequences. The probes of the booster set can include exomes of the GRCh37 and GRCh38 reference sequences with high GC content or low gene homology due to multiple alternative gene sequences or genes with high mutation rates. The nucleic acid molecules captured by the first pool and the nucleic acid molecules captured by the second pool are combined to generate a combined pool of nucleic acids. The two pools are combined and titrated at different ratios to improve the sensitivity of one pool relative to the other. The hybridized capture probes and captured nucleic acids are incubated with magnetic streptavidin beads. A magnetic separator is placed outside of the tube containing the sample and capture probes and it immobilizes the magnetic streptavidin beads. The liquid is poured off and additional buffer is added to wash the beads and remove any unbound nucleic acids. An amplification reaction is performed on the captured nucleic acids to add adapters for further downstream sequencing reactions. The sequences are aligned to reference sequences to identify the presence of specific microorganisms in the microbiome of the subject or to identify any disease-causing pathogens.

[0187] Example 11. Identifying or CAR-T cells A nucleic acid sample is obtained from a subject. Two probe pools are generated, where the first pool consists of biotinylated Agilent Clinical Research Exome Kit probes. A supplemental set of biotinylated probes is added to the first pool, which includes sequences specific to a chimeric antigen receptor found in CAR-T cells. The second biotinylated capture probe pool (e.g., a "booster set" of probes targeting specific sequences or subsets of sequences) includes sequences against CAR sequences with high GC content and / or sequences with low overall sequence homology to the captured nucleic acids due to mutations or recombination. The sample is split into two sample pools, the first sample pool is subjected to the first probe pool and the second sample pool is subjected to the second probe pool to allow the capture probes to capture nucleic acids from each sample pool. The first probe pool and the second probe pool hybridized to the captured nucleic acids are combined. Magnetic streptavidin beads are added to the mixture and the biotinylated probes bind to the beads. A magnetic separator is placed outside of the tube containing the sample and capture probes and it immobilizes the magnetic streptavidin beads. The liquid is poured off and additional buffer is added to wash the beads and remove any unbound nucleic acids. The captured nucleic acids are resuspended in fresh buffer and an amplification reaction is performed to append adapters, then the captured nucleic acids are sequenced. The sequence reads are analyzed and aligned to a CAR-T gene reference sequence to identify the presence of CAR-T related nucleic acids.

[0188] Example 12: Identification of a fragmented transcriptome Samples are obtained from different tissue types from a subject. The samples are treated by enzyme digestion to remove DNA and multiple types of RNA molecules (rRNA, tRNA, miRNA). The mRNA is not digested and is then reverse transcribed to generate cDNA molecules. A set of biotinylated probes with sequence homology to approximately 20,000 genes are generated and mixed with the cDNA samples. Magnetic streptavidin beads are added to the mixture and the biotinylated probes bind to the beads. A magnetic separator is placed outside of the tube containing the sample and capture probes and it immobilizes the magnetic streptavidin beads. The liquid is poured out and additional buffer is added to wash the beads and remove any unbound nucleic acids. The captured nucleic acids are resuspended in fresh buffer and amplification reactions are performed to add adaptors and the captured nucleic acids are sequenced. The sequence reads are analyzed and aligned to reference sequences to determine the identity of the genes in each sample and correlate them to the tissue type of the sample. Multiple replicates are performed on samples where the biotinylated probes are titrated at different ratios. By analyzing the specific signal of the nucleic acids in relation to the amount of probes provided to the sample, the expression level of each gene can be determined. Transcriptome analysis can be performed in parallel with exome or genome analysis. Exome or genome probes can be used as the first probe pool and cDNA capture probes can be used as the second probe pool. In this case, the sample can be split into several parts and DNA specific or mRNA specific reactions are performed. The captured nucleic acids can be pooled together and a sequencing reaction is performed to determine the identity of the nucleic acids as disclosed in previous examples.

[0189] Example 13. Identification and capture of cell-free DNA (cfDNA) associated with cancer Nucleic acid sample extracts are obtained from whole blood or serum. Cells are removed from the blood sample by centrifugation, resulting in a cell-free sample. A pool of biotinylated probes with sequence homology to approximately 20,000 genes is generated and mixed with the cell-free sample. A first pool of biotinylated capture probes is generated by combining the Agilent Clinical Research Exome Kit (based on the exome in the GRCh37 reference genome) with additional probes of interest (including exome regions corresponding to the GRCh38 reference sequence) and probes with homology to genes associated with cancer. Additional probes are titrated into the pool to adjust the relative capture rate of nucleic acids. A second pool of probes (e.g., a “booster set” of probes targeting specific sequences or subsets of sequences) can capture nucleic acids with difficult-to-capture sequences. Probes of the booster set can include exomes of the GRCh37 and GRCh38 reference sequences with high GC content or low gene homology due to multiple alternative gene sequences or genes with high mutation rates. Magnetic streptavidin beads are added to the mixture and the biotinylated probes bind to the beads. A magnetic separator is placed outside of the tube containing the sample and capture probes and it immobilizes the magnetic streptavidin beads. The liquid is poured off and additional buffer is added to wash the beads and remove any unbound nucleic acids. The captured nucleic acids are resuspended in fresh buffer and an amplification reaction is performed to append adapters, and the captured nucleic acids are sequenced. Sequence reads are analyzed and aligned to reference sequences to determine the identity of genes in each sample. cfDNA analysis can be performed in parallel with exome sequencing of genomic DNA by applying genomic DNA to the same or similar set of probes as performed for cfDNA.

[0190] The present invention also provides the following items: 1. A method for processing a biological sample of a subject, comprising: (a) generating a subset of nucleic acid molecules from the biological sample using a pool of nucleic acid probes, wherein the probes comprise (i) a first plurality of nucleic acid probes configured to target elements of a human genome and (ii) a second plurality of nucleic acid probes configured to target elements of one or more non-human genomes; and (b) performing an assay on the subset of nucleic acid molecules to generate sequence information comprising (i) human nucleic acids from the biological sample of the subject and (ii) non-human nucleic acids from the biological sample of the subject.

[0191] 2. The method of item 1, wherein the first plurality of nucleic acid probes of (i) are configured to target elements derived from the human genome.

[0192] 3. The method of item 1, wherein the second plurality of nucleic acid probes are configured to target elements of non-human genomic sequences selected from one or more species selected from the group consisting of viruses, bacteria, bacteriophages, fungi, protists, archaea, amoebas, worms, algae, genetically modified cells, and genetically modified vectors.

[0193] 4. The method of item 1, wherein generating the subset of nucleic acid molecules from the biological sample comprises performing one or more hybridization reactions.

[0194] 5. The method of item 1, further comprising obtaining the biological sample from the subject, wherein the biological sample is derived from a tumor biopsy, whole blood, or plasma.

[0195] 6. The method of item 1, further comprising aligning the sequences of the subset of nucleic acid molecules to one or more reference sequences.

[0196] 7. The method of item 6, wherein the one or more reference sequences comprise a plurality of reference sequences, and wherein the plurality of reference sequences correspond to two or more different species.

[0197] 8. The method of item 6, further comprising identifying the origins of the nucleic acid molecules in the subset based on the alignment.

[0198] 9. The method of item 8, further comprising generating an output comprising the origins of the nucleic acid molecules in the biological sample.

[0199] 10. The method of item 1, wherein the concentration of the second plurality of nucleic acid probes is greater than the concentration of the first plurality of nucleic acid probes in the pool of nucleic acid probes.

[0200] 11. The method of item 1, wherein the concentration of the second plurality of nucleic acid probes in the pool of probes is greater than the concentration of the first plurality of nucleic acid probes in the pool of nucleic acid probes.

[0201] 12. The method of item 1, wherein the first plurality of nucleic acid probes comprises a human exome capture probe set.

[0202] 13. The method of item 1, wherein the first plurality of nucleic acid probes comprises probes configured to target junction sequences resulting from human V(D)J rearrangement or recombination.

[0203] 14. The method of item 1, wherein the second plurality of nucleic acid probes comprises one or more probes configured to target one or more elements of a human papillomavirus E6 gene, an E7 gene, and / or a bacterial 16S ribosomal RNA gene.

[0204] 15. The method according to Project 1, wherein the determination in (b) comprises performing sequencing to produce paired-end reads of 130 to 280 bases in length.

[0205] 16. The method according to Project 1 further includes generating one or more biomedical reports, said one or more biomedical reports containing one or more sets of data selected from: (i) candidate tumor neoantigens, (ii) detected non-human species, (iii) detected complementarity-determining region 3 (CDR3) sequences, and any combination thereof.

[0206] 17. The method according to item 16, wherein the detected non-human species is an antigen, and wherein the CDR3 sequence is capable of binding the antigen.

[0207] 18. The method according to Item 16, wherein the one or more biomedical reports comprise any two of (i)-(iii).

[0208] 19. A method for processing a biological sample of a subject, comprising: (a) A subgroup of nucleic acid molecules generated from the biological sample, wherein the subgroup of nucleic acid molecules includes (i) a first plurality of nucleic acid molecules from the subject and (ii) a second plurality of nucleic acid molecules not from the subject, and wherein the abundance of the first plurality of nucleic acid molecules in the biological sample is greater than the abundance of the second plurality of nucleic acid molecules; and (b) Determining subgroups of the nucleic acid molecules to generate sequence information including: (i) the first plurality of nucleic acid molecules and (ii) the second plurality of nucleic acid molecules.

[0209] 20. The method according to Item 19, wherein the nucleic acid molecule in (i) is derived from the genome of the subject.

[0210] 21. The method according to item 19, wherein the second plurality of nucleic acid molecules not from the subject includes one or more members selected from: viruses, bacteria, bacterial phages, fungi, protozoa, archaea, amoebas, worms, algae, genetically modified cells, and genetically modified vectors.

[0211] 22. The method according to item 19, wherein the subgroup of generating the nucleic acid molecule from the biological sample includes performing one or more hybridization reactions.

[0212] 23. The method according to item 19 further includes obtaining the biological sample from the subject, wherein the biological sample is derived from a tumor biopsy, whole blood, or plasma.

[0213] 24. The method of item 19, further comprising aligning the sequences of the subset of nucleic acid molecules to one or more reference sequences.

[0214] 25. The method of item 24, wherein the one or more reference sequences comprises a plurality of reference sequences, and wherein the plurality of reference sequences correspond to two or more different species.

[0215] 26. The method of item 24, further comprising identifying the origin of the nucleic acid molecules in the subset based on the alignment.

[0216] 27. The method of item 26, further comprising generating an output comprising the identified origins of the nucleic acid molecules in the biological sample.

[0217] 28. The method of item 19, wherein the abundance of the first plurality of nucleic acid molecules is greater than the abundance of the first plurality of nucleic acid molecules in the biological sample.

[0218] 29. The method of item 19, wherein the relative abundance of the second plurality of nucleic acid molecules in the subset is greater than the relative abundance of the first plurality of nucleic acid molecules in the subset.

[0219] 30. A system for processing a biological sample of a subject, comprising: a processing unit comprising one or more computer processors individually or collectively programmed to assay a subset of nucleic acid molecules to generate sequence information comprising sequences of (i) human nucleic acids from the biological sample of the subject and (ii) non-human nucleic acids from the biological sample of the subject, the subset of nucleic acid molecules generated from the biological sample using a pool of nucleic acid probes, wherein the probes comprise (i) a first plurality of nucleic acid probes configured to target elements of a human genome; and (ii) a second plurality of nucleic acid probes configured to target elements of one or more non-human genomes; and a computer memory configured to store the sequence information.

[0220] 31. The system of item 30, wherein the one or more computer processors are programmed to generate an alignment of the sequences to one or more reference sequences.

[0221] 32. The system of item 31, wherein the one or more reference sequences comprises a plurality of reference sequences, and wherein the plurality of reference sequences correspond to two or more different species.

[0222] 33. The system of item 31, wherein the one or more computer processors are programmed to identify, based on the alignment, an origin of the nucleic acid molecules in the subset.

[0223] 34. The system of item 33, wherein the one or more computer processors are programmed to generate an output comprising the origin of the nucleic acid molecules in the biological sample.

[0224] 35. The system of item 30, wherein the one or more computer processors are programmed to produce one or more biomedical reports comprising information selected from the group consisting of: (i) candidate tumor neoantigens, (ii) detected non-human species, (iii) detected complementarity determining region 3 (CDR3) sequences, and any combination thereof.

[0225] 36. A system for processing a biological sample of a subject, comprising: a processing unit comprising one or more computer processors individually or collectively programmed to assay a subset of nucleic acid molecules to produce sequence information comprising: (i) a first plurality of nucleic acid molecules and (ii) a second plurality of nucleic acid molecules, the subset of nucleic acid molecules being produced from the biological sample, wherein the subset of nucleic acid molecules comprises (i) a first plurality of nucleic acid molecules from the subject and (ii) a second plurality of nucleic acid molecules not from the subject, and wherein the abundance of the first plurality of nucleic acid molecules is greater than the abundance of the second plurality of nucleic acid molecules in the biological sample; and a computer memory configured to store the sequence information.

[0226] 37. The system of item 36, wherein the one or more computer processors are programmed to generate an alignment of the sequences to one or more reference sequences.

[0227] 38. The system of item 37, wherein the one or more reference sequences comprise a plurality of reference sequences, and wherein the plurality of reference sequences correspond to two or more different species.

[0228] 39. The system of item 37, wherein the one or more computer processors are programmed to identify, based on the alignment, an origin of the nucleic acid molecules in the subset.

[0229] 40. The system of item 38, wherein the one or more computer processors are programmed to generate an output comprising the origin of the nucleic acid molecules in the biological sample.

[0230] 41. The system of item 36, wherein the one or more computer processors are programmed to produce one or more biomedical reports comprising information selected from the group consisting of: (i) candidate tumor neoantigens, (ii) detected non-human species, (iii) detected complementarity determining region 3 (CDR3) sequences, and any combination thereof.

[0231] While preferred embodiments of the application have been shown and described herein, it will be apparent to those skilled in the art that many more modifications, variations, and alternatives of the present application can be made. It is therefore contemplated to cover any and all modifications, variations, and alternatives that fall within the scope of the present application. It is to be understood that the foregoing description is merely that of a preferred embodiment of the application and that various changes, modifications, and alterations can be made thereto by those skilled in the art without departing from the intended spirit and scope thereof. The description is not intended to limit the scope of the application, but is merely to provide a preferred example of the application.

Claims

1. A system for processing a biological sample of a subject, comprising: a processing unit comprising one or more computer processors individually or collectively programmed to perform an assay on a subset of nucleic acid molecules to produce sequence information comprising sequences of (i) human nucleic acids from the biological sample of the subject and (ii) non-human nucleic acids from the biological sample of the subject, wherein the non-human nucleic acids comprise a subset of non-human nucleic acids derived from one or more oncogenic viruses or bacteria associated with cancer, the subset of nucleic acid molecules being produced from the biological sample using a pool of nucleic acid probes, wherein the probes comprise (i) a first plurality of nucleic acid probes comprising a human exome capture probe set, and probes formulated to target junction sequences resulting from human V(D)J rearrangement or recombination; (ii) a second plurality of nucleic acid probes configured to target one or more elements of non-human genomes, and (iii) wherein, in the pool of nucleic acid probes, the concentration of the first plurality of nucleic acid probes is greater than the concentration of the second plurality of nucleic acid probes, wherein the one or more elements of non-human genomes comprise one or more elements of a human papillomavirus E6 gene, E7 gene, and / or bacterial 16S ribosomal RNA gene; and a computer memory configured to store the sequence information.

2. The system of claim 1, wherein the first plurality of nucleic acid probes of (i) are configured to target elements derived from the human genome.

3. The system of claim 1, wherein producing the subset of nucleic acid molecules from the biological sample comprises performing one or more hybridization reactions.

4. The system of claim 1, wherein the biological sample is derived from a tumor biopsy, whole blood, or plasma obtained from the subject.

5. The system of claim 1, wherein the assay comprises performing sequencing to produce paired-end read sequences of 130 bases to 280 bases in length.

6. The system of claim 1, wherein the one or more computer processors are programmed to generate an alignment of the sequences to one or more reference sequences.

7. The system of claim 6, wherein the one or more reference sequences comprise a plurality of reference sequences, and wherein the plurality of reference sequences correspond to two or more different species.

8. The system of claim 6, wherein the one or more computer processors are programmed to identify, based on the alignment, a source of the nucleic acid molecules in the subset.

9. The system of claim 8, wherein the one or more computer processors are programmed to generate an output comprising the source of the nucleic acid molecules in the biological sample.

10. The system of claim 1, wherein the one or more computer processors are programmed to generate one or more biomedical reports comprising information selected from the group consisting of: (i) candidate tumor neoantigens, (ii) detected non-human species, (iii) detected complementarity determining region 3 (CDR3) sequences, and any combination thereof.