Compositions, methods and systems for processing or analyzing multi-species nucleic acid samples
The method and system for simultaneous sequencing of human and non-human nucleic acids using a pool of probes address the inefficiencies of current sequencing technologies, providing cost-effective and comprehensive genetic data from a single sample.
Patent Information
- Application Number
- JP2025102053
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-08-07
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-02
AI Technical Summary
Current whole genome and/or exome sequencing methods are expensive and inefficient in capturing biomedically important variants, particularly in regions with high CG content and repetitive elements, and often require separate assays for human and non-human genomes.
A method and system for processing biological samples using a pool of nucleic acid probes that target both human and non-human genomes, allowing simultaneous sequencing of human and non-human nucleic acids, including viruses, bacteria, and other organisms, through hybridization and alignment with reference sequences.
Enables cost-effective and comprehensive sequencing of both human and non-human genetic data from a single sample, improving the detection of biomedically relevant variants and reducing the need for multiple sample assays.
Smart Images

Figure 2025128366000002 
Figure 2025128366000003 
Figure 2025128366000004
Abstract
Description
[Technical Field]
[0001] cross reference This application claims priority to U.S. Patent Application No. 16 / 056,982, filed August 7, 2018, and U.S. Provisional Patent Application No. 62 / 678,475, filed May 31, 2018, which are incorporated herein by reference in their entireties. [Background technology]
[0002] Current methods for whole genome and / or exome sequencing are expensive and may not capture many biomedically important variants. For example, commercially available exome enrichment kits (e.g., Illumina's TruSeq Exome Enrichment and Agilent's SureSelect Exome Enrichment) may not be able to target non-exome and exome regions of biomedical interest. In many cases, whole genome and / or exome sequencing using standard sequencing methods performs poorly in content ranges with very high CG content (>70%). Furthermore, whole genome and / or exome sequencing cannot provide adequate and / or cost-effective sequencing of repetitive elements in the genome. Summary of the Invention [Means for solving the problem]
[0003] The methods disclosed herein address these challenges and provide specific sequencing protocols or techniques that extend analysis to human and non-human genomes in a single sample.
[0004] Disclosed herein is a method for processing a biological sample obtained from a subject, the method comprising: (a) generating a subset of nucleic acid molecules from the biological sample using a pool of nucleic acid probes, the probes including (i) a first plurality of nucleic acid probes configured to target elements of the human genome and (ii) a second plurality of nucleic acid probes configured to target elements of one or more non-human genomes; and (b) subjecting the subset of nucleic acid molecules to an assay to obtain sequence information including sequences of (i) human nucleic acids from the biological sample from the subject and (ii) non-human nucleic acids from the subject's biological sample. In some cases, the first plurality of nucleic acid probes in (i) are configured to target elements derived from the human genome. In some cases, the subject may be human. In some cases, the second plurality of nucleic acid probes are configured to target elements derived from the non-human genome sequence of one or more species selected from the group consisting of viruses, bacteria, bacterial phages, fungi, protists, archaea, amoebas, helminths, algae, genetically modified cells, and genetically modified vectors. In embodiments, generating a subset of nucleic acid molecules from a biological sample comprises performing one or more hybridization reactions. In some cases, the method further comprises obtaining a biological sample from the subject. In some cases, the subject's biological sample may be derived from a tumor biopsy, whole blood, or plasma. In some embodiments, the method further comprises aligning the sequences of the subset of nucleic acid molecules with one or more reference sequences. In some cases, the one or more reference sequences comprise a plurality of reference sequences. In some cases, the plurality of reference sequences corresponds to two or more different species. The method may further comprise identifying the origin of the nucleic acid molecules in the subset based on the alignment. In some cases, the method may further comprise generating an output comprising the identified origin of the nucleic acid molecules in the biological sample. In some cases, the concentration of the second plurality of nucleic acid probes is greater than the concentration of the first plurality of nucleic acid probes in the pool of nucleic acid probes. In embodiments, the relative concentration of the second plurality of nucleic acid probes in the pool of probes is greater than the relative concentration of the first plurality of nucleic acid probes in the pool of nucleic acid probes.In embodiments, the first plurality of nucleic acid probes comprises a probe set for human exome capture. In some cases, the first plurality of nucleic acid probes comprises probes configured to target junction sequences created by human V(D)J rearrangement or recombination. In some cases, the second plurality of nucleic acid probes comprises one or more probes configured to target the E6 and / or E7 genes of human papillomavirus. In some cases, the second plurality of nucleic acid probes comprises probes configured to target one or more elements of the bacterial 16S ribosomal RNA gene. In embodiments, the assay (b) comprises performing sequencing to generate paired-end read sequences 130 bases to 280 bases in length. In embodiments, the method may further comprise generating one or more biomedical reports including one or more sets of data selected from the group consisting of (i) candidate tumor neoantigens, (ii) detected non-human species, (iii) detected CDR3 sequences, and any combination thereof. In embodiments, the detected non-human species is an antigen. In some cases, the CDR3 sequence is generated by V(D)J rearrangement or recombination. In some cases, the CDR3 sequence corresponds to an immune response to an antigen. In some cases, the one or more biomedical reports include (i)-(iii).
[0005] Disclosed herein is a method for processing a biological sample obtained from a subject, the method comprising: (a) generating a subset of nucleic acid molecules from the biological sample, the subset of nucleic acid molecules comprising (i) a first plurality of nucleic acid molecules derived from the subject and (ii) a second plurality of nucleic acid molecules not derived from the subject, wherein the abundance of the first plurality of nucleic acid molecules is greater than the abundance of the second plurality of nucleic acid molecules in the biological sample; and (b) subjecting the subset of nucleic acid molecules to an assay to obtain sequence information comprising the sequences of (i) the first plurality of nucleic acid molecules and (ii) the second plurality of nucleic acid molecules. In some cases, the nucleic acid molecules in (i) are derived from the genome of the subject. In some cases, the subject is human. In some cases, the second plurality of nucleic acid molecules not derived from the subject comprise one or more members selected from the group consisting of viruses, bacteria, bacterial phages, fungi, protists, archaea, amoebas, helminths, algae, genetically modified cells, and genetically modified vectors. In embodiments, generating a subset of nucleic acid molecules from a biological sample comprises performing one or more hybridization reactions. In embodiments, the method further comprises obtaining a biological sample from a subject. In some cases, the subject's biological sample is derived from a tumor biopsy, whole blood, or plasma. In embodiments, the method further comprises aligning the sequences of the subset of nucleic acid molecules with one or more reference sequences. In embodiments, the one or more reference sequences comprise a plurality of reference sequences. In embodiments, the plurality of reference sequences corresponds to two or more different species. The method can further comprise identifying the origin of the nucleic acid molecules in the subset based on the alignment. In embodiments, the method further comprises generating an output comprising the identified origin of the nucleic acid molecules in the biological sample. In some cases, the abundance of the first plurality of nucleic acid molecules is greater than the abundance of the first plurality of nucleic acid molecules in the biological sample. In some cases, the relative abundance of the second plurality of nucleic acid molecules in the subset is greater than the relative abundance of the first plurality of nucleic acid molecules in the subset.
[0006] Disclosed herein is a composition comprising a pool of probes configured to hybridize with (i) one or more human sequences from a subject and (ii) one or more non-human sequences from the subject. In an embodiment, the pool of probes is a plurality of capture probes. In an embodiment, the pool of probes is a plurality of amplification probes.
[0007] Disclosed herein is a system for processing a subject's biological sample, the system comprising: a processing unit including one or more computer processors individually or collectively programmed to assay a subset of nucleic acid molecules to obtain sequence information including sequences of (i) human nucleic acids from the biological sample from the subject and (ii) non-human nucleic acids from the subject's biological sample, wherein the subset of nucleic acid molecules is generated from the biological sample using a pool of nucleic acid probes, the probes including (i) a first plurality of nucleic acid probes configured to target elements of the human genome and (ii) a second plurality of nucleic acid probes configured to target elements of one or more non-human genomes; and a computer memory configured to store the sequence information. In some embodiments, the one or more computer processors are programmed to generate an alignment of the sequence with one or more reference sequences. In some embodiments, the one or more reference sequences include multiple reference sequences, where the multiple reference sequences correspond to two or more different species. In some embodiments, the one or more computer processors are programmed to identify the origin of the nucleic acid molecules in the subset based on the alignment. In some embodiments, the one or more computer processors are programmed to generate one or more biomedical reports that include information selected from the group consisting of: (i) candidate tumor neo-antigens, (ii) detected non-human species, (iii) detected complementarity determining region 3 (CDR3) sequences, and any combination thereof.
[0008] Disclosed herein is a system for processing a biological sample from a subject, the system comprising: a processing unit including one or more computer processors individually or collectively programmed to subject a subset of nucleic acid molecules to an assay to obtain sequence information including sequences of (i) a first plurality of nucleic acid molecules and (ii) a second plurality of nucleic acid molecules, wherein the subset of nucleic acid molecules is generated from the biological sample, the subset of nucleic acid molecules including (i) a first plurality of nucleic acid molecules derived from the subject and (ii) a second plurality of nucleic acid molecules not derived from the subject, wherein the abundance of the first plurality of nucleic acid molecules is greater than the abundance of the second plurality of nucleic acid molecules in the biological sample; and a computer memory configured to store the sequence information. In some embodiments, the one or more computer processors are programmed to generate an alignment of the sequence with one or more reference sequences. In some embodiments, the one or more reference sequences include multiple reference sequences, where the multiple reference sequences correspond to two or more different species. In some embodiments, the one or more computer processors are programmed to identify the origin of the nucleic acid molecules in the subset based on the alignment. In some embodiments, the one or more computer processors are programmed to generate one or more biomedical reports that include information selected from the group consisting of: (i) candidate tumor neo-antigens, (ii) detected non-human species, (iii) detected complementarity determining region 3 (CDR3) sequences, and any combination thereof.
[0009] Another aspect of the present disclosure provides a non-transitory computer-readable medium containing machine-executable code that, when executed by one or more computer processors, performs any of the methods described above or elsewhere herein.
[0010] Another aspect of the present disclosure provides a system including one or more computer processors and a computer memory coupled thereto, the computer memory including machine-executable code that, when executed by the one or more computer processors, performs any of the methods described above or elsewhere herein.
[0011] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, which shows and describes only illustrative embodiments of the present disclosure. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive. Incorporation by Reference
[0012] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.
[0013] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention may be obtained by reference to illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also referred to herein as "Figures" and "FIGs."). ) can be obtained by reference to the following detailed description. [Brief explanation of the drawings]
[0014] [Figure 1] FIG. 1 schematically illustrates a method for generating a subset of nucleic acid molecules using a first plurality of nucleic acid probes and a second plurality of nucleic acid probes.
[0015] [Figure 2]2 schematically illustrates a method for generating a subset of nucleic acid molecules using a first plurality of nucleic acid probes and a second plurality of nucleic acid probes, wherein the first plurality of nucleic acid probes can bind to a non-human genome.
[0016] [Figure 3] 3 schematically illustrates a method for generating a subset of nucleic acid molecules using a first plurality of nucleic acid probes and a second plurality of nucleic acid probes, wherein the first plurality of nucleic acid probes can bind to a segmented transcriptome.
[0017] [Figure 4] 4 is a schematic diagram illustrating a method for generating a subset of nucleic acid molecules using a plurality of nucleic acid probes. The plurality of nucleic acid probes can include a probe set that targets the exome of interest and a probe set that targets a non-interest nucleic acid sequence.
[0018] [Figure 5] FIG. 5 illustrates a computer system programmed or otherwise configured to perform the methods of the present disclosure.
[0019] [Figure 6]Figure 6A shows a schematic diagram of a workflow. Preparation 1 and Preparation 2 refer to subsets of nucleic acids. Assay 1, Analysis 1, and Output refer to any assay, analysis, and output described herein. Figure 6B shows a schematic diagram of a workflow. Preparation 1 and Preparation 2 refer to subsets of nucleic acids. Assay 1, Assay 2, Analysis 1, and Output refer to any assay, analysis, and output described herein. Figure 6C shows a schematic diagram of a workflow. Preparation 1 and Preparation 2 refer to subsets of nucleic acids. Assay 1, Assay 2, Analysis 1, Analysis 2, and Output refer to any assay, analysis, and output described herein. Figure 6D shows a schematic diagram of a workflow that includes (1) separation of a nucleic acid sample into several subsets processed by several protocols. These protocols may include enrichment for different genomic or non-genomic regions and may include one or more different amplification operations to prepare libraries of nucleic acid molecules for assay. Some of these libraries may be combined (2) for assay. Results of some assays may be combined (3) for subsequent analysis. Variant calls, or other assessments of sequence or genetic status, may be further combined with (4) to generate a composite assessment at the locus addressed by the assay. Protocols 1-4 refer to any of the methods described herein. Assay 1, Assay 2, Assay 3, Analysis 1, Analysis 2, and Output refer to any of the assays, analyses, and outputs described herein.
[0020] [Figure 7] FIG. 7 depicts an example workflow for the assay described herein.
[0021] [Figure 8] FIG. 8 shows a schematic diagram of the workflow of the present disclosure.
[0022] [Figure 9] FIG. 9 shows a schematic diagram of the workflow of the present disclosure.
[0023] [Figure 10]FIG. 10 shows the effect of shear time on fragment size.
[0024] [Figure 11] FIG. 11 shows the effect of bead ratio on fragment size.
[0025] [Figure 12] FIG. 12 shows the effect of shear time on fragment size.
[0026] [Figure 13] FIG. 13 depicts a schematic of the nucleic acid library construction workflow.
[0027] [Figure 14] FIG. 14 illustrates a method for developing multithreaded assays that address multiple biomedical applications.
[0028] [Figure 15] Figure 15 depicts an example workflow for an assay that includes multiple subsets of DNA enriched for different genomic regions and undergoes some independent processing operations before being combined for a sequencing assay. Read data from two or more subsets are combined either a) on the sequencing device or b) subsequently in silico (e.g., using one or more algorithms) to generate a single test result for the region addressed by the union of the two or more subsets, resulting in a data pool that can be used for one or more biomedical reports. Replenishment pulls may include human and non-human target sequences.
[0029] [Figure 16]Figure 16 depicts an example workflow for an assay that includes multiple subsets of DNA enriched for different genomic regions, which are independently sequenced and subjected to several independent processing operations before being analyzed for variants. Variants from two subsets can be merged to generate results for the region addressed by the union of the two or more subsets, resulting in a data pool that can be used for one or more biomedical reports. Recruitment pulls can include human and non-human target sequences.
[0030] [Figure 17] Figure 17 depicts an example workflow for an assay that contains multiple subsets of DNA enriched for different genomic regions, is independently sequenced, and undergoes some independent processing operations before generating primary data, which may include sequence read data. The primary data from two or more assays is combined and analyzed (e.g., by one or more software programs or algorithms) to generate results for all regions addressed by the union of the two or more subsets, resulting in a data pool that can be used for one or more biomedical reports. The supplemental pulls may include human and non-human target sequences.
[0031] [Figure 18]Figure 18 depicts a multithreaded assay containing two subsets of DNA generated by size selection and further divided into two subsets enriched for different genomic regions based on GC content. The longer molecules can be sequenced using a technique appropriate for the longer molecules. The two shorter subsets can be further prepared, amplified based on a protocol appropriate for the Tm of the subset, and then pooled for sequencing on a high-throughput, short-read sequencer, the HiSeq. The primary data from the sequencing can be merged and analyzed (e.g., by one or more software programs or algorithms) to generate a single best result for all regions addressed by the subsets, resulting in a data pool that can be used for one or more biomedical reports. The supplemental pulls may include human and non-human target sequences. DETAILED DESCRIPTION OF THE INVENTION
[0032] While various embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.
[0033] The human body frequently harbors numerous microbial species, in some cases exceeding 1,000. These have been documented by the Human Microbiome Project and other studies and include numerous viruses, bacteria, bacterial phages, fungi, protists, archaea, and some amoebae, helminths, and algae. While these non-human species can be beneficial, they can also cause or modulate diseases, including cancer. These effects can be direct (e.g., causing mutations in human cells, thus leading to cancer) or indirect (e.g., stimulating the immune system, affecting its ability to fight disease). Microbial species in individuals can be identified and characterized by sequencing nucleic acids extracted from people's samples. Many of these microorganisms live at the interface between the body and its surrounding environment (e.g., skin, saliva, nostrils, digestive tract, intestines, genitals). They can be sampled from these interfaces by swabbing, biopsy, feces or urine, saliva, or similar methods. Multiple methods have been developed for the sequence analysis of microbial nucleic acids from these samples. These include both targeted PCR amplification (e.g., of subsections of the 16S ribosomal RNA gene) and non-targeted (metagenomics) methods using deep sequencing. Many of these methods involve sample types and / or enrichment methods that attempt to minimize the amount of human DNA to optimize sensitivity to the intended microbial target. The human genome (approximately 3 billion bases) is approximately 1,000 times larger than a typical bacterial genome (less than a million bases up to several million) and thousands of times larger than the size of most viral genomes (typically several thousand to tens of thousands of bases). Therefore, without sample collection or assay techniques that avoid or reduce this, human DNA content can dominate the assay's performance.
[0034] The progression of many diseases is also influenced by the genetics of human cells. This can include inherited genetic variants, somatic variants in cancer, VDJ recombination in immune cells, differential gene expression in different cell types, and other properties. These can also be assayed by sequencing nucleic acids, typically from either blood (most commonly PBMCs or cell-free DNA) or diseased tissue (e.g., tumor biopsies). Multiple methods have been developed for sequence analysis of nucleic acids from these samples, including amplicon panels (typically for up to several hundred genes), hybrid capture (typically for large numbers of genes, including the exome), and untargeted methods (whole genome sequencing). Many of these methods involve sample types and / or enrichment methods optimized for their intended human nucleic acid targets.
[0035] Sample types for microbial analysis may differ from those used for human genetic analysis. While fecal samples may be used for analysis of the gastrointestinal microbiome, for example, they may be an insufficient sample for testing for inherited human genetic diseases such as cystic fibrosis. White blood cells from blood (PBMCs) may be sequenced to search for the causes of inherited genetic diseases, but they are typically a poor choice for bacterial analysis because the immune system largely eliminates bacteria from the blood. Assay methods for human and microbial analysis also differ. For example, while both may use PCR, PCR produces very narrowly focused results (i.e., amplicons each generally represent a small portion of the genome of the target species), so human and microbial targets are generally not combined into a single PCR-based assay. Non-targeted assays can be used for both human genetics (i.e., whole human genome sequencing) and microbial metagenomics, but due to the vast differences in genome sizes mentioned above, it is generally best to assay each separately, using source materials and assays optimized for each. For the same reasons, hybrid capture assay technologies (e.g., exome) have been developed for specific (generally mammalian) species (e.g., human exome, or bovine exome).
[0036] Cancer may be a special case. Many tumors use checkpoint genes and other methods to partially or completely suppress the immune system. Therefore, microorganisms that successfully infiltrate tumors may be able to survive and even thrive there. Intact, live bacteria (e.g., Fusobacterium) are found in many cancer cells and live their entire lives there, including the cell division cycle and metastasis. Bacteria can cause cancer (e.g., Helicobacter pylori causes gastric cancer), and they can affect cancer progression and response to cancer therapeutics (e.g., immune checkpoint inhibitors). Viruses can also enter cells and, in some cases, cause cancer and / or integrate their genomes into human chromosomes. Thus, tumor biopsies may contain both human and microbiome species and their nucleic acids. When these cells die, their nucleic acids may be shed into their surroundings and may eventually become detectable as cell-free nucleic acid in plasma.
[0037] The amount of sample from tumor biopsy is often very limited, making it more difficult to perform multiple different assays for human and microbial genetic targets from the same sample.An integrated assay that can simultaneously provide sequence data from both human and microbiome species can be advantageous when the amount of sample is limited.The amount of cell-free DNA and RNA from plasma is also often very limited.An integrated assay that can simultaneously provide sequence data from both human and microbiome species from a single small amount of plasma sample can also be advantageous.
[0038] Disclosed herein is an assay that supports the simultaneous detection of a wide range of human genetic data using microbiome data from the same sample. A single assay can be performed without requiring more samples than would be required for a comparable human-only assay. This assay can use a human exome capture kit (e.g., Agilent Clinical Research Exome v2). A probe pool kit or composition can use hybridization probes complementary to the human sequences it targets. By using over 50,000 capture probes, the method, kit, or composition can target most exons of a given human gene. For example, before performing a hybridization reaction between nucleic acids from a cancer sample and the probes in this kit, we add an additional set of capture probes designed to target non-human sequences. The non-human sequences can be derived from viral, bacterial, fungal, or archaeal genomes, i.e., from the human microbiome. When human exome probes are combined with non-human microbiome probes, the probe mixture can be used in a single hybridization-based capture reaction with nucleic acids extracted from a patient sample (Figure 2). The captured nucleic acids can then be sequenced. In our laboratory, this sequencing is performed using an Illumina NovaSeq-6000 DNA sequencer.
[0039] The mixed (human and non-human) DNA sequences resulting from such a process can be separated by alignment with human and microbiome-species reference sequences. Separation of sequences by alignment is possible because the human genome has diverged so significantly from microbial genomes during evolution.
[0040] The use of capture probes can be advantageous for polymerase chain reaction (PCR)-based assays to enrich and target sequences of interest, due to the difficulty of PCR targeting multiple genes simultaneously.To target multiple genes, PCR may require a large number of primers, for example, potentially up to 100,000 primers, to amplify and target 50,000 sequences, and may require enzymatic manipulation to generate nucleic acid molecules for identification.Optimizing the ratio of PCR primers for multiple genes may also be difficult, and may result in biased amplification, which may not represent the relative amount of target nucleic acid in nucleic acid samples.
[0041] As used in this specification and claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. For example, the term "chimeric transmembrane receptor polypeptide" includes a plurality of chimeric transmembrane receptor polypeptides.
[0042] The term "about" or "approximately" means within an acceptable error range for a particular value as determined by one of ordinary skill in the art, which depends, in part, on how the value is measured or determined, i.e., the limitations of the measurement system. For example, "about" can mean within one or more standard deviations, per practice in the art. Alternatively, "about" can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, for example, with respect to biological systems or processes, the term can mean within an order of magnitude, or within five-fold, or within two-fold of a value. When a particular value is described in this application and claims, unless otherwise specified, the term "about" should be assumed to mean within an acceptable error range for the particular value.
[0043] As used herein, "cell" generally refers to a biological cell. A cell may be the basic structural, functional, and / or biological unit of a living organism. A cell may originate from any organism having one or more cells. Some non-limiting examples include prokaryotic cells, eukaryotic cells, bacterial cells, archaeal cells, cells of unicellular eukaryotes, protozoan cells, cells from plants (e.g., plant crops, fruits, vegetables, grains, soybeans, corn, maize, wheat, seeds, tomatoes, rice, cassava, sugarcane, pumpkins, hay, potatoes, cotton, cannabis, tobacco, flowering plants, conifers, gymnosperms, ferns, club mosses, hornworts, liverworts, mosses), cells of algae (e.g., Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens, etc.). C. Agardh, etc.), seaweed (e.g., kelp), fungal cells (e.g., yeast cells, cells from mushrooms), animal cells, cells from invertebrates (e.g., Drosophila, cnidarians, echinoderms, nematodes, etc.), cells from vertebrates (e.g., fish, amphibians, reptiles, birds, mammals), cells from mammals (e.g., pigs, cows, goats, sheep, rodents, rats, mice, non-human primates, humans, etc.), etc. The cells may not originate from a natural organism (e.g., the cells may be synthetically produced and may be referred to as artificial cells).
[0044] The term "nucleotide," as used herein, generally refers to a base-sugar-phosphate combination. Nucleotides can include synthetic nucleotides. Nucleotides can include synthetic nucleotide analogs. Nucleotides can be nucleic acid sequences of monomeric units (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide can include ribonucleoside triphosphates, adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP), and deoxyribonucleoside triphosphates such as dATP, dCTP, dITP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives can include, for example, [αS]dATP, 7-deaza-dGTP, and 7-deaza-dATP, as well as nucleotide derivatives that confer nuclease resistance to nucleic acid molecules containing them. As used herein, the term nucleotide can refer to dideoxyribonucleoside triphosphates (ddNTPs) and their derivatives. Illustrative examples of dideoxyribonucleoside triphosphates include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP.Nucleotides can be unlabeled or detectably labeled.Labeling can also be performed using quantum dots.Detectable labels can include, for example, radioisotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzyme labels. Fluorescent labels for nucleotides include, but are not limited to, fluorescein, 5-carboxyfluorescein (FAM), 2'7'-dimethoxy-4'5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N',N'-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-(4'dimethylaminophenylazo)benzoic acid (DABCYL), Cascade Blue, Oregon Green, Texas Red, cyanine, and 5-(2'-aminoethyl)aminonaphthalene-l-sulfonic acid (EDANS).Specific examples of fluorescently labeled nucleotides include [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP, available from Perkin Elmer, Foster City, Calif.; FluoroLink deoxynucleotides, FluoroLink Cy3-dCTP, FluoroLink Cy5-dCTP, FluoroLink Fluor X-dCTP, FluoroLink Cy3-dUTP, and FluoroLink Cy5-dUTP, available from Amersham, Arlington Heights, Ill.; and Boehringer Fluorescein-15-dATP, fluorescein-12-dUTP, tetramethyl-rhodamine-6-dUTP, IR770-9-dATP, fluorescein-12-ddUTP, fluorescein-12-UTP, and fluorescein-15-2'-dATP available from Mannheim, Indianapolis, Ind.; and Molecular Chromosomal labeled nucleotides include BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, Fluorescein-12-UTP, Fluorescein-12-dUTP, Oregon Green 488-5-dUTP, Rhodamine Green-5-UTP, Rhodamine Green-5-dUTP, Tetramethylrhodamine-6-UTP, Tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP, available from Probes, Eugene, Oreg. Nucleotides can also be labeled or marked by chemical modification. The chemically modified single nucleotide may be a biotin-dNTP.Some non-limiting examples of biotinylated dNTPs include biotin-dATP (e.g., bio-N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP).
[0045] The terms "genome" and "genomes," as used herein, are generally used to refer to a portion of a subject's genome or the entirety of a subject's genome. For example, a genome can refer to a subject's gene sequence. A genome can refer to a subject's entire genomic sequence.
[0046] The terms "polynucleotide," "oligonucleotide," and "nucleic acid" are used interchangeably to refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or their analogs in either single-, double-, or multiple-stranded form. A polynucleotide can be exogenous or endogenous to a cell. A polynucleotide can be present in a cell-free environment. A polynucleotide can be a gene or a fragment thereof. A polynucleotide can be DNA. A polynucleotide can be RNA. A polynucleotide can have any three-dimensional structure and can perform any function. A polynucleotide can contain one or more analogs (e.g., altered backbones, sugars, or nucleobases). Modifications to the nucleotide structure, if present, can be imparted before or after assembly of the polymer. Some non-limiting examples of analogs include 5-bromouracil, peptide nucleic acids, xenonucleic acids, morpholinos, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein linked to a sugar), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudourdine, dihydrouridine, queusine, and wyosine.Non-limiting examples of polynucleotides include the coding or non-coding regions of genes or gene fragments, loci (locuses) defined by linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), small interfering RNA (siRNA), small hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotides including cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers. The sequence of nucleotides can be interrupted by non-nucleotide components. Any of the aforementioned nucleic acid molecules can be engineered or synthesized.
[0047] The term "gene," as used herein, refers to nucleic acids (e.g., DNA, such as genomic DNA and cDNA) involved in coding for RNA transcripts and their corresponding nucleotide sequences. As used herein with respect to genomic DNA, this term includes regulatory regions, as well as intervening regions and non-coding regions, and can include the 5' and 3' ends. In some uses, this term encompasses transcribed sequences, including 5' and 3' untranslated regions (5'-UTR and 3'-UTR), exons, and introns. In some genes, the transcribed region contains an "open reading frame" that encodes a polypeptide. In some uses of this term, a "gene" includes only the coding sequence (e.g., an "open reading frame" or "coding region") necessary to encode a polypeptide. In some cases, a gene does not encode a polypeptide, such as a ribosomal RNA gene (rRNA) or a transfer RNA (tRNA) gene. In some cases, the term "gene" includes not only the transcribed sequence but also non-transcribed regions, including upstream and downstream regulatory regions, enhancers, and promoters. A gene can refer to an "endogenous gene" or native gene in its natural location in the genome of an organism. A gene can refer to an "exogenous gene" or non-native gene. A non-native gene can refer to a gene that is not normally found in a host organism but is introduced into the host organism by gene transfer. A non-native gene can also refer to a gene that is not in its natural location in the genome of an organism, such as a genetically modified organism. A non-native gene can also refer to a naturally occurring nucleic acid or polypeptide sequence (e.g., a non-native sequence) that contains mutations, insertions, and / or deletions.
[0048] The term "percent identity (%)" as used herein refers to the percentage of amino acid or nucleic acid residues in a candidate sequence that are identical to those in a reference sequence, after aligning the sequences and, if necessary, introducing gaps to achieve the maximum percent identity (i.e., for optimal alignment, gaps can be introduced in one or both of the candidate and reference sequences, and non-homologous sequences can be ignored for comparison purposes). For purposes of determining percent identity, alignment can be achieved in a variety of ways, such as using computer software such as BLAST, ALIGN, or Megalign (DNASTAR) software. The percent identity of two sequences can be calculated using BLAST by aligning a test sequence with a comparison sequence, determining the number of amino acids or nucleotides in the aligned test sequence that are identical to amino acids or nucleotides at the same positions in the comparison sequence, and dividing the number of identical amino acids or nucleotides by the number of amino acids or nucleotides in the comparison sequence.
[0049] The term "subject," as used herein, generally refers to any animal, e.g., a mammal or marsupial. A subject may be a patient. A subject may be symptomatic or asymptomatic with respect to a disease or condition. A subject may be a primate (e.g., a human), a non-human primate (e.g., a rhesus monkey or other species of macaque), a dog, a cat, a mouse, a pig, a horse, a donkey, a cow, a sheep, a rat, and a poultry. Mammals include, but are not limited to, mice, monkeys, humans, farm animals, sport animals, and pets. Also encompassed are tissues, cells, and their progeny of biological entities obtained in vivo or cultured in vitro. A host is an organism capable of harboring a non-host. A subject may be asymptomatic with respect to a disease. Alternatively, a subject is not asymptomatic with respect to a disease.
[0050] The terms "treatment" and "treating," as used herein, refer to an approach to obtaining beneficial or desired results, including, but not limited to, therapeutic benefit and / or preventative benefit. For example, treatment can include administering a system or cell population disclosed herein. Therapeutic benefit refers to any treatment-related improvement in, or effect on, one or more diseases, conditions, or symptoms under treatment. For preventative benefit, the composition can be administered to a subject at risk of developing a particular disease, condition, or symptom, or to a subject reporting one or more physiological symptoms of the disease, even if the disease, condition, or symptom has not yet manifested.
[0051] In some cases, the present disclosure also provides compositions and methods for processing and analyzing biological samples. In some cases, a biological sample from a subject can contain nucleic acids from the subject and nucleic acid molecules not from the subject. In some cases, a biological sample can contain nucleic acids from human and non-human genomes. In some cases, the non-human genome can be derived from viruses, bacteria, bacterial phages, fungi, protists, archaea, amoebas, helminths, algae, or a combination thereof. In some cases, the source of the non-human genome can be beneficial to the human host. In some cases, the source of the non-human genome can be involved in modulating diseases such as cancer. In some cases, the source of the non-human organism containing the non-human genome can act directly by causing mutations in human cells, thereby causing cancer in the subject. In some cases, the source of the non-human genome can act indirectly, for example, by stimulating the immune system of the human host, affecting its ability to fight disease. In some cases, the present disclosure also provides methods including identifying the presence of non-human genomes and human genomes in a sample.
[0052] The sample may be from the skin, saliva, nasal passages, the digestive tract, the intestines, the genitals, or a combination thereof. In some cases, the sample may be obtained by swabbing, biopsy, fecal collection, urine collection, saliva collection, etc.
[0053] The human genome (about 3 billion bases) is about 1,000 times larger than non-human genomes, e.g., bacterial genomes, and thousands of times the size of most viral genomes. In some cases, the methods can include sampling methods to enrich for non-human genomes in a mixed sample.
[0054] The present disclosure also provides a method for analyzing the sequence (or sequencing) of nucleic acids of microorganisms from a sample.Sequencing analysis can include, for example, PCR amplification of a subsection of the 16S ribosomal RNA gene and untargeted metagenomics methods using deep sequencing.The method can include selecting a sample type and / or enrichment method that can minimize the amount of human genomes to optimize sensitivity to non-human genomes.
[0055] In one embodiment, the first plurality of nucleic acid molecules can be derived from a subject. The subject can be a human, and thus the present disclosure also provides human nucleic acid molecules. The present disclosure also provides a first subset or product using a plurality of nucleic acid probes that can be derived from a human genome. The present disclosure also provides a human genome. In some cases, the human genome can include, for example, inherited genetic variants, somatic variants, VDJ recombination in immune cells, differential gene expression in different cell types, and other properties. At least some of the foregoing can contribute to or be associated with disease in the subject (e.g., somatic variants in cancer). The genome can include genes, exons, UTRs, regulatory regions, splice sites, rearranged genes, alternative sequences, rearranged genes, gene phasing, exogenous sequences, etc.
[0056] In some cases, the nucleic acid in biological sample can be analyzed by sequencing.In some cases, typically, the nucleic acid from blood or diseased tissue can be sequenced.In some cases, blood sample can comprise peripheral blood mononuclear cells, cell-free DNA or a combination thereof.In some cases, diseased tissue can comprise cancer.In some cases, the sequence analysis of the nucleic acid from blood sample or diseased tissue sample can comprise non-targeted methods such as amplicon panel generation, hybrid capture, and whole genome sequencing.In some cases, the method can comprise the selection of sample type, human genome enrichment method, and combinations thereof.
[0057] In some cases, the biological sample can include nucleic acids of interest, nucleic acids of non-interest, and combinations thereof. In some cases, the biological sample can include nucleic acids of host, nucleic acids of non-host, and combinations thereof. In some cases, the biological sample can include a genome encoding a receptor. In some cases, the receptor can be derived from an immune cell. In some cases, the receptor derived from an immune cell can be a T cell receptor (TCR), a B cell receptor (BCR), a chimeric antigen receptor (CAR), etc.
[0058] In some cases, the methods provided herein include subjecting nucleic acid molecules to assays to obtain sequence information. The assays can include sequencing nucleic acids containing VDJ rearrangements or VDJ recombinations. In some cases, the VDJ rearrangements or recombinations can refer to cellular receptors. For example, somatic hypermutation can result in antibody-encoding B cell receptor (BCR) sequences due to significant antigen diversity. In some cases, the BCR can be sequenced. The methods provided herein can include sequencing the BCR to elucidate how antibodies develop. For example, the method can include sequence analysis to annotate each base as derived from a particular V, D, or J gene, or as derived from an N addition (also known as a non-templated insertion). In some cases, the VDJ recombination can generate a CDR3 sequence. In some cases, the CDR3 sequence can generate a polypeptide that binds to an antigen, e.g., a tumor antigen. In some cases, the CDR3 sequence can generate a polypeptide that binds to an antigen, e.g., a neoantigen.
[0059] In some cases, the methods provided herein can include sequencing the nucleic acid encoding or associated with the candidate tumor neoantigen. In some cases, the methods provided herein can include sequencing the detected CDR3 sequence. In some embodiments, the target TCR can be identified using various methods. In some cases, the TCR can be identified using whole exome sequencing. For example, the TCR can target the neoantigen or neoepitope identified by whole exome sequencing of target cells. Alternatively, the TCR can be identified from autologous, allogeneic or xenogeneic repertoires. In some cases, genes that may contain mutations that give rise to neoantigens or neoepitopes include ABL1, ACO1 1997, ACVR2A, AFP, AKT1, ALK, ALPPL2, ANAPC1, APC, ARID1A, AR, AR-v7, ASCL2, β2M, BRAF, BTK, C15ORF40, CDH1, CLDN6, CNOT1, CT45A5, CTAG1B, DCT, DKK4, EEF1B2, EEF1DP3, EGFR, EIF2B3, env, EPHB2, ERBB3, ESR1, ESRP1, FAM11 IB, FGFR3, FRG1B, GAGE1, GAGE 10, GATA3, GBP3, HER2, IDH1, JAK1, KIT, KRAS, LMAN1, MABEB 16, MAGEA1, MAGEA10, MAGEA4, MAGEA8, MAGEB 17, MAGEB4, MAGEC1, MEK, MLANA, MLL2, MMP13, MSH3, MSH6, MYC, NDUFC2, NRAS, NY-ESO, PAGE2, PAGE5, PDGFRa, PIK3CA, PMEL, pol protein, POLE, PTEN, RAC1, RBM27, RNF43, RPL22, RUNX1, SEC31A, SEC63, SF3B 1, SLC35F5, SLC45A2, SMAP1, SMAP1, SPOP, TFAM, TGFBR2, THAP5, TP53, TTK, TYR, UBR5, VHL, XPOT.
[0060] The present disclosure also provides a method for obtaining or preparing a nucleic acid sample or a subset of nucleic acid molecules comprising one or more genomes. The method disclosed herein can analyze a subset of nucleic acid molecules generated from a biological sample. The subset can include nucleic acid molecules derived from a subject and nucleic acid molecules not derived from a subject. For example, human and microbial sequences can be enriched using target capture and sequencing. The one or more genomes can include one or more genomic features.
[0061] The genomic features may include the entire genome or a portion thereof. The genomic features may include the entire exome or a portion thereof. The genomic features may include one or more sets of genes. The genomic features may include one or more genes. The genomic features may include one or more sets of regulatory elements. The genomic features may include one or more regulatory elements. The genomic features may include a set of polymorphisms. The genomic features may include one or more polymorphisms. In some cases, a polymorphism refers to a mutation in a genotype. A polymorphism may include one or more base changes, insertions, repeats, or deletions of one or more bases. The genomic features may include copy number variants (CNVs), transversions, other rearrangements, and other forms of genetic variation. In some cases, one or more features of the subset of nucleic acid samples may be polymorphic markers, including restriction fragment length polymorphisms, tandem repeat polymorphisms (VNTRs), hypervariable regions, minisatellites, dinucleotide repeats, trinucleotide repeats, tetranucleotide repeats, simple sequence repeats, and insertion elements such as Alu. In some cases, the differences between the first subset of nucleic acid molecules and the second subset of nucleic acid molecules may be polymorphic markers, including restriction fragment length polymorphisms, tandem repeat polymorphisms (VNTRs), hypervariable regions, minisatellites, dinucleotide repeats, trinucleotide repeats, tetranucleotide repeats, simple sequence repeats, and insertion elements such as Alu. The allelic form occurring most frequently in a selected population may be referred to as the wild-type form. A diploid organism may be homozygous or heterozygous for an allelic form. A biallelic polymorphism has two forms. A triallelic polymorphism has three forms. The polymorphism may include a single nucleotide polymorphism (SNP). In some embodiments of the present disclosure, the one or more polymorphisms include one or more single nucleotide variations, InDels, small insertions, small deletions, structural variant junctions, variable length tandem repeats, flanking sequences, or combinations thereof. The one or more polymorphisms may be located within coding and / or non-coding regions.The one or more polymorphisms may be located within, around, or adjacent to a gene, exon, intron, splice site, untranslated region, or a combination thereof. The one or more polymorphisms may span at least a portion of a gene, exon, intron, or untranslated region. In some cases, the genomic features may be related to the GC content, complexity, and / or mappability of one or more nucleic acid molecules. The genomic features may include one or more simple tandem repeats (STRs), unstable extended repeats, segmental duplications, single and paired read degenerate mapping scores, GRCh38 or GRCh37 patches, or a combination thereof. The genomic features may include one or more low average coverage regions from whole genome sequencing (WGS), zero average coverage regions from WGS, effective compaction, or a combination thereof. The genomic features may include one or more alternative or non-reference sequences.
[0062] The genomic features may include one or more gene phasing and rearrangement genes. Examples of phasing and rearrangement genes include, but are not limited to, one or more of major histocompatibility complex, blood typing, and amylase gene family. In some cases, the gene phasing and / or rearrangement genes may include genes related to blood typing. Blood typing genes may include ABO, RHD, RHCE, or a combination thereof.
[0063] In some cases, the genomic feature may include a rearrangement gene. The rearrangement gene may include a gene involved in the immune response. The gene involved in the immune response may include a gene involved in the major histocompatibility complex, immune response, and cellular function. The one or more major histocompatibility complexes may include one or more HLA class I, HLA class II, or a combination thereof. The HLA class I may be any one of HLA-A, HLA-B, HLA-C, or a combination thereof. The HLA class II may be any one of HLA-DP, HLA-DM, HLA-DOA, HLA-DOB, HLA-DQ, HLA-DR, or a combination thereof. The gene involved in the immune response may be RAG1, RAG2, or a combination thereof. In some cases, the rearrangement gene may include a gene involved in VDJ recombination. For example, to establish diversity in B cell and T cell receptors (BCR and TCR), genes may be generated by recombining existing gene segments. In some cases, different combinations of a finite set of gene segments can generate receptors that can recognize an unlimited number of foreign or non-human genomes. VDJ recombination can include cleaving DNA containing a recombination signal sequence (RSS). In some cases, the fragmented sequence can be reconstructed using cellular repair mechanisms. The present disclosure also provides methods that include sequencing the fragmented section of the genome from VDJ recombination. For example, the present disclosure provides methods that include sequencing at the site of VDJ recombination.
[0064] In some embodiments of the present disclosure, one or more genomic features may not be mutually exclusive. For example, a genomic feature comprising a whole genome or a portion thereof may overlap with additional genomic features, such as a whole exome or a portion thereof, one or more genes, or one or more regulatory elements. Alternatively, one or more genomic features may be mutually exclusive. For example, a genome comprising a non-coding portion of a whole genome may not overlap with a genomic feature such as an exome or a portion thereof, or a coding portion of a gene. Alternatively, or in addition, one or more genomic features may be partially exclusive or partially inclusive. For example, a genome comprising a whole exome or a portion thereof may partially overlap with a genome comprising the exon portion of a gene. However, a genome comprising a whole exome or a portion thereof may not overlap with a genome comprising the intron portion of a gene. Thus, a genomic feature comprising a gene or a portion thereof may partially exclude and / or partially include a genomic feature comprising a whole exome or a portion thereof. In some cases, a gene feature may be associated with a species such that more than one species can be distinguished. In some cases, genetic features can be species-associated, such that human genetic features can be distinguished from bacterial genetic features. In some cases, a first subset of nucleic acid molecules is specific to one species, and a second subset of nucleic acid molecules is specific to a second species.
[0065] Biological samples can include human genomes, non-human genomes, or a combination thereof. In some cases, biological samples can be processed. In some cases, biological samples can be enriched for target (e.g., human) nucleic acid sequences, or for nucleic acid sequences that are not derived from a target (e.g., non-human), simultaneous detection of human genomes and non-human genomes from samples can be performed. Biological samples can include cell-free nucleic acid molecules such as cell-free DNA (cfDNA) or cell-free RNA. Cell-free nucleic acid molecules can be circulating tumor nucleic acid molecules (e.g., circulating tumor DNA). Cell-free DNA can include mutations that can indicate, be associated with, or be related to diseases such as cancer.
[0066] A biological sample may contain nucleic acid molecules containing engineered sequences. For example, the nucleic acids in the biological sample may contain exogenous or surrogate sequences such as tags, exogenous receptors such as chimeric antigen receptors (CARs), plasmid sequences, and neoantigen-specific sequences, to name a few. In some cases, the engineered sequences may be used as diagnostic markers. In some cases, the engineered sequences may be used to determine whether a therapeutic agent is delivered to a target, such as a tumor target. The surrogate sequences may be exogenous or endogenous. In some cases, the surrogate sequences may include those derived from plasmid sequences. The plasmid sequences may be DNA or RNA. In some cases, the plasmid sequences may also be DNA minicircular sequences or doggybone sequences.
[0067] The present disclosure also provides methods that include obtaining or preparing a nucleic acid sample or a subset of nucleic acid molecules comprising one or more transcriptomes. The nucleic acid sample may include mRNA. The mRNA may be derived from different tissues in a subject. The amount of mRNA in a sample can be used to analyze mRNA or protein expression levels in a tissue- and subject-specific manner. For example, the amount of mRNA in a sample may be associated with a particular characteristic or disease (e.g., cancer) in a subject. In addition, the amount of mRNA in a sample may indicate changes or relative differences in expression in a particular tissue type in a subject. In some cases, mRNA may be processed to form cDNA. For example, mRNA may be subjected to reverse transcription using reverse transcription to synthesize cDNA molecules. The cDNA molecules may be subjected to capture, isolation, enrichment, amplification, sequencing, or other reactions that can be performed on nucleic acids as described elsewhere herein.
[0068] Nucleic acids from biological samples containing target and non-target sequences can be enriched and sequenced in a single pool, and can be separated in silico. For example, mixed human and non-human genomes can be separated by aligning with reference sequences of human and non-human species. Because the human genome has diverged from microbial genomes during evolution, sequence separation by alignment may be possible. Genomes, such as DNA, from non-human genomes, such as microbial species, are generally not aligned with the human genome, and vice versa. Non-human genomes, such as microbial genomes, may have segments in common or similar to other non-human genomes. More than one non-human genome in a mixture of human and non-human genomes may require additional alignment to identify the non-human species.
[0069] In some cases, the method can include generating a subset of nucleic acid molecules from the biological sample. In some cases, the method can include generating a subset of nucleic acid molecules from the biological sample by performing one or more hybridization reactions. The hybridization reaction can include enrichment. In some cases, enrichment can be performed. Enrichment can be performed by a variety of methods. In some cases, enrichment can be performed by hybrid capture, array capture, bead capture, etc. In some cases, hybrid capture can be in solution or on a solid support, such as on an array. In some cases, enrichment can be performed by molecular inversion probes (MIPs). Enrichment can be performed by amplification, for example, using PCR. In some cases, the method can include generating a subset of nucleic acids by amplification of a human genome or a non-human genome for the purpose of enrichment. In some cases, amplification can include one of polymerase chain reaction (PCR)-based techniques (e.g., solid-phase PCR, RT-PCR, qPCR, multiplex PCR, touchdown PCR, nano-PCR, nested PCR, hot-start PCR, etc.), helicase-dependent amplification (HDA), loop-mediated isothermal amplification (LAMP), self-sustained sequence replication (3SR), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), rolling circle amplification (RCA), ligase chain reaction (LCR), and any other suitable amplification technique.
[0070] The nucleic acid probe pool can include a human exome capture kit, such as the Agilent Clinical Research Exome v2 (Figure 1). The nucleic acid probe pool can include hybridization probes complementary to the human sequence it targets. The nucleic acid probe pool can include approximately 50,000 capture probes to target a large portion of the genome, such as a given human gene. The capture probes can be designed to target genes, exons, UTRs, regulatory regions, splice sites, rearranged genes, alternative sequences, and additional genomic content. In some cases, the method can include a nucleic acid probe pool designed to target non-human sequences. In some cases, the nucleic acid probe pool can be specific to a non-human genome, such as a viral, bacterial, fungal, or archaeal genome, i.e., from the human microbiome. In some cases, the nucleic acid probe pool can be combined with a second pool of nucleic acid probes that can be specific to a second species compared to the first pool of nucleic acid probes (Figure 2). The pool of nucleic acid probes can be configured to bind to human and non-human sequences, and the pool of nucleic acid probes can bind to sequences from the segmented transcriptome (Figure 3). In some cases, the pool of nucleic acid probes can bind to human sequences, and the pool of nucleic acid probes can bind to non-human sequences (Figure 4). In some cases, the pool of nucleic acid probes can be used for a single hybridization-based capture reaction with nucleic acids extracted from a patient sample. In some cases, the method can include sequencing the pool of nucleic acids being captured. The captured or enriched nucleic acids can be human, non-human, or human and non-human sequences, and the sequencing can include an Illumina NovaSeq-6000 DNA sequencing instrument. In some cases, capture probes targeting microbial species can be designed to target one or more species-different regions or regions of the microbial species sequence immediately adjacent to one or more species-different regions, e.g., non-human or human regions.In some cases, the methods involve targeting regions of difference between non-human, such as microbial, sequences, thereby enabling capture of nucleic acids from a large number of potential non-human microbial species with a small number of capture probes. By using capture probes that may contain shared sequence regions but flank variable regions, many non-human sequences that can be captured can span both.
[0071] In some cases, sequences from non-human genomes can then be assigned to their species of origin by one or more species-specific regions. For example, the 16S ribosomal RNA gene, present in almost all bacteria, has approximately nine regions whose sequences vary from species to species and are interleaved with regions of shared sequence. In some cases, captured molecules whose sequences extend from these shared regions to variable regions can then be assigned to their species of origin based on sequences from the variable region portion. Fungal nucleic acid sequences can be similarly evaluated by using the partially conserved D2 region of the ribosomal RNA gene of the large subunit of the fungal genome. The exome of the genome can be analyzed. The intronic regions of the genome can be analyzed. The exome primarily targets the coding region of the human genome and may represent less than 2% of the entire human genome. By excluding most of the intronic and intergenic portions of the human genome, the amount of human sequence can be reduced by approximately 98%. The exome can be expanded to include non-coding content. Exome sequencing, such as in cancer, can enable deep sequencing, thereby improving the detection of somatic variants at low allele frequencies, as well as improving the detection of non-human sequences co-captured from the sample.
[0072] The present disclosure also provides compositions and methods for processing biological samples.Biological samples can be obtained from subjects such as adults or children.In some cases, the method for processing biological samples includes: (a) using a pool of nucleic acid probes to generate a subset of nucleic acid molecules from the biological sample, wherein the probes include: (i) a first plurality of nucleic acid probes that are configured to target elements of human genome, and (ii) a second plurality of nucleic acid probes that are configured to target elements of one or more non-human genomes; and (b) subjecting the subset of nucleic acid molecules to assay to obtain sequence information, including the sequences of (i) human nucleic acids from the biological sample from the subject and (ii) non-human nucleic acids from the biological sample from the subject.
[0073] The disclosed methods may include detecting, monitoring, quantifying, or assessing one or more diseases or conditions caused by one or more non-human nucleic acid molecules, or one or more non-human or non-host genomes. In some cases, the capture probes can target between different genera. In some cases, the capture probes can target between different species. In some embodiments, the capture probes can target between more than one organism from different orders. In some cases, the capture probes can target between organisms from the kingdoms Plantae, Animalia, Fungi, Protozoa, and the like. ), eubacteria, and / or archaea. In some cases, the capture probe can target viruses, bacteria, bacterial phages, fungi, protists, archaea, amoebas, helminths, algae, genetically modified cells, genetically modified vectors, and combinations thereof. In some cases, the non-human sequence can be bacterial. For example, bacterial sequences include acidiobacteria, actiniobacteria, Aquificae, armatimonadetes, Bacteroidetes, caldiserica, chlamydiae, Chlorobi, chloroflexi, chrysiogenetes, cyanobacteria, deferribacteres, deinococcus-thermus, dictyoglomi, elusimicrobia, fibrob It may be from acteres, firmicutes, fusobacteria, gemmatimonadetes, lentisphaerae, Nitrospirae, planctomycetes, Proteobacteria, spirochaetes, synergistetes, tenericutes, thermosulfobacteria, thermomicrobia, thermotogae, and / or verrucomicrobia. In some cases, the non-human sequence may be derived from, but not limited to, Bordetella, Borrelia, Brucella, Campylobacter, Chlamydia, Chlamydophila, Clostridium, Corynebacterium, Enterococcus, Escherichia, Francisella, Haemophilus, Helicobacter, Legionella, Leptospira, Listeria, Mycobacterium, Mycoplasma, Neisseria, Pseudomonas, Rickettsia, Salmonella, Shigella, Staphylococcus, Streptococcus, Treponema, Vibrio, or Yersinia.Additional pathogens include, but are not limited to, Mycobacterium tuberculosis, Streptococcus, Pseudomonas, Shigella, Campylobacter, and Salmonella. In some cases, the capture probe can target fungi such as blastocladiomycota, chytridiomycota, Glomeromycota, Microsporidia, Neocallimastigomycota, Deuteromycota, Ascomycota, Pezizomycotina, Saccharomycotina, Taphrinomycotina, Basidiomycota, Agaricomycotina, Pucciniomycotina, Ustilaginomycotina, Entomophthoromycotina, Kickxellomycotina, Mucoromycotina, Zoopagomycotina, etc. In some cases, the non-human or non-host may be from a cow, horse, fish, donkey, rabbit, rat, mouse, hamster, dog, cat, pig, snake, sheep, goat, or the like.
[0074] The disease or condition caused by or associated with one or more non-human genomes may include tuberculosis, pneumonia, oral infectious diseases, tetanus, typhoid fever, diphtheria, syphilis, leprosy, bacterial vaginosis, bacterial meningitis, bacterial pneumonia, urinary tract infection, bacterial gastroenteritis, bacterial skin infection, or any combination thereof. Examples of bacterial skin infections include, but are not limited to, impetigo, which may be caused by Staphylococcus aureus or Streptococcus pyogenes; erysipelas, which may be caused by a deep epidermal streptococcus bacterial infection via lymphatic spread; and cellulitis, which may be caused by normal skin flora or exogenous bacteria.
[0075] The non-target nucleic acid sequence may be derived from a fungus such as Candida, Aspergillus, Cryptococcus, Histoplasma, Pneumocystis, and Stachybotrys. Examples of diseases or conditions caused by fungi include, but are not limited to, tinea, candidiasis, ringworm, and athlete's foot.
[0076] In some cases, the non-target or non-host nucleic acid sequence may be derived from a protist. The protist may include a protozoan, a protophyte, a fungus, and a combination thereof. The protist may be an archaeplastida. The archaeplastida may be a Rhodophyta or a Glaucophyta. The protist may be a Sar or a Harosa. The SAR may be a clade including stramenopiles, alveolates, and Rhizaria (SAR). In addition, the clade SAR may include Stramenopiles, Alveolata, Apicomplexa, Ciliophora, Dinoflagellata, Rhizaria, Cercozoa, Foraminifera, Radiolaria, and a combination thereof. In some cases, the protist may be an Excavata. The Excavata may be a Euglenozoa, Percolozoa, Metamonada, and a combination thereof. In some cases, the non-host may be Amoebozoa, Hacrobia, Apusozoa, Opisthokonta, and / or Choanozoa.
[0077] Non-target nucleic acid sequences can be derived from viruses. Examples of viruses include, but are not limited to, adenovirus, Coxsackievirus, Epstein-Barr virus, hepatitis virus (e.g., hepatitis A, B, and C), herpes simplex virus (type 1 and type 2), cytomegalovirus, herpes virus, HIV, influenza virus, measles virus, mumps virus, papillomavirus, parainfluenza virus, poliovirus, respiratory syncytial virus, rubella virus, and varicella-zoster virus. Examples of diseases or conditions caused by viruses include, but are not limited to, cold, influenza, hepatitis, AIDS, chicken pox, rubella, mumps, measles, warts, and poliomyelitis.
[0078] Non-target nucleic acids include Acanthamoeba (e.g., A.astronyxis, A.castellanii, A.culbertsoni, A.hatchetti, A.polyphaga, A.rhysodes, A.healyi, A.divionensis), Brachiola (e.g., B The bacterial species may be derived from protozoa such as Bacillus connori, B. vesicularum), Cryptosporidium (e.g., C. parvum), Cyclospora (e.g., C. cayetanensis), Encephalitozoon (e.g., E. cuniculi, E. hellem, E. intestinalis), Entamoeba (e.g., E. histolytica), Enterocytozoon (e.g., E. bieneusi), Giardia (e.g., G. lamblia), Isospora (e.g., I. belli), Microsporidium (e.g., M. africanum, M. ceylonensis), Naegleria (e.g., N. fowleri), Nosema (e.g., N. algerae, N. ocularum), Pleistophora, Trachipleistophora (e.g., T. anthropophthera, T. hominis), and Vittaforma (e.g., V. corneae).
[0079] Nucleic acids can be extracted and / or isolated from a biological sample from a subject, for example, by isolating a cellular fraction. In a variation, sample processing can therefore include any one or more of: lysing the sample, disrupting membranes in the cells of the sample, separating undesired elements (e.g., RNA, protein) from the sample, purifying nucleic acids (e.g., DNA) in the sample to generate a nucleic acid sample containing the nucleic acid content of the non-human microbiome and the human genome, amplifying nucleic acids from the nucleic acid sample, further purifying the amplified nucleic acids of the nucleic acid sample, sequencing the amplified nucleic acids of the nucleic acid sample, and any combination thereof. In a variation, lysing the sample and / or disrupting membranes in the cells of the sample can include physical methods of cell lysis / membrane disruption (e.g., bead beating, nitrogen vacuum, homogenization, sonication), which remove certain reagents that cause bias in the representation of certain microbial groups during sequencing. Additionally or alternatively, lysis or disruption can include chemical methods (e.g., using detergents, solvents, surfactants, etc.).
[0080] In a variation, separating undesired elements from the sample can include removing RNA using RNase and / or removing proteins using proteases. In a variation, purifying nucleic acids in a sample to produce a nucleic acid sample can include one or more of: precipitation of nucleic acids from a biological sample (e.g., using alcohol-based precipitation), liquid-liquid purification techniques (e.g., phenol-chloroform extraction), chromatography-based purification techniques (e.g., column adsorption), purification techniques including the use of binding moiety-conjugated particles (e.g., magnetic beads, levitating beads, beads with a size distribution, ultrasonically responsive beads, etc.) configured to bind to nucleic acids and release the nucleic acids in the presence of an elution environment (e.g., having an elution solution, providing a pH change, providing a temperature change, etc.), and any other suitable purification techniques.
[0081] Nucleic acids can be extracted and / or isolated from biological samples for extraction and isolation, and / or isolation can be performed in a sterile environment (e.g., a sterile laboratory hood, a sterile room) free of any contaminating substances (e.g., substances that may affect the nucleic acids in the sample or contribute to nucleic acid contamination), and the environment can be controlled for temperature, oxygen content, carbon dioxide content, and / or light exposure (e.g., exposure to ultraviolet light). Extraction can include lysis, which disrupts cell membranes and promotes the release of nucleic acids from cells in the biological sample. In one non-limiting example, lysis can include a bead mill device (e.g., a Tissue Lyser) configured for use with beads that are mixed with the sample and function to agitate the biological content of the sample. In some cases, processing of the biological sample can include one or more combinations of lysis reagents (e.g., proteinases), thermal modules, and any other suitable device(s) for lysis.
[0082] For the isolation of nucleic acids from a lysed sample, non-nucleic acid content of the sample is separated from the nucleic acid content of the sample. The purification module of the sample processing method can include force-based separation, size-based separation, binding moiety-based separation (e.g., by magnetic binding moieties, by buoyant binding moieties, etc.), and / or any other suitable form of separation. For example, the purification operation of the method can include one or more of a centrifuge to facilitate extraction of the supernatant, a filter (e.g., a filtration plate), a liquid delivery module configured to combine the lysed sample with moieties that bind the nucleic acid content of the sample and / or waste, a wash reagent delivery system, an elution reagent delivery system, and any other suitable device for purifying nucleic acid content from a sample.
[0083] A subset of nucleic acid molecules can be assayed to obtain sequence information. Assays that provide sequence information can include a sequencing reaction. In some cases, the sequencing can be RNA sequencing. For example, the sequencing can be RNA transcript sequencing. RNA sequencing can be performed using methods such as chromatin isolation with RNA purification (ChIRP-Seq), global run-on sequencing (GRO-Seq), ribosome profiling sequencing (Ribo-Seq) / ARTseq™, RNA immunoprecipitation sequencing (RIP-Seq), CLIP (Chip-Seq), and others. The sequencing may include any one of high-throughput sequencing of cDNA libraries (HITS-CLIP), crosslinking and immunoprecipitation sequencing (CLIP-Seq), photoactivatable ribonucleoside-enhanced crosslinking and immunoprecipitation (PAR-CLIP), CLIP at individual nucleotide resolution (iCLIP), native elongation transcript sequencing (NET-Seq), targeted purification of polysomal mRNA (TRAP-Seq), crosslinking, ligation, and sequencing of hybrids (CLASH-Seq), parallel analysis of RNA end sequencing (PARE-Seq), genome-wide mapping of uncapped transcripts (GMUCT), transcript isoform sequencing (TIF-Seq), paired-end analysis of transcribed sequences (TEAT), and any combination thereof. In some cases, the sequencing may include RNA structure. Sequencing of RNA structure can include any one of selective 2'-hydroxyl acylation analyzed by primer extension sequencing (SHAPE-Seq), parallel analysis of RNA structure (PARS-Seq), fragmentation sequencing (FRAG-Seq), CXXC affinity purification sequencing (CAP-Seq), alkaline phosphatase, calf intestine-tobacco acid pyrophosphatase sequencing (CIP-TAP), inosine chemical elimination sequencing (ICE), m6A-specific methylated RNA immunoprecipitation sequencing (MeRIP-Seq), and any combination thereof. In some cases, sequencing can include low-level RNA detection.Low-level RNA detection can include digital RNA sequencing, single-cell whole-transcript amplification (Quartz-Seq), designed-primer-based RNA sequencing (DP-Seq), switch mechanism at the 5' end of the RNA template (Smart-Seq), switch mechanism version 2 at the 5' end of the RNA template (Smart-Seq2), unique molecular identifiers (UMI), cellular expression by linear amplification sequencing (CEL-Seq), single-cell tagged reverse transcription sequencing (STRT-Seq), and any combination thereof. In some cases, sequencing can be DNA sequencing. DNA sequencing can include low-level DNA detection. DNA sequencing including low-level DNA detection can include at least one of single-molecule molecular inversion probes (smMIPs), multiple displacement amplification (MDA), multiplex annealing and looping-based amplification cycles (MALBAC), oligonucleotide-selective sequencing (OS-Seq), duplex sequencing (Duplex-Seq), and any combination thereof. In some embodiments, sequencing can include DNA methylation. DNA methylation can include at least one of bisulfite sequencing (BS-Seq), post-bisulfite adapter tagging (PBAT), tagmentation-based whole-genome bisulfite sequencing (T-WGBS), oxidative bisulfite sequencing (oxBS-Seq), Tet-assisted bisulfite sequencing (TAB-Seq), methylated DNA immunoprecipitation sequencing (MeDIP-Seq), methylation capture (MethylCap) sequencing, methyl-binding domain capture (MBDCap) sequencing, reduced presentation bisulfite sequencing (RRBS-Seq), and combinations thereof. In some cases, the sequencing can include DNA-protein interactions.For example, sequencing involving DNA-protein interactions can include DNase I hypersensitive site sequencing (DNase-Seq), MNase-assisted isolation of nucleosomes (MAINE-Seq), chromatin immunoprecipitation sequencing (ChIP-Seq), formaldehyde-assisted isolation of regulatory elements (FAIRE-Seq), assay for transposase-accessible chromatin sequencing (ATAC-Seq), chromatin interaction analysis by paired-end tag sequencing (ChIA-PET), chromatin conformation capture (Hi-C / 3C-Seq), circular chromatin conformation capture (4-C or 4C-Seq), chromatin conformation capture carbon copy (5-C), and combinations thereof. In some cases, sequencing can include rearrangements. Sequencing of the sequence rearrangements can include at least one of retrotransposon capture sequencing (RC-Seq), transposon sequencing (Tn-Seq) or insertion sequencing (INSeq), transposition capture sequencing (TC-Seq), and combinations thereof.
[0084] Sequencing analysis can include, for example, PCR amplification of a subsection of 16S ribosomal RNA gene and untargeted metagenomics methods using deep sequencing. In some embodiments, the method can include a process for next-generation amplification, and sequencing can include simultaneously amplifying the entire 16S region for each set of microorganisms, fragmenting the amplicons of the entire 16S region for each set of microorganisms to generate a set of amplicon fragments, and generating an analysis based on the set of amplicon fragments, where the analysis includes at least one of the characteristics of the microbial population, the identification of microbial species, and the identified target microbial sequence. In some cases, whole exome sequencing can be used.
[0085] In some cases, the method can include aligning the non-human genome at the gene level. For example, the alignment can include aligning the 16S sequence with respect to the 18S sequence, with respect to the ITS sequence, etc. Thus, the output can be used to identify features of interest that can be used to characterize the microorganisms of the biological sample, where the features can be non-human (e.g., the presence of a bacterial genus), genetic (e.g., based on the representation of specific gene regions and / or sequences), and / or based on any other suitable scale.
[0086] In variants, alignment and mapping to a reference non-human genome, e.g., a bacterial genome (e.g., provided by the National Center for Biotechnology Information), can be performed using the Needleman-Wunsch algorithm, which performs a global alignment of two reads (e.g., a sequence read and a reference read) with a stopping condition based on a global alignment score (e.g., in terms of insertions, deletions, matches, mismatches); or the Smith-Waterman algorithm, which performs a local alignment of two reads (e.g., a sequence read and a reference read) with a stopping condition based on a local alignment score (e.g., in terms of insertions, deletions, matches, mismatches). Alignment algorithms can be used, including one or more of: algorithms; basic local alignment search tools (BLAST), which identify regions of local similarity between sequences (e.g., sequence read data and reference read data); FPGA-accelerated alignment tools; BWT indexing using the BWA tool; BWT indexing using the SOAP tool; BWT indexing using the Bowtie tool; Sequence Search and Alignment by Hashing Algorithm (SSAHA2), which maps nucleic acid sequence read data to a genomic reference sequence using word hashing and dynamic programming; and any other suitable alignment algorithm. Mapping of unidentified sequences can further include mapping to a reference viral genome and / or fungal genome to further identify viral and / or fungal components of an individual's microbiome. For example, PCR can be performed in parallel or sequentially using multiple markers (e.g., a first marker, a second marker, a third marker, an Nth marker), and can be associated with one or more bacterial markers, fungal markers, and eukaryotic markers. Furthermore, overlapping read data (e.g., generated by paired-end sequencing) can be constructed based on the output of an alignment algorithm, or aligned sequence read data can be merged with a reference sequence (e.g., using a hidden Markov model banding approach, using the Durbin-Holmes approach).However, alignment and mapping can be performed using any other suitable algorithm or method. In some cases, sequence read data can be coded to facilitate alignment and mapping operations. In one example, each base of a sequence can be coded as a byte according to the following sequence: 0000TGCA, whereby the least significant bit is 1 if the base is sequenced as likely containing base A (e.g., A is represented as 00000001), the next significant bit is 1 if the base is sequenced as likely containing base C (e.g., C is represented as 00000010), the next significant bit is 1 if the base is sequenced as likely containing base G (e.g., G is represented as 00000100), and the next significant bit is 1 if the base is sequenced as likely containing base T (e.g., T is represented as 00001000). In this example, the four most significant bits are set to zero. However, alternative variations of this example can code bases in any other suitable manner. Furthermore, the predetermined sequences of the primers used during amplification can be used to trim the sequence read data to exclude the primer sequences, increasing the efficiency of alignment and mapping.
[0087] The subset of nucleic acid molecules may comprise one or more genomes disclosed herein. The subset of nucleic acid molecules may comprise one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, eleven or more, twelve or more, thirteen or more, fourteen or more, fifteen or more, twenty or more, twenty-five or more, thirty or more, thirty-five or more, forty or more, fifty or more, sixty or more, seventy or more, eighty or more, ninety or more, or one hundred or more genomes. The one or more genomes may be identical, similar, different, or a combination thereof. In some cases, there are two subsets of nucleic acids (Figure 6A).
[0088] The subset of nucleic acid molecules may comprise one or more genomic features disclosed herein. The subset of nucleic acid molecules may comprise one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, eleven or more, twelve or more, thirteen or more, fourteen or more, fifteen or more, twenty or more, twenty-five or more, thirty or more, thirty-five or more, forty or more, fifty or more, sixty or more, seventy or more, eighty or more, ninety or more, or one hundred or more genomic features. The one or more genomic features may be identical, similar, different, or a combination thereof.
[0089] A subset of nucleic acid molecules may contain nucleic acid molecules of different sizes. The length of the nucleic acid molecules in a subset of nucleic acid molecules may be referred to as the size of the nucleic acid molecules. The average length of the nucleic acid molecules in a subset of nucleic acid molecules may be referred to as the average size of the nucleic acid molecules. As used herein, the terms "size of nucleic acid molecules," "average size of nucleic acid molecules," "molecular size," and "average molecular size" may be used interchangeably. The size of nucleic acid molecules may be used to distinguish two or more subsets of nucleic acid molecules. The difference between the average size of nucleic acid molecules in a subset of nucleic acid molecules and the average size of nucleic acid molecules in another subset of nucleic acid molecules may be used to distinguish two subsets of nucleic acid molecules. The average size of nucleic acid molecules in one subset of nucleic acid molecules may be larger than the average size of nucleic acid molecules in at least one other subset of nucleic acid molecules. The average size of nucleic acid molecules in one subset of nucleic acid molecules may be smaller than the average size of nucleic acid molecules in at least one other subset of nucleic acid molecules. The difference in average molecular size between two or more subsets of nucleic acid molecules is at least about 50; 75; 100; 125; 150; 175; 200; 225; 250; 275; 300; 350; 400; 450; 500; 550; 600; 650; 700; 750; 800; 850; 900; 950; 1,000; 1100; 1200; 1300; 1400; 1500 The difference in average molecular size between two or more subsets of nucleic acid molecules may be at least about 200 bases or base pairs. Alternatively, the difference in average molecular size between two or more subsets of nucleic acid molecules may be at least about 300 bases or base pairs.
[0090] The subset of nucleic acid molecules may contain nucleic acid molecules of different sequencing sizes. The length of the nucleic acid molecules in the subset of nucleic acid molecules to be sequenced may be referred to as the sequencing size of the nucleic acid molecules. The average length of the nucleic acid molecules in the subset of nucleic acid molecules may be referred to as the average sequencing size of the nucleic acid molecules. As used herein, the terms "sequencing size of nucleic acid molecules," "average sequencing size of nucleic acid molecules," "molecular sequencing size," and "average molecular sequencing size" may be used interchangeably. The average molecular sequencing size of one or more subsets of nucleic acid molecules is at least about 50; 75; 100; 125; 150; 175; 200; 225; 250; 275; 300; 350; 400; 450; 500; 550; 600; 650; 700; 750; 800; 850; 900; 950; 1,000; 1100; 1200; 1300; 1400; 1500; 1600 The bases or base pairs may be 1,700; 1,800; 1,900; 2,000; 3,000; 4,000; 5,000; 6,000; 7,000; 8,000; 9,000; 10,000; 15,000; 20,000; 30,000; 40,000; 50,000; 60,000; 70,000; 80,000; 90,000; 100,000 or more. The sequencing size of nucleic acid molecules can be used to distinguish between two or more subsets of nucleic acid molecules. The difference between the average sequencing size of nucleic acid molecules in a subset of nucleic acid molecules and the average sequencing size of nucleic acid molecules in another subset of nucleic acid molecules can be used to distinguish between two subsets of nucleic acid molecules. The average sequencing size of the nucleic acid molecules in one subset of nucleic acid molecules may be greater than the average sequencing size of the nucleic acid molecules in at least one other subset of nucleic acid molecules. The average sequencing size of the nucleic acid molecules in one subset of nucleic acid molecules may be smaller than the average sequencing size of the nucleic acid molecules in at least one other subset of nucleic acid molecules.The difference in average molecular sequencing size between two or more subsets of nucleic acid molecules is at least about 50; 75; 100; 125; 150; 175; 200; 225; 250; 275; 300; 350; 400; 450; 500; 550; 600; 650; 700; 750; 800; 850; 900; 950; 1000; 1100; 1200; 1300; 1400; 1500; The difference in average molecular sequencing size between two or more subsets of nucleic acid molecules is at least about 200 bases or base pairs.Alternatively, the difference in average molecular sequencing size between two or more subsets of nucleic acid molecules is at least about 300 bases or base pairs.
[0091] The methods disclosed herein may include one or more capture probes, multiple capture probes, or one or more capture probe sets. Typically, the capture probe includes a nucleic acid binding site. The capture probe may hybridize with the captured nucleic acid. The capture probe may include a nucleic acid sequence complementary to the captured nucleic acid. In some cases, the capture probe may include a nucleic acid sequence that is completely complementary to a portion of the captured nucleic acid. For example, each nucleic acid in the capture probe may be complementary to a base in the captured nucleic acid. The capture probe may be longer than the capture nucleic acid. For example, each base in the captured nucleic acid may be complementary to a base in the capture probe, but not all bases in the capture probe are complementary to bases in the captured nucleic acid. The capture probe may be shorter than the captured nucleic acid. For example, each base in the capture probe may be complementary to a base in the captured nucleic acid, but not all bases in the captured nucleic acid may be complementary to the capture probe.
[0092] The capture probe may perform a capture reaction in solution. The capture probe may be in solution and capture nucleic acid in solution. The captured nucleic acid may then be isolated and / or eluted as described elsewhere herein. The capture probe may capture nucleic acid in solution, and then the capture probe may be attached to a support such as a solid support (e.g., an array or beads). In some situations, the support may be in the form of a semi-solid material (e.g., a gel).
[0093] Attachment to the support may be non-covalent. For example, nucleic acid may be captured by a biotinylated probe, which then binds to avidin / streptavidin beads, thereby attaching the captured nucleic acid complex to the avidin / streptavidin beads. Other binding pairs may also be used to attach the capture probe to the surface. The capture probe may be covalently attached to the support. For example, the support may have a chemically reactive linker that can react with the capture probe so that the capture probe is covalently linked to the support.
[0094] The capture probe may be attached to a support, and a capture reaction may be carried out. For example, the capture probe may be coupled with a support, and then the nucleic acid molecule may be captured. The support may be a solid or semi-solid (e.g., gel) material. Examples of the support include, but are not limited to, beads, slides, and chips. The support may be, for example, glass, silica, silicon, plastic (such as polystyrene), agar, or agarose.
[0095] The capture probe may further comprise one or more linkers. The capture probe may further comprise one or more labels. The one or more linkers may attach one or more labels to the nucleic acid binding site. In some cases, the capture probe may be designed to hybridize with a shared region of the 16S gene sequence. Capture probes that can be designed to target a shared region of the 16S gene sequence can be used to capture nucleic acid molecules from a wide variety of species, even species that have not yet been identified and characterized. In some cases, the method may include a first plurality of nucleic acid probes configured to target elements of a human genome sequence. In some cases, the method may include a second plurality of nucleic acid probes configured to target elements from a genome sequence of a non-human species.
[0096] The methods disclosed herein may involve the detection of one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, twenty or more, thirty or more, forty or more, fifty or more, sixty or more, seventy or more, eighty or more, ninety or more, one hundred or more, one hundred or more, one hundred and twenty-five or more, one hundred and fifty or more, one hundred and seventy-five or more, two hundred or more, two hundred and fifty or more, three hundred and fifty or more, five hundred and fifty or more, sixty or more, seventy or more, eighty or more, ninety or more, one hundred and fif ... or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, 1000 or more, 5000 or more, 10,000 or more, 20,000 or more, 30,000 or more, 40,000 or more, 50,000 or more, 60,000 or more, 70,000 or more, 80,000 or more, 90,000 or more, or 100,0000 or more capture probes or capture probe sets. In some cases, the method can include about 50,000 capture probes. The one or more capture probes or capture probe sets can be different, similar, identical, or a combination thereof. The one or more capture probes or capture probe sets can be at various relative concentrations compared to the other capture probes.For example, some capture probes may be at higher concentrations to capture nucleic acids that may be difficult to capture (e.g., sequences with high GC content, sequences that are the result of sequence recombination, sequences with a large number of mutations, sequences with a high mutation rate), thereby increasing the likelihood of capturing a particular nucleic acid.
[0097] The one or more capture probes may comprise a nucleic acid binding site that hybridizes to at least a portion of one or more nucleic acid molecules, or variants or derivatives thereof, or a subset of nucleic acid molecules in a sample. The capture probes may comprise a nucleic acid binding site that hybridizes to one or more genomes. The capture probes may hybridize to different, similar, and / or identical genomes. The one or more capture probes may be at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99%, or more complementary to one or more nucleic acid molecules, or variants or derivatives thereof.
[0098] A capture probe may comprise one or more nucleotides. A capture probe may comprise 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, or more than 100 nucleotides. The capture probe may contain more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more nucleotides. The capture probe may contain about 100 nucleotides. The capture probe may contain between about 10 and about 500 nucleotides, between about 20 and about 450 nucleotides, between about 30 and about 400 nucleotides, between about 40 and about 350 nucleotides, between about 50 and about 300 nucleotides, between about 60 and about 250 nucleotides, between about 70 and about 200 nucleotides, or between about 80 and about 150 nucleotides. In some embodiments of the present disclosure, the capture probe comprises between about 80 nucleotides and about 100 nucleotides.
[0099] A plurality of capture probes or a capture probe set may include two or more capture probes with identical, similar, and / or different nucleic acid binding site sequences, linkers, and / or labels. For example, two or more capture probes include identical nucleic acid binding sites. In another example, two or more capture probes include similar nucleic acid binding sites. In yet another example, two or more capture probes include different nucleic acid binding sites. Two or more capture probes may further include one or more linkers. Two or more capture probes may further include different linkers. Two or more capture probes may further include similar linkers. Two or more capture probes may further include identical linkers. Two or more capture probes may further include one or more labels. Two or more capture probes may further include different labels. Two or more capture probes may further include similar labels. The two or more capture probes may further comprise the same label.
[0100] Assays can include, but are not limited to, sequencing, amplification, hybridization, enrichment, isolation, elution, fragmentation, detection, and quantification of one or more nucleic acid molecules. Assays can include methods for preparing one or more nucleic acid molecules. Assays can include traditional assays, long read data, high GC content, and hybridization assays (Figure 9). Any number of assays can be performed. The number of assays can be 1, 2, 3, 4, 5, 6, 7, 8, 9, or up to 10 assays. For example, Figure 6B illustrates a schematic showing the use of two assays. Similarly, any number of analyses of data from one or more assays can be performed. The number of analyses can be 1, 2, 3, 4, 5, 6, 7, 8, 9, or up to 10 analyses of data from one or more assays. For example, Figure 6C illustrates two different analyses being performed. In some cases, the analysis is a bioinformatics analysis (Figure 8). Assays can be performed using any number of protocols. For example, one may utilize protocols 1, 2, 3, 4, 5, 6, 7, 8, 9, or up to 10. Figure 6D provides a schematic diagram utilizing four protocols.
[0101] The methods disclosed herein may include performing one or more sequencing reactions on one or more nucleic acid molecules in a sample. The methods disclosed herein may include performing one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, fifteen or more, twenty or more, thirty or more, forty or more, fifty or more, sixty or more, seventy or more, eighty or more, ninety or more, one hundred or more, two hundred or more, three hundred or more, four hundred or more, five hundred or more, six hundred or more, seventy or more, eighty or more, ninety or more, or one thousand or more sequencing reactions on one or more nucleic acid molecules in a sample. The sequencing reactions may be performed simultaneously, sequentially, or a combination thereof. The sequencing reactions may include whole genome sequencing or exome sequencing. The sequencing reactions may include Maxim-Gilbert, chain termination, or high-throughput systems. Alternatively, or in addition, the sequencing reactions may include Helioscope™ single molecule sequencing, nanopore DNA sequencing, Lynx Therapeutics' Massively Parallel Signature Sequencing (MPSS), 454 pyrosequencing, single molecule real-time (RNAP) sequencing, Illumina (Solexa) sequencing, SOLiD sequencing, Ion Torrent™, ion semiconductor sequencing, single molecule SMRT™ sequencing, Polony sequencing, DNA nanoball sequencing, VisiGen bioengineering approaches, or a combination thereof.Alternatively or additionally, the sequencing reaction can include one or more sequencing platforms, including, but not limited to, single molecule real-time (SMRT™) technologies such as the Genome Analyzer IIx, HiSeq, and MiSeq offered by Illumina, the PacBio RS system and Solexa sequencer offered by Pacific Biosciences (California), and true single molecule sequencing (tSMS™) technologies such as the HeliScope™ sequencer offered by Helicos Inc. (Cambridge, Massachusetts). The sequencing reaction can also include electron microscopy or chemically sensitive field effect transistor (chemFET) arrays. In some embodiments of the present disclosure, the sequencing reaction includes capillary sequencing, next-generation sequencing, Sanger sequencing, sequencing by synthesis, sequencing by ligation, sequencing by hybridization, single molecule sequencing, or a combination thereof. Sequencing by synthesis may include reversible terminator sequencing, progressive single molecule sequencing, sequential flow sequencing, or a combination thereof. Sequential flow sequencing may include pyrosequencing, pH-mediated sequencing, semiconductor sequencing, or a combination thereof.
[0102] The methods disclosed herein may include performing at least one long read sequencing reaction and at least one short read sequencing reaction. An example of a method including sequencing long read data and short read data is illustrated in FIG. 18. The long read sequencing reaction and / or the short read sequencing reaction may be performed on at least a portion of a subset of nucleic acid molecules. The long read sequencing reaction and / or the short read sequencing reaction may be performed on at least a portion of two or more subsets of nucleic acid molecules. Both the long read sequencing reaction and the short read sequencing reaction may be performed on at least a portion of one or more subsets of nucleic acid molecules.
[0103] Sequencing of one or more nucleic acid molecules or a subset thereof may be performed for at least about 5; 10; 15; 20; 25; 30; 35; 40; 45; 50; 60; 70; 80; 90; 100; 200; 300; 400; 500; 600; 700; 800; 900; 1,000; 1500; 2,000; 2500; 3,000; 3500; 4,000; 4500; 5,000; 5500; 6,000; 6500; 7,000; 7500; 8,000; It may contain 8500; 9,000; 10,000; 25,000; 50,000; 75,000; 100,000; 250,000; 500,000; 750,000; 10,000,000; 25,000,000; 50,000,000; 100,000,000; 250,000,000; 500,000,000; 750,000,000; 1,000,000,000 or more sequence reads.
[0104] The sequencing reaction may comprise sequencing at least about 50;60;70;80;90;100;110;120;130;140;150;160;170;180;190;200;210;220;230;240;250;260;270;280;290;300;325;350;375;400;425;450;475;500;600;700;800;900;1000;1500;2 This may include sequencing of 10,000; 2500; 3,000; 3500; 4,000; 4500; 5,000; 5500; 6,000; 6500; 7,000; 7500; 8,000; 8500; 9,000; 10,000; 20,000; 30,000; 40,000; 50,000; 60,000; 70,000; 80,000; 90,000; 100,000 or more bases or base pairs. The sequencing reaction may comprise sequencing at least about 50;60;70;80;90;100;110;120;130;140;150;160;170;180;190;200;210;220;230;240;250;260;270;280;290;300;325;350;375;400;425;450;475;500;600;700;800;900;1000;1500;2000; The present invention may involve sequencing of 00; 2500; 3,000; 3500; 4,000; 4500; 5,000; 5500; 6,000; 6500; 7,000; 7500; 8,000; 8500; 9,000; 10,000; 20,000; 30,000; 40,000; 50,000; 60,000; 70,000; 80,000; 90,000; 100,000 or more consecutive bases or base pairs.
[0105] The sequencing techniques used in the methods of the present disclosure may generate at least 100 reads per run, at least 200 reads per run, at least 300 reads per run, at least 400 reads per run, at least 500 reads per run, at least 600 reads per run, at least 700 reads per run, at least 800 reads per run, at least 900 reads per run, at least 1000 reads per run, at least 5,000 reads per run, at least 10,000 reads per run, at least 50,000 reads per run, at least 100,000 reads per run, at least 500,000 reads per run, or at least 1,000,000 reads per run. Alternatively, the sequencing techniques used in the methods of the present disclosure may generate at least 1,500,000 reads per run, at least 2,000,000 reads per run, at least 2,500,000 reads per run, at least 3,000,000 reads per run, at least 3,500,000 reads per run, at least 4,000,000 reads per run, at least 4,500,000 reads per run, or at least 5,000,000 reads per run.
[0106] The sequencing techniques used in the disclosed methods may generate at least about 30 base pairs, at least about 40 base pairs, at least about 50 base pairs, at least about 60 base pairs, at least about 70 base pairs, at least about 80 base pairs, at least about 90 base pairs, at least about 100 base pairs, at least about 110 base pairs, at least about 120 base pairs, at least about 150 base pairs, at least about 200 base pairs, at least about 250 base pairs, at least about 300 base pairs, at least about 350 base pairs, at least about 400 base pairs, at least about 450 base pairs, at least about 500 base pairs, at least about 550 base pairs, at least about 600 base pairs, at least about 700 base pairs, at least about 800 base pairs, at least about 900 base pairs, or at least about 1,000 base pairs per read. Alternatively, the sequencing techniques used in the methods of the present disclosure may generate long sequence read data.In some examples, the sequencing techniques used in the disclosed methods provide a sequencing capability of at least about 1,200 base pairs per read, at least about 1,500 base pairs per read, at least about 1,800 base pairs per read, at least about 2,000 base pairs per read, at least about 2,500 base pairs per read, at least about 3,000 base pairs per read, at least about 3,500 base pairs per read, at least about 4,000 base pairs per read, at least about 4,500 base pairs per read, at least about 5,000 base pairs per read, at least about 6,000 base pairs per read, base pairs per read, at least about 7,000 base pairs per read, at least about 8,000 base pairs per read, at least about 9,000 base pairs per read, at least about 10,000 base pairs per read, 20,000 base pairs per read, 30,000 base pairs per read, 40,000 base pairs per read, 50,000 base pairs per read, 60,000 base pairs per read, 70,000 base pairs per read, 80,000 base pairs per read, 90,000 base pairs per read, or 100,000 base pairs per read.
[0107] High-throughput sequencing system can detect the nucleotides to be sequenced immediately after or when they are incorporated into the growing chain, i.e., detect sequences in real time or substantially real time.In some cases, high-throughput sequencing generates at least 1,000, at least 5,000, at least 10,000, at least 20,000, at least 30,000, at least 40,000, at least 50,000, at least 100,000, or at least 500,000 sequence read data per hour, and each read data is at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 120, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, or at least 500 bases per read data.Sequencing can be performed using the nucleic acid described herein, such as genomic DNA, cDNA derived from RNA transcripts, or RNA as a template.
[0108] The methods disclosed herein may include performing one or more amplification reactions on one or more nucleic acid molecules in a sample. The term "amplification" refers to any process that generates at least one copy of a nucleic acid molecule. The terms "amplicon" and "amplified nucleic acid molecule" refer to copies of a nucleic acid molecule and can be used interchangeably. The amplification reaction may include PCR-based methods, non-PCR-based methods, or a combination thereof. Examples of non-PCR-based methods include, but are not limited to, multiplex displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, rolling circle amplification, or circle-circle amplification. PCR-based methods may include, but are not limited to, PCR, HD-PCR, Next Gen PCR, digital RTA, or any combination thereof. Additional PCR methods include, but are not limited to, linear amplification, allele-specific PCR, Alu PCR, assembly PCR, asymmetric PCR, droplet PCR, emulsion PCR, helicase-dependent amplification (HDA), hot-start PCR, inverse PCR, exponential-then-linear (LATE)-PCR, long PCR, multiplex PCR, nested PCR, semi-nested PCR, quantitative PCR, RT-PCR, real-time PCR, single-cell PCR, and touchdown PCR.
[0109] The methods disclosed herein may include performing one or more hybridization reactions on one or more nucleic acid molecules in a sample. The hybridization reaction may include hybridization of one or more capture probes to one or more nucleic acid molecules or a subset of nucleic acid molecules in the sample. The hybridization reaction may include hybridization of one or more capture probe sets to one or more nucleic acid molecules or a subset of nucleic acid molecules in the sample. The hybridization reaction may include one or more hybridization arrays, multiplex hybridization reactions, hybridization chain reactions, isothermal hybridization reactions, nucleic acid hybridization reactions, or combinations thereof. The one or more hybridization arrays may include hybridization array genotyping, hybridization array proportional sensing, DNA hybridization arrays, macroarrays, microarrays, high-density oligonucleotide arrays, genomic hybridization arrays, comparative hybridization arrays, or combinations thereof. The hybridization reaction may include one or more capture probes, one or more beads, one or more labels, one or more subsets of nucleic acid molecules, one or more nucleic acid samples, one or more reagents, one or more wash buffers, one or more elution buffers, one or more hybridization buffers, one or more hybridization chambers, one or more incubators, one or more separators, or combinations thereof.
[0110] The methods disclosed herein may include performing one or more enrichment reactions on one or more nucleic acid molecules in a sample. The enrichment reaction may include contacting the sample with one or more beads or bead sets. The enrichment reaction may include differential amplification of two or more subsets of nucleic acid molecules based on one or more genomic features. For example, the enrichment reaction may include differential amplification of two or more subsets of nucleic acid molecules based on GC content. Alternatively, or in addition, the enrichment reaction may include differential amplification of two or more subsets of nucleic acid molecules based on methylation status. The enrichment reaction may include one or more hybridization reactions. The enrichment reaction may further include isolating and / or purifying one or more hybridized nucleic acid molecules, one or more bead-bound nucleic acid molecules, one or more free nucleic acid molecules (e.g., nucleic acid molecules without capture probes, nucleic acid molecules without beads), one or more labeled nucleic acid molecules, one or more unlabeled nucleic acid molecules, one or more amplicons, one or more unamplified nucleic acid molecules, or a combination thereof. Alternatively, or in addition, the enrichment reaction may include enriching for one or more cell types in the sample. One or more cell types may be enriched by flow cytometry.
[0111] One or more enrichment reactions can produce one or more enriched nucleic acid molecules. The enriched nucleic acid molecules may include nucleic acid molecules or variants or derivatives thereof. For example, the enriched nucleic acid molecules may include one or more hybridized nucleic acid molecules, one or more bead-bound nucleic acid molecules, one or more free nucleic acid molecules (e.g., nucleic acid molecules without capture probes, nucleic acid molecules without beads), one or more labeled nucleic acid molecules, one or more unlabeled nucleic acid molecules, one or more amplicons, one or more unamplified nucleic acid molecules, or a combination thereof. The enriched nucleic acid molecules may be distinguished from non-enriched nucleic acid molecules by GC content, molecular size, genome, genome features, or a combination thereof. The enriched nucleic acid molecules may be derived from one or more assays, supernatants, eluates, or a combination thereof. The enriched nucleic acid molecules may differ from non-enriched nucleic acid molecules by average size, average GC content, genome, or a combination thereof. In some cases, the enrichment may include multiple subsets of DNA enriched for different genomic regions, which undergo independent processing operations before being combined for sequencing assays (Figure 15). In some cases, the enrichment may include multiple subsets of DNA enriched for different genomic regions, which undergo independent processing operations before being independently sequenced and analyzed (Figure 16).
[0112] The methods disclosed herein may include performing one or more isolation or purification reactions on one or more nucleic acid molecules in a sample. The isolation or purification reaction may include contacting the sample with one or more beads or bead sets. The isolation or purification reaction may include one or more hybridization reactions, enrichment reactions, amplification reactions, sequencing reactions, or combinations thereof. The isolation or purification reaction may include the use of one or more separators. The one or more separators may include a magnetic separator. The isolation or purification reaction may include separating nucleic acid molecules bound to beads from nucleic acid molecules that do not contain beads. The isolation or purification reaction may include separating nucleic acid molecules hybridized to a capture probe from nucleic acid molecules that do not contain the capture probe. The isolation or purification reaction may include separating a first subset of nucleic acid molecules from a second subset of nucleic acid molecules, where the first subset of nucleic acid molecules differs from the second subset of nucleic acid molecules by average size, average GC content, genome, or a combination thereof.
[0113] The methods disclosed herein may include performing one or more elution reactions on one or more nucleic acid molecules in a sample. The elution reaction may include contacting the sample with one or more beads or bead sets. The elution reaction may include separating nucleic acid molecules bound to beads from nucleic acid molecules that do not contain beads. The elution reaction may include separating nucleic acid molecules hybridized to a capture probe from nucleic acid molecules that do not contain a capture probe. The elution reaction may include separating a first subset of nucleic acid molecules from a second subset of nucleic acid molecules, wherein the first subset of nucleic acid molecules differs from the second subset of nucleic acid molecules by average size, average GC content, genome, or a combination thereof.
[0114] The methods disclosed herein may include one or more fragmentation reactions. The fragmentation reaction may involve fragmenting one or more nucleic acid molecules, or a subset of nucleic acid molecules, in a sample to generate one or more fragmented nucleic acid molecules. The one or more nucleic acid molecules may be fragmented by sonication, needle shearing, nebulization, shearing (e.g., acoustic shearing, mechanical shearing, point-sink shearing), passage through a French pressure cell, or enzymatic digestion. Enzymatic digestion may occur by nuclease digestion (e.g., micrococcal nuclease digestion, endonuclease, exonuclease, RNase H, or DNase I). Fragmentation of one or more nucleic acid molecules may result in fragments ranging in size from about 100 base pairs to about 2000 base pairs, from about 200 base pairs to about 1500 base pairs, from about 200 base pairs to about 1000 base pairs, from about 200 base pairs to about 500 base pairs, from about 500 base pairs to about 1500 base pairs, and from about 500 base pairs to about 1000 base pairs. One or more fragmentation reactions may result in fragments ranging in size from about 50 base pairs to about 1000 base pairs. The one or more fragmentation reactions may result in fragments of about 100 base pairs, 150 base pairs, 200 base pairs, 250 base pairs, 300 base pairs, 350 base pairs, 400 base pairs, 450 base pairs, 500 base pairs, 550 base pairs, 600 base pairs, 650 base pairs, 700 base pairs, 750 base pairs, 800 base pairs, 850 base pairs, 900 base pairs, 950 base pairs, 1000 base pairs, or greater in size.
[0115] Fragmentation of one or more nucleic acid molecules may involve mechanical shearing of one or more nucleic acid molecules in a sample for a period of time. The fragmentation reaction may occur for at least about 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500 seconds, or longer.
[0116] Fragmenting the one or more nucleic acid molecules may include contacting the nucleic acid sample with one or more beads.Fragmenting the one or more nucleic acid molecules may include contacting the nucleic acid sample with a plurality of beads, wherein the ratio of the volume of the plurality of beads to the volume of the nucleic acid sample is about 0.10, 0.20, 0.30, 0.40, 0.50, 0.60, 0.70, 0.80, 0.90, 1.00, 1.10, 1.20, 1.30, 1.40, 1.50, 1.60, 1.70, 1.80, 1.90, 2.00, or more. Fragmentation of one or more nucleic acid molecules may include contacting the nucleic acid sample with a plurality of beads, wherein the ratio of the volume of the plurality of beads to the volume of the nucleic acid is about 2.00, 1.90, 1.80, 1.70, 1.60, 1.50, 1.40, 1.30, 1.20, 1.10, 1.00, 0.90, 0.80, 0.70, 0.60, 0.50, 0.40, 0.30, 0.20, 0.10, 0.05, 0.04, 0.03, 0.02, 0.01, or less.
[0117] The methods disclosed herein may include performing one or more detection reactions on one or more nucleic acid molecules in a sample. The detection reaction may include one or more sequencing reactions. Alternatively, performing the detection reaction includes optical sensing, electrical sensing, or a combination thereof. The optical sensing may include optical sensing of photoluminescence photon emission, fluorescence photon emission, pyrophosphate photon emission, chemiluminescence photon emission, or a combination thereof. The electrical sensing may include electrical sensing of ion concentration, ion current modulation, nucleotide electric field, nucleotide tunneling current, or a combination thereof.
[0118] The methods disclosed herein may include performing one or more quantification reactions on one or more nucleic acid molecules in a sample. The quantification reactions may include sequencing, PCR, qPCR, digital PCR, or a combination thereof.
[0119] The method disclosed herein may comprise one or more samples.The method disclosed herein may comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more samples.The sample may be derived from a subject.Two or more samples may be derived from a single subject. The two or more samples may be from 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, or more different subjects. The subjects may be mammals, reptiles, amphibians, birds, and fish. Mammals may be humans, apes, orangutans, monkeys, chimpanzees, cows, pigs, horses, rodents, dogs, cats, or other animals. Reptiles may be lizards, snakes, alligators, turtles, crocodiles, and tortoises. Amphibians may be toads, frogs, newts, and salamanders. Examples of birds include, but are not limited to, ducks, geese, penguins, ostriches, and owls. Examples of fish include, but are not limited to, catfish, eels, sharks, and swordfish. The subject may be a human. The subject may be suffering from a disease or condition.
[0120] The two or more samples may be taken over 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 or more time points. The time points may be present over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 or more hours. The time points may occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 or more days. The time points may occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 or more weeks. The time points may occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 or more months. The time points may occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 or more years.
[0121] In some cases, the method can include obtaining a biological sample from a subject. The subject can be human or non-human. The subject can be an adult or a child. In some cases, an adult subject can be 18 years of age or over 18 years of age. In some cases, the subject's biological sample can be derived from a tumor biopsy, whole blood, or plasma. In some cases, the biological sample can be derived from a bodily fluid, cell, skin, tissue, organ, or a combination thereof. The sample can be blood, plasma, a blood fraction, saliva, sputum, urine, semen, vaginal fluid, cerebrospinal fluid, feces, cell, or tissue biopsy. The sample can be derived from the adrenal gland, appendix, bladder, brain, ear, esophagus, eye, gallbladder, heart, kidney, large intestine, liver, lung, mouth, muscle, nose, pancreas, parathyroid gland, pineal gland, pituitary gland, skin, small intestine, spleen, stomach, thymus, thyroid gland, trachea, uterus, appendix, cornea, skin, heart valve, artery, or vein.
[0122] The sample may contain one or more nucleic acid molecules. The nucleic acid molecule may be a DNA molecule, an RNA molecule (e.g., mRNA, cRNA, or miRNA), or a DNA / RNA hybrid. Examples of DNA molecules include, but are not limited to, double-stranded DNA, single-stranded DNA, single-stranded DNA hairpins, cDNA, and genomic DNA. The nucleic acid may be an RNA molecule, such as double-stranded RNA, single-stranded RNA, ncRNA, RNA hairpins, and mRNA. Examples of ncRNA include, but are not limited to, siRNA, miRNA, snoRNA, piRNA, tiRNA, PASR, TASR, aTASR, TSSa-RNA, snRNA, RE-RNA, uaRNA, x-ncRNA, hY RNA, usRNA, snaR, and vtRNA.
[0123] The methods disclosed herein may include one or more containers. The methods disclosed herein may include one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, twenty or more, thirty or more, forty or more, fifty or more, sixty or more, seventy or more, eighty ... twenty or more, thirty or more, forty or more, fifty or more, sixty or more, seventy or more, eighty or more, nine or more, ten or more, twenty or more, twenty or more, thirty or more, forty or more, fifty or more, sixty or more, seventy or more, eighty or more, nine or more, ten or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more The container may contain 0 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more containers. The one or more containers may be different, similar, identical, or a combination thereof. Examples of containers include, but are not limited to, plates, microplates, PCR plates, wells, microwells, tubes, Eppendorf tubes, vials, arrays, microarrays, and chips.
[0124] The methods disclosed herein may include one or more reagents. The methods disclosed herein may include one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, twenty or more, thirty or more, forty or more, fifty or more, sixty or more, seventy or more, eighty ... twenty or more, thirty or more, forty or more, fifty or more, sixty or more, seventy or more, eighty or more, nine or more, ten or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty The assay may include 0 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more reagents. One or more reagents may be different, similar, identical, or a combination thereof. A reagent may improve the efficiency of one or more assays. A reagent may improve the stability of a nucleic acid molecule, or a variant or derivative thereof. Reagents may include, but are not limited to, enzymes, proteases, nucleases, molecules, polymerases, reverse transcriptases, ligases, and chemical compounds. The methods disclosed herein may include performing an assay with one or more antioxidants. Generally, an antioxidant is a molecule that inhibits the oxidation of another molecule. Examples of antioxidants include, but are not limited to, ascorbic acid (e.g., vitamin C), glutathione, lipoic acid, uric acid, carotene, α-tocopherol (e.g., vitamin E), ubiquinol (e.g., coenzyme Q), and vitamin A.
[0125] The methods disclosed herein may include one or more buffers or solutions. The methods disclosed herein may include one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, twenty or more, thirty or more, forty or more, fifty or more, sixty or more, seventy or more, eighty or more, ninety or more, or ... The assay may include 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more buffers or solutions. One or more buffers or solutions may be different, similar, identical, or a combination thereof. The buffer or solution may improve the efficiency of one or more assays. The buffer or solution may improve the stability of the nucleic acid molecule, or a variant or derivative thereof. Buffers or solutions may include, but are not limited to, wash buffers, elution buffers, and hybridization buffers.
[0126] The methods disclosed herein may include one or more beads, multiple beads, or one or a set of beads. The methods disclosed herein may include one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, twenty or more, thirty or more, forty or more, fifty or more, sixty or more, seventy or more, eighty or more, ninety or more, or ten or more. The beads or bead sets may comprise more than 100, more than 125, more than 150, more than 175, more than 200, more than 250, more than 300, more than 350, more than 400, more than 500, more than 600, more than 700, more than 800, more than 900, or more than 1000 beads or bead sets. One or more beads or bead sets may be different, similar, identical, or a combination thereof. The beads may be magnetic, antibody-coated, protein A-crosslinked, protein G-crosslinked, streptavidin-coated, oligonucleotide-conjugated, silica-coated, or a combination thereof.Examples of beads include, but are not limited to, Ampure beads, AMPure XP beads, streptavidin beads, agarose beads, magnetic beads, Dynabeads®, MACS® microbeads, antibody-conjugated beads (e.g., anti-immunoglobulin microbeads), Protein A-conjugated beads, Protein G-conjugated beads, Protein A / G-conjugated beads, Protein L-conjugated beads, oligo-dT-conjugated beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorescent dye microbeads, and BcMag™ carboxy-terminated magnetic beads. In some embodiments of the present disclosure, one or more beads comprise one or more Ampure beads. Alternatively, or in addition, one or more beads comprise AMPure XP beads.
[0127] The methods disclosed herein may include one or more primers, multiple primers, or one or more primer sets. The primers may further include one or more linkers. The primers may further include one or more labels. The primers may be used in one or more assays. For example, the primers are used in one or more sequencing reactions, amplification reactions, or a combination thereof. The methods disclosed herein may include one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, twenty or more, thirty or more, forty or more, fifty or more, sixty or more, seventy or more, eighty or more, ninety ... twenty or more, twenty or more, twenty or more, seventy or more, eighty or more, ninety or more, ten or more, twenty or more, twenty or more, twenty or more, seventy or more, seventy or more, eighty or more, ninety or more, ten or more, twenty or more, twenty or more, twenty or more, twenty or more, seventy or more, seventy or more, eighty or more, ninety or more, ten or more, twenty or more, twenty or more, twenty or more, seventy or more, seventy or more, eighty or more, ninety or more, ten or more, twenty or more, twenty or more, twenty or more, seventy or more, seventy or more, eighty or more, ninety or more, ten or more, twenty or more, twenty or more, twenty or more, seventy or more, seventy or The primers or primer sets may comprise more than 100, 125, 150, 175, 200, 250, 300, 350, 400, 500, 600, 700, 800, 900, or 1000 primers or primer sets. The primers may comprise about 100 nucleotides. The primers may comprise between about 10 and about 500 nucleotides, between about 20 and about 450 nucleotides, between about 30 and about 400 nucleotides, between about 40 and about 350 nucleotides, between about 50 and about 300 nucleotides, between about 60 and about 250 nucleotides, between about 70 and about 200 nucleotides, or between about 80 and about 150 nucleotides. In some embodiments of the present disclosure, the primer comprises between about 80 nucleotides and about 100 nucleotides.The one or more primers or primer sets may be different, similar, identical, or a combination thereof.
[0128] The primers may hybridize to at least a portion of one or more nucleic acid molecules or variants or derivatives thereof, or a subset of nucleic acid molecules in a sample. The primers may hybridize to one or more genomes. The primers may hybridize to different, similar, and / or identical genomes. The one or more primers may be at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99%, or more complementary to one or more nucleic acid molecules or variants or derivatives thereof.
[0129] A primer may comprise one or more nucleotides. A primer may comprise one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, twenty or more, thirty or more, forty or more, fifty or more, sixty or more, seventy or more, eighty or more, ninety or more, or ten or more nucleotides. The primer may contain 100 or more nucleotides, 100 or more nucleotides, 125 or more nucleotides, 150 or more nucleotides, 175 or more nucleotides, 200 or more nucleotides, 250 or more nucleotides, 300 or more nucleotides, 350 or more nucleotides, 400 or more nucleotides, 500 or more nucleotides, 600 or more nucleotides, 700 or more nucleotides, 800 or more nucleotides, 900 or more nucleotides, or 1000 or more nucleotides. The primer may contain about 100 nucleotides. The primer may contain between about 10 and about 500 nucleotides, between about 20 and about 450 nucleotides, between about 30 and about 400 nucleotides, between about 40 and about 350 nucleotides, between about 50 and about 300 nucleotides, between about 60 and about 250 nucleotides, between about 70 and about 200 nucleotides, or between about 80 and about 150 nucleotides. In some embodiments of the present disclosure, the primer comprises between about 80 nucleotides and about 100 nucleotides.
[0130] A plurality of primers or primer sets may include two or more primers with identical, similar, and / or different sequences, linkers, and / or labels. For example, two or more primers include identical sequences. In another example, two or more primers include similar sequences. In yet another example, two or more primers include different sequences. Two or more primers may further include one or more linkers. Two or more primers may further include different linkers. Two or more primers may further include similar linkers. Two or more primers may further include identical linkers. Two or more primers may further include one or more labels. Two or more primers may further include different labels. Two or more primers may further include similar labels. Two or more primers may further include identical labels.
[0131] In some cases, universal primers can be utilized and can include one or more of the following: 8F primer, 27F primer, CC[F] primer, 357F primer, 515F primer, 533F primer, 16S.1100.F16 primer, 1237F primer, 519R primer, CD[R] primer, 907R primer, 1391R primer, 1492R(I) primer, 1492R(s) primer, U1492R primer, 928F primer, 336R primer, 1100F primer, 1100R primer, 337F primer, 785F primer, 805R primer, 518R primer, and any other suitable universal primers. Alternatively, for samples for which specific primers may be appropriate, amplification can be performed using the specific primers. In examples, specific primers may include the CYA106 primer (for cyanobacteria), the CYA359F primer (for cyanobacteria), the 895F primer (for bacteria excluding plastids and cyanobacteria), the CYA781R primer (for cyanobacteria), the 902R primer (for bacteria excluding plastids and cyanobacteria), the 904R primer (for bacteria excluding plastids and cyanobacteria), the 1100R primer (for bacteria), the 1185mR primer (for bacteria excluding plastids and cyanobacteria), the 1185aR primer (for lichen-associated Rhizobiales), the 1381R primer (for bacteria excluding Asterochloris species plastids), or any other suitable specific primer.
[0132] The capture probe, primer, label, and / or bead may comprise one or more nucleotides that may include RNA, DNA, a mixture of DNA and RNA residues, or modified analogs thereof such as 2'-OMe or 2'-fluoro (2'-F), locked nucleic acid (LNA), or an abasic moiety.
[0133] The methods disclosed herein may include one or more labels. The methods disclosed herein may include one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, twenty or more, thirty or more, forty or more, fifty or more, sixty or more, seventy or more, eighty ... twenty or more, thirty or more, forty or more, fifty or more, sixty or more, seventy or more, eighty or more, nine or more, ten or more, twenty or more, twenty or more, thirty or more, forty or more, fifty or more, sixty or more, seventy or more, eighty or more, nine or more, ten or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or more, twenty or It may contain 0 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more labels. One or more labels may be different, similar, identical, or a combination thereof.
[0134] Examples of labels include, but are not limited to, chemical labels, biochemical labels, biological labels, chromogenic labels, enzymatic labels, fluorescent labels, and luminescent labels. Labels include dyes, photocrosslinkers, cytotoxic compounds, drugs, affinity labels, photoaffinity labels, reactive compounds, antibodies or antibody fragments, biomaterials, nanoparticles, spin labels, fluorophores, metal-containing moieties, radioactive moieties, novel functional groups, groups that interact covalently or non-covalently with other molecules, photocaged moieties, actin radioexcitation moieties, ligands, photoisomerizable moieties, biotin, biotin analogs, heavy atom-incorporating moieties, chemically cleavable groups, photocleavable groups, redox-active agents, isotope-labeled moieties, biophysical probes, phosphorescent groups, chemiluminescent groups, electron-dense groups, magnetic groups, intercalating groups, chromophores, energy transfer agents, bioactive agents, detectable labels, or combinations thereof.
[0135] The label may be a chemical label. Examples of chemical labels include, but are not limited to, biotin and radioisotopes (e.g., iodine, carbon, phosphorus, hydrogen).
[0136] The methods, kits, and compositions disclosed herein may include biological labels, including, but not limited to, metabolic labels, including bioorthogonal azide-modified amino acids, sugars, and other compounds.
[0137] The methods, kits, and compositions disclosed herein may include an enzyme label. Examples of enzyme labels include, but are not limited to, horseradish peroxidase (HRP), alkaline phosphatase (AP), glucose oxidase, and β-galactosidase. The enzyme label may be luciferase.
[0138] The methods, kits, and compositions disclosed herein may include a fluorescent label. The fluorescent label may be an organic dye (e.g., FITC), a biological fluorophore (e.g., green fluorescent protein), or a quantum dot. A non-limiting list of fluorescent labels includes fluorescein isothiocyanate (FITC), DyLight Fluors, fluorescein, rhodamine (tetramethylrhodamine isothiocyanate, TRITC), coumarin, Lucifer Yellow, and BODIPY. The label may be a fluorophore. Examples of fluorophores include, but are not limited to, indocarbocyanine (C3), indodicarbocyanine (C5), Cy3, Cy3.5, Cy5, Cy5.5, Cy7, Texas Red, Pacific Blue, Oregon Green 488, Alexa Fluor®-355, Alexa Fluor 488, Alexa Fluor 532, Alexa Fluor 546, Alexa Fluor-555, Alexa Fluor 568, Alexa Fluor 594, Alexa Fluor 647, Alexa Fluor 660, Alexa Fluor 680, JOE, Lissamine, rhodamine green, BODIPY, fluorescein isothiocyanate (FITC), carboxy-fluorescein (FAM), phycoerythrin, rhodamine, dichlororhodamine (d-rhodamine), carboxytetramethylrhodamine (TAMRA), carboxy-X-rhodamine (ROX™), LIZ™, VIC™, NED™, PET™, SYBR, PicoGreen, RiboGreen, etc. Fluorescent labels can be green fluorescent protein (GFP), red fluorescent protein (RFP), yellow fluorescent protein, phycobiliproteins (e.g., allophycocyanin, phycocyanin, phycoerythrin, and phycoerythrocyanin).
[0139] The methods disclosed herein may include one or more linkers. The methods disclosed herein may include one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, twenty or more, thirty or more, forty or more, fifty or more, sixty or more, seventy or more, eighty or more, ninety or more, ten or more, twenty or more, thirty or more, forty or more, fifty or more, sixty or more, seventy or more, eighty or more, ninety or more, tenty or more, twenty ... or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more linkers. The one or more linkers may be different, similar, identical, or a combination thereof.
[0140] Suitable linkers include any chemical or biological compound capable of attaching to a label, primer, and / or capture probe disclosed herein. When a linker attaches to both a label and a primer or capture probe, a suitable linker may therefore be capable of sufficiently separating the label and the primer or capture probe. A suitable linker may not significantly interfere with the ability of the primer and / or capture probe to hybridize to a nucleic acid molecule, portion thereof, or variant or derivative thereof. A suitable linker may not significantly interfere with the ability of the label to be detected. The linker may be rigid. The linker may be flexible. The linker may be semi-rigid. The linker may be proteolytically stable (e.g., resistant to proteolytic cleavage). The linker may be proteolytically unstable (e.g., susceptible to proteolytic cleavage). The linker may be helical. The linker may be non-helical. The linker may be coiled. The linker may be β-stranded. The linker may include a turn conformation. The linker can be a single chain. The linker can be a long chain. The linker can be a short chain. The linker can contain at least about 5 residues, at least about 10 residues, at least about 15 residues, at least about 20 residues, at least about 25 residues, at least about 30 residues, or at least about 40 residues, or more.
[0141] Examples of linkers include, but are not limited to, hydrazones, disulfides, thioethers, and peptide linkers. The linker can be a peptide linker. The peptide linker can include a proline residue. The peptide linker can include arginine, phenylalanine, threonine, glutamine, glutamate, or any combination thereof. The linker can be a heterobifunctional crosslinker.
[0142] The methods disclosed herein may include performing one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, eleven or more, twelve or more, thirteen or more, fourteen or more, fifteen or more, twenty or more, twenty-five or more, thirty or more, thirty-five or more, forty or more, forty-five or more, or fifty or more assays on a sample containing one or more nucleic acid molecules. The two or more assays may be different, similar, identical, or a combination thereof. For example, the methods disclosed herein may include performing two or more sequencing reactions. In another example, the methods disclosed herein include performing two or more assays, where at least one of the two or more assays includes a sequencing reaction. In yet another example, the methods disclosed herein include performing two or more assays, where at least two of the two or more assays include a sequencing reaction and a hybridization reaction. The two or more assays may be performed sequentially, simultaneously, or a combination thereof. For example, two or more sequencing reactions may be performed simultaneously. In another example, the methods disclosed herein include performing a hybridization reaction followed by a sequencing reaction. In yet another example, the methods disclosed herein include performing two or more hybridization reactions simultaneously, followed by two or more sequencing reactions simultaneously. The two or more assays may be performed by one or more devices. For example, the two or more amplification reactions may be performed by a PCR machine.In another example, two or more sequencing reactions may be performed by two or more sequencers.
[0143] The methods disclosed herein may include one or more devices. The methods disclosed herein may include one or more assays including one or more devices. The methods disclosed herein may include the use of one or more devices to perform one or more operations or assays. The methods disclosed herein may include the use of one or more devices in one or more operations or assays. For example, performing a sequencing reaction may include one or more sequencers. In another example, generating a subset of nucleic acid molecules may include the use of one or more magnetic separators. In yet another example, one or more processors may be used in analyzing one or more nucleic acid samples. Examples of devices include, but are not limited to, sequencers, thermocyclers, real-time PCR machines, magnetic separators, transmission devices, hybridization chambers, electrophoresis instruments, centrifuges, microscopes, imaging devices, fluorometers, luminometers, plate readers, computers, processors, and bioanalyzers.
[0144] The methods disclosed herein may include one or more sequencers. The one or more sequencers may include one or more of HiSeq, MiSeq, HiScan, Genome Analyzer IIx, SOLiD Sequencer, Ion Torrent PGM, 454 GS Junior, PacBio RS, or a combination thereof. The one or more sequencers may include one or more sequencing platforms. The one or more sequencing platforms may include 454 GS FLX by Life Technologies / Roche, Genome Analyzer by Solexa / Illumina, SOLiD by Applied Biosystems, CGA Platform by Complete Genomics, PacBio RS by Pacific Biosciences, or a combination thereof.
[0145] The methods disclosed herein may include one or more thermocyclers. The one or more thermocyclers may be used to amplify one or more nucleic acid molecules. The methods disclosed herein may include one or more real-time PCR devices. The one or more real-time PCR devices may include a thermal cycler and a fluorometer. The one or more thermocyclers may be used to amplify and detect one or more nucleic acid molecules.
[0146] The methods disclosed herein may include one or more magnetic separators. The one or more magnetic separators may be used to separate paramagnetic and ferromagnetic particles from the suspension. The one or more magnetic separators may include one or more LifeStep™ biomagnetic separators, SPHERO™ FlexiMag separators, SPHERO™ MicroMag separators, SPHERO™ HandiMag separators, SPHERO™ MiniTube Mag separators, SPHERO™ UltraMag separators, DynaMag™ magnets, DynaMag™-2 magnets, or combinations thereof.
[0147] The methods disclosed herein may include one or more bioanalyzers. In some cases, the bioanalyzer is a chip-based capillary electrophoresis machine capable of analyzing RNA, DNA, and proteins. The one or more bioanalyzers may include Agilent's 2100 bioanalyzer.
[0148] The methods disclosed herein may include one or more processors. The one or more processors may analyze, compile, store, sort, combine, evaluate, or otherwise process one or more data and / or results from one or more assays, one or more data and / or results based on or derived from one or more assays, one or more outputs from one or more assays, one or more outputs based on or derived from one or more assays, one or more outputs from one or more data and / or results, one or more outputs based on or derived from one or more data and / or results, or combinations thereof. In some cases, the methods disclosed herein may include combining data for analysis, as shown in FIG. 17. The one or more processors may transmit one or more data, results, or outputs from one or more assays, one or more data, results, or outputs based on or derived from one or more assays, one or more outputs from one or more data or results, one or more outputs based on or derived from one or more data or results, or combinations thereof. The one or more processors may receive and / or store requests from a user. The one or more processors may generate or generate one or more data, results, outputs. The one or more processors may generate or generate one or more biomedical reports. The one or more processors may transmit one or more biomedical reports. The one or more processors may analyze, compile, store, sort, combine, evaluate, or otherwise process information from one or more databases, one or more data or results, one or more outputs, or combinations thereof.One or more processors may analyze, compile, store, sort, combine, evaluate, or otherwise process information from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, or more databases. One or more processors may transmit one or more requests, data, results, output, and / or information to one or more users, processors, computers, computer systems, memory locations, devices, databases, or combinations thereof. One or more processors may receive one or more requests, data, results, output, and / or information from one or more users, processors, computers, computer systems, memory locations, devices, databases, or combinations thereof. One or more processors may retrieve one or more requests, data, results, output, and / or information from one or more users, processors, computers, computer systems, memory locations, devices, databases, or combinations thereof. The present disclosure also provides methods that can be utilized for multiple biomedical applications. In some cases, variants, genes, rearranged genes, exons, UTRs, regulatory regions, splice sites, alternative sequences, and other content of interest for human or non-human genomes can be combined from several databases to generate an aggregated set of content applicable to multiple biomedical reports. This content can then be categorized based on local or global genomic context, nucleotide content, sequencing performance, and interpretation needs, and then subsequently grouped into subsets for specific protocols, assay progression, etc. In some cases, variants, genes, exons, UTRs, regulatory regions, splice sites, alternative sequences, and other content of interest can be combined from several databases to generate an aggregated set of content applicable to multiple biomedical reports.This content is then classified based on local or global genomic context, nucleotide content, sequencing performance, and interpretation needs, and then grouped into subsets for subsequent progression through specific protocols and assays (Figure 14). In some cases, protocols and / or assays may include supplemental extraction. The supplemental extraction may include human target sequences, non-human target sequences, and combinations thereof. The supplemental extraction may include nucleic acid molecules derived from the subject (e.g., originating from cells derived from the subject's tissue) and nucleic acid molecules not derived from the subject (e.g., microbial (commensal or parasitic), pathogen, or transplanted). Figures 15-17 provide example workflows for assays that include supplemental extraction for one or more of multiple subsets of DNA enriched for different genomic regions.
[0149] The methods disclosed herein may include one or more memory locations. The one or more memory locations may store information, data, results, output, requests, or combinations thereof. The one or more memory locations may receive information, data, results, output, requests, or combinations thereof from one or more users, processors, computers, computer systems, devices, or combinations thereof.
[0150] The methods described herein can be performed with the aid of one or more computers and / or computer systems. The computer or computer system may include an electronic storage location (e.g., a database, memory) having machine-executable code for performing the methods provided in this disclosure, and one or more processors for executing the machine-executable code.
[0151] The methods disclosed herein may include treating and / or preventing a disease or condition in a subject based on one or more biomedical outputs. The one or more biomedical outputs may recommend one or more therapies. The one or more biomedical outputs may suggest, select, prescribe, recommend, or otherwise determine a course of treatment and / or prevention of a disease or condition. The one or more biomedical outputs may recommend modifying or continuing one or more therapies. Modifying one or more therapies may include administering, initiating, reducing, increasing, and / or terminating one or more therapies. The one or more therapies may include an anti-cancer therapy, an anti-viral therapy, an anti-bacterial therapy, an anti-fungal therapy, an immunosuppressive therapy, or a combination thereof. The one or more therapies may treat, alleviate, or prevent one or more diseases or symptoms.
[0152] Examples of anti-cancer treatments include, but are not limited to, surgery, chemotherapy, radiation therapy, immunotherapy / biological therapy, and photodynamic therapy. Anti-cancer treatments may include chemotherapeutic agents, monoclonal antibodies (e.g., rituximab, trastuzumab), cancer vaccines (e.g., therapeutic vaccines, prophylactic vaccines), gene therapy, or a combination thereof.
[0153] The one or more treatments may include an antibacterial agent. Generally, an antibacterial agent refers to a substance that kills or inhibits the growth of microorganisms, such as bacteria, fungi, viruses, or protozoa. Antibacterial agents either kill microorganisms (bactericidal) or prevent their growth (bacteriostatic). There are two main classes of antibacterial agents: those derived from natural sources (e.g., antibiotics, protein synthesis inhibitors (aminoglycosides, macrolides, tetracyclines, chloramphenicol, polypeptides, etc.)), and synthetic agents (e.g., sulfonamides, cotrimoxazole, quinolones). In some cases, the antibacterial agent is an antibiotic, antiviral, antifungal, antimalarial, antituberculous, antileprosy, or antiprotozoal.
[0154] Antibiotics are generally used to treat bacterial infections. Antibiotics can be divided into two categories: bactericidal antibiotics and bacteriostatic antibiotics. Generally, bactericidal antibiotics can directly kill bacteria, while bacteriostatic antibiotics can prevent them from dividing. Antibiotics may be derived from living organisms or may contain synthetic antibacterial agents such as sulfonamides. Antibiotics may include aminoglycosides such as amikacin, gentamicin, kanamycin, neomycin, netilmicin, tobramycin, and paromomycin. Alternatively, the antibiotic can be an ansamycin (e.g., geldanamycin, herbimycin), a cabacephem (e.g., loracarbef), a carbapenem (e.g., ertapenem, doripenem, imipenem, cilastatin, meropenem), a glycopeptide (e.g., teicoplanin, vancomycin, telavancin), a lincosamide (e.g., clindamycin, lincomycin, daptomycin), a macrolide (e.g., azithromycin, clarithromycin, dirithromycin, erythromycin, roxithromycin, troleandomycin, telithromycin, spectinomycin, spiramycin), a nitrofuran (e.g., furazolidone, nitrofurantoin), and a polypeptide (e.g., bacitracin, colistin, polymyxin B).
[0155] In some instances, antibiotic treatments include cephalosporins such as cefadroxil, cefazolin, cephalothin, cephalexin, cefaclor, cefamandole, cefoxitin, cefprozil, cefuroxime, cefixime, cefdinir, cefditoren, cefoperazone, cefotaxime, cefpodoxime, ceftazidime, ceftibuten, ceftizoxime, ceftriaxone, cefepime, ceftaroline fosamil, and ceftobiprole.
[0156] Antibiotic treatments may also include penicillins. Examples of penicillins include amoxicillin, ampicillin, azlocillin, carbenicillin, cloxacillin, dicloxacillin, flucloxacillin, mezlocillin, methicillin, nafcillin, oxacillin, penicillin G, penicillin V, piperacillin, temocillin, and ticarcillin.
[0157] Alternatively, quinolines may be used to treat bacterial infections. Examples of quinilones include ciprofloxacin, enoxacin, gatifloxacin, levofloxacin, lomefloxacin, moxifloxacin, nalidixic acid, norfloxacin, ofloxacin, trovafloxacin, grepafloxacin, sparfloxacin, and temafloxacin.
[0158] In some cases, antibiotic therapy includes a combination of two or more therapies. For example, amoxicillin and clavulanate, ampicillin and sulbactam, piperacillin and tazobactam, or ticarcillin and clavulanate may be used to treat a bacterial infection.
[0159] Sulfonamides may be used to treat bacterial infections. Examples of sulfonamides include, but are not limited to, mafenide, sulfonamide chrysoidine, sulfacetamide, sulfadiazine, silver sulfadiazine, sulfamethizole, sulfamethoxazole, sulfanilimide, sulfasalazine, sulfisoxazole, trimethoprim, and trimethoprim-sulfamethoxazole (cotrimoxazole) (tmp-smx).
[0160] Tetracyclines are another example of antibiotics. Tetracyclines can bind to the 30S ribosomal subunit in the mRNA translation complex, thereby inhibiting the binding of aminoacyl-tRNA to the mRNA ribosomal complex. Tetracyclines include demeclocycline, doxycycline, minocycline, oxytetracycline, and tetracycline. Additional components that can be used to treat bacterial infections include arsphenamine, chloramphenicol, fosfomycin, fusidic acid, linezolid, metronidazole, mupirocin, platensimycin, quinupristin / dalfopristin, rifaximin, thiamphenicol, tigecycline, tinidazole, clofazimine, dapsone, capreomycin, cycloserine, ethambutol, ethionamide, isoniazid, pyrazinamide, rifampicin, rifamycin, rifabutin, rifapentine, and streptomycin.
[0161] Antiviral therapies are a class of medications specifically used to treat viral infections. Like antibiotics, specific antiviral agents are used for specific viruses. They are relatively harmless to the host and can therefore be used to treat infections. Antiviral therapies can inhibit various stages of the viral life cycle. For example, antiviral therapies may inhibit viral attachment to cellular receptors. Such antiviral therapies may include agents that mimic viral-associated proteins (VAPs) and bind to cellular receptors. Other antiviral therapies may inhibit viral entry, viral uncoating (e.g., amantadine, rimantadine, pleconaril), viral synthesis, viral integration, viral transcription, or viral translation (e.g., fomivirsen). In some cases, antiviral therapies are morpholino antisense. Antiviral therapies must be distinguished from virucidal agents, which actively inactivate viral particles outside the body.
[0162] Many available antiviral drugs are designed to treat infections caused by retroviruses, most commonly HIV. Antiretroviral drugs can include the classes of protease inhibitors, reverse transcriptase inhibitors, and integrase inhibitors. Drugs for treating HIV include protease inhibitors (e.g., Invirase, Saquinavir, Kaletra, Lopinavir, Lexiva, Fosamprenavir, Norvir, Ritonavir, Prizista, Darunavir, Reyataz, Viracept), integrase inhibitors (e.g., Raltegra, vir), transcriptase inhibitors (e.g., abacavir, Ziagen, Agenerase, amprenavir, Aptivus, tipranavir, Crixivan, indinavir, Fortvase, saquinavir, Intelence™, etravirine, Isentress, Viread), reverse transcriptase inhibitors (e.g., delavirdine, efavirenz, Epivir, Hivid, nevirapine, retrovir, AZT, stuvadine, Truvada, Vyde Alternatively, antiretroviral therapy may include atripla (e.g., efavirenz, emtricitabine, and tenofovira disoproxil fumarate) and completer (emtricitabine, rilpivir), fusion inhibitors (e.g., fuzeon, enfuvirtide), chemokine coreceptor antagonists (e.g., clsentri, emtriva, emtricitabine, epzicom, or trizivir). Therapeutic options include combination therapy with antivirals such as aviranthesin (aviran), aviran, and tenofovir disoproxil fumarate (tenofovir disoproxil fumarate). Herpes viruses, which can cause herpes simplex and genital herpes, are typically treated with the nucleoside analog acyclovir. Viral hepatitis (A-E) is caused by five distinct hepatotropic viruses and is generally treated with antiviral medications, depending on the type of infection. Influenza A and B viruses are important targets for the development of new influenza treatments that overcome resistance to existing neuraminidase inhibitors, such as oseltamivir.
[0163] In some cases, the antiviral treatment may include a reverse transcriptase inhibitor. The reverse transcriptase inhibitor may be a nucleoside reverse transcriptase inhibitor or a non-nucleoside reverse transcriptase inhibitor. Nucleoside reverse transcriptase inhibitors may include, but are not limited to, Combivir, Emtriva, Epivir, Epzicom, Hivid, Retrovir, Trizivir, Truvada, Videx ec, Videx, Viread, Zerrit, and Ziagen. Non-nucleoside reverse transcriptase inhibitors may include Edurant, Intelence, Rescriptor, Sustiva, and Viramune (immediate release or sustained release).
[0164] Protease inhibitors are another example of antiviral drugs and may include, but are not limited to, Agenerase, Aptivus, Crixivan, Fortovase, Invirase, Kaletra, Lexiva, Norvir, Prezista, Reyataz, and Viracept. Alternatively, antiviral treatments may include fusion inhibitors (e.g., enfuviride), or entry inhibitors (e.g., maraviroc).
[0165] Additional examples of antiviral drugs include abacavir, acyclovir, adefovir, amantadine, amprenavir, Ampligen, arbidol, atazanavir, atripla, boceprevir, cidofovir, combivir, darunavir, delavirdine, didanosine, docosanol, edoxudine, efavirenz, emtricitabine, enfuvirtide, entecavir, famciclovir, fomivirsen, fosamprenavir, foscarnet, phosphonet, fusion inhibitors, ganciclovir, ibacitabine, immunovir, idoxuridine, imiquimod, indinavir, inosine, integrase inhibitors, interferons (e.g., type I, type II, type III interferons), lamivudine, and loxacin. pinavir, loviride, maraviroc, moroxydine, methisazone, nelfinavir, nevirapine, nexavir, nucleoside analogs, oseltamivir, peg-interferon alfa-2a, penciclovir, peramivir, pleconaril, podophyllotoxin, protease inhibitors, raltegravir, reverse transcriptase inhibitors, ribavirin, rimantadine, ritonavir, pyramidine, saquinavir, stavudine, tea tree oil, tenofovir, tenofovir disoproxil, tipranavir, trifluridine, trizivir, tromantadine, truvada, valacyclovir, valganciclovir, vicriviroc, vidarabine, viramidine, zalcitabine, zanamivir, and zidovudine.
[0166] Antifungal drugs are medicines that can be used to treat fungal infections such as athlete's foot, ringworm, candidiasis (thrush), and serious systemic infections such as cryptococcal meningitis. Antifungal agents work by exploiting the differences between mammalian and fungal cells to kill fungal organisms. Unlike bacteria, fungi and humans are both eukaryotic organisms. Therefore, fungal cells and human cells are similar at the molecular level, which makes it more difficult to find targets for antifungal drugs to attack that are not even present in the infected organism.
[0167] Antiparasitic agents are a class of medications indicated for the treatment of infections by parasitic organisms such as nematodes, cestodes, trematodes, infectious protozoans, and amoebas. Like antifungals, they must kill the infecting pest without serious damage to the host.
[0168] The methods of the present disclosure can be performed by a system, a kit, a library, or a combination thereof. The methods of the present disclosure may include one or more systems. The systems of the present disclosure can be performed by a kit, a library, or both. The systems may include one or more components for performing any of the methods or any of the operations of the methods disclosed herein. For example, the systems may include one or more kits, devices, libraries, or a combination thereof. The systems may include one or more sequencers, processors, memory locations, computers, computer systems, or a combination thereof. The systems may include a transmission device.
[0169] The kit may include various reagents for performing various operations disclosed herein, including sample processing and / or analytical operations. The kit may include instructions for performing at least some of the operations disclosed herein. The kit may include one or more capture probes, one or more beads, one or more labels, one or more linkers, one or more devices, one or more reagents, one or more buffers, one or more samples, one or more databases, or a combination thereof.
[0170] The library may include one or more capture probes. The library may include one or more subsets of nucleic acid molecules. The library may include one or more databases. The library may be generated or produced from any of the methods, kits, or systems disclosed herein. A database library may be generated from one or more databases. A method for generating one or more libraries may include (a) aggregating information from one or more databases to generate an aggregated dataset; (b) analyzing the aggregated dataset; and (c) generating one or more database libraries from the aggregated dataset. Figure 13 provides an example of a library construction workflow. In some cases, libraries may be pooled (Figure 7). Computer Systems
[0171] The present disclosure provides a computer system programmed to perform the methods of the present disclosure. Figure 5 shows a computer system 501 programmed or otherwise configured to map and / or align sequence read data to identify the origin of a nucleic acid molecule (e.g., human or non-human, host or non-host), to identify one or more characteristics (e.g., genetic variants), or any combination thereof. The computer system 501 can control various aspects of processing the sequencing information provided in the present disclosure, such as aligning the sequence read data with one or more reference sequences to identify the origin of a nucleic acid sequence in a biological sample. The computer system 501 can be a user's electronic device or a computer system remotely located relative to the electronic device. The electronic device can be a mobile electronic device.
[0172] The computer system 501 includes a central processing unit (CPU, also referred to herein as "processor" and "computer processor") 505, which may be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system 501 also includes memory or memory locations 510 (e.g., random access memory, read-only memory, flash memory), electronic storage 515 (e.g., hard disk), a communication interface 520 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 525, such as cache, other memory, data storage, and / or electronic display adapters. The memory 510, storage 515, interface 520, and peripheral devices 525 communicate with the CPU 505 through a communication bus (solid lines), such as a motherboard. The storage 515 may be a data storage device (or data repository) for storing data. The computer system 501 may be operatively coupled to a computer network ("network") 530 with the aid of the communication interface 520. Network 530 may be the Internet, an Internet and / or extranet, or an intranet and / or extranet communicating over the Internet. Network 530, in some cases, is a telecommunications and / or data network. Network 530 may include one or more computer servers, which may enable distributed computing such as cloud computing. In some cases, network 530 with the aid of computer system 501 may implement a peer-to-peer network, which may enable devices coupled to computer system 501 to act as clients or servers.
[0173] The CPU 505 can execute a series of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 510. The instructions may be directed to the CPU 505, which may then program or otherwise configure the CPU 505 to perform the methods of the present disclosure. Examples of operations performed by the CPU 505 may include fetch, decode, execute, and writeback.
[0174] The CPU 505 may be part of a circuit, such as an integrated circuit. One or more other components of the system 501 may be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).
[0175] Storage device 515 may store files such as drivers, libraries, and saved programs. Storage device 515 may store user data, such as user preferences and user programs. Computer system 501 may, in some cases, include one or more additional data storage devices external to computer system 501, such as located on a remote server that communicates with computer system 501 through an intranet or the Internet.
[0176] Computer system 501 can communicate with one or more remote computer systems through network 530. For example, computer system 501 can communicate with a remote computer system of a user (e.g., a healthcare provider). Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., an Apple® iPad®, a Samsung® Galaxy Tab), a phone, a smartphone (e.g., an Apple® iPhone®, an Android-enabled device, a Blackberry®), or a personal digital assistant. A user can access computer system 501 through network 1130.
[0177] The methods described herein can be performed by machine (e.g., a computer processor) executable code stored in an electronic storage location of the computer system 501, such as, for example, memory 510 or electronic storage 515. The machine-executable or machine-readable code can be provided in the form of software. During use, the code can be executed by the processor 505. In some cases, the code can be read from storage 515 and stored in memory 510 ready for access by the processor 505. In some situations, the electronic storage 515 can be omitted, and the machine-executable instructions are stored in memory 510.
[0178] The code may be pre-compiled and configured for use by a machine having a processor adapted to execute the code, or may be compiled during run-time. The code may be provided in a programming language that may be selected to allow the code to be executed in a pre-compiled or compiled manner.
[0179] Aspects of the systems and methods provided herein, such as computer system 501, can be embodied in programming. Various aspects of the technology may be considered "products" or "articles of manufacture," typically in the form of machine (or processor) executable code and / or associated data performed or embodied in some type of machine-readable medium. The machine-executable code can be stored in electronic storage, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. "Storage" media can include any or all of the tangible memory of a computer, processor, or the like, or their associated modules, such as various semiconductor memories, tape drives, disk drives, etc., which may provide non-transitory storage for software programming at any time. All or portions of the software may sometimes be communicated over the Internet or various other telecommunications networks. Such communication may, for example, enable loading of the software from one computer or processor to another, e.g., from a management server or host computer to the computer platform of an application server. Thus, another type of medium that may bear software elements includes optical, electrical, and electromagnetic waves used across physical interfaces, e.g., between local devices, over wired and optical terrestrial communication networks, as well as various air links. The physical elements that carry such waves, such as wired or wireless links, optical links, etc., may also be considered as media bearing software. As used herein, unless limited to non-transitory tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.
[0180] As such, machine-readable media such as computer-executable code may take many forms, including, but not limited to, tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer(s), which may be used to execute, for example, databases shown in the figures. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves, such as those generated during radio wave (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punch cards, perforated paper tape, any other physical storage medium with a pattern of holes, RAM, ROM, PROMs and EPROMs, FLASH-EPROMs, any other memory chips or cartridges, carrier waves carrying data or instructions, cables or links carrying such carrier waves, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer-readable media may be involved in carrying a sequence or multiple sequences of one or more instructions to a processor for execution.
[0181] The computer system 501 may include or be in communication with an electronic display 535 that includes a user interface (UI) 540 for providing one or more biomedical reports including one or more sets of data selected from the group consisting of, for example, (i) candidate tumor neo-antigens, (ii) detected non-human species, (iii) detected CDR3 sequences, and any combination thereof. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0182] The methods and systems of the present disclosure can be performed by one or more algorithms. The algorithms can be implemented by software when executed by the central processing unit 505. The algorithms can, for example, map and / or align sequence read data, call variants, annotate sequence information, or any combination thereof. [Example]
[0183] Example 1 Preparation of genomic DNA The following procedure was used to prepare a subset of nucleic acid molecules from a sample containing genomic DNA.
[0184] 1. Shear the sample containing genomic DNA for 15-35 seconds at M220.
[0185] 2. The fragmented gDNA was purified using SPRI beads after ligation (the volume ratio of SPRI beads to DNA sample was 1), and the DNA was eluted in 100 μL of elution buffer (EB).
[0186] 3. 50 μL of SPRI beads were added to 100 μL of DNA.
[0187] 4. The supernatant was transferred to a new tube.
[0188] 5. The DNA was eluted from the remaining bead-bound DNA. This eluted DNA was called the long insert.
[0189] 6. 10 μL of SPRI beads were added to the supernatant from step 4.
[0190] 7. The supernatant from step 6 was transferred to a new tube.
[0191] 8. DNA was eluted from the remaining bead-bound DNA from step 6. This eluted DNA was called the middle insert.
[0192] 9. 20 μL of SPRI beads were added to the supernatant from step 7.
[0193] 10. The supernatant from step 9 was transferred to a new tube.
[0194] 11. DNA was eluted from the remaining bead-bound DNA from step 9. This eluted DNA was called the short insert.
[0195] Example 2 Obtaining biological samples Subjects with evaluable metastatic cancer undergo tumor resection. Tumor-derived lymphocytes, called tumor-infiltrating lymphocytes (TILs), are grown and expanded. Multiple individual fragments or multiple individual cultures of TILs are grown. The individual cultures are expanded separately to obtain a sufficient yield of TILs (approximately 10 8 When 10 ...
[0196] Example 3 Extraction of genomic material from biological samples Genomic DNA (gDNA) and total RNA were purified from various tumors and matched normal apheresis samples using the QIAGEN AllPrep DNA / RNA kit (catalog no. 80204) according to the manufacturer's recommendations. Tumor samples were formalin-fixed, paraffin-embedded (FFPE), and gDNA was extracted using the Covaris truXTRAC™ FFPE DNA kit according to the manufacturer's instructions.
[0197] Example 4 Sequencing analysis of biological samples Construction of a whole-exome library of approximately 20,000 coding genes and exon capture are performed using the Agilent Technologies SureSelectXT Target Enrichment System (catalog number 5190-8646) for paired-end libraries coupled with Human All Exson V6 RNA bait (catalog number 5190-8863) (Agilent Technologies, Santa Clara, CA, USA) and bacterial RNA bait. Whole-exome sequencing (WES) libraries are then sequenced on a NextSeq 500 benchtop sequencer (Illumina, San Diego, CA, USA). Libraries are prepared using 3 μg of gDNA from fresh tumor tissue samples and 200 ng of gDNA from FFPE tumor samples according to the manufacturer's protocol. Paired-end sequencing is performed using the Illumina High-Throughput Flow Cell Kit (300 cycles) (catalog number FC-404-2004). Tu-1, Tu-2A, and Tu-2B samples were initially run with v1 of the reagent / flow cell kit, and subsequent runs of the same library preparation were performed using v2 of the reagent / flow cell kit. Tumor samples were run with the v2 reagent / flow cell kit. The average sequencing depth and tumor percentage (tumor purity) in each sample were determined, as estimated using the bioinformatics program Allele-Specific Copy Number Analysis of Tumors (ASCAT)1. RNA-seq libraries were prepared using 2 μg of total RNA with the Illumina TruSeq RNA Standard Library Preparation Kit according to the manufacturer's protocol. RNA-seq libraries were paired-end sequenced on a NextSeq 500 benchtop sequencer (Illumina, San Diego, CA, USA). Alignment, processing, and variant calling were performed. For WES, alignments were performed against human genome build hg19 using the novocraft (http: / / www.novocraft.com / ) library. This was done using voalign MPI. Duplicates were marked using Picard's MarkDuplicates tool. In / del resequencing and base recalibration were performed according to the GATK best practice workflow (https: / / www.broadinstitute.org / gatk / ). After data cleanup, pileup files were generated using samtools (http: / / samtools.sourceforge.net) and analyzed using Varscan2 (http: / / varscan.sourceforge.net) according to the following criteria: tumor and normal read counts. Somatic variants are called using a β-mutation index of 10 or higher, a variant allele frequency of 10% or higher, and tumor variant read data of 4 or higher. These variants are then annotated using Annovar (http: / / annovar.openbioinformatics.org). For RNA-seq, alignments are annotated using the Human Genome Building Blocks (HGBB). For hg19, we used the STAR (https: / / github.com / alexdobin / STAR) two-pass method. Duplicates are marked using Picard's MarkDuplicates tool. Reads are split and trimmed using the GATK SplitNTrim tool. In / del resequencing and base recalibration are then performed using the GATK toolbox. A pileup file is created using the final recalibrated bam file and samtools mpileup. Finally, variants are called using Varscan2.
[0198] Example 5 Alignment of non-human genome sequences The genomic information extracted from the sequencing analysis is aligned against a genomic reference of the 16S ribosomal RNA gene. Reads that perfectly match the 16S gene are identified. Capture probes designed to hybridize to the shared region of the 16S gene can be used to capture nucleic acid molecules from a wide variety of species, even those that have yet to be identified and characterized. Captured molecules whose sequences extend from these shared regions into the variable region are assigned to their species of origin based on sequences derived from the variable region.
[0199] Example 6 Shear time and fragment size Genomic DNA (gDNA) was sheared using different shearing time settings on the Covaris. The gDNA fragments generated by the various shearing times were then analyzed. The results are shown in Figure 10 and Table 1. [Table 1]
[0200] Example 7 Bead ratio and fragment size The ratio of bead volume to nucleic acid sample volume was changed, and the effect of these ratios on the average fragment size was analyzed. As can be seen in Figure 11, changing the ratio of bead volume to nucleic acid sample volume to 0.8 (line 1), 0.7 (line 2), 0.6 (line 3), 0.5 (line 4) and 0.4 (line 5) resulted in changes in the average size of DNA fragments. In general, the lower the ratio, the larger the average fragment size appeared to be.
[0201] Example 8 Ligation reaction and fragment size Two different shearing times and three different ligation reaction combinations were performed on nucleic acid samples. Sample 1 was sheared for 25 seconds and a ligation reaction was performed on long-insert DNA. Sample 2 was sheared for 32 seconds and a ligation reaction was performed on long-insert DNA. Sample 3 was sheared for 25 seconds and a ligation reaction was performed on medium-insert DNA. Sample 4 was sheared for 32 seconds and a ligation reaction was performed on medium-insert DNA. Sample 5 was sheared for 25 seconds and a ligation reaction was performed on short-insert DNA. Sample 6 was sheared for 32 seconds and a ligation reaction was performed on short-insert DNA. Figure 12 shows the average fragment sizes for the six reactions. Example 9 Capture of target nucleic acid A nucleic acid sample is obtained from a subject. The sample can be divided into two separate samples, each of which can be subjected to a different pool of probes and conditions. A first pool of biotinylated capture probes is generated by combining the Agilent Clinical Research Exome Kit (based on the exome in the GRCh37 reference genome) with additional probes of interest, including exome regions corresponding to the GRCh38 reference sequence, HLA-specific probes, T cell receptor and B cell receptor recombination-specific probes (i.e., regions corresponding to the V(D)J region), regions of microsatellite instability, and oncovirus sequences. The additional probes are titrated into the pool to adjust the relative capture rate of nucleic acids. This can be beneficial for downstream sequencing reactions that increase the sequencing depth or sensitivity of the set of captured nucleic acids. A second pool of probes (e.g., a "boost set" of probes targeting specific sequences, or a subset of sequences) can capture nucleic acids due to sequences that are difficult to capture. The boost set of probes can include exomes of GRCh37 and GRCh38 reference sequences with high GC content, probes related to cancer therapy, and additional T cell receptor and B cell recombination-specific probes. The nucleic acid molecules captured by the first pool and the nucleic acid molecules captured by the second pool are combined to create a combined pool of capture probes and captured nucleic acids. A combined pool can be created by combining two pools of captured molecules at different ratios to adjust the relative amounts of captured nucleic acids corresponding to each pool. This can be beneficial for downstream sequencing reactions to increase the depth or sensitivity of the set of capture nucleic acids. The hybridized capture probes and capture nucleic acids are incubated with magnetic streptavidin beads. A magnetic separator is placed outside the tube containing the sample, allowing the capture probes to immobilize the magnetic streptavidin beads. The liquid is decanted, and additional buffer is added to wash the beads and remove any unbound nucleic acids. The captured nucleic acids are subjected to an amplification reaction to add adapters for further downstream sequencing reactions.
[0202] Example 10 Microbiome and oncovirus identification A nucleic acid sample is obtained from a subject. The sample can be divided into two separate samples, and each separate sample can be subjected to different pools of probes and conditions. A first pool of biotinylated capture probes is generated by combining the Agilent Clinical Research Exome Kit (based on the exome in the GRCh37 reference genome) with additional probes of interest, including exome regions corresponding to the GRCh38 reference sequence, probes with homology to the 16S rRNA region of bacteria, fungi, archaea, and protist genes, including Helicobacter pylori and Fusobacterium genes, probes for viral genes (including human papillomavirus, hepatitis B virus, and hepatitis C virus), and probes for pathogenicity genes. The additional probes are titrated into the pool to adjust the relative capture rate of nucleic acids. A second pool of probes (e.g., a "boost set" of probes targeting specific sequences, or a subset of sequences) can capture nucleic acids due to sequences that are difficult to capture. The probes in the boost set can include exomes of the GRCh37 and GRCh38 reference sequences, which have high GC content or low homology to genes due to multiple alternative gene sequences or genes with high mutation rates. The nucleic acid molecules captured by the first pool and the nucleic acid molecules captured by the second pool are combined to create a combined pool of nucleic acids. These two pools are combined and titrated at different ratios to increase the sensitivity of one pool relative to the other. The hybridized capture probes and capture nucleic acids are incubated with magnetic streptavidin beads. A magnetic separator is placed outside the tube containing the sample and capture probes, allowing the magnetic streptavidin beads to be immobilized. The liquid is decanted, and additional buffer is added to wash the beads and remove any unbound nucleic acids. The captured nucleic acids are subjected to an amplification reaction, followed by the addition of adapters for further downstream sequencing reactions. The sequences are aligned with the reference sequences to identify the presence of specific microorganisms in the target microbiome or to identify any disease-causing pathogens.
[0203] Example 11 Identification of CAR-T cells A nucleic acid sample from a subject is obtained. Two pools of probes are generated, with the first pool consisting of biotinylated probes from the Agilent Clinical Research Exome Kit. A supplementary set of biotinylated probes is added to the first pool, which contains sequences specific to chimeric antigen receptors found in CAR-T cells. The second pool of biotinylated capture probes (e.g., a "boost set" of probes targeting specific sequences, or a subset of sequences) contains sequences directed toward CAR sequences with high GC content and / or sequences that may have lower overall sequence homology with the captured nucleic acids due to mutation or recombination. The sample is divided into two sample pools, with the first sample pool being applied to the first pool of probes and the second sample pool being applied to the second pool of probes, allowing the capture probes to capture nucleic acids from each of the sample pools. The first and second pools of probes hybridized with the captured nucleic acids are combined. Magnetic streptavidin beads are added to the mixture, and the biotinylated probe is bound to the beads. A magnetic separator is placed outside the tube containing the sample and the capture probe, allowing the magnetic streptavidin beads to be immobilized. The liquid is decanted, and additional buffer is added to wash the beads and remove any unbound nucleic acid. The captured nucleic acid is resuspended in fresh buffer and subjected to an amplification reaction to add adapters, and then the captured nucleic acid is sequenced. The sequence reading data is analyzed and aligned with the CAR-T gene reference sequence to identify the presence of CAR-T-related nucleic acid.
[0204] Example 12 Identification of segmented transcriptomes Samples are obtained from subjects from different tissue types. The samples are treated by enzymatic digestion to remove DNA and several types of RNA molecules (rRNA, tRNA, miRNA). The remaining mRNA is subjected to reverse transcription to generate cDNA molecules. A set of biotinylated probes with sequence homology to approximately 20,000 genes is generated and mixed with the cDNA sample. Magnetic streptavidin beads are added to the mixture, and the biotinylated probes are bound to the beads. A magnetic separator is placed outside the tube containing the sample and capture probes, allowing the magnetic streptavidin beads to be immobilized. The liquid is decanted, and additional buffer is added to wash the beads and remove any unbound nucleic acids. The captured nucleic acids are resuspended in fresh buffer and subjected to an amplification reaction to add adapters, after which the captured nucleic acids are sequenced. The sequence read data are analyzed and aligned to a reference sequence to determine the identity of the genes in each sample and associate them with the tissue type of the sample. Multiple replicates of the sample are run, with biotinylated probes titrated at different ratios. By analyzing the specific signal of nucleic acids related to the amount of probe provided in the sample, the expression level of each gene can be determined. Transcriptome analysis can be performed in parallel with exome or genome analysis. Exome or genome probes can be used as the first pool of probes, and cDNA capture probes can be used as the second pool. In this case, the sample can be divided into portions and subjected to DNA-specific or mRNA-specific reactions. As disclosed in the previous example, the captured nucleic acids can be pooled together and subjected to sequencing reactions to determine the identity of the nucleic acids.
[0205] Example 13 Identifying and capturing cancer-related cell-free DNA (cfDNA) A nucleic acid sample extract is obtained from whole blood or serum. Cells from the blood sample are removed by centrifugation to obtain a cell-free sample. A biotinylated probe set with sequence homology to approximately 20,000 genes is generated and mixed with the cell-free sample. A first pool of biotinylated capture probes is generated by combining the Agilent Clinical Research Exome Kit (based on the exome in the GRCh37 reference genome) with additional probes of interest containing exome regions corresponding to the GRCh38 reference sequence and probes with homology to genes associated with cancer. The additional probes are titrated into the pool to adjust the relative capture rates of nucleic acids. A second pool of probes (e.g., a "boost set" of probes targeting specific sequences, or a subset of sequences) can capture nucleic acids with sequences that are difficult to capture. The boost set probes can include exomes of the GRCh37 and GRCh38 reference sequences with high GC content or low homology to genes due to multiple alternative gene sequences or genes with high mutation rates. Magnetic streptavidin beads are added to the mixture, and the biotinylated probes are bound to the beads. A magnetic separator is placed outside the tube containing the sample and the capture probe, allowing the magnetic streptavidin beads to be immobilized. The liquid is decanted, and additional buffer is added to wash the beads and remove any unbound nucleic acids. The captured nucleic acids are resuspended in fresh buffer and subjected to an amplification reaction to add adapters, after which the captured nucleic acids are sequenced. The sequence read data is analyzed and aligned with a reference sequence to determine the identity of the gene in each sample. cfDNA analysis can be performed in parallel with exome sequencing of genomic DNA by applying the same or similar probe set to genomic DNA as is done with cfDNA.
[0206] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the present invention be limited by the specific examples provided herein. While the present invention has been described with reference to the foregoing specification, the description and illustration of the embodiments herein are not meant to be construed in a limiting sense. Numerous modifications, changes, and substitutions will occur to those skilled in the art without departing from the invention. Furthermore, it should be understood that all aspects of the present invention are not limited to the specific expressions, arrangements, or relative proportions set forth herein, which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the present invention described herein may be employed in practicing the present invention. It is therefore intended that the present invention encompass any such alternatives, modifications, variations, or equivalents. The following claims define the scope of the invention, and it is intended that methods and structures within the scope of these claims, and their equivalents, be covered thereby. The present invention provides, for example, the following items. (Item 1) 1. A method for processing a biological sample from a subject, comprising: (a) generating a subset of nucleic acid molecules from the biological sample using a pool of nucleic acid probes, the probes comprising: (i) a first plurality of nucleic acid probes configured to target elements of the human genome, and (ii) a second plurality of nucleic acid probes configured to target elements of one or more non-human genomes; and (b) assaying a subset of said nucleic acid molecules to obtain sequence information comprising the sequences of (i) human nucleic acids from said biological sample from said subject, and (ii) non-human nucleic acids from said biological sample from said subject. A method comprising: (Item 2) 2. The method of claim 1, wherein the first plurality of nucleic acid probes of (i) are configured to target elements derived from the human genome. (Item 3) Item 12. The method of item 1, wherein the second plurality of nucleic acid probes are configured to target elements from a non-human genomic sequence of one or more species selected from the group consisting of viruses, bacteria, bacterial phages, fungi, protists, archaea, amoebas, helminths, algae, genetically modified cells, and genetically modified vectors. (Item 4) 2. The method of claim 1, wherein generating the subset of nucleic acid molecules from the biological sample comprises performing one or more hybridization reactions. (Item 5) 2. The method of claim 1, further comprising obtaining the biological sample from the subject, wherein the biological sample is derived from a tumor biopsy, whole blood, or plasma. (Item 6) 2. The method of claim 1, further comprising aligning the sequences of the subset of nucleic acid molecules to one or more reference sequences. (Item 7) 7. The method of claim 6, wherein the one or more reference sequences comprise a plurality of reference sequences, the plurality of reference sequences corresponding to two or more different species. (Item 8) 7. The method of claim 6, further comprising identifying the origin of the nucleic acid molecules in the subset based on the alignment. (Item 9) 9. The method of claim 8, further comprising generating an output comprising the origin of the nucleic acid molecule in the biological sample. (Item 10) 2. The method of claim 1, wherein the concentration of the second plurality of nucleic acid probes is higher than the concentration of the first plurality of nucleic acid probes in the pool of nucleic acid probes. (Item 11) 2. The method of claim 1, wherein the concentration of the second plurality of nucleic acid probes in the pool of probes is higher than the concentration of the first plurality of nucleic acid probes in the pool of nucleic acid probes. (Item 12) Item 10. The method of item 1, wherein the first plurality of nucleic acid probes comprises a probe set for human exome capture. (Item 13) Item 10. The method of item 1, wherein the first plurality of nucleic acid probes comprises probes configured to target junction sequences generated by human V(D)J rearrangement or recombination. (Item 14) 15. The method of claim 1, wherein the second plurality of nucleic acid probes comprises one or more probes configured to target one or more elements of the E6 gene, E7 gene, and / or the bacterial 16S ribosomal RNA gene of human papillomavirus. 2. The method of claim 1, wherein the assay (b) comprises performing sequencing to generate paired-end read sequences of 130 bases to 280 bases in length. (Item 16) 2. The method of claim 1, further comprising generating one or more biomedical reports comprising one or more sets of data selected from the group consisting of: (i) candidate tumor neoantigens, (ii) detected non-human species, (iii) detected complementarity determining region 3 (CDR3) sequences, and any combination thereof. (Item 17) 17. The method of claim 16, wherein the non-human species to be detected is an antigen and the CDR3 sequence has the ability to bind to the antigen. (Item 18) Item 17. The method according to item 16, wherein the one or more biomedical reports include any two of (i) to (iii). (Item 19) 1. A method for processing a biological sample from a subject, comprising: (a) generating a subset of nucleic acid molecules from the biological sample, the subset of nucleic acid molecules comprising (i) a first plurality of nucleic acid molecules derived from the subject, and (ii) a second plurality of nucleic acid molecules not derived from the subject, wherein the abundance of the first plurality of nucleic acid molecules is greater than the abundance of the second plurality of nucleic acid molecules in the biological sample; and (b) assaying a subset of said nucleic acid molecules to obtain sequence information comprising the sequences of (i) said first plurality of nucleic acid molecules and (ii) said second plurality of nucleic acid molecules. A method comprising: (Item 20) 20. The method of item 19, wherein the nucleic acid molecule of (i) is derived from the genome of the subject. (Item 21) 20. The method of claim 19, wherein the second plurality of nucleic acid molecules not derived from the subject comprises one or more members selected from the group consisting of viruses, bacteria, bacterial phages, fungi, protists, archaea, amoebas, helminths, algae, genetically modified cells, and genetically modified vectors. (Item 22) 20. The method of claim 19, wherein generating the subset of nucleic acid molecules from the biological sample comprises performing one or more hybridization reactions. (Item 23) 20. The method of claim 19, further comprising obtaining the biological sample from the subject, wherein the biological sample is derived from a tumor biopsy, whole blood, or plasma. (Item 24) 20. The method of claim 19, further comprising aligning the sequences of the subset of nucleic acid molecules to one or more reference sequences. (Item 25) 25. The method of claim 24, wherein the one or more reference sequences comprise a plurality of reference sequences, the plurality of reference sequences corresponding to two or more different species. (Item 26) 25. The method of claim 24, further comprising identifying the origin of the nucleic acid molecules in the subset based on the alignment. (Item 27) 27. The method of claim 26, further comprising generating an output comprising the identified origin of nucleic acid molecules in the biological sample. (Item 28) 20. The method of claim 19, wherein the abundance of the first plurality of nucleic acid molecules is greater than the abundance of the first plurality of nucleic acid molecules in the biological sample. (Item 29) 20. The method of claim 19, wherein the relative abundance of the second plurality of nucleic acid molecules in the subset is greater than the relative abundance of the first plurality of nucleic acid molecules in the subset. (Item 30) 1. A system for processing a biological sample from a subject, comprising: a processing unit comprising one or more computer processors individually or collectively programmed to assay a subset of nucleic acid molecules to obtain sequence information comprising sequences of (i) human nucleic acids from said biological sample from said subject, and (ii) non-human nucleic acids from said biological sample from said subject, wherein said subset of nucleic acid molecules is generated from said biological sample using a pool of nucleic acid probes, said probes comprising (i) a first plurality of nucleic acid probes configured to target elements of the human genome, and (ii) a second plurality of nucleic acid probes configured to target elements of one or more non-human genomes; and a computer memory configured to store said sequence information. Including, the system. (Item 31) 31. The system of item 30, wherein the one or more computer processors are programmed to generate an alignment of the sequence with one or more reference sequences. (Item 32) 32. The system of claim 31, wherein the one or more reference sequences comprise a plurality of reference sequences, the plurality of reference sequences corresponding to two or more different species. (Item 33) 32. The system of claim 31, wherein the one or more computer processors are programmed to identify the origin of the nucleic acid molecules in the subset based on the alignment. (Item 34) 34. The system of claim 33, wherein the one or more computer processors are programmed to generate an output comprising the origin of the nucleic acid molecules in the biological sample. (Item 35) 31. The system of claim 30, wherein the one or more computer processors are programmed to generate one or more biomedical reports including information selected from the group consisting of: (i) candidate tumor neoantigens, (ii) detected non-human species, (iii) detected complementarity determining region 3 (CDR3) sequences, and any combination thereof. (Item 36) 1. A system for processing a biological sample from a subject, comprising: a processing unit comprising one or more computer processors individually or collectively programmed to assay a subset of nucleic acid molecules to obtain sequence information comprising sequences of (i) the first plurality of nucleic acid molecules and (ii) the second plurality of nucleic acid molecules, wherein the subset of nucleic acid molecules is generated from the biological sample, the subset of nucleic acid molecules comprising (i) a first plurality of nucleic acid molecules derived from the subject, and (ii) a second plurality of nucleic acid molecules not derived from the subject, wherein the abundance of the first plurality of nucleic acid molecules is greater than the abundance of the second plurality of nucleic acid molecules in the biological sample; and a computer memory configured to store said sequence information. Including, the system. (Item 37) 37. The system of item 36, wherein the one or more computer processors are programmed to generate an alignment of the sequence with one or more reference sequences. (Item 38) 38. The system of item 37, wherein the one or more reference sequences comprise a plurality of reference sequences, the plurality of reference sequences corresponding to two or more different species. (Item 39) 38. The system of claim 37, wherein the one or more computer processors are programmed to identify the origin of the nucleic acid molecules in the subset based on the alignment. (Item 40) 39. The system of claim 38, wherein the one or more computer processors are programmed to generate an output comprising the origin of the nucleic acid molecules in the biological sample. (Item 41) 37. The system of claim 36, wherein the one or more computer processors are programmed to generate one or more biomedical reports including information selected from the group consisting of: (i) candidate tumor neoantigens, (ii) detected non-human species, (iii) detected complementarity determining region 3 (CDR3) sequences, and any combination thereof.
Claims
[Claim 1] The invention as described in the drawings of this application.