Compositions, methods, and systems for processing or analyzing multispecies nucleic acid samples
Patent Information
- Application Number
- CN201980050741.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-08-07
- Filing Date
- 2019-05-24
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2039-05-24
AI Technical Summary
Existing genome-wide and exome sequencing methods are expensive and cannot effectively capture biomedical important non-exome regions, and are poorly performed in high CG-content regions, providing adequate and cost-effective sequencing of duplicate elements in the genome.
A subset of nucleic acid molecules are generated from biological samples using a nucleic acid probe pool, including elements targeting the human and non-human genomes, through hybridization and sequencing technology, combined with computer processing units for sequence alignment and analysis, and biomedical reports are generated.
It has achieved efficient and economical sequencing of human and non-human nucleic acids, and can simultaneously detect candidate tumor neoantigens, non-human species and CDR3 sequences, providing more comprehensive genomic information.
Smart Images

Figure CN112888794B_ABST
Abstract
Description
[0001] Cross-references
[0002] This application claims priority to U.S. Patent Application No. 16 / 056,982, filed on August 7, 2018, and U.S. Provisional Patent Application No. 62 / 678,475, filed on May 31, 2018, which are hereby incorporated by reference in their entireties. Background Art
[0003] Current methods for whole genome and / or exome sequencing can be expensive and fail to capture many biomedically important variants. For example, commercially available exome enrichment kits (e.g., Illumina's TruSeq Exome Enrichment and Agilent's SureSelect Exome Enrichment) may not target non-exomic and exome regions of biomedical interest. Typically, whole genome and / or exome sequencing using standard sequencing methods performs poorly in regions of content with very high CG content (>70%). In addition, whole genome and / or exome sequencing also fails to provide sufficient and / or cost-effective sequencing of repetitive elements in the genome.
[0004] The methods disclosed herein provide specialized sequencing protocols or technologies to address these issues and extend the analysis to both human and non-human genomes in a single sample. Summary of the Invention
[0005] Disclosed herein is a method for processing a biological sample obtained from a subject, comprising (a) using a pool of nucleic acid probes to generate a subset of nucleic acid molecules from the biological sample, wherein the probes include (i) a first plurality of nucleic acid probes configured to target elements of the human genome and (ii) a second plurality of nucleic acid probes configured to target elements of one or more non-human genomes; and (b) assaying the subset of nucleic acid molecules to generate sequence information comprising the following sequences: (i) human nucleic acids from the biological sample of the subject and (ii) non-human nucleic acids from the biological sample of the subject. In some cases, the first plurality of nucleic acid probes of (i) are configured to target elements derived from the human genome. In some cases, the subject can be a human. In some cases, the second plurality of nucleic acid probes are configured to target elements in non-human genome sequences from one or more species, the one or more species being selected from viruses, bacteria, bacteriophages, fungi, protozoa, archaea, amoebas, worms, algae, genetically modified cells, and genetically modified vectors. On the one hand, generating a subset of nucleic acid molecules from a biological sample includes performing one or more hybridization reactions. In some cases, the method further comprises obtaining a biological sample from a subject. In some cases, the biological sample of the subject can be derived from a tumor biopsy, whole blood or plasma. In some aspects, the method also includes comparing the sequence of the subgroup of nucleic acid molecules with one or more reference sequences. In some cases, one or more reference sequences include multiple reference sequences. In some cases, multiple reference sequences correspond to two or more different species. The method can also include identifying the source of nucleic acid molecules in the subgroup based on the comparison. In some cases, the method can also include generating an output including identifying the nucleic acid molecule source in the biological sample. In some cases, in the nucleic acid probe pool, the concentration of the second multiple nucleic acid probes is greater than the concentration of the first multiple nucleic acid probes. On the one hand, the relative concentration of the second multiple nucleic acid probes in the probe pool is greater than the relative concentration of the first multiple nucleic acid probes in the nucleic acid probe pool. On the one hand, the first multiple nucleic acid probes include people's exon group capture probe group. In some cases, the first multiple nucleic acid probes include probes configured to target the junction sequence produced by human V (D) J rearrangement or recombination. In some cases, the second multiple nucleic acid probes include one or more probes configured to target human papillomavirus E6 gene and / or E7 gene. In some cases, the second plurality of nucleic acid probes includes probes configured to target one or more elements of a bacterial 16S ribosomal RNA gene. In one aspect, the determining of (b) includes sequencing to generate paired-end read sequences having a length of 130 to 280 bases.In one aspect, the method can further comprise generating one or more biomedical reports comprising one or more sets of data selected from the group consisting of: (i) candidate tumor neoantigens, (ii) detected non-human species, (iii) detected CDR3 sequences, and any combination thereof. In one aspect, the detected non-human species is an antigen. In some cases, the CDR3 sequence is generated by V(D)J rearrangement or recombination. In some cases, the CDR3 sequence corresponds to an immune response to the antigen. In some cases, the one or more biomedical reports comprise (i)-(iii).
[0006] Disclosed herein is a method for processing a biological sample obtained from a subject, the method comprising (a) generating a subset of nucleic acid molecules from the biological sample, wherein the subset of nucleic acid molecules comprises (i) a first plurality of nucleic acid molecules from the subject and (ii) a second plurality of nucleic acid molecules not from the subject, and wherein the abundance of the first plurality of nucleic acid molecules in the biological sample is greater than the abundance of the second plurality of nucleic acid molecules; and (b) assaying the subset of nucleic acid molecules to generate sequence information comprising the sequences of: (i) the first plurality of nucleic acid molecules and (ii) the second plurality of nucleic acid molecules. In some cases, the nucleic acid molecules in (i) are derived from the genome of the subject. In some cases, the subject is human. In some cases, the second plurality of nucleic acid molecules not from the subject comprises one or more members selected from the group consisting of viruses, bacteria, bacteriophages, fungi, protists, archaea, amoebas, worms, algae, genetically modified cells, and genetically modified vectors. In one aspect, generating the subset of nucleic acid molecules from the biological sample comprises performing one or more hybridization reactions. In one aspect, the method further comprises obtaining the biological sample from the subject. In some cases, the subject's biological sample is derived from a tumor biopsy, whole blood, or plasma. In one aspect, the method further comprises comparing the sequences of the subset of nucleic acid molecules with one or more reference sequences. In one aspect, the one or more reference sequences comprise a plurality of reference sequences. In one aspect, the plurality of reference sequences correspond to two or more different species. The method may further comprise identifying the source of the nucleic acid molecules in the subset based on the comparison. In one aspect, the method further comprises generating an output comprising the source of the nucleic acid molecules identified in the biological sample. In some cases, the abundance of the first plurality of nucleic acid molecules in the biological sample is greater than the abundance of the first plurality of nucleic acid molecules. In some cases, the relative abundance of the second plurality of nucleic acid molecules in the subset is greater than the relative abundance of the first plurality of nucleic acid molecules in the subset.
[0007] Disclosed herein is a composition comprising a pool of probes configured to hybridize to (i) one or more human sequences from a subject and (ii) one or more non-human sequences from a subject. In one aspect, the pool of probes is a plurality of capture probes. In one aspect, the pool of probes is a plurality of amplification probes.
[0008] Disclosed herein is a system for processing a biological sample of a subject, comprising: a processing unit comprising one or more computer processors, the one or more computer processors being programmed individually or collectively to assay a subset of nucleic acid molecules to generate sequence information comprising the following sequences: i) human nucleic acids from the biological sample of the subject and (ii) non-human nucleic acids from the biological sample of the subject, the subset of nucleic acid molecules being generated from the biological sample using a pool of nucleic acid probes, wherein the probes comprise (i) a first plurality of nucleic acid probes configured to target elements of the human genome and (ii) a second plurality of nucleic acid probes configured to target elements of one or more non-human genomes; and a computer memory configured to store sequence information. In some embodiments, the one or more computer processors are programmed to generate an alignment of a sequence with one or more reference sequences. In some embodiments, the one or more reference sequences include multiple reference sequences, and wherein the multiple reference sequences correspond to two or more different species. In some embodiments, the one or more computer processors are programmed to identify the source of the nucleic acid molecules in the subset based on the alignment. In some embodiments, one or more computer processors are programmed to generate one or more biomedical reports comprising information selected from the group consisting of: (i) candidate tumor neoantigens, (ii) detected non-human species, (iii) detected complementarity determining region 3 (CDR3) sequences, and any combination thereof.
[0009] Disclosed herein is a system for processing a biological sample of a subject, comprising: a processing unit comprising one or more computer processors, the one or more computer processors being programmed individually or collectively to assay a subset of nucleic acid molecules to generate sequence information comprising the following sequences: i) a first plurality of nucleic acid molecules and (ii) a second plurality of nucleic acid molecules, the subset of nucleic acid molecules being generated from a biological sample, wherein the subset of nucleic acid molecules comprises (i) a first plurality of nucleic acid molecules from the subject and (ii) a second plurality of nucleic acid molecules not from the subject, wherein the abundance of the first plurality of nucleic acid molecules in the biological sample is greater than the abundance of the second plurality of nucleic acid molecules; and a computer memory configured to store sequence information. In some embodiments, the one or more computer processors are programmed to generate an alignment of a sequence with one or more reference sequences. In some embodiments, the one or more reference sequences include multiple reference sequences, and wherein the multiple reference sequences correspond to two or more different species. In some embodiments, the one or more computer processors are programmed to identify the source of the nucleic acid molecules in the subset based on the alignment. In some embodiments, one or more computer processors are programmed to generate one or more biomedical reports comprising information selected from the group consisting of: (i) candidate tumor neoantigens, (ii) detected non-human species, (iii) detected complementarity determining region 3 (CDR3) sequences, and any combination thereof.
[0010] Another aspect of the present disclosure provides a non-transitory computer-readable medium comprising machine-executable code that, when executed by one or more computer processors, implements any of the methods described above or elsewhere herein.
[0011] Another aspect of the present disclosure provides a system comprising one or more computer processors and a computer memory coupled thereto, wherein the computer memory comprises machine executable code that, when executed by the one or more computer processors, implements any method described above or elsewhere herein.
[0012] Other aspects and advantages of the present disclosure will become apparent to those skilled in the art through the detailed description below, wherein only illustrative embodiments of the present disclosure are shown and described. As will be appreciated, the present disclosure is capable of other and different embodiments, and its many details are capable of modification in various obvious aspects, all without departing from the present disclosure. Therefore, the drawings and description are to be regarded as illustrative in nature, and not restrictive.
[0013] Incorporation by reference
[0014] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The novel features of the present invention are set forth with particularity in the appended claims. The features and advantages of the present invention may be better understood by reference to the following detailed description and accompanying drawings (also referred to herein as "Figures") which illustrate exemplary embodiments in which the principles of the invention are utilized, wherein:
[0016] Figure 1 Schematically illustrating a method of generating a subset of nucleic acid molecules using a first plurality of nucleic acid probes and a second plurality of nucleic acid probes;
[0017] Figure 2 A method for generating a subset of nucleic acid molecules using a first plurality of nucleic acid probes and a second plurality of nucleic acid probes is schematically illustrated. The first plurality of nucleic acid probes can bind to a non-human genome;
[0018] Figure 3 Schematically illustrates a method for generating a subset of nucleic acid molecules using a first plurality of nucleic acid probes and a second plurality of nucleic acid probes. The first plurality of nucleic acid probes can bind to a segmented transcriptome;
[0019] Figure 4 A method for generating a subset of nucleic acid molecules using a plurality of nucleic acid probes is schematically illustrated. The plurality of nucleic acid probes may include a probe set targeting an exome of a subject and a probe set targeting nucleic acid sequences not belonging to the subject;
[0020] Figure 5 A computer system is shown that is programmed or otherwise configured to implement the methods of the present disclosure;
[0021] Figure 6A A schematic diagram of the workflow is shown. Prep 1 and Prep 2 refer to subsets of nucleic acids. Assay 1, Analysis 1, and Output refer to any assay, analysis, and output described herein. Figure 6B A schematic diagram of the workflow is shown. Prep 1 and Prep 2 refer to subsets of nucleic acids. Assay 1, Assay 2, Analysis 1, and Output refer to any assay, analysis, and output described herein. Figure 6C A schematic diagram of the workflow is shown. Prep 1 and Prep 2 refer to subsets of nucleic acids. Assay 1, Assay 2, Analysis 1, Analysis 2, and Output refer to any assay, analysis, and output described herein. Figure 6DA schematic diagram of a workflow is shown, which includes (1) separating a nucleic acid sample into multiple subsets processed using multiple protocols. These protocols may involve enrichment of different genomic regions or non-genomic regions and include one or more different amplification operations to prepare nucleic acid molecule libraries for assay. Some of these libraries can be combined (2) for assay. The results of some assays can be combined (3) for subsequent analysis. Variant calls or other assessments of sequence or genetic state can be further combined (4) to produce a combined assessment at the east locus targeted by the assay. Protocols 1-4 refer to any method described herein. Assay 1, Assay 2, Assay 3, Analysis 1, Analysis 2, and Output refer to any assay, analysis, and output described herein;
[0022] Figure 7 An example of the assay workflow described herein is depicted;
[0023] Figure 8 A schematic diagram showing the workflow of the present disclosure;
[0024] Figure 9 A schematic diagram showing the workflow of the present disclosure;
[0025] Figure 10 The effect of shearing time on fragment size is shown;
[0026] Figure 11 The effect of bead ratio on fragment size is shown;
[0027] Figure 12 The effect of shearing time on fragment size is shown;
[0028] Figure 13 A schematic diagram depicting the workflow of nucleic acid library construction is presented;
[0029] Figure 14 Methods for developing multithreaded assays for various biomedical applications are shown;
[0030] Figure 15 An example of an assay workflow is depicted that includes multiple subsets of DNA enriched for different genomic regions, which undergo several independent processing operations before being combined for sequencing assays. Reads from two or more subsets are combined a) in a sequencing instrument or b) subsequently in silico (e.g., using one or more algorithms) to produce a single test result for the combined addressed region of the two or more subsets, and to generate a data pool that can be used for one or more biomedical reports. Supplemental pullouts can include human target sequences and non-human target sequences;
[0031] Figure 16Depicted is an example of an assay workflow that includes multiple subsets of DNA enriched for different genomic regions, which undergo independent processing operations before being independently sequenced and analyzed for variants. Variants from two subsets can be combined to generate results for the region addressed by the combination of two or more subsets, generating a data pool that can be used in one or more biomedical reports. Supplemental extracts can include human target sequences and non-human target sequences;
[0032] Figure 17 An example of an assay workflow is depicted that includes multiple subsets of DNA enriched for different genomic regions, which undergo independent processing operations before being independently sequenced and generate raw data that can include sequence reads. The raw data from two or more assays can be combined and analyzed (e.g., by one or more software programs or algorithms) to generate results for all regions addressed by the combination of the two or more subsets, thereby generating a data pool that can be used for one or more biomedical reports. The supplemental extracts can include human target sequences and non-human target sequences;
[0033] Figure 18 Depicts a multithreaded assay that involves generating two subsets of DNA by size selection and further dividing them into two subsets enriched for different genomic regions based on GC content. Longer molecules can be sequenced using techniques suitable for longer molecules. m The protocol further prepares and amplifies two shorter molecular subsets, which are then combined for sequencing on a high-throughput short-read sequencer, HiSeq. The raw data from sequencing can be combined and analyzed (e.g., by one or more software programs or algorithms) to produce a single optimal result for all regions addressed by the subsets and generate a data pool that can be used for one or more biomedical reports. The supplemental extracts can include human target sequences and non-human target sequences. DETAILED DESCRIPTION
[0034] Although various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that these embodiments are provided by way of example only. Without departing from the present invention, those skilled in the art may conceive of many variations, changes, and substitutions. It should be understood that various alternatives of the embodiments of the present invention described herein may be employed.
[0035] The human body often carries a large number of microbial species, in some cases exceeding 1,000. These microbial species have been documented by the Human Microbiome Project and other studies and may include many viruses, bacteria, bacteriophages, fungi, protists, archaea, as well as some amoebas, worms, and algae. These non-human species may be beneficial, but they may also cause or modulate diseases, including cancer. Their effects can be direct (e.g., causing mutations in human cells that lead to cancer) or indirect (e.g., stimulating the immune system, which affects its ability to resist disease). Microbial species in individuals can be identified and characterized by sequencing nucleic acids extracted from human samples. Many of these microorganisms live at the interface between the human body and its surrounding environment (e.g., skin, saliva, nasal passages, intestines, bowel movements, genitals). Sampling can be performed from these interfaces by swabbing, biopsy, taking feces or urine, taking saliva, or similar methods. A variety of methods have been developed to sequence the microbial nucleic acids in these samples. These include both targeted PCR amplification (e.g., of subsections of the 16S ribosomal RNA gene) and non-targeted (metagenomic) methods using deep sequencing. Many of these methods involve sample types and / or enrichment methods that attempt to minimize the amount of human DNA, thereby optimizing sensitivity to the intended microbial target. The human genome (approximately 3 billion bases) is about 1,000 times the size of a typical bacterial genome (less than a million bases to up to several million bases) and thousands of times the size of most viral genomes (typically thousands to up to tens of thousands of bases). Therefore, if there are no sampling methods or detection techniques to avoid or reduce it, the human DNA content will take up most of the detection capacity.
[0036] The progression of many diseases is also influenced by human cytogenetics. This can include inherited genetic variants, somatic variants in cancer, VDJ recombination in immune cells, differential gene expression in different cell types, and other characteristics. These can also be determined by sequencing nucleic acids typically from blood (most commonly DNA or cell-free DNA from PBMCs) or from diseased tissues (e.g., tumor biopsies). A variety of methods have been developed for sequence analysis of nucleic acids in these samples, including amplicons (typically for up to hundreds of genes), hybrid capture (typically for a large number of genes, including exomes), and non-targeted methods (whole genome sequencing). Many of these methods involve sample types and / or enrichment methods optimized for their intended human nucleic acid targets.
[0037] The sample types used for microbial analysis can differ from those used for human genetic analysis. For example, while stool samples can be used to analyze the gut microbiome, they can be a poor sample for detecting inherited human genetic diseases, such as cystic fibrosis. White blood cells (PBMCs) in the blood can be sequenced to identify the causes of inherited genetic diseases, but because the immune system largely excludes bacteria from the blood, they are generally not a good choice for bacterial analysis. The assay methods used for human and microbial analysis also differ. For example, PCR can be used in either, but because PCR produces very narrowly focused results (i.e., each amplicon typically represents a small portion of the target species' genome), human and microbial targets are not typically combined in a single PCR-based assay. Non-targeted assays can be used for both human genetics (i.e., whole human genome sequencing) and microbial metagenomics, but because these genome sizes differ significantly, it is generally best to perform each assay separately, optimizing the raw materials and assays for each. For similar reasons, hybridization capture assays (e.g., exome sequencing) have been developed for specific (usually mammalian) species, such as human exomes or bovine exomes.
[0038] Cancer may be a special case. Many tumors use checkpoint genes and other methods to partially or completely exclude the immune system. Therefore, microorganisms that manage to infiltrate the tumor may be able to survive and even thrive there. Complete living bacteria (such as Fusobacterium) may already be cancer cells and live entirely in them, including through cycles of cell division and through metastasis. Bacteria may cause cancer (for example, Helicobacter pylori is a cause of gastric cancer), they can affect the progression of cancer and respond to cancer therapeutics (such as immune checkpoint inhibitors). Viruses may also enter cells and, in some cases, are the cause of cancer, and / or may integrate their genomes into human chromosomes. Therefore, a tumor biopsy can contain human and microbiome species, as well as their nucleic acids. When these cells die, their multispecies nucleic acids may fall into their surrounding environment and eventually be detected as cell-free nucleic acids in the plasma.
[0039] Tumor biopsies are often very limited in sample size, making it more difficult to perform multiple different assays for both human and microbial genetic targets in the same sample. When sample size is limited, integrated assays that can simultaneously provide sequence data from both human and microbiome species can be advantageous. The amount of cell-free DNA and RNA in plasma is also often limited. Integrated assays that can simultaneously provide sequence data from both human and microbiome species from a single, small plasma sample can also be advantageous.
[0040] Disclosed herein is an assay that supports the parallel detection of a wide range of human genetic data as well as microbiome data from the same sample at the same time. A single assay can be performed that does not require any more samples than an equivalent human-only assay might require. This assay can use a human exome capture kit (e.g., Agilent Clinical Research Exome v2). The kit or composition of the probe pool can use hybridization probes that are complementary to the human sequences they target. By using over 50,000 capture probes, the method, kit or composition can target the exons of most predetermined human genes. For example, before the nucleic acids in a cancer sample are subjected to a hybridization reaction with the probes of the kit, we add a set of additional capture probes that have been designed to target non-human sequences. The non-human sequences can be from viral, bacterial, fungal or archaeal genomes, i.e., from the human microbiome. Once the human exome probes are combined with the probes of the non-human microbiome, the probe mixture can be used to perform a hybridization-based capture reaction ( Figure 2 The captured nucleic acids can then be sequenced. In our laboratory, this sequencing is performed using the Illumina NovaSeq-6000 DNA sequencer.
[0041] The mixed (human and non-human) DNA sequences produced by this process can be separated by alignment with human and microbiome species reference sequences. Separating sequences by alignment is possible because the human genome has diverged significantly from microbial genomes during evolution.
[0042] Because PCR is difficult to target multiple genes simultaneously, the use of capture probes can be advantageous over polymerase chain reaction (PCR)-based assays to enrich and target sequences of interest. To target multiple genes, PCR can require a large number of primers, for example, up to potentially 100,000 primers to amplify and target 50,000 sequences, and requires enzymatic manipulation to generate nucleic acid molecules for identification. Optimizing the ratio of PCR primers for multiple genes can also be difficult and can result in biased amplification results that may not represent the relative amounts of target nucleic acids in the nucleic acid sample.
[0043] As used in the specification and claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. For example, the term "a chimeric transmembrane receptor polypeptide" includes a plurality of chimeric transmembrane receptor polypeptides.
[0044] The term "about" or "approximately" means within an acceptable error range for a particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limits of the measurement system. For example, according to practice in the art, "about" can mean within 1 or more standard deviations. Alternatively, "about" can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, for example with respect to biological systems or processes, the term can mean within an order of magnitude of the value, or within 5 times, or within 2 times the value. Where a particular value is described in the application and claims, it should be assumed that the term "about" means within an acceptable error range for that particular value unless otherwise indicated.
[0045] As used herein, "cell" generally refers to a biological cell. A cell can be the basic structural unit, functional unit, and / or biological unit of a living organism. A cell can be derived from any organism having one or more cells. Some non-limiting examples include: prokaryotic cells, eukaryotic cells, bacterial cells, archaeal cells, cells of unicellular eukaryotic organisms, protozoan cells, cells from plants (e.g., cells from plant crops, fruits, vegetables, grains, soy, corn, maize, wheat, seeds, tomatoes, rice, cassava, sugarcane, pumpkin, hay, potatoes, cotton, hemp, tobacco, flowering plants, conifers, gymnosperms, ferns, clubmosses, hornworts, liverworts, mosses), algal cells (e.g., Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens, C. Agardh, etc.), algae (e.g., seaweed), fungal cells (e.g., yeast cells, cells from mushrooms), animal cells, cells from invertebrates (e.g., fruit flies, cnidarians, echinoderms, nematodes, etc.), cells from vertebrates (e.g., fish, amphibians, reptiles, birds, mammals), cells from mammals (e.g., pigs, cows, goats, sheep, rodents, rats, mice, non-human primates, humans, etc.), etc. Sometimes the cells are not from natural organisms (e.g., cells can be artificially synthesized, sometimes also called artificial cells).
[0046] As used herein, the term "nucleotide" generally refers to a combination of base sugar phosphates. Nucleotides may include synthetic nucleotides. Nucleotides may include synthetic nucleotide analogs. Nucleotides may be monomeric units of nucleic acid sequences (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide may include ribonucleoside triphosphates adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP), and deoxyribonucleoside triphosphates, such as dATP, dCTP, dITP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives may include, for example, [αS] dATP, 7-deaza-dGTP, and 7-deaza-dATP, as well as nucleotide derivatives that confer nuclease resistance on nucleic acid molecules comprising them. The term nucleotide used herein may refer to dideoxyribonucleoside triphosphates (ddNTP) and derivatives thereof. Illustrative examples of dideoxyribonucleoside triphosphates may include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. Nucleotide can be unlabeled or detectably labeled.Labeling can also be carried out with quantum dots.Detectable labeling can include, for example, radioisotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzyme labels.The fluorescent labeling of nucleotides can include, but is not limited to, fluorescein, 5-carboxyfluorescein (FAM), 2'7'-dimethoxy-4'5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N, N, N', N'-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxyl-X-rhodamine (ROX), 4-(4'-dimethylaminophenylazo) benzoic acid (DABCYL), Cascade Blue, Oregon Green, Texas Red, cyanine, and 5-(2'-aminoethyl) aminonaphthalene-1-sulfonic acid (EDANS).Specific examples of fluorescently labeled nucleotides can include [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP, available from Perkin Elmer in Foster City, California; and FluoroLink Deoxy Nucleotides, FluoroLink Cy3-dCTP, FluoroLink Cy5-dCTP, FluoroLink Fluor X-dCTP, FluoroLink Cy3-dUTP, and FluoroLink Cy5-dUTP; fluorescein 15-dATP, fluorescein-12-dUTP, tetramethylrhodamine 6-dUTP, IR770-9-dATP, fluorescein-12-ddUTP, fluorescein-12-UTP, and fluorescein-15-2′-dATP, available from Boehringer Mannheim, Indianapolis, IN; and chromosome-labeled nucleotides, BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-dUTP, Cascade Blue-7-dUTP, available from Molecular Probes, Eugene, OR. In some embodiments, the nucleotides may be labeled or marked by chemical modification. The chemically modified single nucleotides may be biotin-dNTPs. Some non-limiting examples of biotinylated dNTPs can include biotin-dATP (e.g., biotin-N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP).
[0047] As used herein, the term "genome" is generally used to refer to a portion of a subject's genome or the subject's entire genome. For example, a genome can refer to a subject's gene sequence. A genome can refer to a subject's entire genome sequence.
[0048] The terms "polynucleotide," "oligonucleotide," and "nucleic acid" are used interchangeably to refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides or their analogs in single-stranded, double-stranded, or multi-stranded form. A polynucleotide can be exogenous or endogenous to a cell. A polynucleotide can be present in a cell-free environment. A polynucleotide can be a gene or a fragment thereof. A polynucleotide can be DNA. A polynucleotide can be RNA. A polynucleotide can have any three-dimensional structure and can perform any function. A polynucleotide can include one or more analogs (e.g., altered backbones, sugars, or nucleobases). If present, the nucleotide structure can be modified before or after polymer assembly. Some non-limiting examples of analogs include: 5-bromouracil, peptide nucleic acids, xenologous nucleic acids, morpholinos, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein linked to a sugar), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouracil, dihydrouridine, q-glycosides, and wyosine. Non-limiting examples of polynucleotides include coding or non-coding regions of genes or gene fragments, loci defined by linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotides (including cell-free DNA (cfDNA) and cell-free RNA (cfRNA)), nucleic acid probes and primers. The sequence of nucleotides may be interrupted by non-nucleotide components. Any of the aforementioned nucleic acid molecules may be engineered or synthetic.
[0049] As used herein, the term "gene" refers to nucleic acids (such as DNA, such as genomic DNA and cDNA) and their corresponding nucleotide sequences related to coding RNA transcripts. As used herein, the term for genomic DNA includes the non-coding region and regulatory region in the middle, and may include 5' end and 3' end. In some applications, the term encompasses transcribed sequences, including 5' and 3' untranslated regions (5'-UTR and 3'-UTR), exons and introns. In some genes, the transcribed region will include an "open reading frame" encoding a polypeptide. In some applications of the term, a "gene" only includes the coding sequence (such as "open reading frame" or "coding region") required for encoding a polypeptide. In some cases, a gene does not encode a polypeptide, such as ribosomal RNA genes (rRNA) and transfer RNA (tRNA) genes. In some cases, the term "gene" not only includes transcribed sequences, but also includes non-transcribed regions (including upstream and downstream regulatory regions), enhancers and promoters. A gene can refer to an "endogenous gene" or a natural gene that is in its natural position in the genome of an organism. Gene can refer to "exogenous gene" or non-natural gene. Non-natural gene can refer to a gene that is not normally present in the host organism but is introduced into the host organism by gene transfer. Non-natural gene can also refer to a gene that is not in its natural position in the genome of an organism (such as a genetically modified organism). Non-natural gene can also refer to a naturally occurring nucleic acid or polypeptide sequence (e.g., non-natural sequence) that comprises a mutation, insertion and / or deletion.
[0050] As used herein, the term "percent (%) identity" refers to the percentage of amino acid or nucleic acid residues in a candidate sequence that are identical with the amino acid or nucleic acid residues in a reference sequence after aligning the sequences and introducing gaps, if necessary, to achieve the maximum percent identity (e.g., gaps can be introduced in one or both of the candidate and reference sequences for optimal alignment and nonhomologous sequences can be ignored for comparison purposes). To determine percent identity, alignment can be achieved in a variety of ways, for example, using computer software such as BLAST, ALIGN, or Megalign (DNASTAR) software. The percent identity of two sequences can be calculated by aligning a test sequence with a comparison sequence using BLAST, determining the number of amino acids or nucleotides in the aligned test sequence that are identical to the amino acid or nucleotide at the same position in the comparison sequence, and dividing the number of identical amino acids or nucleotides by the number of amino acids or nucleotides in the comparison sequence.
[0051] As used herein, the term "subject" generally refers to any animal, such as a mammal or marsupial. A subject can be a patient. For a disease or condition, a subject can be symptomatic or asymptomatic. A subject can be a primate (e.g., human), non-human primate (e.g., rhesus monkey or other types of macaques), dog, cat, mouse, pig, horse, donkey, cattle, sheep, rat, and poultry. Mammals include, but are not limited to, rodents, monkeys, humans, farm animals, sport animals, and pets. Also included are tissues, cells, and progeny of biological entities obtained in vivo or cultured in vitro. A host is an organism that can carry a non-host. A subject may be asymptomatic for a disease. Alternatively, a subject is not asymptomatic for a disease.
[0052] As used herein, the term "treatment" refers to an approach for obtaining a beneficial or desired result, including but not limited to therapeutic benefit and / or prophylactic benefit. For example, treatment can include administering a system or cell population disclosed herein. Therapeutic benefit refers to any therapeutically relevant improvement or effect on one or more diseases, conditions, or symptoms being treated. For prophylactic benefit, a composition can be administered to a subject at risk for a particular disease, condition, or symptom, or to a subject who reports having one or more physiological symptoms of a disease, even though the disease, condition, or symptom may not yet be manifested.
[0053] In some cases, the present disclosure also provides compositions and methods for processing and analyzing biological samples. In some cases, a biological sample from a subject may contain nucleic acids from the subject and nucleic acid molecules that are not from the subject. In some cases, a biological sample may contain nucleic acids from human and non-human genomes. In some cases, the non-human genome may be from a virus, bacteria, bacteriophage, fungi, protists, archaea, amoeba, worms, algae, or a combination thereof. In some cases, the source of the non-human genome may be beneficial to the human host. In some cases, the source of the non-human genome may be involved in regulating diseases, such as cancer. In some cases, the source of the non-human organism comprising the non-human genome may be directly manifested by causing mutations in human cells, thereby causing cancer in the subject. In some cases, the source of the non-human genome may be indirectly manifested, for example, by stimulating the human host's immune system to affect its ability to resist disease. In some cases, the present disclosure also provides a method comprising identifying the presence of a non-human genome and a human genome in a sample.
[0054] The sample can be from the skin, saliva, nasal passages, intestinal tract, intestines, genitals or a combination thereof. In some cases, the sample can be obtained by swabbing, biopsy, collecting stool, collecting urine, collecting saliva, etc.
[0055] The human genome (approximately 3 billion bases) is 1,000 times larger than non-human genomes (e.g., bacterial genomes) and thousands of times larger than most viral genomes. In some cases, the methods can include a sampling method to enrich for non-human genomes in a mixed sample.
[0056] The present disclosure also provides a method comprising performing sequence (or sequencing) analysis on the microbial nucleic acid from the sample. Sequencing analysis can include PCR amplification (e.g., segmentation of 16S ribosomal RNA genes), and non-targeted metagenomics methods using deep sequencing. The method can include selecting sample types and / or enrichment methods that can seek to minimize the amount of the human genome, thereby optimizing sensitivity to non-human genomes.
[0057] On the one hand, the first multiple nucleic acid molecules can be from a subject. The subject can be the human race, so the present disclosure also provides human nucleic acid molecules. The present disclosure also provides the use of a first subset or product of a plurality of nucleic acid probes that can be from the human genome. The present disclosure also provides the human genome. In some cases, the human genome may include, for example, inherited genetic variants, somatic variants, VDJ recombination in immune cells, differential gene expression in different cell types and other characteristics. At least some of the aforementioned can contribute to the subject's disease or be associated therewith (for example, somatic variants in cancer). The genome can include genes, exons, UTRs, regulatory regions, splice sites, recombinant genes, alternative sequences, reassembly genes, gene phasing, exogenous sequences, etc.
[0058] In some cases, nucleic acids in biological samples can be analyzed by sequencing. In some cases, nucleic acids typically from blood or diseased tissues can be sequenced. In some cases, a blood sample can include peripheral blood mononuclear cells, cell-free DNA, or a combination thereof. In some cases, diseased tissues may include cancer. In some cases, sequence analysis of nucleic acids from a blood sample or diseased tissue sample can include the generation of amplicon groups, hybridization capture, and non-targeted methods, such as whole genome sequencing. In some cases, the method can involve the selection of sample types, human genome enrichment methods, and combinations thereof.
[0059] In some cases, biological sample can include subject nucleic acid, non-subject nucleic acid and combinations thereof. In some cases, biological sample can include host nucleic acid, non-host nucleic acid and combinations thereof. In some cases, biological sample can include the genome encoding receptor. In some cases, receptor can be from immune cells. In some cases, the receptor from immune cells can be T cell receptor (TCR), B cell receptor (BCR), chimeric antigen receptor (CAR) etc.
[0060] In some cases, the method provided herein includes measuring nucleic acid molecules to generate sequence information. The determination can include sequencing the nucleic acid comprising VDJ rearrangement or VDJ recombination. In some cases, VDJ rearrangement or recombination can refer to a cell receptor. For example, somatic hypermutation can produce a B cell receptor (BCR) sequence encoding an antibody to achieve significant antigenic diversity. In some cases, BCR can be sequenced. The method provided herein can include sequencing BCR to illustrate how antibodies develop. For example, the method can include sequence analysis to annotate each base as a specific gene from a V, D or J gene, or from N addition (also known as non-templated insertion). In some cases, VDJ recombination can produce a CDR3 sequence. In some cases, a CDR3 sequence can produce a polypeptide that binds an antigen such as a tumor antigen. In some cases, a CDR3 sequence can produce a polypeptide that binds an antigen such as a neoantigen.
[0061] In some cases, the methods provided herein may include sequencing nucleic acids encoding candidate tumor neoantigens or associated with candidate tumor neoantigens. In some cases, the methods provided herein may include sequencing the detected CDR3 sequences. In some aspects, a variety of methods can be used to identify subject TCRs. In some cases, whole-exomic sequencing can be used to identify TCRs. For example, TCRs can target neoantigens or neoepitopes identified by whole-exome sequencing of target cells. Alternatively, TCRs can be identified from autologous, allogeneic, or xenogeneic libraries. In some cases, the gene that can contain a mutation that generates a neoantigen or neoepitope can be ABL1, ACO1 1997, ACVR2A, AFP, AKT1, ALK, ALPPL2, ANAPC1, APC, ARID1A, AR, AR-v7, ASCL2, β2M, BRAF, BTK, C15ORF40, CDH1, CLDN6, CNOT1, CT45A5, CTAG1B, DCT, DKK4, EEF1B2, EEF1DP3, EGFR, EIF2B3, env, EPHB2, ERBB3, ESR1, ESRP1, FAM11 1B, FGFR3, FRG1B, GAGE1, GAGE 10, GATA3, GBP3, HER2, IDH1, JAK1, KIT, KRAS, LMAN1, MABEB 16. MAGEA1, MAGEA10, MAGEA4, MAGEA8, MAGEB 17. MAGEB4, MAGEC1, MEK, MLANA, MLL2, MMP13, MSH3, MSH6, MYC, NDUFC2, NRAS, NY-ESO, PAGE2, PAGE5, PDGFRa, PIK3CA, PMEL, pol protein, POLE, PTEN, RAC1, RBM27, RNF43, RPL22, RUNX1, SEC31A, SEC63, SF3B 1. SLC35F5, SLC45A2, SMAP1, SMAP1, SPOP, TFAM, TGFBR2, THAP5, TP53, TTK, TYR, UBR5, VHL, XPOT.
[0062] The present disclosure also provides methods for obtaining or providing nucleic acid samples or subsets of nucleic acid molecules comprising one or more genomes. The methods disclosed herein can analyze nucleic acid subsets of nucleic acid molecules generated from a biological sample. The subsets can include nucleic acid molecules from a subject and nucleic acid molecules not from a subject. For example, target capture and sequencing can be used to enrich for human and microbial sequences. One or more genomes can include one or more genomic signatures.
[0063] The genomic signature may include the entire genome or a portion thereof. The genomic signature may include the entire exon group or a portion thereof. The genomic signature may include one or more groups of genes. The genomic signature may include one or more genes. The genomic signature may include one or more groups of regulatory elements. The genomic signature may include one or more regulatory elements. The genomic signature may include a set of polymorphisms. The genomic signature may include one or more polymorphisms. In some cases, polymorphism refers to a mutation in a genotype. A polymorphism may include a change in one or more bases, an insertion, duplication, or deletion of one or more bases. The genomic signature may include copy number variants (CNVs), transversions, other rearrangements, and other forms of genetic variation. In some cases, one or more features of a nucleic acid sample subset may be polymorphic markers, including restriction fragment length polymorphisms, variable number tandem repeats (VNTRs), hypervariable regions, minisatellite sequences, dinucleotide repeats, trinucleotide repeats, tetranucleotide repeats, simple sequence repeats, and insertion elements (e.g., Alu). In some cases, the difference between the first subset of nucleic acid molecules and the second subset of nucleic acid molecules can be a polymorphic marker, including restriction fragment length polymorphism, variable number tandem repeats (VNTR), hypervariable region, minisatellite sequence, dinucleotide repeats, trinucleotide repeats, tetranucleotide repeats, simple sequence repeats and insertion elements (e.g., Alu). The allelic form that most frequently occurs in a selected population is sometimes referred to as wild-type form. Diploid organisms can be homozygous or heterozygous for allelic form. Biallelic polymorphism has two forms. Triallelic polymorphism has three forms. Polymorphism can include single nucleotide polymorphism (SNP). In some aspects of the present disclosure, one or more polymorphisms include one or more single nucleotide variations, inDel, small insertions, small deletions, structural variant connections, tandem repeats of variable length, flanking sequences or their combination. One or more polymorphisms can be located within coding and / or non-coding regions. One or more polymorphisms can be located within, around or near a gene, exon, intron, splice site, untranslated region or its combination. One or more polymorphisms can span at least a portion of a gene, exon, intron, or untranslated region. In some cases, a genomic signature may be relevant to the GC content, complexity, and / or mappability of one or more nucleic acid molecules. A genomic signature may include one or more simple tandem repeats (STRs), unstable extended repeats, segmental duplications, single and paired reads, denaturation mapping scores, GRCh38 or GRCh37 patches, or a combination thereof. A genomic signature may include one or more low average coverage regions from whole genome sequencing (WGS), zero average coverage regions from WGS, verified compression, or a combination thereof. A genomic signature may include one or more alternative or non-reference sequences.
[0064] The genomic signature may include one or more gene phasing and reassembly genes. Examples of phasing and reassembly genes include, but are not limited to, one or more major histocompatibility complexes, blood typing, and amylase gene families. In some cases, gene phasing and / or reassembly genes may include genes related to blood typing. Blood typing genes may include ABO, RHD, RHCE, or a combination thereof.
[0065] In some cases, genomic signatures may include reassembly genes. Reassembly genes may include genes involved in immune response. Genes involved in immune response may include genes related to major histocompatibility complexes, immune receptors, and cellular functions. One or more major histocompatibility complexes may include one or more HLA classes I, HLA classes II, or combinations thereof. HLA class I may be any of HLA-A, HLA-B, HLA-C, or a combination thereof. HLA class II may be any of HLA-DP, HLA-DM, HLA-DOA, HLA-DOB, HLA-DQ, HLA-DR, or a combination thereof. Genes involved in immune response may be RAG1, RAG2, and a combination thereof. In some cases, reassembly genes may include genes involved in VDJ recombination. For example, in order to establish diversity in B cells and T cell receptors (BCR and TCR), genes may be generated by recombining pre-existing gene segments. In some cases, different combinations of a limited set of gene segments may produce receptors capable of identifying an unlimited number of foreign genomes or non-human genomes. VDJ recombination may include cutting DNA containing a recombination signal sequence (RSS). In some cases, the fragmented sequences can be reassembled using cellular repair mechanisms. The present disclosure also provides methods comprising sequencing fragmented portions of a genome from VDJ recombination. For example, the present disclosure provides methods comprising sequencing at VDJ recombination sites.
[0066] In some aspects of the present disclosure, one or more genomic features may not be mutually exclusive. For example, the genomic features comprising the entire genome or a part thereof may overlap with other genomic features (such as the entire exon group or a part thereof, one or more genes, one or more regulatory elements, etc.). Or, one or more genomic features may be mutually exclusive. For example, the genome comprising the non-coding portion of the entire genome may not overlap with the genomic features (such as the coding portion of the exon group or its part or gene). Alternatively or additionally, one or more genomic features are partially exclusive or partially included. For example, the genome comprising the entire exon group or a part thereof may overlap with the genome portion comprising the exon portion of the gene. However, the genome comprising the entire exon group or a part thereof may not overlap with the genome comprising the intron portion of the gene. Therefore, the genomic features comprising a gene or a part thereof may partially exclude and / or partially include the genomic features comprising the entire exon group or a part thereof. In some cases, the genetic signature may be associated with species, thereby distinguishing more than one species. In some cases, the genetic signature may be associated with species, thereby distinguishing human genetic features and bacterial genetic features. In some cases, a first subset of nucleic acid molecules is specific for one species, and a second subset of nucleic acid molecules is specific for a second species.
[0067] The biological sample may comprise a human genome, a non-human genome, or a combination thereof. In some cases, the biological sample may be processed. In some cases, the biological sample may be enriched for nucleic acid sequences of a subject (e.g., a human), or may be enriched for nucleic acid sequences not from a subject (e.g., a non-human), while simultaneously detecting the human genome and the non-human genome in the sample. The biological sample may comprise cell-free nucleic acid molecules, such as cell-free DNA (cfDNA) or cell-free RNA. The cell-free nucleic acid molecules may be circulating tumor nucleic acid molecules (e.g., circulating tumor DNA). The cell-free DNA may comprise mutations, which may indicate a disease, be associated with a disease, or be associated therewith, such as cancer.
[0068] The biological sample can have a nucleic acid molecule comprising an engineered sequence. For example, the nucleic acid in the biological sample may include an exogenous or alternative sequence (e.g., a tag), an exogenous receptor (e.g., a chimeric antigen receptor (CAR) receptor), a plasmid sequence, and a neoantigen-specific sequence, to name a few. In some cases, the engineered sequence can be used as a diagnostic marker. In some cases, the engineered sequence can be used to determine whether the therapeutic agent is transported to a target, such as a tumor target. The alternative sequence can be exogenous or endogenous. In some cases, the alternative sequence can include or be derived from a plasmid sequence. The plasmid sequence can be DNA or RNA. In some cases, the plasmid sequence can also be a DNA mini-circle sequence or a dog bone sequence.
[0069] The present disclosure also provides methods including obtaining or providing a subgroup of nucleic acid samples or nucleic acid molecules comprising one or more transcriptomes. The nucleic acid sample can include mRNA. The mRNA can be from different tissues of a subject. The amount of the mRNA in the sample can be used to analyze the expression level of mRNA or protein in the tissue and the subject's specific species. For example, the amount of the mRNA in the sample can be relevant to a specific feature or disease (e.g., cancer) in the subject. In addition, the amount of the mRNA in the sample can indicate a change or relative difference in expression in a specific tissue type in the subject. In some cases, mRNA can be processed to form cDNA. For example, reverse transcriptase can be used to reverse transcribe mRNA to synthesize cDNA molecules. cDNA molecules can be captured, separated, enriched, amplified, sequenced, or other reactions that can be performed on nucleic acids, as described elsewhere herein.
[0070] Nucleic acids from biological samples containing subject and non-subject sequences can be enriched and sequenced in a single pool and separated by computer. For example, mixed human genomes and non-human genomes can be separated by alignment with human and non-human species reference sequences. It is possible to separate sequences by alignment because the human genome has diverged from microbial genomes during evolution. Genomes, such as DNA from non-human genomes (such as microbial species) are generally not aligned with human genomes, and vice versa. Non-human genomes, such as microbial genomes, can have segments that are identical or similar to other non-human genomes. More than one non-human genome in a mixture of human and non-human genomes may require additional alignment to identify non-human species.
[0071] In some cases, the method can include producing a subset of nucleic acid molecules from a biological sample. In some cases, the method can include producing a subset of nucleic acid molecules from a biological sample by performing one or more hybridization reactions. The hybridization reaction can include enrichment. In some cases, enrichment can be carried out. Enrichment can be carried out by various methods. In some cases, enrichment can be carried out by hybridization capture, array capture, bead capture, etc. In some cases, hybridization capture can be in solution or on a solid support, such as on an array. In some cases, enrichment can be carried out by molecular inversion probes (MIPs). Enrichment can be carried out by amplification, such as using PCR. In some cases, for the purpose of enrichment, the method can include producing a subset of nucleic acid by amplifying the human genome or non-human genome. In some cases, amplification can include one of the following: polymerase chain reaction (PCR)-based techniques (e.g., solid-phase PCR, RT-PCR, qPCR, multiplex PCR, touchdown PCR, nanoPCR, nested PCR, hot-start PCR, etc.), helicase-dependent amplification (HDA), loop-mediated isothermal amplification (LAMP), self-sustained sequence replication (3SR), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), rolling circle amplification (RCA), ligase chain reaction (LCR), and any other suitable amplification technology.
[0072] The nucleic acid probe pool may comprise a human exome capture kit, such as the Agilent Clinical Research Exome v2 ( Figure 1 ). The nucleic acid probe pool may comprise hybridization probes that are complementary to the human sequences to which they are targeted. The nucleic acid probe pool may comprise approximately 50,000 capture probes to target the genome of most predetermined human genes, etc. Capture probes can be designed to target genes, exons, UTRs, regulatory regions, splice sites, reassembled genes, alternative sequences, and other genomic contents. In some cases, the method may comprise a nucleic acid probe pool designed to target non-human sequences. In some cases, the nucleic acid probe pool may be specific for non-human genomes, such as genomes from viruses, bacteria, fungi, or archaea (i.e., from the human microbiome). In some cases, the nucleic acid probe pool may be combined with a second nucleic acid probe pool that is specific for a second species compared to the first nucleic acid probe pool ( Figure 2 The nucleic acid probe pool can be configured to bind to human and non-human sequences, and the nucleic acid probe pool can bind to sequences from segmented transcriptomes ( Figure 3 In some cases, a pool of nucleic acid probes can bind to human sequences and a pool of nucleic acid probes can bind to non-human sequences ( Figure 4). In some cases, a pool of nucleic acid probes can be used to perform a hybridization-based capture reaction with nucleic acids extracted from a patient sample. In some cases, the method can include sequencing the captured nucleic acid pool. The captured or enriched nucleic acids can be human sequences, non-human sequences, or human sequences and non-human sequences, and the sequencing can include an Illumina NovaSeq-6000 DNA sequencer. In some cases, capture probes targeting microbial species can be designed to target regions of microbial species sequences that are different from one or more species, or directly adjacent to regions that are different from one or more species, such as non-human or human. In some cases, the method includes targeting different regions between non-human sequences (e.g., microbial sequences), thereby enabling the use of a small number of capture probes to capture nucleic acids from a large number of potential non-human microbiome species. By using capture probes that may contain a consensus sequence region but are adjacent to a variable region, many non-human sequences that can be captured can span both.
[0073] In some cases, sequences from non-human genomes can then be assigned to their species of origin based on one or more species-unique regions. For example, the 16S ribosomal RNA gene, present in nearly all bacteria, has approximately nine regions where the sequence varies between species, interspersed with regions of shared sequence. In some cases, capture molecules whose sequences extend from these shared regions to the variable regions can then be assigned to their species of origin based on sequences from portions of the variable regions. Fungal nucleic acid sequences can be similarly evaluated using the partially conserved D2 region of the large subunit ribosomal RNA gene of fungal genomes. The exome of a genome can be analyzed. The intronic regions of the genome can also be analyzed. The exome primarily targets the coding regions of the human genome and can comprise less than 2% of the entire human genome. By excluding most introns and intergenic regions of the human genome, the amount of human sequence can be reduced by approximately 98%. Exome sequencing can be expanded to include non-coding content. Exome sequencing, for example in cancer, can allow for deep sequencing, thereby improving the detection of somatic variants with low allele frequency, as well as the detection of non-human sequences co-captured from a sample.
[0074] The present disclosure also provides compositions and methods for processing biological samples. Biological samples can be obtained from subjects such as adults or children. In some cases, the method for processing biological samples can include (a) using a pool of nucleic acid probes to generate a subset of nucleic acid molecules from a biological sample, wherein the probes include (i) a first plurality of nucleic acid probes configured to target elements of a human genome; and (ii) a second plurality of nucleic acid probes configured to target elements of one or more non-human genomes; (b) the subset of nucleic acid molecules is measured to generate sequence information including the following sequences: (i) human nucleic acids from the biological sample of the subject and (ii) non-human nucleic acids from the biological sample of the subject.
[0075] The disclosed methods can include detecting, monitoring, quantifying or evaluating one or more non-human nucleic acid molecules or one or more diseases or conditions caused by one or more non-human genomes or non-host genomes. In some cases, the capture probes can target different genera. In some cases, the capture probes can target different species. In some aspects, the capture probes can target different orders of more than one organism. In some cases, the capture probes can target plant kingdom, animal kingdom, fungi, protozoa (protest), true bacteria and / or archaebacteria. In some cases, the capture probes can target viruses, bacteria, bacteriophages, fungi, protozoa, archaebacteria, amoebas, worms, algae, genetically modified cells, genetically modified vectors and combinations thereof. In some cases, the non-human sequence can be bacteria. For example, bacterial sequences can be from acidobacteria, actiniobacteria, Aquificae, armatimonadetes, Bacteroidetes, caldiserica, chlamydiae, Chlorobi, chloroflexi, chrysiogenetes, cyanobacteria, deferribacteres, deinococcus-thermus, dictyoglomi, elusimicrobia, Fibrobacteres, Firmicutes, Fusobacteria, Gemmatimonadetes, Lentisphaerae, Nitrospirae, Planctomycetes, Proteobacteria, Spirochaetes, Synergistetes, Tenericutes, Thermodesulfobacteria, Thermomicrobia, Thermotogae, and / or Verrucomicrobia.In some cases, the non-human sequence can be from, but is not limited to, Bordetella, Borrelia, Brucella, Campylobacter, Chlamydia, Chlamydophila, Clostridium, Corynebacterium, Enterococcus, Escherichia, Francisella, Haemophilus, Helicobacter , Legionella (Legionella), Leptospira (Leptospira), Listeria (Listeria), Mycobacterium (Mycobacterium), Mycoplasma (Mycoplasma), Neisseria (Neisseria), Pseudomonas (Pseudomonas), Rickettsia (Rickettsia), Salmonella (Salmonella), Shigella (Shigella), Staphylococcus (Staphylococcus), Streptococcus (Streptococcus), Treponema (Treponema), Vibrio (Vibrio) or Yersinia (Yersinia). Other pathogens include but are not limited to Mycobacterium tuberculosis (Mycobacterium tuberculosis), Streptococcus, Pseudomonas, Shigella, Campylobacter and Salmonella.In some cases, the capture probe can target fungi, such as blastocladiomycota, chytridiomycota, Glomeromycota, Microsporidia, Neocallimastigomycota, Deuteromycota, Ascomycota, Pezizomycotina, Saccharomycotina, Taphrinomycotina, Basidiomycota, Agaricomycotina, Pucciniomycotina, Ustilaginomycotina, Entomophthoromycotina, Kickxellomycotina, Mucoromycotina, Zoopagomycotina, etc. In some cases, the non-human or non-host can be from a cow, horse, fish, donkey, rabbit, rat, mouse, hamster, dog, cat, pig, snake, sheep, goat, etc.
[0076] Diseases or conditions caused by or associated with one or more non-human genomes may include tuberculosis, pneumonia, foodborne illness, tetanus, typhoid fever, diphtheria, syphilis, leprosy, bacterial vaginosis, bacterial meningitis, bacterial pneumonia, urinary tract infection, bacterial gastroenteritis, bacterial skin infection, or any combination thereof. Examples of bacterial skin infections include, but are not limited to, impetigo, which may be caused by Staphylococcus aureus or Streptococcus pyogenes; erysipelas, which may be caused by streptococcal infection of the deep epidermis with lymphatic spread; and cellulitis, which may be caused by normal skin flora or exogenous bacteria.
[0077] The non-subject nucleic acid sequence can be derived from a fungus, such as Candida, Aspergillus, Cryptococcus, Histoplasma, Pneumocystis, and Stachybotrys. Examples of diseases or conditions caused by fungi include, but are not limited to, tinea cruris, yeast infection, ringworm, and tinea pedis.
[0078] In some cases, the nucleic acid sequence of the non-subject or non-host can be from a protist. Protists can include protozoa, protophytes, fungi, and combinations thereof. Protists can be primitive plastid organisms. Primitive plastid organisms can be Rhodophyta or Glaucophyta. Protists can be Sar or Harosa. SAR can be a clade including stramenopiles, alveolates, and Rhizaria (SAR). In addition, the clade SAR can include stramenopiles, alveolates, Apicomplexa, Ciliophora, Dinoflagellata, Foraminifera, Cercozoa, Foraminifera, Radiolaria, and combinations thereof. In some cases, protists can be Excavata. The archaeal worm kingdom can be Euglenozoa, Percolozoa, Metamonada, and combinations thereof. In some cases, the non-host can be Amoebozoa, Hacrobia, Apusozoa, Opisthokonta, and / or Choanozoa.
[0079] Non-subject nucleic acid sequences can be derived from viruses. Examples of viruses include, but are not limited to, adenovirus, coxsackievirus, Epstein-Barr virus, hepatitis viruses (e.g., hepatitis A, B, and C), herpes simplex virus (type 1 and type 2), cytomegalovirus, herpes virus, HIV, influenza virus, measles virus, mumps virus, papillomavirus, parainfluenza virus, poliovirus, respiratory syncytial virus, rubella virus, and varicella-zoster virus. Examples of diseases or conditions caused by viruses include, but are not limited to, colds, influenza, hepatitis, AIDS, chickenpox, rubella, mumps, measles, warts, and polio.
[0080] The non-subject nucleic acid can be derived from a protozoan, such as Acanthamoeba (e.g., A. astronyxis, A. castellanii, A. culbertsoni, A. hatchetti, A. polyphaga, A. rhysodes, A. healyi, A. divionensis), Brachiola (e.g., B. conneri), or a protozoan. connori), B. vesicularum), Cryptosporidium (e.g., C. parvum), Cyclospora (e.g., C. cayetanensis), Encephalitozoon (e.g., E. cuniculi, E. hellem, E. intestinalis), Entamoeba (e.g., E. oeba) (e.g., Entamoeba histolytica), Enterocytozoon (e.g., E. bieneusi), Giardia (e.g., G. lamblia), Isospora (e.g., I. belli), Microsporidium (e.g., M. africanum, M. ceylonensis), Nasella Naegleria (e.g., N. fowleri), Nosema (e.g., N. algerae, N. ocularum), Pleistophora, Trachipleistophora (e.g., T. anthropophthera, T. hominis), and Vittaforma (e.g., V. corneae).
[0081] Can for example extract and / or separate nucleic acid from the biological sample of experimenter by the separation of carrying out cell fraction.In modification, therefore sample processing can comprise one or more of following: lysis sample, the film in destruction sample cell, separate unwanted element (such as RNA, protein) from sample, the nucleic acid (such as DNA) in purification sample to produce the nucleic acid sample of the non-human microorganism group nucleic acid content and human genome nucleic acid content that comprises sample, amplify nucleic acid from this nucleic acid sample, further purify the nucleic acid after the amplification of nucleic acid sample, the nucleic acid after the amplification of nucleic acid sample is sequenced, and any combination thereof.In modification, lysis sample and / or the film in destruction sample cell can comprise the physical method (such as, bead beating, nitrogen decompression, homogenization, ultrasonic treatment) of cell lysis / membrane destruction, it omits some reagent that produces deviation aspect some microorganism species when expressing when sequencing.In addition or alternatively, cracking or destruction can relate to chemical method (such as, use detergent, use solvent, use surfactant etc.).
[0082] In a variation, separating unwanted elements from the sample can include removing RNA using RNase and / or removing proteins using proteases. In a variation, purifying nucleic acids in a sample to produce a nucleic acid sample can include one or more of the following: precipitating nucleic acids from a biological sample (e.g., using an alcohol-based precipitation method); liquid-liquid-based purification techniques (e.g., phenol-chloroform extraction); chromatography-based purification techniques (e.g., column adsorption); purification techniques involving the use of particles configured to bind nucleic acids and bound to binding moieties configured to release nucleic acids (e.g., magnetic beads, buoyant beads, beads having a size distribution, ultrasound-responsive beads, etc.) in the presence of an elution environment (e.g., with an eluent, providing a pH change, providing a temperature change, etc.), and any other suitable purification techniques.
[0083] Can extract and / or separate nucleic acid from biological sample, so that can carry out extraction and separation and / or separation, it can be carried out in the environment (such as sterile laboratory fume hood, sterile room) that any contaminant (such as may affect the nucleic acid in sample or can promote the material of contaminating nucleic acid) is disinfected, can control ambient temperature, control oxygen content, control carbon dioxide content and / or control light exposure (such as, be exposed to ultraviolet light).Extraction can include cracking to destroy cell membrane and promote nucleic acid to release from the cell in biological sample.In a non-limiting example, cracking can include bead milling equipment (such as, tissue lysis instrument), it is configured to be mixed with sample and be used for the bead of the biological content of agitating sample.In some cases, the processing of biological sample can include one or more combination in the following: lysis reagent (such as protease), heating module and any other suitable cracking device.
[0084] In order to separate nucleic acid from the sample of cracking, the non-nucleic acid content of sample is separated with the nucleic acid content of sample.The purification module of sample treatment method can comprise the separation based on power, the separation based on size, the separation based on binding part (for example, have magnetic binding part, have buoyancy binding part etc.), and / or the separation of any other suitable form.For example, the purification operation of method can comprise one or more of following: be convenient to the centrifuge that supernatant extracts, filter (for example filter plate), be configured to the fluid delivery module, washing reagent delivery system, elution reagent delivery system and any other suitable equipment that is used for nucleic acid content in the purified sample that the sample of cracking is combined with the part of binding nucleic acid content and / or sample waste.
[0085] A subset of nucleic acid molecules can be assayed to generate sequence information. The assay that generates sequence information can induce a sequencing reaction. In some cases, sequencing can be directed to RNA. For example, sequencing can be directed to RNA transcripts. RNA sequencing can include any of the following: chromatin isolation by RNA purification (ChlRP-Seq), global run sequencing (GRO-Seq), ribosome profiling sequencing (Ribo-Seq) / ARTseq TM, RNA immunoprecipitation sequencing (RIP-Seq), high-throughput sequencing of CLIP cDNA libraries (HITS-CLIP), cross-linking and immunoprecipitation sequencing (CLIP-Seq), photoactivatable ribonucleoside enhanced cross-linking and immunoprecipitation (PAR-CLIP), single nucleotide resolution CLIP (iCLIP), native extended transcript sequencing (NET-Seq), targeted purification of polysomal mRNA (TRAP-Seq), cross-linking, ligation and hybridization sequencing (CLASH-Seq), parallel analysis of RNA end sequencing (PARE-Seq), genome-wide mapping of uncapped transcripts (GMUCT), transcript isoform sequencing (TIF-Seq), paired-end analysis of TSS (PEAT), and any combination thereof. In some cases, sequencing can include RNA structure. Sequencing of RNA structure may include any of the following: selective 2'-hydroxyl acylation by primer extension sequencing (SHAPE-Seq), parallel analysis of RNA structure (PARS-Seq), fragmented sequencing (FRAG-Seq), CXXC affinity purification sequencing (CAP-Seq), calf intestinal alkaline phosphatase-tobacco acid pyrophosphatase sequencing (CIP-TAP), inosine chemical cleanup sequencing (ICE), m6A-specific methylated RNA immunoprecipitation sequencing (MeRIP-Seq), and any combination thereof. In some cases, sequencing may include low-level RNA detection. Low-level RNA detection may include: digital RNA sequencing, whole transcriptome amplification of single cells (Quartz-Seq), designed primer-based RNA sequencing (DP-Seq), switching mechanism of the 5' end of RNA templates (Smart-Seq), switching mechanism of the 5' end of RNA template version 2 (Smart-Seq2), unique molecular identifiers (UMIs), cellular expression by linear amplification sequencing (CEL-Seq), single-cell labeled reverse transcription sequencing (STRT-Seq), and any combination thereof. In some cases, sequencing can be directed to DNA. DNA sequencing can include low-level DNA detection. DNA sequencing including low-level DNA detection can include at least one of single-molecule molecular inversion probe (smMIP), multiple displacement amplification (MDA), multiple annealing and ring-forming amplification cycles (MALBAC), oligonucleotide selective sequencing (OS-Seq), dual sequencing (Duplex-Seq) and any combination thereof. In some aspects, sequencing can include DNA methylation.DNA methylation may comprise at least one of bisulfite sequencing (BS-Seq), post-bisulfite adapter tagging (PBAT), tagmentation-based whole genome bisulfite sequencing (T-WGBS), oxidative sulfite sequencing (oxBS-Seq), Tet-assisted bisulfite sequencing (TAB-Seq), methylated DNA immunoprecipitation sequencing (MeDIP-Seq), methylation capture (MethylCap) sequencing, methyl binding domain capture (MBDCap) sequencing, reduced representation bisulfite sequencing (RRBS-Seq), and combinations thereof. In some cases, sequencing may include DNA-protein interactions. For example, sequencing including DNA-protein interactions may include: DNase1 hypersensitive site sequencing (DNase-Seq), MNase-assisted nucleosome separation sequencing (MAINE-Seq), chromatin immunoprecipitation sequencing (ChIP-Seq), formaldehyde-assisted regulatory element separation (FAIRE-Seq), assay for transposase accessible chromatin sequencing (ATAC-Seq), chromatin interaction analysis by double-end tag sequencing (ChlA-PET), chromatin conformation capture (Hi-C / 3C-Seq), circular chromatin conformation capture (4-C or 4C-Seq), chromatin conformation capture carbon copy (5-C) and combinations thereof. In some cases, sequencing may include rearrangement. Sequencing of sequence rearrangements may include at least one of: retrotransposon capture sequencing (RC-Seq), transposon sequencing (Tn-Seq) or insertion sequencing (INSeq), translocation capture sequencing (TC-Seq) and combinations thereof.
[0086] Sequencing analysis can include PCR amplification, such as segmentation of the 16S ribosomal RNA gene, and non-targeted metagenomics methods using deep sequencing. In some embodiments, the method can include a method for next-generation amplification and sequencing, comprising: simultaneously amplifying the entire 16S region of each of a group of microorganisms, fragmenting the amplicons of the entire 16S region of each of the group of microorganisms to produce a set of amplicon fragments, and generating an analysis based on the set of amplicon fragments, wherein the analysis includes at least one of microbial population characteristics, microbial species identification, and identified target microbial sequences. In some cases, whole exome sequencing can be utilized.
[0087] In some cases, the method can include performing an alignment at the genetic level of a non-human genome. For example, the alignment can include aligning a 16S sequence relative to an 18S sequence, relative to an ITS sequence, etc. Thus, the output can be used to identify features of interest that can be used to characterize the microbiome of a biological sample, where the feature can be non-human (e.g., the presence of a bacterial genus), genetic (e.g., based on the expression of a specific genetic region and / or sequence), and / or based on any other suitable level.
[0088] In variations, alignment and mapping to a reference non-human genome (e.g., a bacterial genome) (e.g., provided by the National Center for Biotechnology Information) can be performed using an alignment algorithm comprising one or more of the following: the Needleman-Wunsch algorithm, which performs a global alignment of two reads (e.g., a sequenced read and a reference read) with a termination condition based on a score for the global alignment (e.g., in terms of insertions, deletions, matches, mismatches); the Smith-Waterman algorithm, which performs a local alignment of two reads (e.g., a sequenced read and a reference read) and scores the local alignment (e.g., in terms of insertions, deletions, matches, mismatches); the Basic Local Alignment Search tool; Tool (BLAST), which identifies regions of local similarity between sequences (e.g., sequence reads and reference reads); FPGA accelerated alignment tool; BWT indexing using the BWA tool; BWT indexing using the SOAP tool; BWT indexing using the Bowtie tool; sequence search and alignment by a hashing algorithm (SSAHA2) (SSAHA2), which uses word hashing and dynamic programming to map nucleic acid sequencing reads to a genomic reference sequence; and any other suitable alignment algorithm. Mapping of unidentified sequences may also include mapping to reference viral genomes and / or fungal genomes to further identify the viral and / or fungal components of an individual's microbiome. For example, PCR can be performed with multiple markers (e.g., a first marker, a second marker, a third marker, an Nth marker) in parallel or in series and associated with one or more of a bacterial marker, a fungal marker, and a eukaryotic marker. In addition, overlapping reads can be assembled based on the output of the alignment algorithm (e.g., reads generated by paired-end sequencing), or aligned sequence reads can be merged with a reference sequence (e.g., using hidden Markov model banding techniques, using Durbin-Holmes techniques). However, alignment and mapping can implement any other suitable algorithm or technique. In some cases, sequence reads can be encoded to facilitate alignment and mapping operations. In one example, each base of the sequence can be encoded as a byte according to the arrangement 0000TGCA, where the least significant bit is 1 if the base is sequenced as likely to contain base A (e.g., A is represented as 00000001); the next significant bit is 1 if the base is sequenced as likely to contain base C (e.g., C is represented as 00000010); the next significant bit is 1 if the base is sequenced as likely to contain base G (e.g., G is represented as 00000100); and the next significant bit is 1 if the sequence is sequenced as likely to contain base T (e.g., T is represented as 00001000). In this example, the four most significant bits are set to zero.However, alternative variations of the examples may encode the bases in any other suitable manner. Additionally, the predetermined primer sequences used during amplification may be used to trim sequence reads to omit the primer sequences to improve the efficiency of alignment and mapping.
[0089] The subset of nucleic acid molecules may comprise one or more genomes disclosed herein. The subset of nucleic acid molecules may comprise 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 20 or more, 25 or more, 30 or more, 35 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, or 100 or more genomes. The one or more genomes may be the same, similar, different, or a combination thereof. In some cases, there are two subsets of nucleic acid molecules, Figure 6A .
[0090] A subset of nucleic acid molecules may comprise one or more genomic features disclosed herein. A subset of nucleic acid molecules may comprise 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 20 or more, 25 or more, 30 or more, 35 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, or 100 or more genomic features. One or more genomic features may be the same, similar, different, or a combination thereof.
[0091] Nucleic acid molecule subsets may contain nucleic acid molecules of different sizes. The length of the nucleic acid molecules in a nucleic acid molecule subset may be referred to as the size of the nucleic acid molecules. The average length of the nucleic acid molecules in a nucleic acid molecule subset may be referred to as the average size of the nucleic acid molecules. As used herein, the terms "size of the nucleic acid molecule," "average size of the nucleic acid molecules," "molecular size," and "average molecular size" may be used interchangeably. The size of the nucleic acid molecules may be used to distinguish between two or more nucleic acid molecule subsets. The difference between the average size of the nucleic acid molecules in a nucleic acid molecule subset and the average size of the nucleic acid molecules in another nucleic acid molecule subset may be used to distinguish between the two nucleic acid molecule subsets. The average size of the nucleic acid molecules in one nucleic acid molecule subset may be greater than the average size of the nucleic acid molecules in at least one other nucleic acid molecule subset. The average size of the nucleic acid molecules in one nucleic acid molecule subset may be less than the average size of the nucleic acid molecules in at least one other nucleic acid molecule subset. The difference in average molecular size between two or more subsets of nucleic acid molecules may be at least about 50; 75; 100; 125; 150; 175; 200; 225; 250; 275; 300; 350; 400; 450; 500; 550; 600; 650; 700; 750; 800; 850; 900; 950; 1,000; 1100; 1200; 1300; 1400; 1500; 16 In some aspects of the present disclosure, the difference in average molecular size between two or more subsets of nucleic acid molecules is at least about 200 bases or base pairs. Alternatively, the difference in average molecular size between two or more subsets of nucleic acid molecules is at least about 300 bases or base pairs.
[0092] The nucleic acid molecule subsets may comprise nucleic acid molecules of different sequencing sizes. The length of the nucleic acid molecules in the nucleic acid molecule subset to be sequenced may be referred to as the sequencing size of the nucleic acid molecule. The average length of the nucleic acid molecules in the nucleic acid molecule subset may be referred to as the average sequencing size of the nucleic acid molecule. As used herein, the terms "sequencing size of a nucleic acid molecule," "average sequencing size of a nucleic acid molecule," "molecular sequencing size," and "average molecular sequencing size" may be used interchangeably. The average molecular sequencing size of one or more nucleic acid molecule subsets may be at least about 50; 75; 100; 125; 150; 175; 200; 225; 250; 275; 300; 350; 400; 450; 500; 550; 600; 650; 700; 750; 800; 850; 900; 950; 1,000; 1100; 1200; 1300; 1400; 1500; 160 0; 1700; 1800; 1900; 2,000; 3,000; 4,000; 5,000; 6,000; 7,000; 8,000; 9,000; 10,000; 15,000; 20,000; 30,000; 40,000; 50,000; 60,000; 70,000; 80,000; 90,000; 100,000 or more bases or base pairs. The sequenced size of the nucleic acid molecules can be used to distinguish between two or more subgroups of nucleic acid molecules. The difference between the average sequenced size of the nucleic acid molecules in a subgroup of nucleic acid molecules and the average sequenced size of the nucleic acid molecules in another subgroup of nucleic acid molecules can be used to distinguish between the two subgroups of nucleic acid molecules. The average sequenced size of the nucleic acid molecules in one subgroup of nucleic acid molecules can be greater than the average sequenced size of the nucleic acid molecules in at least one other subgroup of nucleic acid molecules. The average sequenced size of the nucleic acid molecules in one subset of nucleic acid molecules can be less than the average sequenced size of the nucleic acid molecules in at least one other subset of nucleic acid molecules. The difference in average sequenced size between two or more subsets of nucleic acid molecules can be at least about 50; 75; 100; 125; 150; 175; 200; 225; 250; 275; 300; 350; 400; 450; 500; 550; 600; 650; 700; 750; 800; 850; 900; 950; 1,000; 1100; 1200; 1300; 1400; 1500; 1600; 1700; 1800; 1900; 2,000; 3,000; 4,000; 5,000; 6,000; 7,000; 8,000; 9,000; 10,000; 15,000; 20,000; 30,000; 40,000; 50,000; 60,000; 70,000; 80,000; 90,000; 100,000 or more bases or base pairs.In some aspects of the present disclosure, the difference in average molecular sequencing size between two or more subsets of nucleic acid molecules is at least about 200 bases or base pairs. Alternatively, the difference in average molecular sequencing size between two or more subsets of nucleic acid molecules is at least about 300 bases or base pairs.
[0093] The methods disclosed herein may include one or more capture probes, multiple capture probes, or one or more capture probe groups. Typically, the capture probe comprises a nucleic acid binding site. The capture probe can hybridize with the captured nucleic acid. The capture probe can comprise a nucleic acid sequence that is complementary to the captured nucleic acid. In some cases, the capture probe can comprise a nucleic acid sequence that is completely complementary to a portion of the captured nucleic acid. For example, each nucleic acid in the capture probe can be complementary to a base in the captured nucleic acid. The capture probe can be longer than the captured nucleic acid. For example, each base in the captured nucleic acid can be complementary to a base in the capture probe, but not all bases in the capture probe are complementary to bases in the captured nucleic acid. The capture probe can be shorter than the captured nucleic acid. For example, each base in the capture probe can be complementary to a base in the captured nucleic acid, but not all bases in the captured nucleic acid are complementary to the capture probe.
[0094] Capture probe can carry out capture reaction in solution.Capture probe can be in solution and can capture nucleic acid in solution.As described elsewhere herein, the nucleic acid captured can be separated and / or eluted subsequently.Capture probe can capture nucleic acid in solution, and then capture probe can be subsequently connected to support, such as solid support (for example, array or beads).In some cases, support can be formed by semi-solid material (for example, gel).
[0095] The attachment to the support can be non-covalent attachment. For example, nucleic acid can be captured by a biotinylated probe, which is then combined with avidin / streptavidin beads to attach the captured nucleic acid complex to the avidin / streptavidin beads. Other binding pairs can be used to attach the capture probe to the surface. The capture probe can be covalently attached to the support. For example, the support can have a chemically reactive linker that can react with the capture probe so that the capture probe is covalently linked to the carrier.
[0096] Capture probe can be attached to support and carry out capture reaction.For example, capture probe can be connected with support and capture nucleic acid molecule subsequently.Support can be solid or semisolid (for example gel) material.The example of support includes but is not limited to bead, slide and fragment.Support can be for example glass, silicon dioxide, silicon, plastic (for example polystyrene), agar or agarose.
[0097] The capture probe can also include one or more linkers. The capture probe can also include one or more labels. One or more linkers can attach one or more labels to the nucleic acid binding site. In some cases, the capture probe can be designed to hybridize with the shared region of the 16S gene sequence. The capture probe that can be designed to target the shared region of the 16S gene sequence can be used to capture nucleic acid molecules from a variety of species, even species that have not yet been identified and characterized. In some cases, the method can include a first plurality of nucleic acid probes configured to target an element of a human genome sequence. In some cases, the method can include a second plurality of nucleic acid probes configured to target an element of a genome sequence of a non-human species.
[0098] The methods disclosed herein can include 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, In some embodiments, the method comprises the steps of: a) using a plurality of capture probes or a plurality of capture probes; b) using a plurality of capture probes or a plurality of capture probe groups; c) using a plurality of capture probes or a plurality of capture probe groups; and d) using a plurality of capture probes or a plurality of capture probe groups. In some embodiments, the method comprises the steps of: a) using a plurality of capture probes or a plurality of capture probe groups; and b) using a plurality of capture probes or a plurality of capture probe groups. In some embodiments, the method comprises the steps of: a) using a plurality of capture probes or a plurality of capture probe groups; and c) using a plurality of capture probes or a plurality of capture probe groups. For example, the concentration of certain capture probes can be higher to capture nucleic acids that may be difficult to capture (e.g., sequences with high GC content, sequences resulting from sequence recombination, sequences with a large number of mutations, sequences with a high mutation rate), thereby increasing the likelihood of capturing a specific nucleic acid.
[0099] The one or more capture probes may comprise a nucleic acid binding site that hybridizes to at least a portion of one or more nucleic acid molecules, or variants thereof, or derivatives thereof in a sample of nucleic acid molecules or a subset of nucleic acid molecules. The capture probe may comprise a nucleic acid binding site that hybridizes to one or more genomes. The capture probe may hybridize to different, similar, and / or identical genomes. The one or more capture probes may be at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99% or more complementary to one or more nucleic acid molecules, or variants thereof, or derivatives thereof.
[0100] The capture probe can comprise one or more nucleotides. The capture probe can comprise 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more or 1000 or more nucleotides. The capture probe can comprise about 100 nucleotides. The capture probe may comprise from about 10 to about 500 nucleotides, from about 20 to about 450 nucleotides, from about 30 to about 400 nucleotides, from about 40 to about 350 nucleotides, from about 50 to about 300 nucleotides, from about 60 to about 250 nucleotides, from about 70 to about 200 nucleotides, or from about 80 to about 150 nucleotides. In some aspects of the present disclosure, the capture probe comprises from about 80 nucleotides to about 100 nucleotides.
[0101] Multiple capture probes or capture probe groups may comprise two or more capture probes having identical, similar, and / or different nucleic acid binding site sequences, linkers, and / or labels. For example, two or more capture probes comprise the same nucleic acid binding site. In another example, two or more capture probes comprise similar nucleic acid binding sites. In another example, two or more capture probes comprise different nucleic acid binding sites. Two or more capture probes may also comprise one or more linkers. Two or more capture probes may also comprise different linkers. Two or more capture probes may also comprise similar linkers. Two or more capture probes may also comprise the same linker. Two or more capture probes may also comprise one or more labels. Two or more capture probes may also comprise different labels. Two or more capture probes may also comprise similar labels. Two or more capture probes may also comprise the same label.
[0102] Assays may include, but are not limited to, sequencing, amplification, hybridization, enrichment, separation, elution, fragmentation, detection, quantification of one or more nucleic acid molecules. Assays may include methods for preparing one or more nucleic acid molecules. Assays may include conventional assays, long reads, high GC content, and hybridization assays. Figure 9 Any number of assays may be performed. The plurality of assays may be 1, 2, 3, 4, 5, 6, 7, 8, 9 or up to 10 assays. For example, Figure 6B A schematic diagram showing the use of two assays is shown. Similarly, any number of analyses can be performed on data from one or more assays. The number of analyses can be 1, 2, 3, 4, 5, 6, 7, 8, 9, or up to 10 analyses of data from one or more assays. For example, Figure 6C Two different analyses are shown being performed. In some cases, the analysis is a bioinformatics analysis, Figure 8 Any number of protocols can be used for the assay. For example, 1, 2, 3, 4, 5, 6, 7, 8, 9, or up to 10 protocols can be used. Figure 6D Schematic diagrams utilizing four schemes are provided.
[0103] Method disclosed herein can include carrying out one or more sequencing reactions to one or more nucleic acid molecules in a sample. Method disclosed herein can include carrying out 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 15 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 200 or more, 300 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more or 1000 or more sequencing reactions to one or more nucleic acid molecules in a sample. Sequencing reactions can be carried out simultaneously, sequentially or in combination. Sequencing reactions can include whole genome sequencing or exome sequencing. Sequencing reactions can include Maxim-Gilbert, chain termination or high throughput systems. Alternatively or additionally, sequencing reactions can include Helioscope TM Single-molecule sequencing, nanopore DNA sequencing, Lynx Therapeutics' massively parallel signature sequencing (MPSS), 454 pyrosequencing, single-molecule real-time (RNAP) sequencing, Illumina (Solexa) sequencing, SOLiD sequencing, Ion Torrent TM, ion semiconductor sequencing, single molecule SMRT (TM) sequencing, Polony sequencing, DNA nanoball sequencing, VisiGen Biotechnologies method or a combination thereof. Alternatively or additionally, the sequencing reaction may comprise one or more sequencing platforms, including but not limited to the Genome Analyzer IIx, HiSeq and MiSeq provided by Illumina, single molecule real-time (SMRT) sequencing ... TM ) technologies (e.g., the PacBioRS system provided by Pacific Biosciences (California) and Solexa Sequencer), true single-molecule sequencing (tSMS TM ) technologies (e.g., HeliScope provided by Helicos Inc. (Cambridge, MA) TM Sequencer). The sequencing reaction may also include an electron microscope or a chemically sensitive field effect transistor (chemFET) array. In some aspects of the present disclosure, the sequencing reaction includes capillary sequencing, next generation sequencing, Sanger sequencing, sequencing by synthesis, sequencing by ligation, sequencing by hybridization, single molecule sequencing, or a combination thereof. Sequencing by synthesis may include reversible terminator sequencing, progressive single molecule sequencing, sequential flow sequencing, or a combination thereof. Sequential flow sequencing may include pyrosequencing, pH-mediated sequencing, semiconductor sequencing, or a combination thereof.
[0104] The methods disclosed herein can include performing at least one long-read sequencing reaction and at least one short-read sequencing reaction. Figure 18 Examples of methods including long-read and short-read sequencing are shown in . A long-read sequencing reaction and / or a short-read sequencing reaction can be performed on at least a portion of a subset of nucleic acid molecules. A long-read sequencing reaction and / or a short-read sequencing reaction can be performed on at least a portion of two or more subsets of nucleic acid molecules. A long-read sequencing reaction and a short-read sequencing reaction can be performed on at least a portion of one or more subsets of nucleic acid molecules.
[0105] Sequencing of one or more nucleic acid molecules or a subset thereof may comprise at least about 5; 10; 15; 20; 25; 30; 35; 40; 45; 50; 60; 70; 80; 90; 100; 200; 300; 400; 500; 600; 700; 800; 900; 1,000; 1,500; 2,000; 2,500; 3,000; 3,500; 4,000; 4,500; 5,000; 5500; 6,000; 6500; 7,000; 7500; 8 or more sequencing reads.
[0106] The sequencing reaction may comprise at least about 50; 60; 70; 80; 90; 100; 110; 120; 130; 140; 150; 160; 170; 180; 190; 200; 210; 220; 230; 240; 250; 260; 270; 280; 290; 300; 325; 350; 375; 400; 425; 450; 475; 500; 600; 700; 800; 900; 1,000 ; 1500; 2,000; 2500; 3,000; 3500; 4,000; 4500; 5,000; 5500; 6,000; 6500; 7,000; 7500; 8,000; 8500; 9,000; 10,000; 20,000; 30,000; 40,000; 50,000; 60,000; 70,000; 80,000; 90,000; 100,000 or more bases or base pairs are sequenced. The sequencing reaction may comprise at least about 50;60;70;80;90;100;110;120;130;140;150;160;170;180;190;200;210;220;230;240;250;260;270;280;290;300;325;350;375;400;425;450;475;500;600;700;800;900;1,000; 1500; 2,000; 2500; 3,000; 3500; 4,000; 4500; 5,000; 5500; 6,000; 6500; 7,000; 7500; 8,000; 8500; 9,000; 10,000; 20,000; 30,000; 40,000; 50,000; 60,000; 70,000; 80,000; 90,000; 100,000 or more consecutive bases or base pairs are sequenced.
[0107] The sequencing technology used in the methods of the present disclosure can produce at least 100 reads per run, at least 200 reads per run, at least 300 reads per run, at least 400 reads per run, at least 500 reads per run, at least 600 reads per run, at least 700 reads per run, at least 800 reads per run, at least 900 reads per run, at least 1000 reads per run, at least 5,000 reads per run, at least 10,000 reads per run, at least 50,000 reads per run, at least 100,000 reads per run, at least 500,000 reads per run, or at least 1,000,000 reads per run. Alternatively, the sequencing technology used in the methods of the present disclosure can produce at least 1,500,000 reads per run, at least 2,000,000 reads per run, at least 2,500,000 reads per run, at least 3,000,000 reads per run, at least 3,500,000 reads per run, at least 4,000,000 reads per run, at least 4,500,000 reads per run, or at least 5,000,000 reads per run.
[0108] The sequencing technology used in the methods of the present disclosure can generate at least about 30 base pairs, at least about 40 base pairs, at least about 50 base pairs, at least about 60 base pairs, at least about 70 base pairs, at least about 80 base pairs, at least about 90 base pairs, at least about 100 base pairs, at least about 110, at least about 120 base pairs, at least about 150 base pairs, at least about 200 base pairs, at least about 250 base pairs, at least about 300 base pairs, at least about 350 base pairs, at least about 400 base pairs, at least about 450 base pairs, at least about 500 base pairs, at least about 550 base pairs, about 600 base pairs, at least about 700 base pairs, at least about 800 base pairs, at least about 900 base pairs, or at least about 1,000 base pairs per read. Alternatively, the sequencing technology used in the methods of the present disclosure can generate long sequencing reads. In some cases, the sequencing technology used in the methods of the present disclosure can generate at least about 1200 base pairs / read, at least about 1500 base pairs / read, at least about 1800 base pairs / read, at least about 2000 base pairs / read, at least about 2500 base pairs / read, at least about 3,000 base pairs / read, at least about 3500 base pairs / read, at least about 4,000 base pairs / read, at least about 4,500 base pairs / read, at least about 5,000 base pairs / read, at least about 6,000 base pairs / read, or at least about 7,000 base pairs / read. reads, at least about 7,000 base pairs / read, at least about 8,000 base pairs / read, at least about 9,000 base pairs / read, at least about 10,000 base pairs / read, 20,000 base pairs / read, 30,000 base pairs / read, 40,000 base pairs / read, 50,000 base pairs / read, 60,000 base pairs / read, 70,000 base pairs / read, 80,000 base pairs / read, 90,000 base pairs / read, or 100,000 base pairs / read.
[0109] High-throughput sequencing systems can allow for the detection of sequenced nucleotides immediately after or when the sequenced nucleotides are incorporated into the growing chain, i.e., real-time or substantially real-time sequence detection. In some cases, high-throughput sequencing generates at least 1,000, at least 5,000, at least 10,000, at least 20,000, at least 30,000, at least 40,000, at least 50,000, at least 100,000, or at least 500,000 sequence reads per hour; wherein each read is at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 120, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, or at least 500 bases / read. Sequencing can be performed using nucleic acids described herein, such as genomic DNA, cDNA or RNA derived from RNA transcripts as templates.
[0110] Method disclosed herein can include performing one or more amplification reactions on one or more nucleic acid molecules in a sample.Term "amplification" refers to any process of producing at least one copy of a nucleic acid molecule.Term "amplicon" and "amplified nucleic acid molecule" refer to copies of nucleic acid molecules and can be used interchangeably.Amplification reactions can include PCR-based methods, non-PCR-based methods, or combinations thereof.Examples based on non-PCR methods include, but are not limited to, multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, rolling circle amplification, or ring-ring amplification.PCR-based methods can include, but are not limited to, PCR, HD-PCR, next-generation PCR, digital RTA, or any combination thereof. Other PCR methods include, but are not limited to, linear amplification, allele-specific PCR, Alu PCR, assembly PCR, asymmetric PCR, droplet PCR, emulsion PCR, helicase-dependent amplification HDA, hot-start PCR, inverse PCR, linear-after-the-exponential (LATE)-PCR, long PCR, multiplex PCR, nested PCR, semi-nested PCR, quantitative PCR, RT-PCR, real-time PCR, single-cell PCR, and touchdown PCR.
[0111] The methods disclosed herein can include performing one or more hybridization reactions on one or more nucleic acid molecules in a sample. The hybridization reaction can include hybridization of one or more capture probes with one or more nucleic acid molecules in a sample or a subset of nucleic acid molecules. The hybridization reaction can include hybridizing one or more capture probe groups with one or more nucleic acid molecules in a sample or a subset of nucleic acid molecules. The hybridization reaction can include one or more hybridization arrays, multiple hybridization reactions, hybridization chain reactions, isothermal hybridization reactions, nucleic acid hybridization reactions, or a combination thereof. The one or more hybridization arrays can include hybridization array genotyping, hybridization array ratio sensing, DNA hybridization arrays, macroarrays, microarrays, high-density oligonucleotide arrays, genomic hybridization arrays, comparative hybridization arrays, or a combination thereof. The hybridization reaction can include one or more capture probes, one or more beads, one or more labels, one or more nucleic acid molecule subsets, one or more nucleic acid samples, one or more reagents, one or more wash buffers, one or more elution buffers, one or more hybridization buffers, one or more hybridization chambers, one or more incubators, one or more separators, or a combination thereof.
[0112] Methods disclosed herein can include carrying out one or more enrichment reactions to one or more nucleic acid molecules in a sample. The enrichment reaction can include contacting the sample with one or more beads or bead groups. The enrichment reaction can include differential amplification of two or more nucleic acid molecule subgroups based on one or more genomic features. For example, the enrichment reaction includes differential amplification of two or more nucleic acid molecule subgroups based on GC content. Alternatively or additionally, the enrichment reaction includes differential amplification of two or more nucleic acid molecule subgroups based on methylation status. The enrichment reaction can include one or more hybridization reactions. The enrichment reaction can also include separation and / or purification of one or more hybridized nucleic acid molecules, one or more bead-bound nucleic acid molecules, one or more free nucleic acid molecules (e.g., nucleic acid molecules without capture probes, nucleic acid molecules without beads), one or more labeled nucleic acid molecules, one or more unlabeled nucleic acid molecules, one or more amplicons, one or more unamplified nucleic acid molecules, or a combination thereof. Alternatively or additionally, the enrichment reaction can include one or more cell types in the enriched sample. The one or more cell types can be enriched by flow cytometry.
[0113] One or more enrichment reactions can produce one or more enriched nucleic acid molecules. The enriched nucleic acid molecules can comprise nucleic acid molecules or their variants or derivatives. For example, the enriched nucleic acid molecules include one or more hybridized nucleic acid molecules, one or more bead-bound nucleic acid molecules, one or more free nucleic acid molecules (e.g., nucleic acid molecules without capture probes, nucleic acid molecules without beads), one or more labeled nucleic acid molecules, one or more unlabeled nucleic acid molecules, one or more amplicons, one or more unamplified nucleic acid molecules, or a combination thereof. The enriched nucleic acid molecules can be distinguished from non-enriched nucleic acid molecules by GC content, molecular size, genome, genomic characteristics, or a combination thereof. The enriched nucleic acid molecules can be derived from one or more assays, supernatants, eluates, or a combination thereof. The enriched nucleic acid molecules can be different from non-enriched nucleic acid molecules in terms of average size, average GC content, genome, or a combination thereof. In some cases, enrichment can include multiple DNA subsets enriched for different genomic regions, which undergo independent processing operations before being combined for sequencing assays, Figure 15 In some cases, enrichment can include multiple DNA subsets enriched for different genomic regions that undergo independent processing operations before being independently sequenced and analyzed. Figure 16 .
[0114] The methods disclosed herein may include performing one or more separation or purification reactions on one or more nucleic acid molecules in a sample. The separation or purification reaction may include contacting the sample with one or more beads or bead groups. The separation or purification reaction may include one or more hybridization reactions, enrichment reactions, amplification reactions, sequencing reactions, or a combination thereof. The separation or purification reaction may include the use of one or more separators. The one or more separators may include a magnetic separator. The separation or purification reaction may include separating bead-bound nucleic acid molecules from bead-free nucleic acid molecules. The separation or purification reaction may include separating capture probe-hybridized nucleic acid molecules from capture probe-free nucleic acid molecules. The separation or purification reaction may include separating a first subset of nucleic acid molecules from a second subset of nucleic acid molecules, wherein the first subset of nucleic acid molecules differs from the second subset of nucleic acid molecules in average size, average GC content, genome, or a combination thereof.
[0115] The methods disclosed herein can include performing one or more elution reactions on one or more nucleic acid molecules in a sample. The elution reactions can include contacting the sample with one or more beads or sets of beads. The elution reactions can include separating bead-bound nucleic acid molecules from bead-free nucleic acid molecules. The elution reactions can include separating capture probe-hybridized nucleic acid molecules from capture probe-free nucleic acid molecules. The elution reactions can include separating a first subset of nucleic acid molecules from a second subset of nucleic acid molecules, wherein the first subset of nucleic acid molecules differs from the second subset of nucleic acid molecules in average size, average GC content, genome, or a combination thereof.
[0116] Method disclosed herein can include one or more fragmentation reactions.Fragmentation reaction can include making one or more nucleic acid molecule fragmentation in the sample of nucleic acid molecule or subgroup to produce one or more fragmented nucleic acid molecules.One or more nucleic acid molecules can be by ultrasonic treatment, pin shearing, atomization, shearing (such as acoustic shearing, mechanical shearing, point-groove shearing (point-sink shearing)), by the passage or enzymatic digestion of French pressure cell and carry out fragmentation.Enzymatic digestion can be carried out (such as, micrococcal nuclease digestion, nuclease endonuclease, exonuclease, RNase H or DNA enzyme I) by nuclease digestion.The fragmentation of one or more nucleic acid molecules can produce the fragment size of about 100 base pairs to about 2000 base pairs, about 200 base pairs to about 1500 base pairs, about 200 base pairs to about 1000 base pairs, about 200 base pairs to about 500 base pairs, about 500 base pairs to about 1500 base pairs and about 500 base pairs to about 1000 base pairs. The one or more fragmentation reactions can produce fragment sizes of about 50 base pairs to about 1000 base pairs. The one or more fragmentation reactions can produce fragment sizes of about 100 base pairs, 150 base pairs, 200 base pairs, 250 base pairs, 300 base pairs, 350 base pairs, 400 base pairs, 450 base pairs, 500 base pairs, 550 base pairs, 600 base pairs, 650 base pairs, 700 base pairs, 750 base pairs, 800 base pairs, 850 base pairs, 900 base pairs, 950 base pairs, 1000 base pairs or more.
[0117] Fragmenting one or more nucleic acid molecules can include mechanically shearing one or more nucleic acid molecules in a sample over a period of time. The fragmentation reaction can occur for at least about 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500 or more seconds.
[0118] Fragmenting the one or more nucleic acid molecules can comprise contacting the nucleic acid sample with one or more beads. Fragmenting the one or more nucleic acid molecules can comprise contacting the nucleic acid sample with a plurality of beads, wherein the ratio of the volume of the plurality of beads to the volume of the nucleic acid sample is about 0.10, 0.20, 0.30, 0.40, 0.50, 0.60, 0.70, 0.80, 0.90, 1.00, 1.10, 1.20, 1.30, 1.40, 1.50, 1.60, 1.70, 1.80, 1.90, 2.00 or more. Fragmenting one or more nucleic acid molecules can comprise contacting the nucleic acid sample with a plurality of beads, wherein the ratio of the volume of the plurality of beads to the volume of the nucleic acid is about 2.00, 1.90, 1.80, 1.70, 1.60, 1.50, 1.40, 1.30, 1.20, 1.10, 1.00, 0.90, 0.80, 0.70, 0.60, 0.50, 0.40, 0.30, 0.20, 0.10, 0.05, 0.04, 0.03, 0.02, 0.01 or less.
[0119] The methods disclosed herein can include performing one or more detection reactions on one or more nucleic acid molecules in a sample. The detection reactions can include one or more sequencing reactions. Alternatively, performing the detection reactions can include optical sensing, electrical sensing, or a combination thereof. Optical sensing can include optical sensing of photoluminescent photon emission, fluorescent photon emission, pyrophosphate photon emission, chemiluminescent photon emission, or a combination thereof. Electrical sensing can include electrical sensing of ion concentration, ion current modulation, nucleotide electric field, nucleotide tunneling current, or a combination thereof.
[0120] The methods disclosed herein can include performing one or more quantitative reactions on one or more nucleic acid molecules in a sample. The quantitative reactions can include sequencing, PCR, qPCR, digital PCR, or a combination thereof.
[0121] Methods disclosed herein can include one or more samples. Methods disclosed herein can include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more samples. The sample can be derived from a subject. Two or more samples can be derived from a single subject. Two or more samples can be derived from 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more different subjects. The subject can be a mammal, a reptile, an amphibian, a bird, and a fish. Mammals can be humans, apes, gorillas, monkeys, chimpanzees, cows, pigs, horses, rodents, dogs, cats, or other animals. Reptiles can be lizards, snakes, alligators, turtles, crocodiles, and tortoises. Amphibians can be toads, frogs, newts, and salamanders. Examples of birds include, but are not limited to, ducks, geese, penguins, ostriches, and owls. Examples of fish include, but are not limited to, catfish, eels, sharks, and swordfish. The subject can be human. The subject can have a disease or condition.
[0122] Two or more samples can be collected at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 or more time points. The time points can occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 or more hours. 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 or more days. The time points can occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 or more weeks. 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 or more months. The time points can occur over a period of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 or more years.
[0123] In some cases, the method can include obtaining a biological sample from a subject. The subject can be human or non-human. The subject can be an adult or a child. In some cases, an adult subject can be 18 years old or older. In some cases, the subject's biological sample can be derived from a tumor biopsy, whole blood, or plasma. In some cases, the biological sample can be from body fluids, cells, skin, tissue, organs, or a combination thereof. The sample can be blood, plasma, blood fractions, saliva, sputum, urine, semen, vaginal fluid, cerebrospinal fluid, feces, cells, or a tissue biopsy. The sample can be from the adrenal gland, appendix, bladder, brain, ear, esophagus, eye, gallbladder, heart, kidney, large intestine, liver, lung, mouth, muscle, nose, pancreas, parathyroid gland, pineal gland, pituitary gland, skin, small intestine, spleen, stomach, thymus, thyroid gland, trachea, uterus, worm-like appendix, cornea, skin, heart valve, artery, or vein.
[0124] The sample may comprise one or more nucleic acid molecules. The nucleic acid molecule may be a DNA molecule, an RNA molecule (e.g., mRNA, cRNA, or miRNA), and a DNA / RNA hybrid. Examples of DNA molecules include, but are not limited to, double-stranded DNA, single-stranded DNA, single-stranded DNA hairpins, cDNA, and genomic DNA. The nucleic acid may be an RNA molecule, such as double-stranded RNA, single-stranded RNA, ncRNA, RNA hairpins, and mRNA. Examples of ncRNA include, but are not limited to, siRNA, miRNA, snoRNA, piRNA, tiRNA, PASR, TASR, aTASR, TSSa-RNA, snRNA, RE-RNA, uaRNA, x-ncRNA, hYRNA, usRNA, snaR, and vtRNA.
[0125] Method disclosed herein can comprise one or more containers. Method disclosed herein can comprise 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more or 1000 or more containers. One or more containers can be different, similar, identical, or a combination thereof. The example of container includes but is not limited to plate, microplate, PCR plate, hole, microwell, tube, Eppendorf tube, bottle, array, microarray and fragment.
[0126] The methods disclosed herein can include one or more reagents. The methods disclosed herein can include 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more or 1000 or more reagents. One or more reagents can be different, similar, identical, or a combination thereof. Reagents can improve the efficiency of one or more assays. Reagents can improve the stability of nucleic acid molecules or their variants or derivatives. Reagents can include, but are not limited to, enzymes, proteases, nucleases, molecules, polymerases, reverse transcriptases, ligases, and compounds. Methods disclosed herein may include performing an assay comprising one or more antioxidants. Typically, an antioxidant is a molecule that inhibits the oxidation of another molecule. Examples of antioxidants include, but are not limited to, ascorbic acid (e.g., vitamin C), glutathione, lipoic acid, uric acid, carotene, alpha-tocopherol (e.g., vitamin E), ubiquinol (e.g., coenzyme Q), and vitamin A.
[0127] Method disclosed herein can comprise one or more buffers or solutions. Method disclosed herein can comprise 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more or 1000 or more buffers or solutions. One or more buffers or solutions can be different, similar, identical, or a combination thereof. Buffers or solutions can improve the efficiency of one or more assays. Buffers or solutions can improve the stability of nucleic acid molecules or their variants or derivatives. Buffers or solutions may include, but are not limited to, wash buffers, elution buffers, and hybridization buffers.
[0128] Methods disclosed herein can include one or more beads, multiple beads, or one or more bead groups. Methods disclosed herein can include 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more one or more beads or bead groups. One or more beads or bead groups can be different, similar, identical, or a combination thereof. The beads can be magnetic, antibody coated, protein A cross-linked, protein G cross-linked, streptavidin coated, oligonucleotide conjugated, silica coated, or a combination thereof. Examples of beads include, but are not limited to, Ampure beads, AMPure XP beads, streptavidin beads, agarose beads, magnetic beads, microbeads, antibody-conjugated beads (e.g., anti-immunoglobulin microbeads), protein A-conjugated beads, protein G-conjugated beads, protein A / G-conjugated beads, protein L-conjugated beads, oligo-dT-conjugated beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorochrome microbeads, and BcMag TM Carboxyl-terminated magnetic beads. In some aspects of the present disclosure, the one or more beads include one or more Ampure beads. Alternatively or additionally, the one or more beads include AMPure XP beads.
[0129] The methods disclosed herein can include one or more primers, multiple primers, or one or more primer sets. Primers can also include one or more linkers. Primers can also include one or more labels. Primers can be used in one or more assays. For example, primers are used in one or more sequencing reactions, amplification reactions, or combinations thereof. The methods disclosed herein can include 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more, or 1000 or more one or more primers or primer sets. The primer can comprise about 100 nucleotides. The primer can comprise about 10 to about 500 nucleotides, about 20 to about 450 nucleotides, about 30 to about 400 nucleotides, about 40 to about 350 nucleotides, about 50 to about 300 nucleotides, about 60 to about 250 nucleotides, about 70 to about 200 nucleotides or about 80 to about 150 nucleotides. In some aspects of the present disclosure, the primer comprises about 80 nucleotides to about 100 nucleotides. One or more primers or primer sets can be different, similar, identical, or its combination.
[0130] Primers can hybridize to at least a portion of one or more nucleic acid molecules, variants thereof, or derivatives thereof in a sample or subset of nucleic acid molecules. Primers can hybridize to one or more genomes. Primers can hybridize to different, similar, and / or identical genomes. One or more primers can be complementary to one or more nucleic acid molecules, variants thereof, or derivatives thereof by at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99%, or more.
[0131] The primer can comprise one or more nucleotides. The primer can comprise 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more or 1000 or more nucleotides. The primer can comprise about 100 nucleotides. The primer may comprise from about 10 to about 500 nucleotides, from about 20 to about 450 nucleotides, from about 30 to about 400 nucleotides, from about 40 to about 350 nucleotides, from about 50 to about 300 nucleotides, from about 60 to about 250 nucleotides, from about 70 to about 200 nucleotides, or from about 80 to about 150 nucleotides. In some aspects of the present disclosure, the primer comprises from about 80 nucleotides to about 100 nucleotides.
[0132] Multiple primers or primer sets may comprise two or more primers having identical, similar, and / or different sequences, linkers, and / or tags. For example, two or more primers may comprise the same sequence. In another example, two or more primers may comprise similar sequences. In yet another example, two or more primers may comprise different sequences. Two or more primers may also comprise one or more linkers. Two or more primers may also comprise different linkers. Two or more primers may also comprise similar linkers. Two or more primers may also comprise the same linker. Two or more primers may also comprise one or more tags. Two or more primers may also comprise different tags. Two or more primers may also comprise similar tags. Two or more primers may also comprise the same tag.
[0133] In some cases, universal primers can be used, and the universal primers can include one or more of the following: 8F primer, 27F primer, CC[F] primer, 357F primer, 515F primer, 533F primer, 16S.1100.F16 primer, 1237F primer, 519R primer, CD[R] primer, 907R primer, 1391R primer, 1492R(I) primer, 1492R(s) primer, U1492R primer, 928F primer, 336R primer, 1100F primer, 1100R primer, 337F primer, 785F primer, 805R primer, 518R primer, and any other suitable universal primers. Alternatively, for samples for which specific primers may be suitable, specific primers can be used for amplification. In an example, specific primers may include: CYA106 primer (for cyanobacteria), CYA359F primer (for cyanobacteria), 895F primer (for bacteria excluding plastids and cyanobacteria), CYA781R primer (for cyanobacteria), 902R primer (for bacteria excluding plastids and cyanobacteria), 904R primer (for bacteria excluding plastids and cyanobacteria), 1100R primer (for bacteria), 1185mR primer (for bacteria excluding plastids and cyanobacteria), 1185aR primer (for lichen-associated Rhizobiales), 1381R primer (for bacteria excluding plastids of Asterochloris species), or any other suitable specific primers.
[0134] The capture probe, primer, label and / or bead may comprise one or more nucleotides. The one or more nucleotides may comprise RNA, DNA, a mixture of DNA and RNA residues, or modified analogs thereof, such as 2'-OMe or 2'-fluorine (2'-F), locked nucleic acid (LNA), or abasic sites.
[0135] Method disclosed herein can include one or more marks. Method disclosed herein can include 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more or 1000 or more one or more marks. One or more marks can be different, similar, identical, or a combination thereof.
[0136] Examples of labels include, but are not limited to, chemical labels, biochemical labels, biological labels, colorimetric labels, enzymatic labels, fluorescent labels, and luminescent labels. Labels include dyes, photocrosslinkers, cytotoxic compounds, drugs, affinity labels, photoaffinity labels, reactive compounds, antibodies or antibody fragments, biomaterials, nanoparticles, spin labels, fluorophores, metal-containing moieties, radioactive moieties, novel functional groups, groups that covalently or non-covalently interact with other molecules, photocaged moieties, actinic radiation excitable moieties, ligands, photoisomerizable moieties, biotin, biotin analogs, heavy atom-binding moieties, chemically cleavable groups, photocleavable groups, redox-active agents, isotopically labeled moieties, biophysical probes, phosphorescent groups, chemiluminescent groups, electron-dense groups, magnetic groups, intercalating groups, chromophores, energy transfer agents, biologically active agents, detectable labels, or combinations thereof.
[0137] The label can be a chemical label. Examples of chemical labels can include, but are not limited to, biotin and radioactive isotopes (e.g., iodine, carbon, phosphate, hydrogen).
[0138] The methods, kits, and compositions disclosed herein can include biomarkers. Biomarkers can include metabolic markers, including but not limited to bioorthogonal azide-modified amino acids, sugars, and other compounds.
[0139] Methods, kits, and compositions disclosed herein can include enzyme labels. Enzyme labels can include, but are not limited to, horseradish peroxidase (HRP), alkaline phosphatase (AP), glucose oxidase, and beta-galactosidase. The enzyme label can be luciferase.
[0140] The methods, kits, and compositions disclosed herein may include fluorescent labels. The fluorescent label can be an organic dye (e.g., FITC), a biofluorophore (e.g., green fluorescent protein), or a quantum dot. A non-limiting list of fluorescent labels includes fluorescein isothiocyanate (FITC), DyLight Fluors, fluorescein, rhodamine (tetramethylrhodamine isothiocyanate, TRITC), coumarin, Lucifer Yellow, and BODIPY. The label can be a fluorophore. Examples of fluorophores include, but are not limited to, indocarbocyanine (C3), indodicarbocyanine (C5), Cy3, Cy3.5, Cy5, Cy5.5, Cy7, Texas Red, Pacific Blue, Oregon Green 488, Alexa Fluor ... -355, Alexa Fluor 488, Alexa Fluor 532, Alexa Fluor 546, Alexa Fluor-555, Alexa Fluor 568, Alexa Fluor 594, Alexa Fluor 647, Alexa Fluor 660, Alexa Fluor 680, JOE, Lissamine, Rhodamine Green, BODIPY, Fluorescein isothiocyanate (FITC), Carboxyfluorescein (FAM), Phycoerythrin, Rhodamine, Dichlororhodamine (dRhodamine), Carboxytetramethylrhodamine (TAMRA), Carboxy-X-rhodamine (ROX TM ), LIZ TM 、VIC TM 、NED TM PET TM , SYBR, PicoGreen, RiboGreen, etc. The fluorescent label can be green fluorescent protein (GFP), red fluorescent protein (RFP), yellow fluorescent protein, phycobiliprotein (such as allophycocyanin, phycocyanin, phycoerythrin and phycoerythrocyanin).
[0141] The methods disclosed herein may comprise one or more linkers. The methods disclosed herein may comprise 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, 100 or more, 125 or more, 150 or more, 175 or more, 200 or more, 250 or more, 300 or more, 350 or more, 400 or more, 500 or more, 600 or more, 700 or more, 800 or more, 900 or more or 1000 or more of one or more linkers. The one or more linkers may be different, similar, identical, or a combination thereof.
[0142] Suitable linkers include any compound or biological compound capable of being connected to the labels, primers and / or capture probes disclosed herein. If the linker is simultaneously connected to the label and the primer or capture probe, the suitable linker may be able to fully separate the label and the primer or capture probe. The suitable linker may not significantly interfere with the ability of the primer and / or capture probe to hybridize with the nucleic acid molecule, its portion or variant or derivative thereof. The suitable linker may not significantly interfere with the ability of the label to be detected. The linker can be rigid. The linker can be flexible. The linker can be semi-rigid. The linker can be proteolytically stable (e.g., resistant to proteolytic cleavage). The linker can be proteolytically unstable (e.g., sensitive to proteolytic cleavage). The linker can be helical. The linker can be non-helical. The linker can be coiled. The linker can be β-stranded. The linker can contain a turn conformation. The linker can be single-stranded. The linker can be long-chain. The linker can be short-chain. The linker can comprise at least about 5 residues, at least about 10 residues, at least about 15 residues, at least about 20 residues, at least about 25 residues, at least about 30 residues, or at least about 40 residues or more.
[0143] Examples of linkers include, but are not limited to, hydrazones, disulfide bonds, thioethers, and peptide linkers. The linker can be a peptide linker. The peptide linker can comprise a proline residue. The peptide linker can comprise arginine, phenylalanine, threonine, glutamine, glutamic acid, or any combination thereof. The linker can be a heterobifunctional cross-linker.
[0144] The methods disclosed herein may comprise one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, eleven or more, twelve or more, thirty or more, thirty or more, thirty or more, forty or more, forty or more, forty or more, fifty or more assays performed on a sample comprising one or more nucleic acid molecules. The two or more assays may be different, similar, identical, or a combination thereof. For example, the methods disclosed herein comprise performing two or more sequencing reactions. In another example, the methods disclosed herein comprise performing two or more assays, wherein at least one of the two or more assays comprises a sequencing reaction. In yet another example, the methods disclosed herein comprise performing two or more assays, wherein at least two of the two or more assays comprise a sequencing reaction and a hybridization reaction. The two or more assays may be performed sequentially, simultaneously, or a combination thereof. For example, the two or more sequencing reactions may be performed simultaneously. In another example, the methods disclosed herein comprise performing a hybridization reaction followed by a sequencing reaction. In yet another example, the method disclosed herein comprises performing two or more hybridization reactions simultaneously, followed by performing two or more sequencing reactions simultaneously. Two or more assays can be performed by one or more devices. For example, two or more amplification reactions can be performed by a PCR instrument. In another example, two or more sequencing reactions can be performed by two or more sequencers.
[0145] The methods disclosed herein may include one or more devices. The methods disclosed herein may include one or more assays comprising one or more devices. The methods disclosed herein may include using one or more devices to perform one or more operations or assays. The methods disclosed herein may include using one or more devices in one or more operations or assays. For example, performing a sequencing reaction may include one or more sequencers. In another example, generating subsets of nucleic acid molecules may include using one or more magnetic separators. In yet another example, one or more processors may be used in the analysis of one or more nucleic acid samples. Examples of devices include, but are not limited to, sequencers, thermal cyclers, real-time PCR instruments, magnetic separators, transfer devices, hybridization chambers, electrophoresis instruments, centrifuges, microscopes, imagers, fluorometers, photometers, microplate readers, computers, processors, and bioanalyzers.
[0146] The methods disclosed herein may include one or more sequencers. The one or more sequencers may include one or more HiSeq, MiSeq, HiScan, Genome Analyzer IIx, SOLiD sequencer, Ion Torrent PGM, 454GSJunior, Pac Bio RS, or a combination thereof. The one or more sequencers may include one or more sequencing platforms. The one or more sequencing platforms may include 454Life Technologies / Roche's GS FLX, Solexa / Illumina's Genome Analyzer, Applied Biosystems' SOLiD, Complete Genomics' CGA Platform, Pacific Biosciences' PacBio RS, or a combination thereof.
[0147] The methods disclosed herein can include one or more thermal cyclers. One or more thermal cyclers can be used to amplify one or more nucleic acid molecules. The methods disclosed herein can include one or more real-time PCR instruments. One or more real-time PCR instruments can include a thermal cycler and a fluorometer. One or more thermal cyclers can be used to amplify and detect one or more nucleic acid molecules.
[0148] The methods disclosed herein may include one or more magnetic separators. The one or more magnetic separators may be used to separate paramagnetic and ferromagnetic particles from a suspension. The one or more magnetic separators may include one or more LifeStep TM Biomagnetic separator, SPHERO TM FlexiMag separator, SPHERO TM MicroMag separator, SPHERO TM HandiMag separator, SPHERO TM MiniTube Mag separator, SPHERO TM UltraMag separator, DynaMag TM Magnet, DynaMag TM -2 magnets, or a combination thereof.
[0149] The methods disclosed herein can include one or more bioanalyzers. In some cases, a bioanalyzer is a chip-based capillary electrophoresis instrument that can analyze RNA, DNA, and protein. The one or more bioanalyzers can include an Agilent 2100 bioanalyzer.
[0150] The methods disclosed herein may include one or more processors. The one or more processors may analyze, compile, store, classify, combine, evaluate, or otherwise process one or more data and / or results from one or more assays, one or more data and / or results based on or derived from one or more assays, one or more outputs from one or more assays, one or more outputs based on or derived from one or more assays, one or more outputs from one or more data and / or results, one or more outputs based on or derived from one or more data and / or results, or a combination thereof. In some cases, the methods disclosed herein may include combining data for analysis, such as Figure 17As shown. One or more processors may send one or more data, results, or outputs from one or more assays; one or more data, results, or outputs based on or derived from one or more assays; one or more outputs from one or more data or results; one or more outputs based on or derived from one or more data or results, or a combination thereof. One or more processors may receive and / or store requests from users. One or more processors may generate or produce one or more data, results, or outputs. One or more processors may generate or produce one or more biomedical reports. One or more processors may send one or more biomedical reports. One or more processors may analyze, compile, store, sort, combine, evaluate, or otherwise process information from one or more databases, one or more data or results, one or more outputs, or a combination thereof. One or more processors may analyze, compile, store, sort, combine, evaluate, or otherwise process information from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, or more databases. One or more processors can send one or more requests, data, results, outputs and / or information to one or more users, processors, computers, computer systems, memory locations, devices, databases or combinations thereof. One or more processors can receive one or more requests, data, results, outputs and / or information from one or more users, processors, computers, computer systems, memory locations, devices, databases or combinations thereof. One or more processors can retrieve one or more requests, data, results, outputs and / or information from one or more users, processors, computers, computer systems, memory locations, devices, databases or combinations thereof. The present disclosure also provides a method that can be used for a variety of biomedical applications. In some cases, variants, genes, reassembled genes, exons, UTRs, regulatory regions, splice sites, alternative sequences and other contents of the human or non-human genome of interest can be combined from multiple databases to generate a summary content set suitable for multiple biomedical reports. This content can then be classified based on local or overall genome context, nucleotide content, sequencing performance and interpretation requirements and then grouped into subgroups for specific protocols, assay development, etc. In some cases, variants, genes, exons, UTRs, regulatory regions, splice sites, alternative sequences, and other content of interest are combined from multiple databases to generate aggregated content sets that can be applied to multiple biomedical reports. This content is then categorized based on local or global genomic context, nucleotide content, sequencing performance, and interpretation requirements, and then subsequently grouped into subgroups for specific protocols and assay development. Figure 14In some cases, the protocol and / or assay may include a supplemental extract. The supplemental extract may comprise a human target sequence, a non-human target sequence, and combinations thereof. The supplemental extract may comprise nucleic acid molecules from the subject (e.g., derived from cells from a tissue from the subject) and nucleic acid molecules not from the subject (e.g., from microorganisms (symbiotic or parasitic), pathogens, or transplants). Figure 15-17 Examples of assay workflows are provided that include supplemental extracts for one or more of multiple DNA subsets enriched for different genomic regions.
[0151] The methods disclosed herein may include one or more memory locations. The one or more memory locations may store information, data, results, output, requests, or a combination thereof. The one or more memory locations may receive information, data, results, output, requests, or a combination thereof from one or more users, processors, computers, computer systems, devices, or a combination thereof.
[0152] The methods described herein can be implemented with the aid of one or more computers and / or computer systems. The computers or computer systems may include an electronic storage location (e.g., a database, a memory) having machine executable code for performing the methods provided in the present disclosure, and one or more processors for executing the machine executable code.
[0153] The methods disclosed herein can include treating and / or preventing a disease or condition in a subject based on one or more biomedical outputs. The one or more biomedical outputs can recommend one or more therapies. The one or more biomedical outputs can suggest, select, specify, recommend, or otherwise determine a course of treatment and / or prevention for a disease or condition. The one or more biomedical outputs can recommend modifying or continuing one or more therapies. Modifying one or more therapies can include administering, initiating, reducing, increasing, and / or terminating one or more therapies. The one or more therapies can include anticancer, antiviral, antibacterial, antifungal, immunosuppressive therapy, or a combination thereof. The one or more therapies can treat, alleviate, or prevent one or more diseases or indications.
[0154] Examples of anticancer therapies include, but are not limited to, surgery, chemotherapy, radiation therapy, immunotherapy / biotherapy, photodynamic therapy. Anticancer therapies may include chemotherapy, monoclonal antibodies (e.g., rituximab, trastuzumab), cancer vaccines (e.g., therapeutic vaccines, prophylactic vaccines), gene therapy, or a combination thereof.
[0155] One or more therapies may include antimicrobials. Generally, antimicrobials refer to substances that kill or inhibit the growth of microorganisms (e.g., bacteria, fungi, viruses, or protozoa). Antimicrobials kill microorganisms (microbicidal) or prevent the growth of microorganisms (microbial inhibitory). There are mainly two types of antimicrobials, which are obtained from natural sources (e.g., antibiotics, protein synthesis inhibitors (e.g., aminoglycosides, macrolides, tetracyclines, chloramphenicols, polypeptides)) and synthetic agents (e.g., sulfonamides, sulfamethoxazole, quinolones). In some cases, antimicrobials are antibiotics, antivirals, antifungals, antimalarials, antituberculosis drugs, antileprosy drugs, or antiprotozoal drugs.
[0156] Antibiotics are commonly used to treat bacterial infections. Antibiotics can be divided into two categories: bactericidal and bacteriostatic. Generally, bactericides kill bacteria directly, while bacteriostatics prevent bacteria from dividing. Antibiotics can be derived from living organisms or include synthetic antibacterial agents such as sulfonamides. Antibiotics can include aminoglycosides, such as amikacin, gentamicin, kanamycin, neomycin, netilmicin, tobramycin, and paromomycin. Alternatively, the antibiotic can be ansamycins (e.g., geldanamycin, herbimycin), cabacephems (e.g., loracarbef), carbapenems (e.g., ertapenem, doripenem, imipenem, cilastatin, meropenem), glycopeptides (e.g., teicoplanin, vancomycin, telavancin), lincosamides (e.g., clindamycin, lincomycin, daptomycin), macrolides (e.g., azithromycin, clarithromycin, dirithromycin, erythromycin, roxithromycin, troleandomycin, telithromycin, spectinomycin, spiramycin), nitrofurans (e.g., furazolidone, nitrofurantoin), and polypeptides (e.g., bacitracin, colistin, polymyxin B).
[0157] In some cases, antibiotic therapy includes cephalosporins, such as cefadroxil, cefazolin, cephalothin, cephalexin, cefaclor, cefamandole, cefoxitin, cefprozil, cefuroxime, cefixime, cefdinir, cefditoren, cefoperazone, cefotaxime, cefpodoxime, ceftazidime, ceftibuten, ceftizoxime, ceftriaxone, cefepime, ceftaroline fosamil, and ceftobiprole.
[0158] Antibiotic treatment may also include penicillins. Examples of penicillins include amoxicillin, ampicillin, azlocillin, carbenicillin, cloxacillin, dicloxacillin, flucloxacillin, mezlocillin, methicillin, nafcillin, oxacillin, penicillin G, penicillin V, piperacillin, temocillin, and ticarcillin.
[0159] Alternatively, quinolines can be used to treat bacterial infections. Examples of quinolines include ciprofloxacin, enoxacin, gatifloxacin, levofloxacin, lomefloxacin, moxifloxacin, nalidixic acid, norfloxacin, ofloxacin, trovafloxacin, grepafloxacin, sparfloxacin, and temafloxacin.
[0160] In some cases, antibiotic therapy includes a combination of two or more therapies. For example, amoxicillin and clavulanate, ampicillin and sulbactam, piperacillin and tazobactam, or ticarcillin and clavulanate can be used to treat bacterial infections.
[0161] Sulfonamides can also be used to treat bacterial infections. Examples of sulfonamides include, but are not limited to, mafenide, sulfonamidochrysoidine, sulfacetamide, sulfadiazine, silver sulfadiazine, sulfamethizole, sulfamethoxazole, sulfadimethoxazole, sulfasalazine, sulfisoxazole, trimethoprim, and trimethoprim-sulfamethoxazole combination (tmp-smx).
[0162] Tetracyclines are another example of antibiotics. Tetracyclines can inhibit the binding of aminoacyl-tRNA to mRNA ribosome complexes by binding to the 30S ribosomal subunit in the mRNA translation complex. Tetracyclines include demeclocycline, doxycycline, minocycline, oxytetracycline and tetracycline. Other antibiotics that can be used to treat bacterial infections include arsphenamine, chloramphenicol, fosfomycin, fusidic acid, linezolid, metronidazole, mupirocin, plate mycin, quinupristin / dalfopristin, rifaximin, thiamphenicol, tigecycline, tinidazole, clofazimine, dapsone, capreomycin, cycloserine, ethambutol, ethionamide, isoniazid, pyrazinamide, rifampicin, rifamycin, rifabutin, rifapentine and streptomycin.
[0163] Antiviral therapy is a class of drugs specifically used to treat viral infections. Like antibiotics, specific antivirals are also used for specific viruses. They are relatively harmless to the host and therefore can be used to treat infections. Antiviral therapy can inhibit various stages of the viral life cycle. For example, antiviral therapy can inhibit the attachment of viruses to cell receptors. Such antiviral therapies may include agents that mimic virus-associated proteins (VAPs) and bind to cell receptors. Other antiviral therapies can inhibit viral entry, viral uncoating (e.g., amantadine, rimantadine, pleconaril), viral synthesis, viral integration, viral transcription, or viral translation (e.g., fomivirsen). In some cases, antiviral therapy is morpholino antisense therapy. Antiviral therapy should be distinguished from virucidal agents, which can effectively inactivate viral particles in vitro.
[0164] Many available antiviral drugs are designed to treat infection with retroviruses, primarily HIV. Antiretroviral drugs may include classes of protease inhibitors, reverse transcriptase inhibitors, and integrase inhibitors. Drugs for the treatment of HIV may include protease inhibitors (e.g., Indinavir, Saquinavir, Kaletra, Lopinavir, Fosamprenavir, Fosamprenavir, Novovir, Ritonavir, Dextromethorphan, Duranavir, Retaz, Panroce), integrase inhibitors (e.g., Raltegravir), transcriptase inhibitors (e.g., Abacavir, Cytoxan, Agenerase, Amprenavir, Aptivus, Tipranavir, Gaxivir, Indinavir, Fudeve, Saquinavir, Intelence TM , etravirine, aspirin, viread), reverse transcriptase inhibitors (e.g., delavirdine, efavirenz, ypervir, havid, nevirapine, zidovudine, AZT, stuvadine, truvada, huituz), fusion inhibitors (e.g., enfuvirtide, enfuvirtide), chemokine co-receptor antagonists (e.g., selzentry, emtriva, emtricitabine, epzicom or santoxaban). Alternatively, antiretroviral therapy can be a combination therapy, such as atripla (e.g., efavirenz, emtricitabine and tenofovir disoproxil fumarate) and complete agent (embricitabine, rilpivirine and tenofovir disoproxil fumarate). Herpes viruses that may cause cold sores and genital herpes are usually treated with the nucleoside analog acyclovir. Viral hepatitis (AE) is caused by five unrelated hepatotropic viruses and is also usually treated with antiviral drugs depending on the type of infection. Influenza A and B viruses are important targets for the development of novel influenza treatments to overcome resistance to existing neuraminidase inhibitors, such as oseltamivir.
[0165] In some cases, antiviral therapy may include reverse transcriptase inhibitors. Reverse transcriptase inhibitors may be nucleoside reverse transcriptase inhibitors or non-nucleoside reverse transcriptase inhibitors. Nucleoside reverse transcriptase inhibitors may include but are not limited to dapoxetine, emtriva, yipingwei, epzicom, havid, ritova, sanxiewei, truvada, huitozi ec, huitozi, weiruide, serite and saijin. Non-nucleoside reverse transcriptase inhibitors may include edurant, intelence, rescriptor, sustiva and vitamin (immediate release or extended release).
[0166] Protease inhibitors are another example of antiviral drugs and may include, but are not limited to, agenerase, aptivus, gaxifen, formosan, involucrata, kaletra, fosamprenavir, edivafur, valproate, rituximab, and vilaset. Alternatively, antiviral therapy may include a fusion inhibitor (e.g., enfuviride) or an entry inhibitor (e.g., maraviroc).
[0167] Other examples of antiviral drugs include abacavir, acyclovir, adefovir, amantadine, amprenavir, ampligen, arbidol, atazanavir, atripla, boceprevir, cidofovir, diazepam, darunavir, delavirdine, didanosine, docosanol, edoxuridine, efavirenz, emtricitabine, enfuvirtide, entecavir, famciclovir, fomivirsen, fosamprenavir, foscarnet, fosfoacetic acid, fusion inhibitors, ganciclovir, ibacitabine, iminovir, idoxuridine, imiquimod, indinavir, inosine, integrase inhibitors, interferons (e.g., type I, II, III interferons), lamivudine, lopinavir, loviride, maraviroc, morphine, methisothiazone , nelfinavir, nevirapine, nexavir, nucleoside analogs, oseltamivir, peginterferon alfa-2a, penciclovir, peramivir, pleconaril, podophyllotoxin, protease inhibitors, raltegravir, reverse transcriptase inhibitors, ribavirin, amantadine, ritonavir, pyramidine, saquinavir, stavudine, tea tree oil, tenofovir, tenofovir disoproxil, tipranavir, trifluridine, trixevir, tromantanamide, Truvada, valacyclovir, valganciclovir, vicriviroc, vidarabine, viramidine, zalcitabine, zanamivir, and zidovudine.
[0168] Antifungal drugs are drugs that can be used to treat fungal infections such as tinea pedis, ringworm, candidiasis (thrush), and serious systemic infections such as cryptococcal meningitis. Antifungals work by exploiting differences between mammalian and fungal cells to kill fungal organisms. Unlike bacteria, fungi and humans are both eukaryotic. Therefore, fungal and human cells are similar at the molecular level, making it more difficult to find a target for antifungal drugs to attack, which is also not present in the infected organism.
[0169] Antiparasitic drugs are a class of drugs used to treat parasites such as nematodes, tapeworms, flukes, infective protozoa, and amoebas. Like antifungals, they must kill the infecting pests without seriously harming the host.
[0170] The methods of the present disclosure may be implemented by a system, a kit, a library, or a combination thereof. The methods of the present disclosure may include one or more systems. The systems of the present disclosure may be implemented by a kit, a library, or both. The systems may include one or more components that perform any of the methods disclosed herein or the operations of any of the methods. For example, the systems may include one or more kits, devices, libraries, or a combination thereof. The systems may include one or more sequencers, processors, memory locations, computers, computer systems, or a combination thereof. The systems may include a transmission device.
[0171] The kit may comprise various reagents for performing the various operations disclosed herein, including sample processing and / or analytical operations. The kit may comprise instructions for performing at least some of the operations disclosed herein. The kit may comprise one or more capture probes, one or more beads, one or more labels, one or more linkers, one or more devices, one or more reagents, one or more buffers, one or more samples, one or more databases, or a combination thereof.
[0172] A library may comprise one or more capture probes. A library may comprise one or more subsets of nucleic acid molecules. A library may comprise one or more databases. A library may be generated or produced by any of the methods, kits, or systems disclosed herein. A database library may be generated from one or more databases. Methods for generating one or more libraries may comprise: (a) aggregating information from one or more databases to generate an aggregated dataset; (b) analyzing the aggregated dataset; (c) generating one or more database libraries from the aggregated dataset. Figure 13 An example of a library construction workflow is provided. In some cases, libraries can be pooled, Figure 7 .
[0173] Computer system
[0174] The present disclosure provides computer systems programmed to implement the methods of the present disclosure. Figure 5 A computer system 501 is shown that is programmed or otherwise configured to map and / or compare sequence reads to identify the source (e.g., human or non-human, host or non-host) of a nucleic acid molecule, identify one or more characteristics (e.g., genetic variants), or any combination thereof. The computer system 501 can regulate various aspects of processing sequencing information as provided in the present disclosure, for example, comparing sequence reads with one or more reference sequences to identify the source of nucleic acid sequences in a biological sample. The computer system 501 can be an electronic device of a user or a computer system remotely located relative to an electronic device. The electronic device can be a mobile electronic device.
[0175] Computer system 501 includes a central processing unit (CPU, also referred to herein as a "processor" and "computer processor") 505, which can be a single-core or multi-core processor, or multiple processors for parallel processing. Computer system 501 also includes memory or memory locations 510 (e.g., random access memory, read-only memory, flash memory), electronic storage 515 (e.g., a hard disk), a communication interface 520 (e.g., a network adapter) for communicating with one or more other systems, and peripherals 525 (e.g., cache, other memory, data storage, and / or an electronic display adapter). Memory 510, storage 515, interface 520, and peripherals 525 communicate with CPU 505 via a communication bus (solid lines), such as a motherboard. Storage 515 can be a data storage unit (or data repository) for storing data. Computer system 501 can be operatively coupled to a computer network ("network") 530 via communication interface 520. Network 530 can be the Internet, an internetwork, and / or an extranet, or an intranet and / or extranet in communication with the Internet. In some cases, network 530 is a telecommunications and / or data network. Network 530 may include one or more computer servers, which may enable distributed computing such as cloud computing. In some cases, network 530, with the assistance of computer system 501, may implement a peer-to-peer network that enables devices coupled to computer system 501 to act as either clients or servers.
[0176] The CPU 505 can execute a series of machine-readable instructions, which can be embodied in a program or software. The instructions can be stored in a memory location such as the memory 510. The instructions can be directed to the CPU 505, which can then program the CPU 505 or otherwise configure the CPU 505 to implement the methods of the present disclosure. Examples of operations performed by the CPU 505 can include fetching, decoding, executing, and writing back.
[0177] CPU 505 may be part of a circuit, such as an integrated circuit, which may include one or more other components of system 501. In some cases, the circuit is an application-specific integrated circuit (ASIC).
[0178] Storage unit 515 can store files, such as drivers, libraries, and saved programs. Storage unit 515 can also store user data, such as user preferences and user programs. In some cases, computer system 501 may include one or more other data storage units external to computer system 501, such as on a remote server that communicates with computer system 501 via an intranet or the Internet.
[0179] Computer system 501 can communicate with one or more remote computer systems via network 530. For example, computer system 501 can communicate with a remote computer system of a user (e.g., a healthcare provider). Examples of remote computer systems include personal computers (e.g., portable PCs), tablet computers or tablet computers (e.g., iPad, Galaxy Tab), phones, smartphones (e.g. iPhone, Android-enabled devices, ) or personal digital assistant. Users can access computer system 501 through network 1130.
[0180] The methods described herein may be implemented by means of machine (e.g., computer processor) executable code stored in an electronic storage location (e.g., memory 510 or electronic storage unit 515) of computer system 501. The machine executable or machine readable code may be provided in the form of software. During use, the code may be executed by processor 505. In some cases, the code may be retrieved from storage unit 515 and stored in memory 510 for ready access by processor 505. In some cases, electronic storage unit 515 may be omitted and the machine executable instructions may be stored in memory 510.
[0181] The code may be precompiled and configured for use with a machine having a processor suitable for executing the code, or may be compiled at runtime. The code may be provided in a selectable programming language so that the code can be executed in a precompiled or as-compiled manner.
[0182] The various aspects of the system and method (e.g., computer system 501) provided herein can be embodied in programming. Various aspects of technology can be considered as "products" or "manufactures" in the form of machine (or processor) executable code and / or associated data generally carried on a machine-readable medium type or implemented with a machine-readable medium type. Machine executable code can be stored on an electronic storage unit, such as a memory (e.g., read-only memory, random access memory, flash memory) or a hard disk. "Storage" type media can include any or all tangible memories, processors, etc. of a computer, or its related modules, such as various semiconductor memories, tape drives, disk drives, etc. that can provide non-transitory storage for software programming at any time. All or part of the software can be communicated at any time through the Internet or various other telecommunications networks. For example, such communication can enable software to be loaded from one computer or processor to another computer or processor, such as from a management server or mainframe to the computer platform of an application server. Therefore, another type of medium that can carry software elements includes, such as, a physical interface between local devices, light waves, electric waves, and electromagnetic waves used by wired and optical landline networks and by various air links. The physical elements that carry such waves, such as wired or wireless links, optical links, etc., can also be considered to be the medium that carries the software. As used herein, unless restricted to non-transitory, tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.
[0183] Thus, a machine-readable medium such as computer executable code may take many forms, including but not limited to tangible storage media, carrier media, or physical transmission media. Non-volatile storage media include, for example, optical or magnetic disks, such as any storage device in any computer, such as may be used to implement the databases shown in the accompanying drawings. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that make up a bus within a computer system. Carrier transmission media may take the form of electrical or electromagnetic signals, or acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example, a floppy disk, a diskette, a hard disk, magnetic tape, any other magnetic medium, a CD-ROM, a DVD or DVD-ROM, any other optical medium, punched card tape, any other physical storage medium having a pattern of holes, a RAM, a ROM, a PROM and EPROM, a FLASH-EPROM, any other memory chip or memory cartridge, a carrier wave that transmits data or instructions, a cable or link that transmits such a carrier wave, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer-readable media may be involved in loading one or more sequences of one or more instructions to a processor for execution.
[0184] The computer system 501 may include or communicate with an electronic display 535 including a user interface (UI) 540 for providing, for example, one or more biomedical reports comprising one or more data sets selected from the group consisting of: i) candidate tumor neoantigens, (ii) detected non-human species, (iii) detected CDR3 sequences, and any combination thereof. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0185] The methods and systems of the present disclosure may be implemented by one or more algorithms. The algorithms may be implemented by software when executed by the central processing unit 505. The algorithms may, for example, map and / or align sequence reads, call variants, annotate sequence information, or any combination thereof.
[0186] Example
[0187] Example 1. Preparation of genomic DNA
[0188] Prepare a subset of nucleic acid molecules from a sample containing genomic DNA using the following procedure:
[0189] 1. Shear the sample containing genomic DNA using M220 for 15-35 seconds.
[0190] 2. After ligation, the fragmented gDNA was purified using SPRI beads (the volume ratio of SPRI beads to DNA sample was 1) and the DNA was eluted into 100 μL of elution buffer (EB).
[0191] 3. Add 50 μL of SPRI beads to 100 μL of DNA.
[0192] 4. Transfer the supernatant to a new tube.
[0193] 5. Elute the DNA from the remaining bound DNA beads. This eluted DNA is called the long insert.
[0194] 6. Add 10 μL of SPRI beads to the supernatant from step 4.
[0195] 7. Transfer the supernatant from step 6 to a new tube.
[0196] 8. Elute DNA from the remaining DNA-bound beads from step 6. This eluted DNA is called the mid-insert.
[0197] 9. Add 20 μL of SPRI beads to the supernatant from step 7.
[0198] 10. Transfer the supernatant from step 9 to a new tube.
[0199] 11. Elute DNA from the remaining DNA-bound beads from step 9. This eluted DNA is called the short insert.
[0200] Example 2. Obtaining biological samples
[0201] Subjects with evaluable metastatic cancer undergo tumor resection. Lymphocytes from the tumor, tumor infiltrating lymphocytes (TILs), are grown and expanded. Multiple individual fragments or multiple individual cultures of TILs are grown. Individual cultures are expanded separately, and when sufficient TIL yields (approximately 10 8 When the TILs are isolated from the original tumor (cells), they are cryopreserved and aliquots are taken for immunological testing. Aliquots of the original tumor are sequenced for exome and transcriptome to identify mutations that are unique to the tumor compared with normal cells. Sequencing also identifies the presence and identity of non-human genomes, including microorganisms.
[0202] Example 3. Extraction of genomic material from biological samples
[0203] Genomic DNA (gDNA) and total RNA were purified from aliquots of various tumors and matched normal blood components using the QIAGEN AllPrep DNA / RNA Kit (Cat. No. 80204) according to the manufacturer's recommendations. Tumor samples were formalin-fixed, paraffin-embedded (FFPE), and gDNA was extracted using the Covaris truXTRAC™ FFPE DNA Kit according to the manufacturer's instructions.
[0204] Example 4. Sequencing analysis of biological samples
[0205] Agilent Technologies' SureSelectXT target enrichment system (catalog number 5190-8646) was used to prepare a double-ended library of human all-exon V6 RNA bait (catalog number 5190-8863) (Agilent Technologies, Santa Clara, California, USA) and bacterial RNA bait for exon capture of approximately 20,000 coding genes. Whole exome sequencing (WES) libraries were subsequently sequenced on a NextSeq 500 desktop sequencer (Illumina, San Diego, California, USA). Libraries were prepared using 3 μg of gDNA from fresh tumor tissue samples and 200 ng of gDNA from FFPE tumor samples according to the manufacturer's protocol. Double-ended sequencing was completed using the Illumina high-output flow cell kit (300 cycles) (catalog number FC-404-2004). Tu-1, Tu-2A and Tu-2B samples were initially run on v1 of the reagent / flow cell kit, followed by subsequent runs of the same library preparation using v2 of the reagent / flow cell kit. Tumor samples were run on the v2 reagent / flow cell kit. The average sequencing depth and percentage (tumor purity) of the tumor in each sample were determined as estimated by the tumor bioinformatics program Allele-Specific Copy Number Analysis (ASCAT) 1. RNA-seq libraries were prepared using 2 μg of total RNA and the Illumina TruSeq RNA Strand Library Preparation Kit according to the manufacturer's protocol. The RNA-seq libraries were sequenced on a NextSeq 500 desktop sequencer (Illumina, San Diego, California, USA). WES was aligned, processed and variants were called, and the human genome build hg19 was aligned using novoalign MPI from novocraft (http: / / www.novocraft.com / ). Duplicates were marked using Picard's MarkDuplicates tool. Indel realignment and base realignment were performed according to the GATK best-practice workflow (https: / / www.broadinstitute.org / gatk / ). After data cleaning, pileup files were created with samtools (http: / / samtools.sourceforge.net) and somatic variants were called using Varscan2 (http: / / varscan.sourceforge.net) using the following criteria: tumor and normal read counts of 10 or more, variant allele frequency of 10% or more, and tumor variant read counts of 4 or more.These variants were then annotated using Annovar (http: / / annovar.openbioinformatics.org). For RNA-seq, a two-pass alignment was performed using STAR (https: / / github.com / alexdobin / STAR) against the human genome build hg19. Duplicates were marked using Picard's MarkDuplicates tool. Reads were split and trimmed using the GATK SplitNTrim tool. In / del realignment and base realignment were then performed using the GATK toolbox. The final realigned bam file was used with samtools mpileup to create a stacking file. Finally, variants were called using Varscan2.
[0206] Example 5: Sequence alignment of non-human genomes
[0207] The genomic information extracted from the sequencing analysis is compared with the ribosomal RNA gene 16S (genomic reference). Reads that fully match the 16S gene are identified. Capture probes designed to hybridize to the shared region of the 16S gene can be used to capture nucleic acid molecules from a variety of species, even species that have not yet been identified and characterized. Based on the sequence from the variable region portion, the capture molecules whose sequences extend from these shared regions to the variable region are assigned to their source species.
[0208] Example 6. Shearing time and fragment size
[0209] Genomic DNA (gDNA) was sheared by changing the shearing time set by Covaris. The gDNA fragments generated by various shearing times were then analyzed. The results are shown in Figure 10 and Table 1.
[0210] Table 1. Shearing time and average fragment size
[0211] serial number Cutting time (seconds) Average fragment size (base pairs) 1 375 150 2 175 200 3 80 200 4 40 400 5 32 500 6 25 800
[0212] Example 7. Bead Ratio and Fragment Size
[0213] The ratio of bead volume to nucleic acid sample volume was varied and the effect of these ratios on the average fragment size was analyzed. Figure 11 As shown, changing the volume ratio of the bead volume to the nucleic acid sample volume from 0.8 (line 1), 0.7 (line 2), 0.6 (line 3), 0.5 (line 4) and 0.4 (line 5) results in changes in the average size of the DNA fragments. In general, it appears that the lower the ratio, the larger the average fragment size.
[0214] Example 8. Ligation reaction and fragment size
[0215] Two different shearing times and three different ligation combinations were performed on the nucleic acid samples. Sample 1 was sheared for 25 seconds and ligated on long-insert DNA. Sample 2 was sheared for 32 seconds and ligated on long-insert DNA. Sample 3 was sheared for 25 seconds and ligated on medium-insert DNA. Sample 4 was sheared for 32 seconds and ligated on medium-insert DNA. Sample 5 was sheared for 25 seconds and ligated on short-insert DNA. Sample 6 was sheared for 32 seconds and ligated on short-insert DNA. Figure 12 Average fragment sizes of six reactions are shown.
[0216] Example 9. Capture of nucleic acids of interest
[0217] A nucleic acid sample is obtained from a subject. The sample can be divided into two separate samples, and each separate sample can be subjected to different probe pools and conditions. By combining the Agilent clinical research exome kit (based on the exome in the GRCh37 reference genome) with additional probes of interest (including exome regions corresponding to the GRCh38 reference sequence, HLA-specific probes, T cell receptor and B cell receptor recombination-specific probes (i.e., regions corresponding to V(D)J regions), microsatellite instability regions, and tumor virus sequences), a first biotinylated capture probe pool is generated. Additional probes are titrated into the pool to adjust the relative capture rate of nucleic acid. This can be beneficial for downstream sequencing reactions to increase the sequencing depth or sensitivity of a group of captured nucleic acids. The second probe pool (e.g., a "boosted set" of probes targeting a specific sequence or sequence subset) can capture nucleic acids with sequences that are difficult to capture. The probes of the boosted set can include exomes of the GRCh37 and GRCh38 reference sequences with high GC content, probes relevant to cancer therapy, additional T cell receptors, and B cell recombination-specific probes. The nucleic acid molecules captured by the first pool and the nucleic acid molecules captured by the second pool are combined to generate a combined pool of capture probes and captured nucleic acids. Combination pools can be created by combining two capture molecule pools in different ratios to adjust the relative amount of captured nucleic acids corresponding to each pool. This can be beneficial for downstream sequencing reactions to increase the depth or sensitivity of a set of captured nucleic acids. The hybridized capture probes and captured nucleic acids are incubated with magnetic streptavidin beads. A magnetic separator is placed outside the tube containing the sample and allows the capture probes to fix the magnetic streptavidin beads. The liquid is poured out and additional buffer is added to wash the beads and remove any unbound nucleic acids. The captured nucleic acids are subjected to an amplification reaction to attach adapters for further downstream sequencing reactions.
[0218] Example 10. Identification of Microbiome and Tumor Viruses
[0219] Nucleic acid sample is obtained from experimenter.This sample can be divided into two independent samples, and each independent sample can be made to experience different probe pools and conditions.By using Agilent clinical research exome test kit (based on the exome in GRCh37 reference genome) and other interested probes (comprising the exome region corresponding to GRCh38 reference sequence), with the 16S rRNA region of bacteria, fungi, archaea and protists gene (comprising the gene of Helicobacter pylori (Helicobacter pylori) and clostridium (Fusobacterium)), there is a probe with homology, the probe of viral gene (comprising human papillomavirus, hepatitis B virus, hepatitis C virus) and the probe of the gene related to pathogenicity are combined together, generate the first biotinylated capture probe pool.Other probe is titrated into the library to adjust the relative capture rate of nucleic acid.The second probe pool (such as " strengthening group " of the probe targeting specific sequence or sequence subset) can capture the nucleic acid with the sequence that is difficult to capture. The probes for the enhanced set can include exons of the GRCh37 and GRCh38 reference sequences with high GC content or low gene homology due to multiple alternative gene sequences or genes with high mutation rates. The nucleic acid molecules captured by the first pool and the nucleic acid molecules captured by the second pool are combined to generate a combined pool of nucleic acids. The two pools are merged and titrated in different ratios to increase the sensitivity of one pool relative to the other. The hybridized capture probes and captured nucleic acids are incubated with magnetic streptavidin beads. A magnetic separator is placed outside the tube containing the sample and capture probes and fixed to the magnetic streptavidin beads. The liquid is poured out and additional buffer is added to wash the beads and remove any unbound nucleic acids. The captured nucleic acids are subjected to an amplification reaction to attach adapters for further downstream sequencing reactions. The sequence is compared with the reference sequence to identify the presence of specific microorganisms in the subject's microbiome or to identify any disease-causing pathogens.
[0220] Example 11. Identification of CAR-T cells
[0221] Nucleic acid samples are obtained from subjects. Two probe pools are generated, wherein the first pool consists of biotinylated Agilent clinical research exome kit probes. The supplemented biotinylated probe group is added to the first pool, which includes sequences specific for chimeric antigen receptors found in CAR-T cells. The second biotinylated capture probe pool (e.g., a "boosted set" of probes targeting a specific sequence or sequence subset) contains sequences for CAR sequences with high GC content and / or sequences with low total sequence homology to the captured nucleic acid due to mutation or recombination. The sample is divided into two sample pools, the first sample pool is subjected to the first probe pool, and the second sample pool is subjected to the second probe pool to allow the capture probe to capture nucleic acids from each sample pool. The first probe pool and the second probe pool hybridized with the captured nucleic acid are merged. Magnetic streptavidin beads are added to the mixture, and the biotinylated probe is bound to the beads. A magnetic separator is placed outside the tube containing the sample and capture probe, and the magnetic streptavidin beads are fixed to it. The liquid is poured out, and additional buffer is added to wash the beads and remove any unbound nucleic acids. The captured nucleic acids are resuspended in fresh buffer and subjected to an amplification reaction to attach adapters, after which the captured nucleic acids are sequenced. Sequence reads are analyzed and aligned to a CAR-T gene reference sequence to identify the presence of CAR-T-associated nucleic acids.
[0222] Example 12: Identification of segmented transcriptomes
[0223] Samples are obtained from different tissue types of the subject. The samples are treated with enzymatic digestion to remove DNA and various types of RNA molecules (rRNA, tRNA, miRNA). The mRNA is not digested and then reverse transcribed to generate cDNA molecules. A biotinylated probe set with sequence homology to approximately 20,000 genes is generated and mixed with the cDNA sample. Magnetic streptavidin beads are added to the mixture, and the biotinylated probe is bound to the beads. A magnetic separator is placed outside the tube containing the sample and capture probe and fixed to the magnetic streptavidin beads. The liquid is poured out and additional buffer is added to wash the beads and remove any unbound nucleic acids. The captured nucleic acid is resuspended in fresh buffer and an amplification reaction is performed to attach adapters, and the captured nucleic acid is subsequently sequenced. The sequence reads are analyzed and compared with the reference sequence to determine the identity of the genes in each sample and associate them with the tissue type of the sample. Multiple replicates are performed on samples in which the biotinylated probes are titrated in different proportions. By analyzing the specific signal of the nucleic acid related to the amount of the probe provided to the sample, the expression level of each gene can be determined. Transcriptome analysis can be carried out in parallel with exon group or genome analysis. Exon group or genome probe can be used as the first probe pool, and cDNA capture probe can be used as the second probe pool. In this case, the sample can be divided into several parts and DNA specific or mRNA specific reaction can be carried out. As disclosed in the previous embodiment, the captured nucleic acid can be merged together and subjected to sequencing reaction to determine the identity of the nucleic acid.
[0224] Example 13. Identification and Capture of Cancer-Associated Cell-Free DNA (cfDNA)
[0225] Nucleic acid sample extracts are obtained from whole blood or serum. Cells are removed from the blood sample by centrifugation to obtain a cell-free sample. A biotinylated probe set with sequence homology to approximately 20,000 genes is generated and mixed with the cell-free sample. A first biotinylated capture probe pool is generated by combining the Agilent Clinical Research Exome Kit (based on the exome in the GRCh37 reference genome) with additional probes of interest (including exome regions corresponding to the GRCh38 reference sequence) and probes with homology to genes associated with cancer. Additional probes are titrated into the pool to adjust the relative capture rate of nucleic acids. A second probe pool (e.g., a "boosted set" of probes targeting specific sequences or sequence subsets) can capture nucleic acids with sequences that are difficult to capture. The boosted set of probes can include exomes of the GRCh37 and GRCh38 reference sequences with high GC content or low gene homology due to multiple alternative gene sequences or genes with high mutation rates. Magnetic streptavidin beads are added to the mixture, and the biotinylated probes are bound to the beads. A magnetic separator is placed outside the tube containing the sample and capture probes and immobilizes magnetic streptavidin beads. The liquid is decanted, and additional buffer is added to wash the beads and remove any unbound nucleic acids. The captured nucleic acids are resuspended in fresh buffer and subjected to an amplification reaction to attach adapters, after which the captured nucleic acids are sequenced. Sequence reads are analyzed and aligned to a reference sequence to determine the identity of genes in each sample. CFDNA analysis can be performed in parallel with exome sequencing of genomic DNA by applying the same or similar probe sets to genomic DNA.
[0226] Although preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that these embodiments are provided by way of example only. It is not intended that the present invention be limited by the specific examples provided in the specification. Although the present invention has been described with reference to the foregoing description, the description and illustration of the embodiments herein are not intended to be interpreted in a restrictive sense. Those skilled in the art will now envision various variations, changes, and substitutions without departing from the present invention. In addition, it should be understood that all aspects of the present invention are not limited to the specific description, configuration, or relative proportions depending on various conditions and variables set forth herein. It should be understood that various alternatives to the embodiments of the present invention described herein can be used to implement the present invention. Therefore, it is contemplated that the present invention should also encompass any such alternatives, modifications, variations, or equivalents. The appended claims are intended to define the scope of the present invention and thus encompass methods and structures within the scope of these claims and their equivalents.
Claims
1. A system for processing a biological sample from a subject, comprising: a processing unit comprising one or more computer processors programmed individually or collectively to assay a subset of nucleic acid molecules to generate sequence information comprising the following sequences: (i) human nucleic acids from the biological sample of the subject and (ii) non-human nucleic acids from the biological sample of the subject, wherein the non-human nucleic acids include a subset of non-human nucleic acids derived from one or more oncogenic viruses or bacteria associated with cancer, the subset of nucleic acid molecules being generated from the biological sample using a pool of nucleic acid probes, wherein the probes comprise (i) a first plurality of nucleic acid probes comprising a set of human exome capture probes and probes formulated to target junctional sequences generated by human V(D)J rearrangement or recombination; (ii) a second plurality of nucleic acid probes configured to target one or more non-human genomic elements, and (iii) wherein in the pool of nucleic acid probes, the concentration of the second plurality of nucleic acid probes is greater than the concentration of the first plurality of nucleic acid probes, wherein the one or more non-human genomic elements include one or more elements of a human papillomavirus E6 gene, an E7 gene, and / or a bacterial 16S ribosomal RNA gene; and A computer memory configured to store the sequence information.
2. The system of claim 1, wherein the first plurality of nucleic acid probes of (i) are configured to target elements derived from the human genome.
3. The system of claim 1, wherein generating a subset of the nucleic acid molecules from the biological sample comprises performing one or more hybridization reactions.
4. The system of claim 1, wherein the biological sample is derived from a tumor biopsy, whole blood, or plasma obtained from the subject.
5. The system of claim 1, wherein the determining comprises performing sequencing to generate paired-end read sequences having a length of 130 bases to 280 bases.
6. The system of claim 1, wherein the one or more computer processors are programmed to generate an alignment of the sequence with one or more reference sequences.
7. The system of claim 6, wherein the one or more reference sequences comprise a plurality of reference sequences, and wherein the plurality of reference sequences correspond to two or more different species.
8. The system of claim 6, wherein the one or more computer processors are programmed to identify the source of the nucleic acid molecules in the subset based on the alignment.
9. The system of claim 8, wherein the one or more computer processors are programmed to generate an output comprising the source of the nucleic acid molecules in the biological sample.
10. The system of claim 1, wherein the one or more computer processors are programmed to generate one or more biomedical reports comprising information selected from the group consisting of: (i) candidate tumor neoantigens, (ii) detected non-human species, (iii) detected complementarity determining region 3 (CDR3) sequences, and any combination thereof.
Citation Information
Patent Citations
Fetal Genomic Analysis From A Maternal Biological Sample
US20110105353A1
Methods and systems for genomic analysis
US20160019341A1