Compositions and methods for assessing microbial populations
By using specific nucleic acid primers and probes for nucleic acid amplification and detection, the problem of difficulty in accurately identifying microbial populations in samples in the prior art is solved, and high sensitivity and specific microbial detection and identification are achieved.
Patent Information
- Application Number
- CN202080082093.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-06
- Filing Date
- 2020-10-09
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2040-10-09
AI Technical Summary
It is difficult to accurately and comprehensively identify microbial populations in samples, especially in multifactorial conditions and causes, complications and diagnostic studies of diseases.
A composition and method are provided, comprising specific nucleic acid primers and probes for amplification, detection, characterization, evaluation, profiling and measurement of nucleic acids. These primers and probes are able to bind, hybridize and amplify the target nucleic acid of the microorganism, specifically amplifying the predetermined unique nucleic acid sequences in the microorganism genome.
It achieves highly sensitive, specific, accurate and repeatable detection and identification of microbial populations in the sample, and can accurately determine the relative and absolute levels or abundance of different microorganisms, and is suitable for a variety of research and diagnostic applications.
Smart Images

Figure CN115176032B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to and the benefit of U.S. Provisional Application No. 62 / 914,366, filed on October 11, 2019, and U.S. Provisional Application No. 62 / 914,368, filed on October 11, 2019, and U.S. Provisional Application No. 62 / 944,877, filed on December 6, 2019, each of which is incorporated herein by reference in its entirety.
[0003] Sequence Listing
[0004] This application hereby incorporates by reference in its entirety the materials in the electronic sequence listing filed concurrently herewith. The materials in the electronic sequence listing are filed in the form of a text (.txt) document titled “LT01495_ST25.txt” created on October 8, 2020, with a file size of 408 kilobytes. Background Art
[0005] As the scientific and medical communities become more aware of the important role of microbiota in ecosystems and in the health of individuals and populations, the diversity of microbiota in various environments has become an area of intensive research. In one example, the intestinal microbiota (also known as the gut microbiome) consists of trillions of bacteria, fungi and other microorganisms. One-third of the human gut microbiota is common to most people, while two-thirds is unique to each person. The healthy human intestine has a variety of coexisting or symbiotic bacteria that live in a relatively stable state. When microbial imbalance or maladaptation occurs, the composition and proportion of the normal flora are changed, and the intestine enters a state of dysbiosis. Dysbiosis usually leads to inflammation of the intestinal cell wall, destruction of the mucus barrier, epithelial barrier and immune-sensitive cells in the gastrointestinal tract. Imbalance in the intestinal microbiota is associated with disease, chronic health conditions and response to immuno-oncology treatment. For example, imbalance in the intestinal microbiota is associated with intestinal diseases such as irritable bowel syndrome (IBS), inflammatory bowel disease (IBD) and obesity, as well as autoimmune conditions such as celiac disease, lupus and rheumatoid arthritis (RA). Furthermore, the composition of the gut microbiota may influence susceptibility to neoplastic pathologies such as cancer and responsiveness to cancer therapies. As the gut microbiota is implicated in a wide range of conditions and diseases across animal species including humans, animals, and insects, characterization and study of the gut microbiota has become a major research focus to advance the understanding of health and disease and to develop therapies for associated pathologies.
[0006] Several different techniques have been used to attempt to identify microorganisms in various environmental and biological samples. The initial techniques relied on microbial culture processes, which were time-consuming and provided limited information, in part because of the different growth conditions required to obtain different microbial cultures. Many of the latest techniques that do not require culture involve analyzing the genetic composition of microbial cells contained in the sample using nucleic acid analysis methods, which include, for example, nucleic acid amplification (e.g., PCR) and / or sequencing. Typically, such methods involve the amplification and analysis of microbial 16S rRNA gene segments. Although the analysis of 16S rRNA sequences reduces the time and labor required for some other methods of assessing the microbial composition of samples, the comprehensiveness, accuracy, quality, and depth of information obtained by methods based on 16S rRNA gene sequence analysis may vary and be limited, for example, by the target amplicons and primers used in the method. Therefore, more sensitive and comprehensive methods are needed to accurately characterize the entire microbial population in a sample by identifying and distinguishing microbial species and their levels in samples containing multiple species. These methods will play an important role in many research areas, including research on the causes, complications, and diagnosis of multifactorial disorders and diseases, as well as in promoting the study and understanding of intestinal microbiota in health and disease. Summary of the invention
[0007] Provided herein are compositions and methods, and combinations, kits and systems comprising the compositions and methods for amplification, detection, characterization, evaluation, analysis and / or measurement of nucleic acids. In certain embodiments, provided herein are compositions comprising nucleic acids used as primers and / or probes, such as single-stranded nucleic acids. In certain embodiments, provided herein are compositions comprising a combination of multiple nucleic acids. In certain embodiments, the primers and / or probes can be combined with, hybridized, amplified and / or detected target nucleic acids of microorganisms (e.g., bacteria), such as may be present in samples (e.g., biological samples), such as samples of animal digestive tract contents. Such nucleic acid provided herein comprises a predetermined unique nucleic acid sequence of amplifying microorganism genome specifically or selectively, a primer and probe for combining, hybridizing and / or detecting the predetermined unique nucleic acid sequence, and amplifying a nucleic acid sequence in one or more genes, a primer and probe for combining, hybridizing and / or detecting the nucleic acid sequence, wherein the one or more genes are homologous in most or substantially all members of the taxonomic category (e.g., domain, kingdom, phylum, class, order, genus, species) of an organism (e.g., microorganism), but different between different organisms. In certain embodiments, such nucleic acid contains one or more modifications that promote the operation and / or multiple amplification of nucleic acid. For example, such modifications include modifications that increase the susceptibility of nucleic acid to cutting relative to nucleic acid that does not include modification. In certain embodiments, the nucleic acid includes one or more pairs of nucleic acids used as a primer (e.g., primer pair) for amplifying target nucleic acid, such as, for example, one or more or more nucleic acids contained in a specific nucleic acid unique to a microorganism species or a homologous gene (e.g., 16S ribosomal RNA (rRNA) gene) shared by a variety of different microorganisms. For example, in certain embodiments, the nucleic acid comprises one or more primer pairs, which amplify two or more regions in the prokaryotic 16S rRNA gene individually, such as hypervariable regions. In certain embodiments, the nucleic acid comprises a combination of multiple primer pairs. In certain embodiments, the combination of multiple primer pairs is designed to amplify nucleic acids in one, some, most or substantially all microorganisms (such as, for example, bacteria) in a sample in a manner that is targeted and / or bounded. Also provided herein is a composition containing a nucleic acid mixture, wherein most or substantially all nucleic acids contain a sequence of a part of a microorganism (e.g., bacteria) genome. In certain embodiments, the sequence length of the microorganism genome portion is less than 250 nucleotides or is about 250 nucleotides. In certain embodiments, the nucleic acid comprises nucleotides containing uracil nucleobases. In certain embodiments, the composition contains one or more or more primers, such as nucleic acids and / or primer pairs of any embodiment described herein. In certain embodiments, the composition comprises DNA polymerase, DNA ligase and / or at least one uracil cutting or modifying enzyme.
[0008] In some embodiments of the amplification methods provided herein, nucleic acids described herein are used as amplification primers to amplify nucleic acids. In some embodiments, the nucleic acid amplification is multiple amplification. In some embodiments, the amplification method comprises a plurality of nucleic acid primers, such as primer pairs, which amplify two or more regions in one or more genes individually, and the one or more genes are homologous in most or substantially all members of the taxonomic category (e.g., domain, kingdom, phylum, class, order, genus, species) of the organism. For example, in some embodiments, a plurality of nucleic acid primers comprise primers or primer pairs that amplify one or more or more hypervariable regions in the prokaryotic 16SrRNA gene individually. In some embodiments, the amplification method comprises one or more or more nucleic acid primers, such as primer pairs, which amplify specific nucleic acids unique to species of organisms (e.g., microorganisms such as bacteria). In some embodiments, the amplification method comprises: a plurality of nucleic acid primers, such as primer pairs, which comprise a combination of primers that amplify two or more regions in one or more genes individually, and the one or more genes are homologous in most or substantially all members of the taxonomic category of the organism; and one or more or more nucleic acid primers, such as primer pairs, which amplify specific nucleic acids unique to species of organisms. In some embodiments, primers for use in amplification methods comprise a nucleic acid containing or consisting of a nucleic acid provided herein and / or a nucleic acid capable of amplifying a nucleic acid containing or consisting essentially of a target sequence provided herein.
[0009] In some embodiments of the method for detecting and / or measuring nucleic acid provided herein, nucleic acid as described herein is used as primer and / or probe. For example, in some methods for detecting and / or measuring nucleic acid, nucleic acid as described herein is used as amplification primer to carry out nucleic acid amplification to nucleic acid, and detect whether there are one or more nucleic acid amplification products. In some embodiments, the amplification is carried out using multiple nucleic acid primers and carried out in a single multiplex amplification reaction mixture. In some embodiments, according to the amplification method provided herein, any one or more primers or primers or primer pairs as described herein are used to amplify. In some embodiments, nucleic acid is contacted with a probe containing nucleic acid as described herein under hybridization conditions, and detect whether there is the hybridization probe. In some embodiments, one or more nucleic acids as provided herein are used as probes (for example, detectable or labeled probes) to detect whether there are one or more nucleic acid amplification products. In some embodiments, the nucleotide sequence information of one or more nucleic acid amplification products is obtained to detect whether there are one or more nucleic acid amplification products. In some embodiments, the level (absolute or relative) of the amplification product detected is measured and determined. In some embodiments, the level (absolute or relative) of the hybridization probe detected is measured and determined. In some embodiments, the nucleic acid detected and / or measured is the nucleic acid of a microorganism (for example, bacteria). In some embodiments, the nucleic acid being detected and / or measured is a nucleic acid in or from a sample, for example, a sample of the contents of the digestive tract of an organism.
[0010] Compositions and methods are also provided herein, as well as combinations, kits and systems comprising the compositions and methods, for characterizing, evaluating, dissecting and / or measuring a microorganism (e.g., bacteria) population and / or its components or composition in a sample (e.g., biological sample), such as a sample of an animal digestive tract content. In certain embodiments, the method for characterizing, evaluating, dissecting and / or measuring a microorganism population in a sample comprises a combination of a nucleic acid primer pair that specifically amplifies a predetermined unique nucleic acid sequence of a microorganism genome and / or a primer pair that amplifies a nucleic acid sequence that appears in a homologous gene or genomic region to carry out nucleic acid amplification of the nucleic acid in or from the sample, wherein the homologous gene or genomic region is common to a variety of microorganisms, but there are differences between different microorganisms. In a specific embodiment, the primer pair that amplifies the nucleic acid sequence that appears in a homologous gene or genomic region comprises one or more primer pairs, which amplify the nucleic acid of the nucleotide sequence of one or more hypervariable regions of a prokaryotic 16SrRNA gene. In certain embodiments, the amplification using the combination of nucleic acid primer pairs is carried out in a single multiplex amplification reaction mixture. In some embodiments, methods for characterizing, evaluating, profiling and / or measuring a microbial population comprise obtaining sequence information from amplified nucleic acid products using a combination of primer pairs and / or determining the levels (e.g., relative and / or absolute levels) of amplified nucleic acid products and using the sequence information and / or level determination to identify the genus of microorganisms in the sample and the species of one or more microorganisms in the sample, and optionally their relative and / or absolute levels, to characterize, evaluate, profile and / or measure the microbial population and / or its components or compositions in the sample.
[0011] Also provided herein are compositions and methods, and combinations, kits and systems comprising the compositions and methods, for diagnosis and / or treatment, symptom relief or prevention of microbial (e.g., bacterial) imbalance and / or dysbiosis in a subject and conditions, disorders and diseases associated therewith. For example, in some cases, the microbial imbalance and / or dysbiosis is in the digestive tract or gastrointestinal tract of the subject. In some embodiments, the diagnosis and / or treatment, symptom relief or prevention of microbial imbalance and / or dysbiosis in a subject comprises: nucleic acid amplification of nucleic acid in or from one or more samples from the subject; obtaining sequence information of the nucleic acid amplification product; detecting whether there are one or more microorganisms in the sample; and detecting whether there are one or more microorganisms at disproportionate levels in the sample, wherein the presence of one or more microorganisms at disproportionate levels indicates microbial imbalance and / or dysbiosis in the subject. In some embodiments of treating a subject with microbial imbalance and / or dysbiosis, a subject with disproportionate levels of one or more microorganisms is treated to establish a balance of microorganisms or biosis in the subject. In some embodiments, a plurality of nucleic acid primers are used to perform the amplification. In some embodiments, the presence or absence of one or more microorganisms in the detection sample comprises the genus of one or more microorganisms in the sample. In some embodiments, the presence or absence of one or more microorganisms in the detection sample comprises the genus of one or more microorganisms in the sample and the one or more species of microorganisms in the sample. In some embodiments, the combination of a nucleic acid primer pair that specifically amplifies a predetermined unique nucleic acid sequence of a microorganism genome and / or amplifies a primer pair that appears in a homologous gene or genomic region to amplify, the homologous gene or genomic region is shared by a variety of microorganisms, but there are differences between different microorganisms. In a specific embodiment, the primer pair that amplifies the nucleic acid sequence that appears in a homologous gene or genomic region comprises one or more primer pairs, and the nucleic acid that amplifies the sequence of one or more hypervariable regions of a prokaryotic 16S rRNA gene is included. In some embodiments, the amplification is carried out in a single multiple amplification reaction mixture. In some embodiments, obtaining the nucleotide sequence information of the nucleic acid amplification product comprises using the nucleic acid provided herein as a probe (e.g., a detectable or labeled probe) to detect the nucleotide sequence. In some embodiments, obtaining the nucleotide sequence information of the nucleic acid amplification product comprises sequencing the amplified product. In some embodiments, detecting whether there are one or more microorganisms of disproportionate levels in the sample comprises determining the relative level of one or more microorganisms in the sample.In some embodiments, treating a subject having disproportionate levels of one or more microorganisms comprises administering to the subject one or more microorganisms and / or one or more compositions that reduce the levels of or eliminate certain microorganisms, e.g., a composition comprising an antibiotic.
[0012] According to the teachings and principles embodied in the present application, novel methods, systems and non-transitory machine-readable storage media are provided to compress reference sequence databases used to map sequence reads for analysis and profiling of microbial populations. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The accompanying drawings, which are incorporated in and form a part of the specification, illustrate one or more exemplary embodiments and serve to explain the principles of various exemplary embodiments. The drawings are exemplary and explanatory only and should not be construed as limiting or restrictive in any way.
[0014] Figure 1 1 is a diagram depicting the structure of a prokaryotic 16S ribosomal RNA (rRNA) gene showing nine hypervariable regions 101 (boxes labeled V1-V9) interspersed between conserved regions of the gene (unlabeled white boxes). The arrows above the gene depict forward and reverse primers that hybridize to eight sequences targeting conserved hypervariable segment regions 102 at designated locations to amplify the hypervariable regions between the arrows.
[0015] Figure 2 A workflow for analyzing nucleotide sequence information generated in the methods provided herein is described.
[0016] Figure 3 is a block diagram of an exemplary workflow for processing sequence read data obtained from sequencing of amplified nucleic acids produced by amplifying microbial nucleic acids using 16S rRNA gene primers.
[0017] Figure 4 is a block diagram of an exemplary workflow for processing sequence read data obtained from sequencing of amplified nucleic acids produced by amplifying microbial nucleic acids using species-specific nucleic acid primers.
[0018] Figure 5A and Figure 5B Each is a graphical representation of the results of an analysis using Spearman's rho comparing the results of the analysis using a pool of 16S rRNA gene primers for amplification ( Figure 5A ) or using a pool of species-specific primers for amplification ( Figure 5B ) Sequencing data of four replicate aliquots of a DNA amplicon library generated from six bacterial samples.
[0019] Fig. 6Ais a bar graph showing the results of read analysis from sequencing of a DNA amplicon library generated from a mixed bacterial sample using a 16S rRNA gene primer pool for nucleic acid amplification. The number of reads mapped to different bacterial genera is shown. Figure 6B The analytics from the analysis are depicted.
[0020] Figure 7 Shown is a graph of the Spearman correlation coefficient analysis of sequencing results of four replicate aliquots of a library generated from a bacterial DNA mixture (sample No. 1 (MSA1002)) using a pool of 16S primers for amplification.
[0021] Fig. 8A is a bar graph showing the results of read analysis from sequencing of a DNA amplicon library generated from a mixed bacterial sample using a pool of species-specific gene primers for nucleic acid amplification. The number of reads mapped to different bacterial species is shown. Figure 6B The analytics from the analysis are depicted.
[0022] Fig. 9 Shown is a graph of the Spearman correlation coefficient analysis of sequencing results of four replicate aliquots of a library generated from a bacterial DNA mixture (sample No. 1 (MSA1002)) using a species-specific primer pool for amplification.
[0023] Fig.10 is a block diagram depicting various embodiments of a nucleic acid sequencing platform, for example, a sequencing instrument 200 may include a fluid delivery and control unit 202 , a sample processing unit 204 , a signal detection unit 206 , and a data acquisition, analysis, and control unit 208 . DETAILED DESCRIPTION
[0024] The following description of various exemplary embodiments is exemplary and illustrative only and should not be interpreted as limited or restrictive in any way. Other embodiments, features, objects and advantages of the present teaching will be apparent from the specification and drawings and claims. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those of ordinary skill in the art to which the present invention belongs.
[0025] Provided herein are compositions and methods, and combinations, kits and systems comprising the compositions and methods, for amplification, detection, characterization, evaluation, analysis and / or measurement of nucleic acids, such as nucleic acids of microorganisms (including microorganisms, such as bacteria). Provided herein are compositions and methods that can be highly sensitive, specific, accurate, and repeatable for detecting and identifying one or more microorganisms in a sample containing a complex microbial population and other biological materials (e.g., cells that are not microorganisms). The compositions and methods further provide accurate determination of the relative and / or absolute levels or abundance of different microorganisms in such samples. These and other aspects of the compositions and methods provided herein make them ideally suitable for, for example, many methods, including but not limited to accurate and comprehensive methods for assessing or characterizing a microbial population in a sample (e.g., a biological sample), or for diagnosing and / or treating microbial imbalances and / or dysbiosis in a subject, alleviating the symptoms of the microbial imbalances and / or dysbiosis, or preventing the microbial imbalances and / or dysbiosis, including such methods described and provided herein. In some embodiments, the compositions and methods are further capable of achieving multiple (including highly multiple) amplification of microbial nucleic acids in a single amplification reaction mixture, thereby providing rapid, high-throughput, sensitive and easily identifiable amplification of nucleic acids from a large number of different microorganisms, which can be found, for example, in many different samples, such as from food, water, soil and animal (e.g., human) samples, such as microbiota of biological fluids (e.g., saliva, sputum, mucus, blood, urine, semen), tissues, skin, respiratory tract, urogenital tract and animal digestive tract (e.g., intestine). In some embodiments, the methods provided herein include multiple next-generation sequencing workflows for accurate, sensitive, high-throughput assessment, characterization or analysis of microbial populations, which are used, for example, to associate the microbial composition of a subject (e.g., the microbiota of the digestive tract, gastrointestinal tract, digestive tract or part thereof of a subject) with health status and disease or condition.
[0026] definition
[0027] As used herein, the terms "comprises / comprising," "includes / including," "has / having," or any other variation thereof, are intended to cover a non-exclusive inclusion. For example, a process, method, article, or apparatus that comprises a list of features is not necessarily limited to only those features but may include other features not expressly listed or inherent to such process, method, article, or apparatus. Further, unless expressly stated to the contrary, "or" refers to an inclusive or and not an exclusive or.
[0028] As used herein, "organism" refers to a life form or organism. Examples of organisms include microorganisms, unicellular organisms, multicellular organisms, plants, and animals. Examples of animals include insects, fish, birds, and mammals, including humans and non-human mammals.
[0029] As used herein, "subject" refers to an organism, typically an animal, such as a human or a non-human animal, such as a mammal, that is the focus of research, investigation, treatment, and / or from which information and / or material (e.g., a sample or specimen) is sought and / or obtained. In some cases, a subject may be a patient.
[0030] As used herein, "microorganism", which is used interchangeably with "microbe" herein, refers to an organism having a microscopic or submicroscopic size. Examples of microorganisms include bacteria, archaea, protists, and fungi. Many microorganisms are unicellular and are capable of division and proliferation. Microorganisms include prokaryotes, such as bacteria, and non-prokaryotic organisms, such as eukaryotic organisms.
[0031] As used herein, "microbiome" refers to a collection, colony or community of microorganisms that inhabit a specific biological niche or ecosystem. The environment in which the microbiome is found includes soil, water, hydrothermal vents and hosts, such as animal hosts. For example, the human microbiome is composed of an array of microorganisms that settle in humans (such as on or within human tissues and biological fluids). There are several habitats in the human microbiome, such as skin, oral mucosa, respiratory tract, conjunctiva, urogenital tract and digestive tract or digestive tube, or gastrointestinal tract, commonly referred to as "intestinal" microbiome. The genetic components (e.g., genes and genomes) of all microbial cells in the microbiome are referred to as "microbiome" in this article.
[0032] As used herein, "sensitivity" with respect to detection and / or identification of microorganisms (e.g., bacteria) in a sample is a performance measure of a method for detecting or identifying microorganisms, such as at the genus and / or species level, i.e., based on calculating the true positive rate, i.e., the proportion of actual positives that are correctly identified. For example, one method of determining the sensitivity of nucleic acid sequencing and an analytical method for detecting or identifying microorganisms is to perform the method on a known control sample of microorganisms and then determine the percentage of sequence reads that are correctly and unambiguously assigned to a specific genus or species in the sample. The higher the sensitivity of the detection or identification, the fewer the number of failures to detect the actual presence of a specific genus or species in the sample.
[0033] As used herein, "specificity" with respect to detection and / or identification of a microorganism (e.g., bacteria) in a sample is a performance measure of a method for detecting or identifying a microorganism, such as at the genus and / or species level, i.e., based on calculating the true negative rate, i.e., the proportion of actual negatives that are correctly identified. For example, one method for determining the specificity of nucleic acid sequencing and an analytical method for detecting or identifying a microorganism is to perform the method on a known control sample of microorganisms known not to contain a specific microorganism, and then determine the percentage of sequence reads that are incorrectly assigned to a specific genus or species that is not present in the sample. The higher the specificity of the detection or identification, the lower the number of errors in identifying a specific genus or species in the sample.
[0034] As used herein, the term "nucleic acid" refers to natural nucleic acids, artificial nucleic acids, analogs thereof, or combinations thereof, including polynucleotides and oligonucleotides. As used herein, the terms "polynucleotide" and "oligonucleotide" are used interchangeably and refer to single-stranded, double-stranded, partially double-stranded polymers of nucleotides, including but not limited to 2'-deoxyribonucleotides (DNA) and ribonucleotides (RNA) connected by internucleotide phosphodiester bonds such as 3'-5' and 2'-5', reverse connections such as 3'-3' and 5'-5', branched structures or analog nucleic acids. Examples of partially double-stranded nucleic acids include, for example, double-stranded molecules with 5' and / or 3' single-stranded overhangs. Polynucleotides have associated counterions, such as H + NH 4+ , trialkylammonium, Mg 2+ 、Na + Etc. Oligonucleotides can be composed entirely of deoxyribonucleotides, entirely of ribonucleotides, or of chimeric mixtures thereof. Oligonucleotides can include nucleobases and sugar analogs. The size of polynucleotides is generally in the range of several, such as 5-40 monomer units (when it is more commonly referred to as oligonucleotides in the art) to thousands of monomer nucleotide units (when it is more commonly referred to as polynucleotides in the art); however, for the purposes of this disclosure, both oligonucleotides and polynucleotides can have any suitable length. Unless otherwise indicated, whenever an oligonucleotide sequence is represented, it is understood that nucleotides are in 5' to 3' order from left to right, and "A" represents deoxyadenosine, "C" represents deoxycytidine, "G" represents deoxyguanosine, "T" represents thymidine, and "U" represents deoxyuridine. As standard in the art, the letters A, C, G, and T can be used to refer to bases themselves, nucleosides, or nucleotides including bases. Oligonucleotides are referred to as having a "5' end" and a "3' end" because a single nucleotide is typically reacted to form an oligonucleotide by linking (optionally via a phosphodiester or other suitable linkage) the 5' phosphate or equivalent of one nucleotide to the 3' hydroxyl or equivalent of its adjacent nucleotide.
[0035] As used herein, the term "nucleotide" and its variants include any compound, including but not limited to any naturally occurring nucleotide or its analog, which can hybridize with another nucleotide and / or can be bound to a polymerase, or can be polymerized by a polymerase. Typically but not necessarily, the selective binding of nucleotides to polymerases is followed by the polymerization of nucleotides into nucleic acid chains by polymerases; however, occasionally, nucleotides can be dissociated from polymerases without becoming incorporated into nucleic acid chains, which is referred to herein as "non-productive" events. Such nucleotides include not only naturally occurring nucleotides, but also any analogs (regardless of their structure) that can selectively bind to polymerases or can be polymerized by polymerases. Although naturally occurring nucleotides generally include bases, sugars, and phosphate moieties, the nucleotides of the present disclosure may include compounds that do not have any, some, or all of such moieties. In some embodiments, nucleotides may optionally include phosphorus atom chains including three, four, five, six, seven, eight, nine, ten, or more phosphorus atoms. In some embodiments, the phosphorus chain may be connected to any carbon of a sugar ring, such as a 5' carbon. The phosphorus chain may be connected to sugars with an intermediate O or S. In one embodiment, one or more phosphorus atoms in the chain can be part of a phosphate group having P and O. In another embodiment, the phosphorus atoms in the chain can be replaced by intermediate O, NH, S, methylene, substituted methylene, ethylene, substituted ethylene, CNH 2 , C(O), C(CH 2 ), CH 2 CH 2 , or C(OH)CH 2 R (where R can be 4-pyridine or 1-imidazole) are connected together. In one embodiment, the phosphorus atoms in the chain can have O, BH 3or S side groups. In the phosphorus chain, the phosphorus atom with a side group other than O can be a substituted phosphate group. In the phosphorus chain, the phosphorus atom with an intermediate atom other than O can be a substituted phosphate group. Some examples of nucleotide analogs are described in U.S. Pat. No. 7,405,281 to Xu. In some embodiments, the nucleotide includes a label and is referred to herein as a "labeled nucleotide"; the label of the labeled nucleotide is referred to herein as a "nucleotide label". In some embodiments, the label may be in the form of a fluorescent dye attached to the terminal phosphate group (i.e., the phosphate group farthest from the sugar). Some examples of nucleotides that can be used in the disclosed methods and compositions include, but are not limited to, ribonucleotides, deoxyribonucleotides, modified ribonucleotides, modified deoxyribonucleotides, ribonucleotide polyphosphates, deoxyribonucleotide polyphosphates, modified ribonucleotide polyphosphates, modified deoxyribonucleotide polyphosphates, peptide nucleotides, modified peptide nucleotides, metal nucleosides, phosphonic acid nucleosides, and modified phosphate-sugar backbone nucleotides, analogs, derivatives or variants of the aforementioned compounds, etc. In some embodiments, nucleotides may include a non-oxygen moiety (such as a thiol or borane moiety) in place of an oxygen moiety that bridges the α-phosphate and sugar of a nucleotide, or the α and β-phosphate of a nucleotide, or the β and γ-phosphate of a nucleotide, or between any other two phosphates of a nucleotide, or any combination thereof. "Nucleotide 5'-triphosphate" refers to a nucleotide having a triphosphate group at the 5' position and sometimes represented as "NTP" or "dNTP" and "ddNTP" to specifically indicate the structural features of ribose. The triphosphate group may include sulfur to replace various oxygens, such as α-thiol-nucleotide 5'-triphosphate. For a review of nucleic acid chemistry, see: Shabarova, Z. and Bogdanov, A., Advanced Organic Chemistry of Nucleic Acids, German Chemical Society Press (VCH), New York, 1994.
[0036] As used herein, the term "hybridization" is consistent with its use in the art and refers to a method for making two nucleic acid molecules perform base pairing interactions. When any part of a nucleic acid molecule is base paired with any part of another nucleic acid molecule, the two nucleic acid molecules are called hybridized; it is not necessary for the two nucleic acid molecules to hybridize across their entire corresponding lengths and in some embodiments, at least one nucleic acid molecule may include a portion that does not hybridize with another nucleic acid molecule. "Hybridization conditions" are conditions (e.g., temperature, ionic strength, etc.) suitable for hybridization of two nucleic acids containing nucleotide sequences capable of base pairing interactions. The phrase "hybridization under stringent conditions" and its variants refer to conditions under which two nucleic acid sequences (e.g., target-specific primers and target sequences) hybridize in the presence of high hybridization temperatures and low ionic strengths. In an exemplary embodiment, stringent hybridization conditions include an aqueous environment containing about 30mM magnesium sulfate, about 300mMTris-sulfate (pH8.9) and about 90mM ammonium sulfate at about 60-68°C or its equivalent. As used herein, the phrase "standard hybridization conditions" and its variants refer to conditions under which two nucleic acids hybridize in the presence of high hybridization temperatures and low ionic strengths. In one exemplary embodiment, standard hybridization conditions comprise an aqueous environment containing about 100 mM magnesium sulfate, about 500 mM Tris-sulfate (pH 8.9), and about 200 mM ammonium sulfate at about 50-55°C, or its equivalent.
[0037] As used herein, the terms "identity" and "identical" and variations thereof, when used to refer to two or more nucleic acid sequences, refer to the sequence similarity of two or more sequences (e.g., nucleotides or polypeptide sequences). In the case of two or more homologous sequences, the identity, similarity or homology percentage of a sequence or its subsequence indicates the same (i.e., about 70% identity or more, about 75%, 80%, 85%, 90%, 95%, 98% or 99% identity) percentage of all monomeric units (e.g., nucleotides or amino acids). The identity percentage can be in a specified region, when compared and aligned for maximum correspondence in a comparison window, or in a specified region, as measured using BLAST or BLAST 2.0 sequence comparison algorithms with default parameters described below or by manual alignment and visual inspection. When there is at least 85% identity at the amino acid level or nucleotide level, the sequence is referred to as "substantially identical". Preferably, identity is present in a region of at least about 25, 50 or 100 residues long or over the entire length of at least one compared sequence. Typical algorithms for determining percentages of sequence identity and sequence similarity are BLAST and BLAST 2.0 algorithms, which are described in Altschul et al., Nuc. Acids Res. 25:3389-3402 (1977). Other methods include the algorithms described by Smith and Waterman, Adv. Appl. Math. 2:482 (1981) and Needleman and Wunsch, J. Mol. Biol. 48:443 (1970), among others. Another indication that two nucleic acid sequences are substantially identical is that the two molecules or their complements hybridize to each other under stringent hybridization conditions.
[0038] As used herein, the terms "complementary" and "complement" and their derivatives refer to any two or more nucleic acid sequences (e.g., part or all of a template nucleic acid molecule, a target sequence, and / or a primer) that undergo cumulative base pairing at two or more independent corresponding positions in an antiparallel orientation (e.g., in a hybridized double helix). Such base pairing can be performed according to any existing set of rules, such as according to the Watson-Crick base pairing rules or according to some other base pairing paradigm. Optionally, there may be "complete" or "overall" complementarity between the first and second nucleic acid sequences, wherein each nucleotide in the first nucleic acid sequence may perform a stabilizing base pairing interaction with a nucleotide in a corresponding antiparallel position on the second nucleic acid sequence. "Partial" complementarity describes a nucleic acid sequence in which at least 20% but less than 100% of the residues of a nucleic acid sequence are complementary to residues in another nucleic acid sequence. In some embodiments, at least 50% but less than 100% of the residues of a nucleic acid sequence are complementary to residues in another nucleic acid sequence. In some embodiments, at least 70%, 80%, 90%, 95% or 98% but less than 100% of the residues of a nucleic acid sequence are complementary to the residues in another nucleic acid sequence. When at least 85% of the residues in a nucleic acid sequence are complementary to the residues in another nucleic acid sequence, the sequence is referred to as "substantially complementary". In some embodiments, two complementary or substantially complementary sequences can hybridize to each other under standard or stringent hybridization conditions. "Non-complementary" describes a nucleic acid sequence in which less than 20% of the residues of one nucleic acid sequence are complementary to the residues in another nucleic acid sequence. When less than 15% of the residues in a nucleic acid sequence are complementary to the residues in another nucleic acid sequence, the sequence is referred to as "substantially non-complementary". In some embodiments, two non-complementary or substantially non-complementary sequences cannot hybridize to each other under standard or stringent hybridization conditions. "Mismatch" is present in any position of two non-complementary relative nucleotides. Complementary nucleotides are included in nucleotides that are effectively incorporated by DNA polymerases relative to each other during DNA replication under physiological conditions. In a typical embodiment, between the nucleobases of nucleotides and / or polynucleotides that are located in antiparallel positions to each other, complementary nucleotides can form base pairs with each other, such as AT / U and GC base pairs formed by specific Watson-Crick type hydrogen bonds, or base pairs formed by some other type of base pairing paradigm. The complementarity of other artificial base pairs can be based on other types of hydrogen bonds and / or hydrophobicity of the bases and / or shape complementarity between the bases.
[0039] As used herein, "sample" and its derivatives are used in their broadest sense and include any sample, culture, etc. that may contain a target composition (such as a target). In some embodiments, the sample includes cDNA, RNA, PNA, LNA, chimeric, hybrid or multiple forms of nucleic acid. The sample may include any biological, clinical, surgical, agricultural, atmospheric or water-based sample containing one or more organisms and / or nucleic acids. An example of a biological or clinical sample is a sample of the contents of an animal's digestive tract. The digestive tract is a continuous passage that starts from the mouth and ends at the anus, through which food and liquid are ingested, digested and absorbed, and waste is processed and eliminated. The digestive tract or digestive tract is also referred to as the gastrointestinal tract and intestinal tract in this article, and includes multiple organs. An example of a sample from the digestive tract is a fecal sample. In some cases, at least some of the nucleic acids in the sample may be contained in a cell. In some cases, nucleic acids can be extracted from one or more cells in the sample. In some cases, the term "nucleic acid sample" may refer to a sample containing nucleic acids that are or are not in a cell or organism and / or nucleic acids extracted from a sample. The term also encompasses any isolated nucleic acid sample, such as expressed RNA, fresh frozen or formalin fixed paraffin embedded nucleic acid specimens.
[0040] As used herein, "homologous" or "homologs" and their derivatives, when used in reference to a portion of a genome or gene, refer to a genomic segment or gene that shows a conserved sequence with substantial sequence similarity but also differs in sequence in multiple organisms (e.g., multiple organisms of a domain, kingdom, phylum, class, order, family, genus, and / or species). Examples of homologous genes include, but are not limited to, 16S rRNA genes, 18S rRNA genes, 23S rRNA genes, and ABC transporter genes.
[0041] As used herein, "unique", when used in reference to a nucleic acid sequence in an organism or group of organisms, refers to a nucleotide sequence of a nucleic acid (e.g., a segment or portion of a genome) in an organism or group of organisms that is sufficiently different from sequences in the genomes of other organisms or other groups of organisms that it can be used to selectively detect or identify members of the organism or group of organisms, and / or to distinguish members of the organism or group of organisms from some, most, most, or substantially all different organisms or organisms that do not belong to the group of organisms. Such unique sequences are also referred to herein as "signature sequences" or "signature regions" of the nucleic acid of the organism or group of organisms. For example, a nucleic acid sequence of nucleotides can be unique to an individual organism, unique to members of a strain of an organism species, unique to members of an organism species, unique to members of an organism genus, unique to members of an organism family, unique to members of an organism order, unique to members of an organism class, unique to members of an organism phylum, unique to members of an organism kingdom, and / or unique to members of an organism domain. Typically, the difference in a unique sequence is the identity and / or order of consecutive nucleotides or nucleobases in the sequence. In certain embodiments, a unique sequence is unique to an organism compared to or with respect to some specific groups of organisms (e.g., organisms in the same kingdom, phylum, class, order, family, genus, species), but may not be unique to the organism compared to the sum of all other organisms or all other organisms outside a specified group. A unique nucleotide sequence can be any length, e.g., between about 20 and 1000 nucleotides, 30 and 750 nucleotides, 40 and 500 nucleotides, 50 and 400 nucleotides, 50 and 350 nucleotides, 50 and 300 nucleotides, 50 and 250 nucleotides, 50 and 200 nucleotides, 50 and 150 nucleotides, or 50 and 100 nucleotides. In some embodiments, the unique nucleotide sequence can be about 1000 nucleotides or less, about 750 nucleotides or less, about 500 nucleotides or less, about 400 nucleotides or less, about 350 nucleotides or less, about 300 nucleotides or less, about 250 nucleotides or less, about 200 nucleotides or less, about 150 nucleotides or less, about 100 nucleotides or less, or about 50 nucleotides or less in length.In some embodiments, the length of a unique nucleotide sequence can be greater than about 25 nucleotides, greater than about 40 nucleotides, greater than about 50 nucleotides, greater than about 60 nucleotides, greater than about 70 nucleotides, greater than about 75 nucleotides, greater than about 90 nucleotides, greater than about 95 nucleotides, greater than about 100 nucleotides, greater than about 150 nucleotides, greater than about 175 nucleotides, greater than about 200 nucleotides, greater than about 250 nucleotides, greater than about 275 nucleotides, greater than about 300 nucleotides, greater than about 325 nucleotides, greater than about 350 nucleotides, or greater than about 400 nucleotides. In some embodiments, a unique sequence is useful for selectively detecting, identifying, and / or distinguishing an organism or a member of a group of organisms, particularly in the presence of nucleic acids from other organisms or organisms that are not members of the group of organisms, by binding to, hybridizing to, and / or being amplified by specific nucleic acid probes and / or primers that specifically or selectively or uniquely bind to, hybridize to, and / or amplify the unique sequence. For example, in some embodiments, a unique or signature sequence of an organism (e.g., a microorganism such as a bacterium) or a group of organisms is a sequence that has less than 60%, less than 65%, less than 70%, less than 75%, less than 80%, less than 81%, less than 82%, less than 83%, less than 84%, less than 85%, less than 86%, less than 87%, less than 88%, less than 89%, less than 90%, less than 91%, less than 92%, less than 93%, less than 94%, or less than 95% identity to a nucleotide sequence in a different organism or a specific group of organisms. In some embodiments, a unique sequence has less than 90% identity to a nucleotide sequence in a different organism or a specific group of organisms. In some embodiments, a unique or signature sequence of an organism (e.g., a microorganism such as a bacterium) or a group of organisms has less than 25%, less than 20%, less than 19%, less than 18%, less than 17%, less than 16%, less than 15%, less than 14%, less than 13%, less than 12%, less than 10% of the nucleotides match nucleotides in a nucleotide sequence of similar length in a different organism or a specific group of organisms. In some embodiments, a unique sequence has less than 17% of the nucleotides match nucleotides in a nucleotide sequence in a different organism or a specific group of organisms. In some embodiments, a unique sequence has less than 90% identity with a nucleotide sequence in a different organism or a specific group of organisms and has less than 17% of the nucleotides match nucleotides in a nucleotide sequence in a different organism or a specific group of organisms.In some embodiments, a unique sequence within a group of organisms (e.g., a bacterial species) is at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical in most or substantially all members of the group (e.g., strains of the species). In some embodiments, a unique sequence within a group of organisms is at least 95%, at least 96%, at least 97% identical in most or substantially all members of the group (e.g., strains of the species). In some embodiments, a specific identity of a unique sequence within a group of organisms is at least or greater than 75%, at least or greater than 80%, at least or greater than 85%, at least or greater than 90%, or at least or greater than 95% of the members of the group. In some embodiments, the nucleotide sequence of a unique sequence within a group of organisms (e.g., a bacterial species) has at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of matching nucleotides in the members of the group. In some embodiments, the nucleotide sequence of a unique sequence within a group of organisms has at least 95% nucleotide matching in the members of the group. In some embodiments, a unique sequence within a group of organisms (e.g., a bacterial species) is at least 95% identical in at least or greater than 90% of the members of the group and has at least 95% nucleotide matching.
[0042] As used herein, "synthesis" and its derivatives refer to reactions involving nucleotide polymerization by a polymerase, optionally in a template-dependent manner. The polymerase synthesizes oligonucleotides by transferring a nucleoside monophosphate from a nucleoside triphosphate (NTP), a deoxynucleoside triphosphate (dNTP), or a dideoxynucleoside triphosphate (ddNTP) to the 3' hydroxyl of an extended oligonucleotide chain. For purposes of this disclosure, synthesis comprises continuously extending hybridized adapters or target-specific primers via transfer of a nucleoside monophosphate from a deoxynucleoside triphosphate.
[0043] As used herein, "polymerase" and its derivatives refer to any enzyme that can catalyze the polymerization of nucleotides (including analogs thereof) into nucleic acid chains. Typically, but not necessarily, such nucleotide polymerization can be carried out in a template-dependent manner. Such polymerases may include, but are not limited to, naturally occurring polymerases and any subunits and truncations thereof, mutant polymerases, variant polymerases, recombinant, fused or otherwise engineered polymerases, chemically modified polymerases, synthetic molecules or assemblies, and any analogs, derivatives, or fragments thereof that retain the ability to catalyze such polymerization. Optionally, the polymerase may be a mutant polymerase comprising one or more mutations, the mutations involving replacement of one or more amino acids with other amino acids, insertion or deletion of one or more amino acids from the polymerase, or connection of two or more parts of the polymerase. Typically, the polymerase includes one or more active sites that can perform nucleotide binding and / or nucleotide polymerization catalysis. Some exemplary polymerases include, but are not limited to, DNA polymerases and RNA polymerases. As used herein, the term "polymerase" and its variations also refer to a fusion protein comprising at least two parts connected to each other, wherein the first part comprises a peptide that can catalyze the polymerization of nucleotides into a nucleic acid chain and the first part is connected to the second part comprising a second polypeptide. In some embodiments, the second polypeptide may comprise a reporter enzyme or a domain that enhances continuous synthesis. Optionally, the polymerase may have 5' exonuclease activity or terminal transferase activity. In some embodiments, the polymerase may optionally be reactivated, for example, by using heat, chemicals or by adding a new amount of the polymerase to the reaction mixture. In some embodiments, the polymerase may comprise a heat start polymerase that may optionally be reactivated or an aptamer-based polymerase.
[0044] As used herein, "amplify / amplifying" or "amplification reaction" and its derivatives refer to any action or process of replicating or copying at least a portion of a nucleic acid molecule (referred to as a template nucleic acid molecule, which may contain a target sequence) to at least one additional nucleic acid molecule. The additional nucleic acid molecule optionally comprises a sequence that is substantially identical or substantially complementary to at least some portions of the template nucleic acid molecule. The template nucleic acid molecule may be single-stranded or double-stranded and the other nucleic acid molecule may be independently single-stranded or double-stranded. In some embodiments, amplification comprises a template-dependent in vitro enzyme-catalyzed reaction for preparing at least one copy of at least a portion of a nucleic acid molecule or preparing at least one copy of a nucleic acid sequence complementary to at least a portion of a nucleic acid molecule. Amplification optionally comprises linear or exponential replication of nucleic acid molecules. In some embodiments, such amplification is performed using isothermal conditions; in other embodiments, such amplification may comprise thermal cycling. In some embodiments, amplification is multiplex amplification, which comprises simultaneously amplifying multiple target sequences in a single amplification reaction. At least some target sequences may be located on the same nucleic acid molecule or different target nucleic acid molecules contained in a single amplification reaction. In some embodiments, "amplification" comprises amplification of at least a portion of DNA- and RNA-based nucleic acids, either alone or in combination. The amplification reaction may comprise a single-stranded or double-stranded nucleic acid substrate, and may further comprise any amplification technology process known to those of ordinary skill in the art. In some embodiments, the amplification reaction comprises a polymerase chain reaction (PCR).
[0045] As used herein, "amplification conditions" and their derivatives refer to conditions suitable for amplifying one or more nucleic acid sequences. Such amplification can be linear or exponential. In some embodiments, the amplification conditions may include isothermal conditions or alternatively may include thermal cycling conditions or a combination of isothermal and thermal cycling conditions. In some embodiments, conditions suitable for amplifying one or more nucleic acid sequences include polymerase chain reaction (PCR) conditions. Typically, amplification conditions refer to a reaction mixture sufficient to amplify nucleic acids such as one or more target sequences or to amplify amplified target sequences (e.g., amplified target sequences connected to adapters) connected to one or more adapters. Amplification conditions include catalysts for amplification or for nucleic acid synthesis, such as polymerases; primers having a certain degree of complementarity with the nucleic acid to be amplified; and nucleotides, such as deoxyribonucleotide triphosphates (dNTPs), to promote extension of primers after hybridization with nucleic acids. Amplification conditions may require hybridization or annealing of primers to nucleic acids, extension of primers, and dissociation steps, such as denaturation, wherein the extended primers are separated from the nucleic acid sequences undergoing amplification. Typically, but not necessarily, amplification conditions may include thermal cycling; in some embodiments, amplification conditions include multiple cycles of repeated annealing, extension, and separation amplification steps. Typically, amplification conditions contain cations such as Mg ++ or Mn ++ (e.g. MgCl2 etc.), and may also contain a variety of ionic strength regulators.
[0046] As defined herein, "multiplex amplification" refers to the selective and non-random amplification of two or more target sequences in a sample using at least one specific primer. In certain embodiments, multiple amplification is performed so that some or all of the target sequences are amplified in a single reaction vessel. "Plexity" or "plex" of a given multiple amplification refers to the number of different target-specific sequences amplified during the single multiple amplification. In certain embodiments, the multiplex number can be about 12-fold, 24-fold, 48-fold, 74-fold, 96-fold, 120-fold, 144-fold, 168-fold, 192-fold, 216-fold, 240-fold, 264-fold, 288-fold, 312-fold, 336-fold, 360-fold, 384-fold or 398-fold.
[0047] The term "polymerase chain reaction" ("PCR") as used herein refers to the method of K.B. Mullis U.S. Pat. Nos. 4,683,195 and 4,683,202, which are incorporated herein by reference, describing a method for increasing the concentration of a segment of a target polynucleotide in a mixture of genomic RNA or cDNA without cloning or purification. This process for amplifying a target polynucleotide consists of introducing a large excess of two oligonucleotide primers into a DNA mixture containing the desired target polynucleotide, followed by a precise sequence of thermal cycles in the presence of a DNA polymerase. The two primers are complementary to their corresponding strands of the target double-stranded polynucleotide. In order to achieve amplification, the mixture is denatured, and then the primers are annealed to their complementary sequences within the target polynucleotide molecule. After annealing, the primers are extended with a polymerase to form a pair of new complementary strands. The steps of denaturation, primer annealing and polymerase extension can be repeated multiple times (i.e., denaturation, annealing and extension constitute one "cycle"; there can be many "cycles") to obtain a high concentration of amplified segments of the desired target polynucleotide. The length of the amplified segment (amplicon) of the desired target polynucleotide is determined by the relative position of the primers relative to each other, and therefore this length is a controllable parameter. By means of repeating the process, the method is called "polymerase chain reaction" (hereinafter referred to as "PCR"). Because the desired amplified segment of the target polynucleotide becomes the main nucleic acid sequence in the mixture (in terms of concentration), it is called "PCR amplified". As defined herein, the target nucleic acid molecules in a sample containing multiple target nucleic acid molecules are amplified via PCR. In an improvement of the method discussed above, the target nucleic acid molecules can be PCR amplified using multiple different primer pairs, and in some cases, one or more primer pairs for each target target nucleic acid molecule, thereby forming a multiple PCR reaction. Using multiple PCR, it is possible to simultaneously amplify multiple target nucleic acid molecules from a sample to form an amplified target sequence. It is also possible to hybridize with a labeled probe by several different methods (e.g., quantification with a bioanalyzer or qPCR; incorporation of biotinylated primers followed by avidin-enzyme conjugate detection; 32 The amplified target sequence can be detected by incorporating a P-labeled deoxynucleotide triphosphate (such as dCTP or dATP) into the amplified target sequence. Any oligonucleotide sequence can be amplified with an appropriate primer set to achieve amplification of target nucleic acid molecules from RNA, cDNA, formalin-fixed paraffin-embedded DNA, fine-needle biopsies, and a variety of other sources. In particular, the amplified target sequence generated by the multiplex PCR process as disclosed herein is itself an effective substrate for subsequent PCR amplification or a variety of downstream assays or operations.
[0048] As used herein, "reamplifying" or "reamplification" and its derivatives refer to any process (referred to in some embodiments as "secondary" amplification or "reamplification") by further amplifying at least a portion of an amplified nucleic acid molecule via any suitable amplification process, thereby producing a re-amplified nucleic acid molecule. The secondary amplification need not be the same as the original amplification process that produced the amplified nucleic acid molecule; nor need the re-amplified nucleic acid molecule be identical or completely complementary to the amplified nucleic acid molecule; all that is required is that the re-amplified nucleic acid molecule contains at least a portion of the amplified nucleic acid molecule or its complement. For example, re-amplification may involve the use of different amplification conditions and / or different primers, including target-specific primers that are different from those of the primary amplification.
[0049] As used herein, when used with respect to a given primer, the term "extension" and variations thereof include any in vivo or in vitro enzymatic activity characteristic of a given polymerase involving the polymerization of one or more nucleotides at the end of an existing nucleic acid molecule. Typically, but not necessarily, such primer extension is performed in a template-dependent manner; during template-dependent extension, the ordering and selection of bases are performed by a given base pairing rule, which may include a Watson-Crick type base pairing rule, or alternatively (and particularly in the case of an extension reaction involving nucleotide analogs) by some other type of base pairing paradigm. In a non-limiting example, extension is performed by a polymerase via the polymerization of nucleotides at the 3'OH end of a nucleic acid molecule.
[0050] As used herein, the term "portion" and variations thereof, when used in reference to a given nucleic acid molecule (e.g., a primer or template nucleic acid molecule), includes any number of contiguous nucleotides within the length of the nucleic acid molecule (including a portion or the entire length of the nucleic acid molecule).
[0051] As used herein, "target sequence" or "target sequence of interest" and derivatives thereof refer to any single-stranded or double-stranded nucleic acid sequence that can be combined, hybridized, amplified and / or synthesized according to the present disclosure, including, for example, any nucleic acid sequence suspected to be present, expected to be present, or likely to be present in a sample. In some embodiments, the target sequence exists in a double-stranded form and includes at least a portion of a specific nucleotide sequence to be combined, hybridized, amplified and / or synthesized, or its complement, before adding a specific primer or an attached adapter. In some embodiments, the target sequence is part of a target. For example, a target nucleic acid sequence can be a sequence located in a target gene, a target genome, and / or a target organism (e.g., bacteria, or a specific family, genus, or species of a target organism, such as Ruminococcus family, Ruminococcus genus, and R. gnavus species). The target sequence can include a nucleic acid to which a primer for an amplification or synthesis reaction can hybridize before being extended by a polymerase. In some cases, a target sequence is a sequence adjacent to and contiguous with a sequence hybridized with a primer for amplifying a target sequence. In some embodiments, the terms refer to a nucleic acid sequence, the sequence identity, order or position of nucleotides of which is determined by one or more of the methods of the present disclosure.
[0052] As used herein, "amplified target sequence" and its derivatives refer to the nucleic acid sequence produced by using specific primers and method amplification / amplification target sequence provided herein. The amplified target sequence can be any of synonymous (positive strand produced in the second round and subsequent even-numbered rounds of amplification) or antisense (i.e., negative strand produced during the first round and subsequent odd-numbered rounds of amplification) relative to the target sequence. In certain embodiments, the amplified target sequence and any part of another amplified target sequence in the reaction generally have less than 50% complementarity. As used herein, "amplicon" refers to the total nucleic acid produced by amplification using primers and methods as provided herein. In some cases, the amplicon may be identical to the target sequence. In some cases, when the target nucleic acid sequence is defined as not comprising a primer sequence, the amplicon comprises the amplified target sequence and the primers for amplifying the target sequence located at each end of the amplified target sequence. In this case, the target sequence can be referred to as the "insert" of the amplicon.
[0053] As used herein, the terms "primer", "probe" and derivatives thereof refer to any polynucleotides that can hybridize with a target target sequence. In some embodiments, primers can also be used to initiate nucleic acid synthesis. Typically, primers serve as substrates on which nucleotides can be polymerized by a polymerase; however, in some embodiments, primers can become incorporated into the synthesized nucleic acid chain and provide a site to which another primer can hybridize to initiate the synthesis of a new chain complementary to the synthesized nucleic acid molecule. A primer or probe can include any combination of nucleotides or their analogs, which can optionally be connected to form a linear polymer of any suitable length. In some embodiments, a primer is a single-stranded oligonucleotide or polynucleotide. (For purposes of this disclosure, the terms "polynucleotide" and "oligonucleotide" can be used interchangeably herein and may not necessarily indicate any difference in length between two nucleotides). In some embodiments, a primer or probe is single-stranded, but it can also be double-stranded. A primer or probe is optionally produced naturally, such as in a purified restriction digest, or can be produced synthetically. In some embodiments, when exposed to amplification or synthesis conditions, the primer serves as a starting point for amplification or synthesis; such amplification or synthesis can be performed in a template-dependent manner and optionally forms a primer extension product complementary to at least a portion of the target sequence. Exemplary amplification or synthesis conditions may include contacting the primer with a polynucleotide template (e.g., a template comprising a target sequence), nucleotides and an inducing agent (such as a polymerase) at a suitable temperature and pH to induce polymerization of nucleotides on the end of the target-specific primer. If double-stranded, the primer or probe may be optionally treated to separate its chain before being used to prepare the primer extension product. In some embodiments, the primer probe is an oligodeoxynucleotide or an oligoribonucleotide. In some embodiments, the primer or probe may include one or more nucleotide analogs. The precise length and / or composition (including sequence) of the primer or probe may affect many properties, including melting temperature (Tm), GC content, formation of secondary structure, repeated nucleotide motifs, predicted length of primer extension products, coverage across target nucleic acid molecules, the number of primers present in a single amplification or synthesis reaction, the presence of nucleotide analogs or modified nucleotides in the primer, etc. In certain embodiments, primers can be paired with compatible primers in amplification or synthesis reaction to form a primer pair consisting of a forward primer and a reverse primer. In certain embodiments, the forward primer of the primer pair comprises a sequence substantially complementary to at least a portion of a chain of a nucleic acid molecule, and the reverse primer of the primer pair comprises a sequence substantially identical to at least a portion of the chain. In certain embodiments, the forward primer and the reverse primer can hybridize with the relative strands of the nucleic acid double helix. Optionally, the forward primer triggers the synthesis of the first nucleic acid strand, and the reverse primer triggers the synthesis of the second nucleic acid strand, wherein the first and second strands are substantially complementary to each other, or can hybridize to form a double-stranded nucleic acid molecule. In certain embodiments, one end of the amplification or synthesis product is defined by the forward primer and the other end of the amplification or synthesis product is defined by the reverse primer.In some embodiments, when it is necessary to amplify or synthesize lengthy primer extension products (such as amplification exons, coding regions or genes), several primer pairs spanning the desired length can be generated to achieve sufficient amplification of the region. In some embodiments, primers or probes can include one or more cleavable groups. Primers and probes can be of any length. In some embodiments, the length of the probe can be about 200 nucleotides or less, 175 nucleotides or less, 150 or less nucleotides, 125 nucleotides or less, 100 or less nucleotides, 90 nucleotides or less, 80 or less nucleotides, 75 nucleotides or less, 70 or less nucleotides, 60 nucleotides or less, 55 nucleotides or less, 50 nucleotides or less, 40 nucleotides or less, 35 nucleotides or less, 30 nucleotides or less, 25 nucleotides or less, 20 nucleotides or less, 15 nucleotides or less or 10 nucleotides or less. In certain embodiments, the primer length is within the range of about 10 to about 60 nucleotides, about 12 to about 50 nucleotides, and about 15 to about 40 nucleotides in length. Typically, when exposed to amplification conditions in the presence of dNTPs and a polymerase, the primer can hybridize with the corresponding target sequence and perform primer extension. In some cases, a part of a specific nucleotide sequence or primer is known at the beginning of the amplification reaction or can be determined by one or more of the methods disclosed herein. In certain embodiments, the primer comprises one or more cleavable groups at one or more positions within the primer. In certain embodiments, the mixture of primers can be a degenerate primer. A degenerate primer is a primer with a similar sequence but different at one or more nucleotide positions, so that one primer may have A at the position, another may have G at the same position, another may have T at the same position, and a fourth primer may have C at the same position. Probes and / or primers can be labeled. Labeling is often used to detect a primer or probe that has been combined or hybridized with another nucleic acid, for example, in order to detect a specific sequence that a primer or probe specifically binds. Compositions and methods for labeling nucleic acids for use as detectable probes are known in the art and include attaching a reporter gene or a signal generating moiety to the probe. Examples of detectable labels include, but are not limited to, fluorescent, luminescent, chemiluminescent, chromogenic, radioactive, and colorimetric moieties. The label may be directly detectable or may be part of a system for generating a detectable signal.
[0054] As used herein, when referring to processes such as amplification, binding or hybridization, "capable" refers to the ability of a nucleic acid (e.g., primer or primer pair) to interact with another nucleic acid (e.g., target nucleic acid, target sequence, template) in a manner that performs, participates in the performance and / or completes the process. For example, a nucleic acid that can bind to another nucleic acid or other molecule by intermolecular force or bond can form a stable connection with another nucleic acid or molecule. A nucleic acid that can hybridize with another nucleic acid can interact with another nucleic acid by base pairing. In certain embodiments, nucleic acids can hybridize under low or high stringency conditions. A nucleic acid that can amplify another nucleic acid can act as a primer in a polymerization reaction that causes nucleic acid extension and produces a complement of a template nucleic acid chain, which can be a copy of the relative chain of a template nucleic acid chain. If a nucleic acid can bind to a certain target molecule, hybridize with a certain target nucleic acid and / or amplify a certain target nucleic acid and does not substantially bind, hybridize and / or substantially amplify a molecule or nucleic acid that is not a target molecule or nucleic acid, then nucleic acid can specifically or selectively bind, hybridize and / or amplify. In some cases, such binding, hybridization, and / or amplification is referred to as "uniquely" binding, hybridization, and / or amplifying a target molecule or nucleic acid.
[0055] As used herein, the term "individually" when used to refer to amplifying nucleic acids refers to primers or primer pairs that are used to amplify a specific defined region of a nucleic acid (e.g., a gene) without amplifying another region of the nucleic acid. For example, primer pairs that individually amplify different hypervariable regions of a 16SrRNA gene each amplify only a single hypervariable region, thereby generating a separate amplicon for each different region, and do not generate an amplicon containing multiple hypervariable regions.
[0056] As defined herein, "cleavable group" refers to any part that can be cut under appropriate conditions after being incorporated into nucleic acid. For example, a cleavable group can be incorporated into a target-specific primer, an amplified sequence, an adapter or a nucleic acid molecule of a sample. In an exemplary embodiment, a target-specific primer can include a cleavable group, which becomes incorporated into the product of amplification and is subsequently cut after amplification, thereby removing a portion or all of the target-specific primers from the amplified product. A cleavable group can be cut or otherwise removed from a target-specific primer, an amplified sequence, an adapter or a nucleic acid molecule of a sample by any acceptable means. For example, a cleavable group can be removed from a target-specific primer, an amplified sequence, an adapter or a nucleic acid molecule of a sample by enzymatic, thermal, photo-oxidation or chemical treatment. In one aspect, a cleavable group can include a non-naturally occurring nucleobase. For example, an oligodeoxynucleotide can include one or more RNA nucleobases, such as uracil that can be removed by uracil glycosylase. In some embodiments, the cleavable group may comprise one or more modified nucleobases (such as 7-methylguanine, 8-oxo-guanine, xanthine, hypoxanthine, 5,6-dihydrouracil or 5-methylcytosine) or one or more modified nucleosides (i.e., 7-methylguanosine, 8-oxo-deoxyguanosine, xanthine nucleoside, inosine, dihydrouridine or 5-methylcytosine). The modified nucleobase or nucleotide may be removed from the nucleic acid by enzymatic, chemical or thermal means. In one embodiment, the cleavable group may be included in a portion that can be removed from the primer after amplification (or synthesis) when exposed to ultraviolet light (i.e., bromodeoxyuridine). In another embodiment, the cleavable group may comprise methylated cytosine. Typically, methylated cytosine may be cleaved by a primer, such as after induced amplification (or synthesis), when treated with sodium bisulfite. In some embodiments, the cleavable portion may comprise a restriction site. For example, the primer or target sequence can include a nucleic acid sequence that is specific to one or more restriction enzymes, and after amplification (or synthesis), the primer or target sequence can be treated with one or more restriction enzymes so that the cleavable group is removed. Typically, one or more cleavable groups and a target-specific primer, amplified sequence, adapter or nucleic acid molecule of a sample can be included at one or more positions.
[0057] As used herein, "cleavage step" and its derivatives refer to any process by which a cleavable group is cleaved or otherwise removed from a target-specific primer, amplified sequence, adapter, or nucleic acid molecule of a sample. In some embodiments, the cleavage step may involve a chemical, thermal, photooxidative, or digestive process.
[0058] In some embodiments, a primer is a single-stranded or double-stranded polynucleotide, typically an oligonucleotide, comprising at least one sequence that is at least 50% complementary, typically at least 75% complementary, or at least 85% complementary, more typically at least 90% complementary, more typically at least 95% complementary, more typically at least 98% or at least 99% complementary, or 100% complementary or identical to at least a portion of a nucleic acid molecule comprising a target sequence. In this case, the primer and the target sequence are described as "corresponding" to each other, and in some cases, the primer may be referred to as "directed against" the target sequence. In some embodiments, the primer is capable of hybridizing to at least a portion of its corresponding target sequence (or the complement of the target sequence); such hybridization may optionally be performed under standard hybridization conditions or under stringent hybridization conditions. In some embodiments, the primer cannot hybridize to the target sequence or its complement, but is capable of hybridizing to a portion of the nucleic acid chain comprising the target sequence or its complement, for example, a sequence upstream or downstream of the target sequence or adjacent to the target sequence. In some embodiments, the primer comprises at least one sequence that is at least 75% complementary, typically at least 85% complementary, more typically at least 90% complementary, more typically at least 95% complementary, more typically at least 98% complementary, or more typically at least 99% complementary to at least a portion of the target sequence itself; in some embodiments, the primer comprises at least one sequence that is at least 75% complementary, typically at least 85% complementary, more typically at least 90% complementary, more typically at least 95% complementary, more typically at least 98% complementary, or more typically at least 99% complementary to at least a portion of a nucleic acid molecule other than the target sequence. In some embodiments, such primers are referred to as "specific primers" or "selective primers" that are substantially non-complementary to a target sequence or its corresponding nucleic acid portion (including the target sequence) other than its corresponding target sequence; optionally, a specific primer or a selective primer is substantially non-complementary to other nucleic acid molecules that may be present in a nucleic acid mixture (e.g., in a sample). In some embodiments, nucleic acid molecules that do not contain or correspond to a target sequence (or the complement of a target sequence) present in a sample are referred to as "non-specific" sequences or "non-specific nucleic acids". In some embodiments, specific primers or selective primers are designed to include a nucleotide sequence that is substantially complementary to at least a portion of its corresponding target sequence. In some embodiments, specific primers or selective primers are at least 95% complementary or at least 99% complementary, 100% complementary or identical to at least a portion of the nucleic acid molecule that includes its corresponding target sequence across its entire length. In some embodiments, specific primers or selective primers can be at least 90%, at least 95% complementary, at least 98% complementary or at least 99% complementary, 100% complementary or identical to at least a portion of its corresponding target sequence across its entire length. In some embodiments, forward specific primers and reverse specific primers define specific primer pairs (or selective primer pairs) that can be used to amplify target sequences by template-dependent primer extension.Typically, each primer in a specific primer pair comprises at least one sequence that is substantially complementary to at least a portion of a nucleic acid molecule comprising a corresponding target sequence but is less than 50% complementary to at least one other target sequence in a mixture or sample. In some embodiments, multiple specific primer pairs can be used in a single amplification reaction for amplification, wherein each primer pair comprises a forward specific primer and a reverse specific primer, each of which comprises at least one sequence that is substantially complementary to or substantially identical to a corresponding target sequence in a mixture or sample, and each specific primer pair has a different corresponding target sequence. In some embodiments, a specific primer may be substantially non-complementary to any other specific primer present in the amplification reaction at its 3' end or its 5' end. In some embodiments, a specific primer may comprise minimal cross hybridization with other specific primers in the amplification reaction. In some embodiments, a specific primer comprises minimal cross hybridization with a non-specific sequence in an amplification reaction mixture. In some embodiments, a specific primer comprises minimal self-complementarity. In some embodiments, a specific primer may comprise one or more cleavable groups located at the 3' end. In some embodiments, a specific primer may comprise one or more cleavable groups located near or around the central nucleotide of the specific primer. In some embodiments, one or more specific primers comprise only non-cleavable nucleotides at the 5' end of the specific primer. In certain embodiments, the specific primer comprises a minimum nucleotide sequence overlap at the 3' end or 5' end of the primer compared to one or more different specific primers, optionally in the same amplification reaction. In certain embodiments, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more specific primers in a single reaction mixture comprise one or more of the above embodiments. In certain embodiments, a plurality of specific primers in a single reaction mixture substantially all comprise one or more of the above embodiments.
[0059] As used herein, the term "ligating / ligation" and its derivatives refer to an action or method for covalently linking two or more molecules together (e.g., covalently linking two or more nucleic acid molecules to each other). In some embodiments, the connection comprises a connection incision between adjacent nucleotides of a nucleic acid. In some embodiments, the connection comprises a covalent bond formed between the end of a first nucleic acid molecule and the end of a second nucleic acid molecule. In some embodiments, such as embodiments in which the nucleic acid molecules to be connected comprise conventional nucleotide residues, the connection may comprise a covalent bond formed between a 5' phosphate group of a nucleic acid and a 3' hydroxyl group of a second nucleic acid, thereby forming a connected nucleic acid molecule. In some embodiments, any method for connecting or combining 5' phosphate with 3' hydroxyl between adjacent nucleotides may be used. In an exemplary embodiment, an enzyme such as a ligase may be used. For purposes of this disclosure, an amplified target sequence may be connected to an adaptor to produce an amplified target sequence connected to an adaptor.
[0060] As used herein, "ligase" and its derivatives refer to any reagent capable of catalyzing the connection of two substrate molecules. In some embodiments, the ligase comprises an enzyme capable of catalyzing the connection of nicks between adjacent nucleotides of a nucleic acid. In some embodiments, the ligase comprises an enzyme capable of catalyzing the formation of a covalent bond between the 5' phosphate of one nucleic acid molecule and the 3' hydroxyl of another nucleic acid molecule, thereby forming a connected nucleic acid molecule. Suitable ligases may include, but are not limited to, T4 DNA ligase, T4 RNA ligase, and E. coli DNA ligase.
[0061] As used herein, "connection conditions" and their derivatives refer to conditions suitable for connecting two molecules to each other. In certain embodiments, connection conditions are suitable for sealing the nicks or gaps between nucleic acids. As defined herein, "nicks" or "gap" refer to nucleic acid molecules that lack the 5' phosphate of a mononucleotide pentose ring in the internal nucleotides of a nucleic acid sequence and are directly bound to the 3' hydroxyl of an adjacent mononucleotide pentose ring. As used herein, the term nick or gap is consistent with the use of the terminology in the art. Typically, nicks or gaps can be connected at a suitable temperature and pH in the presence of an enzyme (such as a ligase). In certain embodiments, T4 DNA ligase can connect nicks between nucleic acids at a temperature of about 70-72°C.
[0062] As used herein, "flat end connection" and its derivatives refer to the connection of two flat-ended double-stranded nucleic acid molecules to each other. "Flat end" refers to the end of a double-stranded nucleic acid molecule in which substantially all nucleotides in one chain end of a nucleic acid molecule are base-paired with the opposite nucleotides in the other chain of the same nucleic acid molecule. If a nucleic acid molecule has an end (referred to as an "overhang" herein) containing a single-stranded portion greater than two nucleotides in length, it is not a flat end. In some embodiments, the end of the nucleic acid molecule does not contain any single-stranded portion, so that each nucleotide in one chain of the end is base-paired with the opposite nucleotide in the other chain of the same nucleic acid molecule. In some embodiments, the ends of the two flat-ended nucleic acid molecules that become connected to each other do not contain any overlapping, shared or complementary sequences. Typically, the flat-end connection does not include the connection of the target sequence of the double-stranded amplification with the double-stranded adapter using other oligonucleotide adapters, such as the patch oligonucleotides described in Mitra and Varley, US2010 / 0129874, published on May 27, 2010. In some embodiments, the flat-end connection includes a nick translation reaction for sealing the nicks generated during the connection process.
[0063] As used herein, the term "adapter" or "adapter and its complement" and its derivatives refer to any linear oligonucleotide that can be connected to the nucleic acid molecules of the present disclosure. Optionally, the adaptor comprises a nucleic acid sequence that is substantially non-complementary to the 3' end or 5' end of at least one target sequence in the sample. In some embodiments, the adaptor is substantially non-complementary to the 3' end or 5' end of any target sequence in the sample. In some embodiments, the adaptor comprises any single-stranded or double-stranded linear oligonucleotide that is substantially non-complementary to the amplified target sequence. In some embodiments, the adaptor is substantially non-complementary to at least one, some or all of the nucleic acid molecules of the sample. In some embodiments, the suitable adaptor length is within the range of about 10-100 nucleotides, about 12-60 nucleotides, and about 15-50 nucleotides in length. The adaptor may comprise any combination of nucleotides and / or nucleic acids. In some aspects, the adaptor may comprise one or more cleavable groups at one or more positions. In another aspect, the adaptor may comprise a sequence that is substantially identical or substantially complementary to at least a portion of a primer (e.g., a universal primer). In some embodiments, the adapter may comprise a barcode or tag to assist in downstream cataloguing, identification or sequencing. In some embodiments, when ligated to an amplified target sequence, specifically, in the presence of a polymerase and dNTPs at a suitable temperature and pH, the single-stranded adapter may serve as a substrate for amplification.
[0064] As used herein, "DNA barcode" or "DNA tag sequence" and its derivatives refer to a unique short (6-14 nucleotide) nucleic acid sequence within an adaptor that can act as a 'key' for distinguishing or separating multiple amplified target sequences in a sample. For the purposes of this disclosure, a DNA barcode or DNA tag sequence can be incorporated into the nucleotide sequence of an adaptor.
[0065] As used herein, "GC content" and its derivatives refer to the cytosine and guanine content of a nucleic acid molecule. In some embodiments, the GC content of a specific primer (or adapter) is 85% or less. In some embodiments, the GC content of a specific primer or adapter is between 15-85%.
[0066] Composition
[0067] The compositions provided herein include compositions containing one or more nucleic acids, and the one or more nucleic acids include, for example, but not limited to, double-stranded, partially double-stranded, single-stranded, modified and unmodified nucleic acids. In some embodiments, the nucleic acid is single-stranded, for example, a single-stranded oligonucleotide that can be used as a primer and / or a probe. In some embodiments, the compositions provided herein contain two nucleic acids that can amplify specific nucleic acids in a nucleic acid amplification process or reaction, for example, a nucleic acid primer pair. Also provided herein are compositions containing multiple nucleic acids (e.g., primers and / or probes, including, for example, multiple primer pairs) or composed of multiple nucleic acids. In some embodiments, the nucleic acids and / or nucleic acid pairs (e.g., primer pairs) in the compositions provided herein can be combined with, hybridized and / or amplified with nucleic acids contained in the genome of one or more microorganisms (such as, for example, bacteria or archaea). In some embodiments, one or more nucleic acids (e.g., primer pairs) in the compositions provided herein are capable of binding to, hybridizing with, and / or amplifying a nucleic acid (e.g., a nucleic acid from a microorganism such as a bacterium) containing a nucleotide sequence as shown in SEQ ID NOs: 1605-1979 in Table 17, or SEQ ID NOs: 1605-1826 in Table 17, or SEQ ID NOs: 1605-1820 in Table 17A, or SEQ ID NOs: 1827-1979 in Table 17C, or SEQ ID NOs: 1827-1976 in Table 17C, or sequences substantially identical or similar to any of these sequences, or specifically binding to, hybridizing with, and / or specifically amplifying the nucleic acid.In some embodiments, one or more nucleic acids (e.g., primer pairs) in the compositions provided herein are capable of amplifying or specifically amplifying a nucleic acid containing a nucleotide sequence shown in SEQ ID NOs: 1605-1979 in Table 17, or SEQ ID NOs: 1605-1826 in Table 17, or SEQ ID NOs: 1605-1820 in Table 17A, or SEQ ID NOs: 1827-1979 in Table 17C, or SEQ ID NOs: 1827-1976 in Table 17C, such as a nucleic acid from a microorganism (e.g., a bacterium) to produce an amplicon sequence of less than about 500, less than about 475, less than about 450, less than about 400, less than about 375, less than about 350, less than about 300, less than about 275, less than about 250, less than about 200, less than about 175, less than about 150, or less than about 100 nucleotides in length, or consisting essentially of a nucleotide sequence selected from a SEQ ID NO in Table 17. NO: 1605-1979, or SEQ ID NO: 1605-1826 in Table 17, or SEQ ID NO: 1605-1820 in Table 17A, or SEQ ID NO: 1827-1979 in Table 17C, or SEQ ID NO: 1827-1976 in Table 17C, or substantially identical or similar sequences and optionally an amplicon sequence consisting of a sequence containing nucleic acid primer sequences at the 5' and 3' ends of the sequence. In some embodiments, the composition contains multiple nucleic acids, which are capable of binding, hybridizing and / or amplifying the multiple nucleic acids, or specifically binding, hybridizing and / or specifically amplifying the multiple nucleic acids, each of which contains a nucleotide sequence selected from SEQ ID NO: 1605-1979 in Table 17, or SEQ ID NO: 1605-1826 in Table 17, or SEQ ID NO: 1605-1820 in Table 17A, or SEQ ID NO: 1827-1979 in Table 17C, or SEQ ID NO: 1827-1976 in Table 17C.In some embodiments, the composition contains a plurality of nucleic acids (e.g., primer pairs) capable of amplifying or specifically amplifying a plurality of nucleic acids (e.g., nucleic acids from microorganisms (e.g., bacteria)), each of the plurality of nucleic acids containing SEQ ID NOs: 1605-1979 in Table 17, or SEQ ID NOs: 1605-1826 in Table 17, or SEQ ID NOs: 1605-1820 in Table 17A, or SEQ ID NOs: 1827-1979 in Table 17C, or SEQ ID NOs: 1827-1979 in Table 17C. NO: 1827-1976, or a nucleotide sequence of a substantially identical or similar sequence to produce an amplicon sequence of less than about 500, less than about 475, less than about 450, less than about 400, less than about 375, less than about 350, less than about 300, less than about 275, less than about 250, less than about 200, less than about 175, less than about 150, or less than about 100 nucleotides in length, or an amplicon sequence consisting essentially of a sequence selected from SEQ ID NO: 1605-1979 in Table 17, or SEQ ID NO: 1605-1826 in Table 17, or SEQ ID NO: 1605-1820 in Table 17A, or SEQ ID NO: 1827-1979 in Table 17C, or SEQ ID NO: 1827-1976 in Table 17C, and optionally containing nucleic acid primer sequences at the 5' and 3' ends of the sequence. In some embodiments, the compositions provided herein contain a plurality of nucleic acids, each of the plurality of nucleic acids comprising a nucleotide sequence selected from SEQ ID NOs: 1605-1979 in Table 17, or SEQ ID NOs: 1605-1826 in Table 17, or SEQ ID NOs: 1605-1820 in Table 17A, or SEQ ID NOs: 1827-1979 in Table 17C, or SEQ ID NOs: 1827-1976 in Table 17C, or substantially identical or similar sequences. In some embodiments, the composition contains a plurality of nucleic acids, each of the plurality of nucleic acids containing or consisting essentially of a nucleotide sequence selected from SEQ ID NOs: 1605-1826 in Table 17, or SEQ ID NOs: 1605-1820 in Table 17A, or substantially identical or similar sequences, and optionally containing nucleic acid primer sequences at the 5' and 3' ends of the sequence, and the nucleotide sequence is less than about 500, less than about 475, less than about 450, less than about 400, less than about 375, less than about 350, less than about 300, less than about 275, less than about 250, less than about 200, less than about 175, less than about 150, or less than about 100 nucleotides in length.
[0068] In some embodiments, the nucleic acids in the compositions provided herein comprise or consist essentially of a nucleotide sequence in Table 15 or Table 16, wherein one or more thymine bases are substituted with uracil bases. In some embodiments, the nucleic acids provided herein comprise or consist essentially of a nucleotide sequence selected from the group consisting of SEQ ID NOs: 11-16, 23, and 24 of Table 15; SEQ ID NOs: 35-40, 47, and 48 of Table 15; SEQ ID NOs: 49-480 of Table 16A; SEQ ID NOs: 49-452 and 457-472 of Table 16A; SEQ ID NOs: 521-820 of Table 16C; SEQ ID NOs: 827-1258 of Table 16D; SEQ ID NOs: 827-1230 and 1235-1250 of Table 16D; or SEQ ID NOs: 1299-1598 of Table 16F; or substantially identical or similar sequences.In some embodiments, the composition contains or consists essentially of a plurality of nucleic acids, each of the plurality of nucleic acids containing or consisting essentially of a sequence selected from the group consisting of: sequences in Table 15; SEQ ID NOs: 1-24 of Table 15; SEQ ID NOs: 11-16, 23, and 24 of Table 15; SEQ ID NOs: 25-48 of Table 15; SEQ ID NOs: 35-40, 47, and 48 of Table 15; sequences in Table 16; SEQ ID NOs: 49-520 of Table 16; SEQ ID NOs: 49-452, 457-472, and 481-520 of Table 16; SEQ ID NOs: 49-492 of Table 16; SEQ ID NOs: 49-452, 457-472, and 481-492 of Table 16; SEQ ID NOs: 49-480 of Table 16A; SEQ ID NOs: 49-520 of Table 16; SEQ ID NOs: 49-452, 457-472, and 481-520 of Table 16; SEQ ID NOs: 49-452 and 457-472; SEQ ID NOs: 521-826 of Table 16C; SEQ ID NOs: 521-820 of Table 16C; SEQ ID NOs: 827-1298 of Table 16; SEQ ID NOs: 827-1230, 1235-1250 and 1259-1298 of Table 16; SEQ ID NOs: 827-1270 of Table 16; SEQ ID NOs: 827-1230, 1235-1250 and 1259-1270 of Table 16; SEQ ID NOs: 827-1258 of Table 16D; SEQ ID NOs: 827-1230 and 1235-1250 of Table 16D; SEQ ID NOs: 1299-1604 of Table 16F; SEQ ID NOs: 1299-1604 of Table 16F NO:1299-1598; substantially the same or similar sequence, and / or any of the above-mentioned nucleotide sequences, wherein one or more thymine bases are replaced by uracil bases. In certain embodiments, the nucleic acid in the compositions provided herein comprises one or more pairs of nucleic acids (e.g., primer pairs). Primer pairs comprise paired (i.e., 2) nucleic acids (polynucleotides) that can be used for amplification of nucleic acids. Examples of primer pairs are shown in Tables 15 and 16 as "primer 1" and "primer 2" in each row of the table, which can amplify nucleic acid sequences contained in the corresponding regions (hypervariable regions) of prokaryotic (e.g., bacteria) 16S rRNA genes (Table 15) or contained in corresponding microbial species (Table 16).In some embodiments, the nucleic acids in the compositions provided herein comprise or consist essentially of one or more pairs of nucleic acids containing or consisting essentially of the nucleotide sequences of Table 15 or Table 16; one or more pairs of nucleotide sequences selected from the group consisting of the sequence pairs shown in Table 15: SEQ ID NOs: 1-24; SEQ ID NOs: 25-48; SEQ ID NOs: 11-16, 23, and 24; SEQ ID NOs: 35-40, 47, and 48; SEQ ID NOs: 49-520; SEQ ID NOs: 49-452, 457-472, and 481-520; SEQ ID NOs: 49-492; SEQ ID NOs: 49-452, 457-472, and 481 ... NO:49-480; SEQ ID NO:49-452 and 457-472 of Table 16A; SEQ ID NO:521-826 of Table 16C; SEQ ID NO:521-820 of Table 16C; SEQ ID NO:827-1298 of Table 16; SEQ ID NO:827-1230, 1235-1250 and 1259-1298 of Table 16; SEQ ID NO:827-1270 of Table 16; SEQ ID NO:827-1230, 1235-1250 and 1259-1270 of Table 16; SEQ ID NO:827-1258 of Table 16D; SEQ ID NO:827-1230 and 1235-1250 of Table 16D; SEQ ID NO:827-1298 of Table 16 NO:1299-1604; or SEQ ID NO:1299-1598 of Table 16F; substantially the same or similar sequence, and / or the nucleotide sequence of any of the above primer pairs, wherein one or more thymine bases are replaced by uracil bases.In some embodiments, the composition contains or consists essentially of multiple pairs of nucleic acids (e.g., primer pairs) containing or consisting essentially of the nucleotide sequences of the following: two or more pairs of nucleotide sequences in Table 15 or Table 16; two or more pairs of nucleotide sequences selected from the sequence pairs shown in the following: SEQ ID NOs: 1-24 of Table 15; SEQ ID NOs: 25-48 of Table 15; SEQ ID NOs: 11-16, 23 and 24 of Table 15; SEQ ID NOs: 35-40, 47 and 48 of Table 15; SEQ ID NOs: 49-520 of Table 16; SEQ ID NOs: 49-452, 457-472 and 481-520 of Table 16; SEQ ID NOs: 49-492 of Table 16; SEQ ID NOs: 49-452, 457-472 and 481-492 of Table 16; SEQ ID NOs: 49-452, 457-472 and 481-492 of Table 16; SEQ ID NOs: 49-452, 457-472 and 481-492 of Table 16; SEQ ID NOs: 49-452, 457-472 and 481-492 of Table 16A NO:49-480; SEQ ID NO:49-452 and 457-472 of Table 16A; SEQ ID NO:521-826 of Table 16C; SEQ ID NO:521-820 of Table 16C; SEQ ID NO:827-1298 of Table 16; SEQ ID NO:827-1230, 1235-1250 and 1259-1298 of Table 16; SEQ ID NO:827-1270 of Table 16; SEQ ID NO:827-1230, 1235-1250 and 1259-1270 of Table 16; SEQ ID NO:827-1258 of Table 16D; SEQ ID NO:827-1230 and 1235-1250 of Table 16D; SEQ ID NO:827-1230 and 1235-1250 of Table 16F ID NO: 1299-1604; or SEQ ID NO: 1299-1598 of Table 16F; or a substantially identical or similar sequence, and / or the nucleotide sequence of any of the above primer pairs, wherein one or more thymine bases are replaced by uracil bases.
[0069] In certain embodiments, nucleic acid or nucleic acid pair and optionally its degenerate sequence and specific microorganism (for example, bacterial species) unique specific nucleic acid sequence combination, hybridization and / or amplification described specific nucleic acid sequence.This nucleic acid or nucleic acid pair, and optionally degenerate sequence, are referred to as "microorganism-specific" or "species-specific" in this article, and amplify nucleic acid in a microorganism-specific or species-specific manner to have in the amplification reaction the situation of nucleic acid from microorganism, in microorganism (for example, bacterium) or microorganism group, produce the single amplification product with unique sequence.The limiting examples of the sequence of this type of nucleic acid and nucleic acid primer pair are provided in Table 16.
[0070] In certain embodiments, nucleic acid pairs, and optionally its degenerate sequence, can increase the sequence in the homologous gene or genomic region that is common but can be different between different microorganisms for multiple, major part, most, substantially all or all microorganisms in the taxonomy group.Taxonomy group comprises kingdom, domain, phylum, class, order, section and kind.In one embodiment, taxonomy group is bacterial kingdom.Can increase the sequence in the homologous gene or genomic region that is common but can be different between the microorganisms of different kingdoms for multiple, major part, most, substantially all or all microorganisms in the kingdom or primer pair, and optionally degenerate sequence, referred to as "kingdom encompasses" in this article, and increase the nucleic acid in the microorganism in the kingdom in the mode encompassed by the kingdom, with in the case of the nucleic acid from microorganism (for example, bacterium) in the amplification reaction, produce the multiple amplification products with the different nucleotide sequences of different microorganisms (for example, different bacteria) in the kingdom.The conserved sequence of nucleic acid can be found in the genome of different organisms or microorganisms. Such sequences may be identical or have substantial similarity in different genomes (see, for example, Isenbarger et al. (2008) Origin of Life and Evolution of the Biosphere (Orig Life Evol Biosph) doi:10.1007 / s11084-008-9148-z). In many cases, conserved sequences are located in essential genes, such as housekeeping genes, which encode elements required for a class or group of organisms or microorganisms to perform basic survival biochemical functions. However, through the evolution and adaptation of organisms and microorganisms to different conditions, even homologous genes have diverged and contain sequences that vary between different organisms and microorganisms, which may be so different that they are unique for a specific organism or microorganism, so that they can be used to identify a single organism or microorganism or a related group (e.g., species) of an organism or microorganism. Homologous genes include, for example, some essential genes required for the basic functions and survival of microorganisms. In some embodiments, homologous genes are 16S rRNA genes, 18S rRNA genes or 23S rRNA genes common to a variety of different organisms or microorganisms (e.g., a variety of different bacteria). For example, in certain embodiments, the nucleic acid comprises one or more primer pairs that individually amplify two or more regions in a prokaryotic (e.g., bacterial) 16S rRNA gene, such as a hypervariable region. Non-limiting examples of such nucleic acid primer pairs are provided in Table 15.
[0071] Variable region analysis has been used for taxonomic classification of prokaryotes, for example, in methods using nucleic acid primers that hybridize to conserved sequences on either side of the variable region. Homologous genes containing multiple variable regions interspersed between conserved regions are particularly useful in such methods because they provide multiple sequences that can be analyzed to more accurately and unambiguously identify individual components of a population of target elements. An example of such a gene is the prokaryotic 16S rRNA gene that encodes ribosomal RNA, which is the major structural and catalytic component of the ribosome. The 16S ribosomal RNA (rRNA) gene of bacteria and archaea is approximately 1500-1700 base pairs long and contains nine hypervariable regions of varying conservation interspersed between conserved regions, commonly referred to as V1-V9 ( Figure 1) (see, e.g., Wang and Qian (2009) PloS ONE 4:e7401 and Kim (2011) J Microbiol Meth 84:81-87). Exemplary 16S rRNA gene sequences are known and include sequences contained in the Greengenes database (http: / / greengenes.lbl.gov), the SILVA database (www.arb-silva.de), and the GRD-genome-based 16 ribosomal RNA database (https: / / metasystems.riken.jp / grd / ). The hypervariable region sequences of the 16S rRNA gene that differ in different microorganisms can be used to identify the microorganisms in a sample. Instead of specifically amplifying the hypervariable region of each possible microorganism that may be present in the sample using many oligonucleotide primers, each primer being specific to the hypervariable region of each organism, a conserved, highly similar or identical sequence located on both sides of the hypervariable region can be used as a primer binding sequence, to which one or a small number of primer pairs will bind, and amplify the hypervariable region in substantially all microorganisms (e.g., bacteria) in the sample. This allows specific nucleic acids that can be used to identify microorganisms to be amplified from substantially all microorganisms, which can then be sequenced to effectively profile the population. However, the results of these methods are often inconsistent and are often incomplete in determining most or all of the microorganisms present in a sample (specifically, a sample containing a variety of different microorganisms). In addition, these methods are generally unable to reliably or accurately distinguish between species of microorganisms, if they can distinguish between species. Most of these methods use primers that are intended to amplify one or a few (rather than all) hypervariable regions of the 16S rRNA gene. If more than a limited number of hypervariable regions are targeted for amplification in these methods, the method typically requires multiple separate amplification reactions for different primers due to overlap of primer sequences, which results in inefficient resource use and time for the method. In addition, such methods typically include primer pairs designed to amplify two or more hypervariable regions (e.g., V2-V3 or V3-V4) into a single amplicon, which results in longer amplicons for sequencing.
[0072] Nucleic Acids
[0073] Provided herein are nucleic acid primer pairs, which individually amplify nucleic acids including sequences in multiple hypervariable regions of prokaryotic 16S rRNA genes. In certain embodiments, the nucleotide sequences of any two 16s rRNA gene primers of nucleic acids including sequences in multiple hypervariable regions amplified individually overlap slightly (e.g., less than or equal to 7 nucleotides) or do not overlap. In some aspects, primer pairs amplification length is less than or equal to about 200 nucleotides, such as 16s rRNA gene sequences with a length between about 125 and 200 nucleotides. In certain embodiments, provided herein are compositions containing multiple nucleic acid primer pairs, which include at least 2, at least 3, at least 4, at least 5, at least 6 or at least 7 separate primer pairs, and optionally degenerate variants thereof, which individually amplify nucleic acids including sequences of a different hypervariable region in 2, 3, 4, 5, 6 or 7 different hypervariable regions in prokaryotic 16s rRNA genes in a nucleic acid amplification reaction. In certain embodiments, compositions provided herein contain multiple nucleic acid primer pairs, and it comprises at least 8 independent primer pairs, and optionally its degenerate variant, and it separately amplifies the nucleic acid of the sequence of one of 8 different hypervariable regions in prokaryotic 16S rRNA gene in nucleic acid amplification reaction.In certain embodiments, compositions comprise the combination of primer pairs, wherein the primer pairs in the primer pair combination separately amplify the nucleic acid of the sequence including being positioned at 3 or more hypervariable regions of prokaryotic 16S rRNA gene, and wherein one of 3 or more regions is V5 district.Degenerate primer variants are included in some compositions, for example, 1 or 2 positions in the primer sequence contain different nucleotides, to ensure that amplification contains the 16S rRNA gene of slight variation in the conserved region.Non-limiting examples of the nucleotide sequence of the primer pairs of 8 hypervariable regions (V2, V3, V4, V5, V6, V7, V8 and V9) of prokaryotic 16SrRNA gene separately are listed in Table 15. In some embodiments, the compositions provided herein comprise or consist essentially of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24 nucleic acids or primer pairs, wherein the nucleic acids or primer pairs contain or consist essentially of a sequence selected from the group consisting of: a sequence in Table 15 or a sequence from SEQ ID NOs: 1-24 of Table 15, SEQ ID NOs: 25-48 of Table 15, SEQ ID NOs: 11-16, 23 and 24 of Table 15, or SEQ ID NOs: 35-40, 47 and 48 of Table 15, or a substantially identical or similar sequence.In some embodiments, the compositions provided herein contain or consist essentially of nucleic acids or primer pairs, wherein the nucleic acids or primer pairs contain or consist essentially of all of the sequences of SEQ ID NOs: 1-24 of Table 15 and / or all of the sequences of SEQ ID NOs: 25-48 of Table 15, respectively. In some embodiments of the compositions provided herein containing a plurality of nucleic acid primer pairs that individually amplify nucleic acids comprising sequences located in a plurality of hypervariable regions of a prokaryotic 16S rRNA gene, the plurality of primer pairs provide at least 85%, or at least 90%, or at least 92%, or at least 95%, or at least 98%, or at least 99%, or 100% coverage of different bacterial 16S rRNA gene sequences in a given database containing bacterial 16S rRNA gene sequences. In some embodiments of the compositions provided herein containing a plurality of nucleic acid primer pairs that individually amplify nucleic acids comprising sequences located in multiple hypervariable regions of a prokaryotic 16S rRNA gene, the plurality of primer pairs are capable of amplifying all or substantially all microbial (e.g., bacterial) nucleic acids in a sample containing a mixture of at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95 or at least 100 or more different genera of microorganisms (e.g., bacteria).
[0074] Species / microorganism specific nucleic acids
[0075] Provided herein are nucleic acids and nucleic acid pairs (e.g., primer pairs), which are combined with, hybridized and / or amplified specific nucleic acid sequences unique to specific microorganisms (e.g., species, subspecies or strains of bacteria). Such microorganism-specific (e.g., bacteria-specific or species-specific) nucleic acids can be used as, for example, specific, selective probes and / or primers, to greatly improve the depth and accuracy of detection and identification of microorganisms in samples, and significantly enhance the characterization, evaluation, measurement and / or analysis of microbial colonies or communities and their components or compositions. Such information is needed to fully understand the biodiversity of microbial communities, e.g., the microbial population of animal digestive tracts. The exemplary nucleic acid sequences provided in Table 16 are combined with, hybridized and / or amplified specific nucleic acid sequences unique to more than 40 different microorganisms (bacteria) genera or at least 43 different microorganisms genera, or more than 70, or at least 73, or at least 74, or at least 75 different microorganisms (bacteria) species.
[0076] Microbe-specific nucleic acids provide many advantages, for example, in completely and accurately assessing, characterizing, measuring and / or analyzing the composition of a microbial population (e.g., microbiota) and determining the relationship of individual microorganisms, as well as associating and / or linking microbial communities, and the state of the subject and / or environment (e.g., health, degree of balance, susceptibility to certain conditions, responsiveness to treatment). The microbiota (i.e., microorganisms, including bacteria) of humans associated with different regions of human subjects contain more than 10 times more microbial cells than human cells. In addition to the appearance of pathogenic microorganisms, the microbiota also contains symbiotic microorganisms. Although the significance of identifying pathogenic microorganisms in animals is relatively clear, analyzing the complex composition of all types of microorganisms in the microbiota is also of great significance for understanding the health of animals and potential disorders and disease treatment interventions. For example, microorganisms present in the digestive tract of animals, commonly referred to as "gut microbiome", contribute to the metabolism of animals, and evidence supports the role of the gut microbiome in inflammatory bowel disease, autoimmune disorders, cardiometabolic disorders, cancer, and neuropsychiatric disorders and diseases.
[0077] Compositions and methods provided herein, comprising microorganism-specific and kingdom-covered nucleic acids, and their purposes in sample analysis, can not only comprehensively investigate the overall and relative levels of microorganisms (for example, bacteria) genus, can also identify the species of microorganisms in detail, the species of the microorganism can be customized to focus on one or more specific target microorganisms, these microorganisms may be very important in some health and disease or imbalanced states. For example, provided herein is a nucleic acid and / or nucleic acid primer pair that is combined with a specific nucleic acid sequence unique to a particular microorganism (for example, species, subspecies or strain of bacteria), hybridized and / or amplified the specific nucleic acid sequence, that is, microorganism-specific nucleic acid. In certain embodiments, nucleic acid can be in a mixture comprising the nucleic acid of a plurality of different microorganisms, for example, in a mixture comprising the nucleic acid of the genome of different microorganisms of the same genus with the microorganism containing the target nucleic acid sequence, specifically combined with and / or hybridized with the target nucleic acid sequence contained in the genome of the microorganism. In certain embodiments, the nucleic acid specifically combined with and / or hybridized with the nucleic acid sequence contained in the genome of the microorganism is not combined with or hybridized with the nucleic acid contained in any other microorganism genus or any other microorganism species. In certain embodiments, in the amplification reaction mixture of nucleic acid including the genome of multiple different microorganisms, and in a specific embodiment, in the amplification reaction mixture of nucleic acid including the genome of different microorganisms from the same genus with the microorganism containing the target nucleic acid sequence, primer pairs specifically amplify the specific target nucleic acid sequence unique to a specific microorganism. In certain embodiments, primer pairs do not amplify the nucleic acid sequence contained in any other microorganism genus or any other organism species. In certain embodiments, the combination of nucleic acid comprises microorganism-specific nucleic acid and / or primer pairs, which specifically bind, hybridize and / or specifically amplify the nucleic acid sequence with the nucleic acid sequence contained in the genome of one or more microorganisms (e.g., bacteria), and the one or more microorganisms are related to one or more symptom, illness and / or disease. In a specific embodiment of the compositions provided herein, the compositions comprises nucleic acid and / or primer pairs, which specifically bind, hybridize and / or specifically amplify the target nucleic acid sequence with the target nucleic acid sequence contained in the genome of the microorganism selected from the microorganism in Table 1. In a specific embodiment, the target nucleic acid sequence contained in the genome of the microorganism selected from the microorganism in Table 1 is unique to the microorganism. In some embodiments, the composition comprises or consists essentially of a plurality of nucleic acids and / or primer pairs, comprising at least one nucleic acid that specifically binds and / or hybridizes to a target nucleic acid of each microorganism in Table 1 and / or at least one primer pair that specifically amplifies a genomic target nucleic acid of each microorganism in Table 1.In some embodiments, the composition comprises or consists essentially of a plurality of nucleic acids and / or primer pairs, comprising at least one nucleic acid that specifically binds and / or hybridizes to a target nucleic acid of each microorganism in Table 1, except or excluding Actinomyces viscosus and / or Blautia coccoides, or except or excluding Actinomyces viscosus, Blautia coccoides, and / or Helicobacter salomonis. In some embodiments, the composition comprises or consists essentially of a plurality of nucleic acids and / or primer pairs, comprising at least one primer pair that specifically amplifies a genomic target nucleic acid of each microorganism in Table 1, except or excluding Actinomyces viscosus and / or Blautia coccoides, or except or excluding Actinomyces viscosus, Blautia coccoides, and / or Helicobacter salomonis. In some embodiments, a plurality of primer pairs comprise or consist essentially of different primer pairs, and the different primer pairs specifically and individually amplify different genomic target nucleic acids contained in or at least in 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 70, 71, 72, 73, 74, 75 or more microorganisms in Table 1. In some embodiments, a plurality of nucleic acid primer pairs comprise or consist essentially of a nucleic acid primer pair set, wherein each different nucleic acid primer pair specifically amplifies different unique nucleic acid sequences contained in different genomes of the microbial groups in Table 1 or each genome of the microbial groups in Table 1, except or not including Actinomycetes viscosus and / or Blautia sphaeroides, or except or not including Actinomycetes viscosus, Blautia sphaeroides and / or Helicobacter zascherii. In a specific embodiment, the target nucleic acid sequences contained in the different microbial genomes are unique to each microorganism.
[0078]
[0079]
[0080] In some embodiments, the nucleic acids and / or nucleic acid primer pairs provided herein bind to, hybridize to, and / or amplify a nucleic acid, such as a nucleic acid from a microorganism (e.g., bacteria), or specifically bind to, hybridize to, and / or specifically amplify a nucleic acid, wherein the nucleic acid contains a nucleotide sequence selected from SEQ ID NOs: 1605-1979 in Table 17, or SEQ ID NOs: 1605-1826 in Table 17, or 1605-1806, 1809-1816, and 1821-1826 in Table 17, or SEQ ID NOs: 1605-1820 in Table 17A, or SEQ ID NOs: 1605-1806 and 1809-1816 in Table 17A, or SEQ ID NOs: 1827-1979 in Table 17C, or SEQ ID NOs: 1827-1976 in Table 17C, and / or a substantially identical or similar nucleotide sequence, or consists essentially of the nucleotide sequence.In some embodiments, the nucleic acid primer pairs provided herein are capable of amplifying or specifically amplifying a nucleic acid (e.g., a nucleic acid from a microorganism (e.g., bacteria)) containing SEQ ID NOs: 1605-1979 in Table 17, or SEQ ID NOs: 1605-1826 in Table 17, or SEQ ID NOs: 1605-1806, 1809-1816, and 1821-1826 in Table 17, or SEQ ID NOs: 1605-1820 in Table 17A, or SEQ ID NOs: 1605-1806 and 1809-1816 in Table 17A, or SEQ ID NOs: 1827-1979 in Table 17C, or SEQ ID NOs: 1827-1979 in Table 17C. 1827-1976 to produce an amplicon sequence of less than about 500, less than about 475, less than about 450, less than about 400, less than about 375, less than about 350, less than about 300, less than about 275, less than about 250, less than about 200, less than about 175, less than about 150, or less than about 100 nucleotides in length, or consisting essentially of SEQ ID NOs: 1605-1979 in Table 17, or SEQ ID NOs: 1605-1826 in Table 17, or SEQ ID NOs: 1605-1806, 1809-1816 and 1821-1826 in Table 17, or SEQ ID NOs: 1605-1820 in Table 17A, or SEQ ID NOs: 1605-1979 in Table 17, or SEQ ID NOs: 1605-1826 in Table 17, or SEQ ID NOs: 1605-1806, 1809-1816 and 1821-1826 in Table 17, or SEQ ID NOs: 1605-1820 in Table 17A, or SEQ ID NOs: NOs: 1605-1806 and 1809-1816, or SEQ ID NOs: 1827-1979 in Table 17C, or SEQ ID NOs: 1827-1976 in Table 17C, and optionally containing primer sequences linked at the 5' and 3' ends.In some embodiments, the compositions provided herein contain a combination of multiple microorganism-specific nucleic acids and / or primer pairs, wherein within the multiple nucleic acids and / or primer pairs, there are at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, at least 110, at least 111, at least 112, at least 90, at least 95, at least 100, at least 125, at least 150, at least 175, at least 200, at least 225, at least 250, at least 275, or at least 300 or more different nucleic acids, different nucleic acids and / or primer pairs that bind, hybridize, and / or amplify (specifically bind to, specifically hybridize to, and / or specifically amplify) the different nucleic acids, the different nucleic acids containing or consisting essentially of different sequences in the following sequences: SEQ ID NO: 17 ID NO: 1605-1979, or SEQ ID NO: 1605-1826 in Table 17, or SEQ ID NO: 1605-1806, 1809-1816 and 1821-1826 in Table 17, or SEQ ID NO: 1605-1820 in Table 17A, or SEQ ID NO: 1605-1806 and 1809-1816 in Table 17A, or SEQ ID NO: 1827-1979 in Table 17C, or SEQ ID NO: 1827-1976 in Table 17C, or substantially identical or similar sequences, and optionally contain primers attached at the 3' end and the 5' end.In some embodiments, such different nucleic acids and / or primer pairs amplify different nucleic acids containing different sequences among the following sequences: SEQ ID NOs: 1605-1979 in Table 17, or SEQ ID NOs: 1605-1826 in Table 17, or SEQ ID NOs: 1605-1806, 1809-1816, and 1821-1826 in Table 17, or SEQ ID NOs: 1605-1820 in Table 17A, or SEQ ID NOs: 1605-1806 and 1809-1816 in Table 17A, or SEQ ID NOs: 1827-1979 in Table 17C, or SEQ ID NOs: 1827-1979 in Table 17C. NOs: 1827-1976 to produce an amplicon sequence of less than about 500, less than about 475, less than about 450, less than about 400, less than about 375, less than about 350, less than about 300, less than about 275, less than about 250, less than about 200, less than about 175, less than about 150, or less than about 100 nucleotides in length, or consisting essentially of SEQ ID NOs: 1605-1979 in Table 17, or SEQ ID NOs: 1605-1826 in Table 17, or SEQ ID NOs: 1605-1806, 1809-1816, and 1821-1826 in Table 17, or SEQ ID NOs: 1605-1820 in Table 17A, or SEQ ID NOs: 1605-1979 in Table 17, or SEQ ID NOs: 1605-1826 in Table 17, or SEQ ID NOs: 1605-1806, 1809-1816, and 1821-1826 in Table 17A NOs: 1605-1806 and 1809-1816, or SEQ ID NOs: 1827-1979 in Table 17C, or SEQ ID NOs: 1827-1976 in Table 17C, and optionally an amplicon sequence consisting of a nucleotide sequence comprising primer sequences linked to the 5' and 3' ends of the sequence.In some embodiments, the combination of microorganism-specific nucleic acids or nucleic acid primer pairs comprises or consists essentially of two or more nucleic acids or primer pairs containing or consisting essentially of a nucleotide sequence or sequence pair selected from the group consisting of Table 16; or SEQ ID NOs: 49-520 of Table 16; or SEQ ID NOs: 49-452, 457-472, and 481-520 of Table 16; or SEQ ID NOs: 49-492 of Table 16; or SEQ ID NOs: 49-452, 457-472, and 481-492 of Table 16; or SEQ ID NOs: 49-480 of Table 16A; or SEQ ID NOs: 49-452 and 457-472 of Table 16A; or SEQ ID NOs: 521-826 of Table 16C; or SEQ ID NOs: 521-826 of Table 16C. or SEQ ID NOs:1299-1598 of Table 16F; or one or more sequences substantially identical or similar thereto, or the nucleotide sequence of any of the foregoing nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases.In some embodiments, the combination of microorganism-specific nucleic acids or nucleic acid primer pairs comprises or consists essentially of a plurality of nucleic acids or primer pairs, wherein there is at least one nucleic acid or primer pair containing or consisting essentially of each nucleotide sequence or sequence pair (for primer pairs) of the following sequences, respectively: SEQ ID NOs: 49-520 of Table 16; or SEQ ID NOs: 49-452, 457-472, and 481-520 of Table 16; or SEQ ID NOs: 49-492 of Table 16; or SEQ ID NOs: 49-452, 457-472, and 481-492 of Table 16; or SEQ ID NOs: 49-480 of Table 16A; or SEQ ID NOs: 49-452 and 457-472 of Table 16A; or SEQ ID NOs: 521-826 of Table 16C; or SEQ ID NOs: 521-820 of Table 16C; or SEQ ID NOs: 521-826 of Table 16C; or SEQ ID NOs:1299-1598 of Table 16F; or substantially the same or similar sequences, and / or the nucleotide sequences of any of the above nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases.In some embodiments, the compositions provided herein contain or consist essentially of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 125, at least 150, at least 175, at least 200, at least 225, at least 250, at least 275, at least 300 or more, or all nucleic acids, or all primer pairs, comprising or consisting essentially of a sequence selected from the group consisting of: Table 16; or SEQ ID NO: 16. or SEQ ID NOs:49-520; or SEQ ID NOs:49-452, 457-472 and 481-520 of Table 16; or SEQ ID NOs:49-492 of Table 16; or SEQ ID NOs:49-452, 457-472 and 481-492 of Table 16; or SEQ ID NOs:49-480 of Table 16A; or SEQ ID NOs:49-452 and 457-472 of Table 16A; or SEQ ID NOs:521-826 of Table 16C; or SEQ ID NOs:521-820 of Table 16C; or SEQ ID NOs:827-1298 of Table 16; or SEQ ID NOs:827-1230, 1235-1250 and 1259-1298 of Table 16; or SEQ ID NOs:827-1270; or SEQ ID NOs:827-1271 of Table 16 NO:827-1230, 1235-1250 and 1259-1270; or SEQ ID NO:827-1258 of Table 16D; or SEQ ID NO:827-1230 and 1235-1250 of Table 16D; or SEQ ID NO:1299-1604 of Table 16F; or SEQ ID NO:1299-1598 of Table 16F; or one or more sequences substantially identical or similar thereto, or the nucleotide sequence of any of the foregoing nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases.
[0081] In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises two or more nucleic acids and / or nucleic acid primer pairs that bind to, hybridize to, and / or amplify a unique nucleic acid sequence contained in the genome of one or more of the following: Akkermansia muciniphila, Bacteroides vulgaris, Bifidobacterium adolescentis, Campylobacter jejuni, Campylobacter conciseness, Clostridium difficile, Escherichia coli, coli), Eubacterium rectum, Helicobacter cholerae, Helicobacter hepatica, Lactobacillus delbrueckii, Parabacteroides distachyon, Ruminococcus brucelli, Streptococcus gallolyticus, and Streptococcus infantum (referred to herein as "Group A" microorganisms; see Table 2A), which are species implicated in playing a role in a variety of pathologies, diseases, and / or conditions, including, for example, neoplastic pathologies (including, for example, responses to immuno-oncology therapy and cancer), gastrointestinal disorders (including, for example, irritable bowel syndrome, inflammatory bowel disease, and celiac disease), and autoimmune diseases (including, for example, lupus and rheumatoid arthritis). In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises a set of nucleic acid primer pairs, wherein each different nucleic acid primer pair specifically amplifies a different unique nucleic acid sequence contained in a different genome of each genome of a different microorganism in Group A. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of nucleic acids and / or nucleic acid primer pairs that bind, hybridize, and / or amplify nucleic acids (e.g., nucleic acids from microorganisms (e.g., bacteria)), or specifically bind, hybridize, and / or specifically amplify nucleic acids, wherein the nucleic acids contain a SEQ ID NO: 17. 1899, 1900, 1932, 1933, 1968, and 1972, or a sequence substantially identical or similar to any of the foregoing sequences.In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of nucleic acids and / or nucleic acid primer pairs that bind to, hybridize to, and / or amplify a nucleic acid, such as a nucleic acid from a microorganism (e.g., bacteria), or specifically bind to, hybridize to, and / or specifically amplify a nucleic acid comprising a sequence selected from the group consisting of SEQ ID NOs: 1607-1609, 1619, 1620, 1635-1637, 1643-1645, 1663-1670, 1679-1684, 1699-1701, 1728-1730, 1752, 1753, 1801, 1802, 1809, and 1810 of Table 17, or a sequence substantially identical or similar to any of the foregoing sequences.In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of primers and / or primer pairs capable of amplifying or specifically amplifying a nucleic acid, such as a nucleic acid from a microorganism (e.g., bacteria), wherein the nucleic acid comprises SEQ ID NOs: 1607-1609, 1619, 1620, 1635-1637, 1643-1645, 1663-1670, 1679-1684, 1699-1701, 1728-1730, 1752, 1753, 1801, 1802, 1809, 1810, 1827, 1828, 1831-1833, 1844, 1845, 1852-1856, 1864, 1876-1885, 1889-1891, 1899 , 1900, 1932, 1933, 1968, and 1972 (or sequences substantially identical or similar to any of the foregoing sequences) to produce an amplicon sequence of less than about 500, less than about 475, less than about 450, less than about 400, less than about 375, less than about 350, less than about 300, less than about 275, less than about 250, less than about 200, less than about 175, less than about 150, or less than about 100 nucleotides in length, or consisting essentially of a SEQ ID NO: 17 ID NO: 1607-1609, 1619, 1620, 1635-1637, 1643-1645, 1663-1670, 1679-1684, 1699-1701, 1728-1730, 1752, 1753, 1801, 1802, 1809, 1810, 1827, 1828, 1831-1833, 1844, 1845, 1852-1856, 1864, 1876-1885, 1889-1891, 1899, 1900, 1932, 1933, 1968 and 1972, or an amplicon sequence consisting of a nucleotide sequence substantially identical or similar to any of the foregoing sequences and optionally containing nucleic acid primer sequences at the 5' and 3' ends of the sequence.In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of primers and / or primer pairs capable of amplifying or specifically amplifying a nucleic acid, such as a nucleic acid from a microorganism (e.g., bacteria), wherein the nucleic acid comprises a SEQ ID NO: 17 or a nucleic acid selected from Table 17. NO:1607-1609, 1619, 1620, 1635-1637, 1643-1645, 1663-1670, 1679-1684, 1699-1701, 1728-1730, 1752, 1753, 1801, 1802, 1809 and 1810 (or a sequence substantially identical or similar to any of the foregoing sequences) to produce an amplicon sequence of less than about 500, less than about 475, less than about 450, less than about 400, less than about 375, less than about 350, less than about 300, less than about 275, less than about 250, less than about 200, less than about 175, less than about 150 or less than about 100 nucleotides in length, or consisting essentially of a SEQ ID NO selected from Table 17 NO:1607-1609, 1619, 1620, 1635-1637, 1643-1645, 1663-1670, 1679-1684, 1699-1701, 1728-1730, 1752, 1753, 1801, 1802, 1809 and 1810 or an amplicon sequence consisting of a nucleotide sequence substantially identical or similar to any of the above sequences and optionally containing nucleic acid primer sequences at the 5' and 3' ends of the sequence.In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of the following nucleic acids and / or nucleic acid primer pairs, which contain or consist essentially of one or more nucleotide sequences selected from the group consisting of: SEQ ID NO: 1 in Table 16 ID NO: 53-58, 77-80, 109-114, 125-130, 165-180, 197-208, 237-242, 295-300, 343-346, 441-444, 457-460, 493-498, 511-520, 521-524, 529-534, 555-558, 571-580, 595, 596, 619-638, 645-650, 665-668, 731-734, 803, 804, 811, 812 and / or SEQ ID NO: NO:831-836、855-858、887-892、903-908、943-958、975-986、1015-1020、1073-1078 ,1121-1124,1219-1222,1235-1238,1271-1276,1289-1298,1299-1302,1307-1312, 1333-1336, 1349-1358, 1373, 1374, 1397-1416, 1423-1428, 1443-1446, 1509-1512, 1581, 1582, 1589 and 1590, or substantially identical or similar sequences, or the nucleotide sequence of any of the foregoing nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases.In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of the following nucleic acids and / or nucleic acid primer pairs, which contain or consist essentially of one or more nucleotide sequences (in the case of primer pairs) selected from the group consisting of: SEQ ID NOs: 53-58, 77-80, 109-114, 125-130, 165-180, 197-208, 237-242, 295-300, 343-346, 441-444, 457-460, 493-498 and 511-520 in Table 16 and / or SEQ ID NOs: NO:831-836, 855-858, 887-892, 903-908, 943-958, 975-986, 1015-1020, 1073-1078, 1121-1124, 1219-1222, 1235-1238, 1271-1276, 1289-1298, or a substantially identical or similar sequence, or the nucleotide sequence of any of the foregoing nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of the following nucleic acids and / or nucleic acid primer pairs, which contain or consist essentially of one or more nucleotide sequences (in the case of primer pairs) selected from the group consisting of: SEQ ID NO: 53-58, 77-80, 109-114, 125-130, 165-180, 197-208, 237-242, 295-300, 343-346, 441-444, 457-460 in Table 16A and / or SEQ ID NO: 53-58, 77-80, 109-114, 125-130, 165-180, 197-208, 237-242, 295-300, 343-346, 441-444, 457-460 in Table 16D NO:831-836, 855-858, 887-892, 903-908, 943-958, 975-986, 1015-1020, 1073-1078, 1121-1124, 1219-1222, 1235-1238, or a substantially identical or similar sequence, or the nucleotide sequence of any of the foregoing nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases.In some embodiments, the combination comprises or consists essentially of different nucleic acids or primers, each of which contains or consists essentially of each of the following different sequences: SEQ ID NO: 53-58, 77-80, 109-114, 125-130, 165-180, 197-208, 237-242, 295-300, 343-346, 441-444, 457-460, 493-498 and 511-520 in Table 16 and / or SEQ ID NO: NO:831-836, 855-858, 887-892, 903-908, 943-958, 975-986, 1015-1020, 1073-1078, 1121-1124, 1219-1222, 1235-1238, 1271-1276, 1289-1298, or a substantially identical or similar sequence, or the nucleotide sequence of any of the foregoing nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases. In some embodiments, the combination comprises or consists essentially of different nucleic acids or primers, each of which contains or consists essentially of each of the following different sequences: SEQ ID NO: 53-58, 77-80, 109-114, 125-130, 165-180, 197-208, 237-242, 295-300, 343-346, 441-444, 457-460 in Table 16A and / or SEQ ID NO: NO:831-836, 855-858, 887-892, 903-908, 943-958, 975-986, 1015-1020, 1073-1078, 1121-1124, 1219-1222, 1235-1238, or a substantially identical or similar sequence, or the nucleotide sequence of any of the foregoing nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases.
[0082]
[0083]
[0084]
[0085]
[0086]
[0087]
[0088]
[0089]
[0090]
[0091]
[0092]
[0093]
[0094] In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises two or more nucleic acids and / or nucleic acid primer pairs that specifically bind to, hybridize with, and / or specifically amplify a unique nucleic acid sequence contained in the genome of one or more of: Akkermansia muciniphila, Anaerobic cocci vaginalis, Miracidium minutissima, Bacteroides nordicus, Bacteroides thetaiotaomicron, Bacteroides vulgaris, Bifidobacterium adolescentis, Bifidobacterium longum, Collinsella aerogenes, Collinsella faecalis, Desmodus alaskaensis Thiovibrio, formigenic Dorrella, Enterococcus faecium, Eubacterium rectum, Faecalibacterium prausnitzii, Gardnerella vaginalis, Bacillus formicum, filamentous Holdermanella, Klebsiella pneumoniae, Parabacteroides dissimilar, Parabacteroides faecium, Pseudomonas aeruginosa, Prevotella tissue-dwelling, Roseburia intestinalis, Ruminococcus brucei, Slackia sclerotiorum, Streptococcus infantis, and Veillonella parvum (referred to herein as "Group B" microorganisms; see Table 2B), which are species believed to be associated with immuno-oncology treatment response. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises a set of nucleic acid primer pairs, wherein each different nucleic acid primer pair specifically amplifies a different unique nucleic acid sequence in a different genome of each genome of a different microorganism in Group B. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of nucleic acids and / or nucleic acid primer pairs that bind, hybridize, and / or amplify nucleic acids (e.g., nucleic acids from microorganisms (e.g., bacteria)), or specifically bind, hybridize, and / or specifically amplify nucleic acids, wherein the nucleic acids contain a SEQ ID NO: 17. NO:1605, 1606, 1643-1645, 1648-1650, 1659-1667, 1682-1684, 1689-1694, 1702-1704, 1718-1723 ,1728-1730,1735-1742,1748-1751,1754-1766,1780-1783,1791,1792,1801,1802,1809-1816,18 21-1826, 1829, 1830, 1864, 1869-1871, 1874-1882, 1890-1896, 1901-1903, 1910-1915, 1920-1922, 1930-1931, 1934-1939, 1954, 1955, 1961-1964, 1968, 1972-1974 and 1977-1979 sequences, and / or substantially identical or similar sequences.In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of the following nucleic acids and / or nucleic acid primer pairs, which bind, hybridize and / or amplify, or specifically bind, hybridize and / or specifically amplify a nucleic acid (e.g., a nucleic acid from a microorganism (e.g., bacteria)), wherein the nucleic acid contains a SEQ ID NO: NO:1605, 1606, 1643-1645, 1648-1650, 1659-1667, 1682-1684, 1689-1694, 1702-1704, 1718-1723, 1728-1730, 1735-1742, 1748-1751, 1754-1766, 1780-1783, 1791, 1792, 1801, 1802, 1809-1816 and 1821-1826, and / or substantially identical or similar sequences. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of nucleic acids and / or nucleic acid primer pairs that bind, hybridize, and / or amplify, or specifically bind, hybridize, and / or specifically amplify, a nucleic acid (e.g., a nucleic acid from a microorganism (e.g., bacteria)) comprising a SEQ ID NO: NO:1605, 1606, 1643-1645, 1648-1650, 1659-1667, 1682-1684, 1689-1694, 1702-1704, 1718-1723, 1728-1730, 1735-1742, 1748-1751, 1754-1766, 1780-1783, 1791, 1792, 1801, 1802 and 1809-1816, and / or substantially identical or similar sequences.In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of primers and / or primer pairs capable of amplifying or specifically amplifying a nucleic acid, such as a nucleic acid from a microorganism (e.g., bacteria), wherein the nucleic acid comprises a SEQ ID NO: 17 or a nucleic acid selected from Table 17. NO:1605, 1606, 1643-1645, 1648-1650, 1659-1667, 1682-1684, 1689-1694, 1702-1704, 1718-1723, 1728-1730, 1735-1742, 1748-1751, 1754 -1766, 1780-1783, 1791, 1792, 1801, 1802, 1809-1816, 1821-1826, 18 29, 1830, 1864, 1869-1871, 1874-1882, 1890-1896, 1901-1903, 1910-1 915, 1920-1922, 1930-1931, 1934-1939, 1954, 1955, 1961-1964, 1968, 1972-1974, and 1977-1979 (or sequences substantially identical or similar to any of the foregoing sequences) to produce an amplicon sequence of less than about 500, less than about 475, less than about 450, less than about 400, less than about 375, less than about 350, less than about 300, less than about 275, less than about 250, less than about 200, less than about 175, less than about 150, or less than about 100 nucleotides in length, or consisting essentially of a SEQ ID NO: 17 selected from Table 17 NO:1605, 1606, 1643-1645, 1648-1650, 1659-1667, 1682-1684, 1689-1694, 1702-1704, 1718-1723, 1728-1730 ,1735-1742,1748-1751,1754-1766,1780-1783,1791,1792,1801,1802,1809-1816,1821-1826,1829,1830,18 64, 1869-1871, 1874-1882, 1890-1896, 1901-1903, 1910-1915, 1920-1922, 1930-1931, 1934-1939, 1954, 1955, 1961-1964, 1968, 1972-1974 and 1977-1979, or a sequence substantially identical or similar to any of the foregoing sequences and optionally comprising a nucleotide sequence at the 5' and 3' ends of the sequence.In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of primers and / or primer pairs capable of amplifying or specifically amplifying a nucleic acid, such as a nucleic acid from a microorganism (e.g., bacteria), wherein the nucleic acid comprises SEQ ID NOs: 1605, 1606, 1643-1645, 1648-1650, 1659-1667, 1682-1684, 1689-1694, 1702-1704, 1718-1723, 1728-1730, 1735-1742, 1748-1751, 1754-1766, 1780-1783, 1791, 1792, 1801, 1802, 1809-1816, and 1817 of Table 17. 821-1826 (or a sequence substantially identical or similar to any of the foregoing sequences) to produce an amplicon sequence of less than about 500, less than about 475, less than about 450, less than about 400, less than about 375, less than about 350, less than about 300, less than about 275, less than about 250, less than about 200, less than about 175, less than about 150, or less than about 100 nucleotides in length, or consisting essentially of a SEQ ID NO: 17 or a sequence selected from Table 17. ID NO: 1605, 1606, 1643-1645, 1648-1650, 1659-1667, 1682-1684, 1689-1694, 1702-1704, 1718-1723, 1728-1730, 1735-1742, 1748-1751, 1754-1766, 1780-1783, 1791, 1792, 1801, 1802, 1809-1816 and 1821-1826, or an amplicon sequence consisting of a nucleotide sequence substantially identical or similar to any of the above sequences and optionally containing nucleic acid primer sequences at the 5' and 3' ends of the sequence.In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of primers and / or primer pairs capable of amplifying or specifically amplifying a nucleic acid, such as a nucleic acid from a microorganism (e.g., bacteria), wherein the nucleic acid comprises SEQ ID NOs: 1605, 1606, 1643-1645, 1648-1650, 1659-1667, 1682-1684, 1689-1694, 1702-1704, 1718-1723, 1728-1730, 1735-1742, 1748-1751, 1754-1766, 1780-1783, 1791, 1792, 1801, 1802, and 1809-1810 of Table 17A. 16 (or a sequence substantially identical or similar to any of the foregoing sequences) to produce an amplicon sequence of less than about 500, less than about 475, less than about 450, less than about 400, less than about 375, less than about 350, less than about 300, less than about 275, less than about 250, less than about 200, less than about 175, less than about 150, or less than about 100 nucleotides in length, or consisting essentially of a SEQ ID NO: 17A selected from the group consisting of ID NO: 1605, 1606, 1643-1645, 1648-1650, 1659-1667, 1682-1684, 1689-1694, 1702-1704, 1718-1723, 1728-1730, 1735-1742, 1748-1751, 1754-1766, 1780-1783, 1791, 1792, 1801, 1802 and 1809-1816, or an amplicon sequence consisting of a nucleotide sequence substantially identical or similar to any of the above sequences and optionally containing nucleic acid primer sequences at the 5' and 3' ends of the sequence.In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of the following nucleic acids and / or nucleic acid primer pairs, which contain or consist essentially of one or more nucleotide sequences (in the case of primer pairs) selected from the group consisting of: SEQ ID NO: 49-52, 125-130, 135-140, 157-174, 203-208, 217-228, 243-248, 275-286, 295-300, 309-324, 335-342, 347-372, 399-406, 421-424, 441-444, 457-472, 462-473, 474-487, 481-489, 490-501, 511-524, 512-526, 513-527, 514-528, 515-529, 520-529, 521-529, 522-521 81-492, 525-528, 595, 596, 605-610, 615-632, 647-660, 669-674, 687-698, 707-712, 727-730, 735-746, 775-778, 789-796, 803, 804, 811-816, 821-826 and / or SEQ ID NO:827-830、903-908、913-918、935-952、981-986、995-1006、1021-1026、1053-1064、1073-1078、1087-1102 ,1113-1120,1125-1150,1177-1184,1199-1202,1219-1222,1235-1250,1259-1270,1303-1306,1373,1374, 1383-1388, 1393-1410, 1425-1438, 1447-1452, 1465-1476, 1485-1490, 1505-1508, 1513-1524, 1553-1556, 1567-1574, 1581, 1582, 1589-1594, 1599-1604, or a substantially identical or similar sequence, or the nucleotide sequence of any of the foregoing nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases.In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of nucleic acids and / or nucleic acid primer pairs containing or consisting essentially of one or more nucleotide sequences (in the case of primer pairs) selected from the group consisting of: SEQ ID NO: 49-52, 125-130, 135-140, 157-174, 203-208, 217-228, 243-248, 275-286, 295-300, 309-324, 335-342, 347-372, 399-406, 421-424, 441-444, 457-472, 481-492 and / or SEQ ID NO: NO:827-830, 903-908, 913-918, 935-952, 981-986, 995-1006, 1021-1026, 1053-1064, 1073-1078, 1087-1102, 1113-1120, 1125-1150, 1177-1184, 1199-1202, 1219-1222, 1235-1250, 1259-1270, or a substantially identical or similar sequence, or the nucleotide sequence of any of the foregoing nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of nucleic acids and / or nucleic acid primer pairs containing or consisting essentially of one or more nucleotide sequences (in the case of primer pairs) selected from the group consisting of: SEQ ID NO: 49-52, 125-130, 135-140, 157-174, 203-208, 217-228, 243-248, 275-286, 295-300, 309-324, 335-342, 347-372, 399-406, 421-424, 441-444, 457-472 in Table 16A and / or SEQ ID NO: 49-52, 125-130, 135-140, 157-174, 203-208, 217-228, 243-248, 275-286, 295-300, 309-324, 335-342, 347-372, 399-406, 421-424, 441-444, 457-472 in Table 16D NO:827-830, 903-908, 913-918, 935-952, 981-986, 995-1006, 1021-1026, 1053-1064, 1073-1078, 1087-1102, 1113-1120, 1125-1150, 1177-1184, 1199-1202, 1219-1222, 1235-1250, or a substantially identical or similar sequence, or the nucleotide sequence of any of the foregoing nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases.In some embodiments, the combination comprises or consists essentially of different nucleic acids or primers, each of which contains or consists essentially of each of the following different sequences: 49-52, 125-130, 135-140, 157-174, 203-208, 217-228, 243-248, 275-286, 295-300, 309-324, 335-342, 347-372, 399-406, 421-424, 441-444, 457-472, 481-492 and / or SEQ ID NO:827-830, 903-908, 913-918, 935-952, 981-986, 995-1006, 1021-1026, 1053-1064, 1073-1078, 1087-1102, 1113-1120, 1125-1150, 1177-1184, 1199-1202, 1219-1222, 1235-1250, 1259-1270, or a substantially identical or similar sequence, or the nucleotide sequence of any of the foregoing nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases. In some embodiments, the combination comprises or consists essentially of different nucleic acids or primers, each of which contains or consists essentially of each of the following different sequences: SEQ ID NO: 49-52, 125-130, 135-140, 157-174, 203-208, 217-228, 243-248, 275-286, 295-300, 309-324, 335-342, 347-372, 399-406, 421-424, 441-444, 457-472 in Table 16A and / or SEQ ID NO: ID NO: 827-830, 903-908, 913-918, 935-952, 981-986, 995-1006, 1021-1026, 1053-1064, 1073-1078, 1087-1102, 1113-1120, 1125-1150, 1177-1184, 1199-1202, 1219-1222, 1235-1250, or a substantially identical or similar sequence, or the nucleotide sequence of any of the foregoing nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases.
[0095] In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises two or more nucleic acids and / or nucleic acid primer pairs that specifically bind to, hybridize to, and / or specifically amplify a unique nucleic acid sequence contained in the genome of one or more of: Bacteroides fragilis, Campylobacter jejuni, Propionibacterium acnes, Escherichia coli, Fusobacterium nucleatum, Helicobacter cholerae, Helicobacter bieneusi, Helicobacter hepatica, Helicobacter pylori, Helicobacter zabbixii, Peptostreptococcus oralis, and Streptococcus gallolyticus (referred to herein as "Group C" microorganisms; see Table 2C), which are species believed to be associated with cancer. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises two or more nucleic acids and / or nucleic acid primer pairs that specifically bind to, hybridize with, and / or specifically amplify a unique nucleic acid sequence contained in the genome of one or more of: Bacteroides fragilis, Campylobacter jejuni, Propionibacterium acnes, Escherichia coli, Fusobacterium nucleatum, Helicobacter bilis, Helicobacter bieneusi, Helicobacter hepatis, Helicobacter pylori, Peptostreptococcus oralis, and Streptococcus gallolyticus (referred to herein as "subgroup 1" of Group C microorganisms). In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises a set of nucleic acid primer pairs, wherein each different nucleic acid primer pair specifically amplifies a different unique nucleic acid sequence in a different genome of each genome of a different microorganism in Group C or Group C that does not contain Helicobacter zabbii (i.e., Subgroup 1 of Group C).In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of the following nucleic acids and / or nucleic acid primer pairs, which bind, hybridize and / or amplify, or specifically bind, hybridize and / or specifically amplify a nucleic acid (e.g., a nucleic acid from a microorganism (e.g., bacteria)), wherein the nucleic acid contains a SEQ ID NO: NO:1616, 1619, 1620, 1625-1628, 1635-1640, 1699, 1700, 1705-1708, 1752, 1753, 1784-1786, 1817-1820, 1827, 1828, 1840, 1841, 1844, 1845, 1852-1859, 1899, 1900, 1904, 1905, 1932, 1933, 1956-1958, 1975, 1976 and / or a substantially identical or similar sequence, or the nucleic acid contains a SEQ ID NO selected from Table 17 NO:1616, 1619, 1620, 1625-1628, 1635-1640, 1699, 1700, 1705-1708, 1752, 1753, 1784-1786, 1827, 1828, 1840, 1841, 1844, 1845, 1852-1859, 1899, 1900, 1904, 1905, 1932, 1933, 1956, 1957, 1958 and / or substantially identical or similar sequences. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of nucleic acids and / or nucleic acid primer pairs that bind, hybridize, and / or amplify, or specifically bind, hybridize, and / or specifically amplify, a nucleic acid, such as a nucleic acid from a microorganism (e.g., bacteria), wherein the nucleic acid comprises a sequence selected from SEQ ID NO: 1616, 1619, 1620, 1625-1628, 1635-1640, 1699, 1700, 1705-1708, 1752, 1753, 1784-1786, 1817-1820 of Table 17A and / or a substantially identical or similar sequence, or the nucleic acid comprises a sequence selected from SEQ ID NO: 1616, 1619, 1620, 1625-1628, 1635-1640, 1699, 1700, 1705-1708, 1752, 1753, 1784-1786, 1817-1820 of Table 17A and / or a substantially identical or similar sequence. NO:1616, 1619, 1620, 1625-1628, 1635-1640, 1699, 1700, 1705-1708, 1752, 1753, 1784-1786 and / or substantially identical or similar sequences.In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of primers and / or primer pairs capable of amplifying or specifically amplifying a nucleic acid, such as a nucleic acid from a microorganism (e.g., bacteria), wherein the nucleic acid comprises SEQ ID NOs: 1616, 1619, 1620, 1625-1628, 1635-1640, 1699, 1700, 1705-1708, 1752, 1753, 1784-1786, 1817-1820, 1827, 1828, 1840, 1841, 1844, 1845, 1852-1859, 1899, 1900, 1904, 1905, 1932, 1933, 1956-1957 of Table 17. 8, 1975, 1976 (or a sequence substantially identical or similar to any of the foregoing sequences) to produce an amplicon sequence of less than about 500, less than about 475, less than about 450, less than about 400, less than about 375, less than about 350, less than about 300, less than about 275, less than about 250, less than about 200, less than about 175, less than about 150, or less than about 100 nucleotides in length, or consisting essentially of a SEQ ID NO: 17 selected from Table 17 ID NO: 1616, 1619, 1620, 1625-1628, 1635-1640, 1699, 1700, 1705-1708, 1752, 1753, 1784-1786, 1817-1820, 1827, 1828, 1840, 1841, 1844, 1845, 1852-1859, 1899, 1900, 1904, 1905, 1932, 1933, 1956-1958, 1975, 1976, or an amplicon sequence consisting of a nucleotide sequence substantially identical or similar to any of the above sequences and optionally containing nucleic acid primer sequences at the 5' and 3' ends of the sequence.In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of primers and / or primer pairs capable of amplifying or specifically amplifying a nucleic acid, such as a nucleic acid from a microorganism (e.g., bacteria), wherein the nucleic acid comprises a SEQ ID NO: 10 selected from Table 17A. NO:1616, 1619, 1620, 1625-1628, 1635-1640, 1699, 1700, 1705-1708, 1752, 1753, 1784-1786, 1817-1820 (or a sequence substantially identical or similar to any of the foregoing sequences) to produce an amplicon sequence of less than about 500, less than about 475, less than about 450, less than about 400, less than about 375, less than about 350, less than about 300, less than about 275, less than about 250, less than about 200, less than about 175, less than about 150, or less than about 100 nucleotides in length, or consisting essentially of a SEQ ID NO selected from Table 17A NO:1616, 1619, 1620, 1625-1628, 1635-1640, 1699, 1700, 1705-1708, 1752, 1753, 1784-1786, 1817-1820 or an amplicon sequence consisting of a nucleotide sequence substantially identical or similar to any of the above sequences and optionally containing a nucleic acid primer sequence at the 5' end and 3' end of the sequence.In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of nucleic acids and / or nucleic acid primer pairs containing or consisting essentially of one or more nucleotide sequences (in the case of primer pairs) selected from the group consisting of: SEQ ID NO: 71, 72, 77-80, 89-96, 109-120, 237-242, 249-256, 343-346, 407-412, 473-480, 493-496, 511-520, 521-524, 547-550, 555-558, 561-568, 571-586, 665-668, 675-678, 731-734, 779-784, 817-820 and / or SEQ ID NO: NO:849, 850, 855-858, 867-874, 887-898, 1012-1020, 1025-1034, 1121-1124, 1185-1190, 1251-1258, 1271-1276, 1289-1298, 1299-1302, 1325-1328, 1333-1336, 1339-1346, 1349-1364, 1443-1446, 1453-1456, 1509-1512, 1557-1562, 1595-1598, or a substantially identical or similar sequence, or the nucleotide sequence of any of the foregoing nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of the following nucleic acids and / or nucleic acid primer pairs, which contain or consist essentially of one or more nucleotide sequences (in the case of primer pairs) selected from the group consisting of: SEQ ID NO: 71, 72, 77-80, 89-96, 109-120, 237-242, 249-256, 343-346, 407-412, 473-480, 493-496, 511-520 and / or SEQ ID NO: NO:849, 850, 855-858, 867-874, 887-898, 1012-1020, 1025-1034, 1121-1124, 1185-1190, 1251-1258, 1271-1276, 1289-1298, or a substantially identical or similar sequence, or the nucleotide sequence of any of the foregoing nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases.In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of the following nucleic acids and / or nucleic acid primer pairs, which contain or consist essentially of one or more nucleotide sequences selected from the following (in the case of primer pairs): SEQ ID NO: 71, 72, 77-80, 89-96, 109-120, 237-242, 249-256, 343-346, 407-412, 473-480 and / or SEQ ID NO: 71, 72, 77-80, 89-96, 109-120, 237-242, 249-256, 343-346, 407-412, 473-480 and / or SEQ ID NO: NO:849, 850, 855-858, 867-874, 887-898, 1012-1020, 1025-1034, 1121-1124, 1185-1190, 1251-1258, or a substantially identical or similar sequence, or the nucleotide sequence of any of the foregoing nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases. In some embodiments, the combination comprises or consists essentially of different nucleic acids or primers, each of which contains or consists essentially of each of the following different sequences: SEQ ID NO: 71, 72, 77-80, 89-96, 109-120, 237-242, 249-256, 343-346, 407-412, 473-480, 493-496, 511-520, 521-524, 547-550, 555-558, 561-568, 571-586, 665-668, 675-678, 731-734, 779-784, 817-820 and / or SEQ ID NO: NO:849, 850, 855-858, 867-874, 887-898, 1012-1020, 1025-1034, 1121-1124, 1185-1190, 1251-1258, 1271-1276, 1289-1298, 1299-1302, 1325-1328, 1333-1336, 1339-1346, 1349-1364, 1443-1446, 1453-1456, 1509-1512, 1557-1562, 1595-1598, or a substantially identical or similar sequence, or the nucleotide sequence of any of the foregoing nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases.In some embodiments, the combination comprises or consists essentially of different nucleic acids or primers, each of which contains or consists essentially of each of the following different sequences: SEQ ID NO: 71, 72, 77-80, 89-96, 109-120, 237-242, 249-256, 343-346, 407-412, 493-496, 511-520, 521-524, 547-550, 555-558, 561-568, 571-586, 665-668, 675-678, 731-734, 779-784 and / or SEQ ID NO: NO:849, 850, 855-858, 867-874, 887-898, 1012-1020, 1025-1034, 1121-1124, 1185-1190, 1271-1276, 1289-1298, 1299-1302, 1325-1328, 1333-1336, 1339-1346, 1349-1364, 1443-1446, 1453-1456, 1509-1512, 1557-1562, or a substantially identical or similar sequence, or the nucleotide sequence of any of the foregoing nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases. In some embodiments, the combination comprises or consists essentially of different nucleic acids or primers, each of which contains or consists essentially of each of the following different sequences: SEQ ID NO: 71, 72, 77-80, 89-96, 109-120, 237-242, 249-256, 343-346, 407-412, 473-480, 493-496, 511-520 and / or SEQ ID NO: IDNO: 849, 850, 855-858, 867-874, 887-898, 1012-1020, 1025-1034, 1121-1124, 1185-1190, 1251-1258, 1271-1276, 1289-1298, or a substantially identical or similar sequence, or the nucleotide sequence of any of the foregoing nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases.In some embodiments, the combination comprises or consists essentially of different nucleic acids or primers, each of which contains or consists essentially of each of the following different sequences: SEQ ID NO: 71, 72, 77-80, 89-96, 109-120, 237-242, 249-256, 343-346, 407-412, 493-496, 511-520 and / or SEQ ID NO: NO:849, 850, 855-858, 867-874, 887-898, 1012-1020, 1025-1034, 1121-1124, 1185-1190, 1271-1276, 1289-1298, or a substantially identical or similar sequence, or the nucleotide sequence of any of the foregoing nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases. In some embodiments, the combination comprises or consists essentially of different nucleic acids or primers, each of which contains or consists essentially of SEQ ID NOs: 71, 72, 77-80, 89-96, 109-120, 237-242, 249-256, 343-346, 407-412, 473-480 and / or SEQ ID NOs: 849, 850, 855-858, 867-874, 887-898, 1012-1020, 1025-1034, 1121-1124, 1185-1190, 1251-1258 in Table 16, or a substantially identical or similar sequence, or the nucleotide sequence of any of the foregoing nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases. In some embodiments, the combination comprises or consists essentially of different nucleic acids or primers, each of which contains or consists essentially of SEQ ID NOs: 71, 72, 77-80, 89-96, 109-120, 237-242, 249-256, 343-346, 407-412 and / or SEQ ID NOs: 849, 850, 855-858, 867-874, 887-898, 1012-1020, 1025-1034, 1121-1124, 1185-1190 in Table 16, or a substantially identical or similar sequence, or the nucleotide sequence of any of the foregoing nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases.
[0096] In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises two or more nucleic acids and / or nucleic acid primer pairs that specifically bind to, hybridize to, and / or specifically amplify a unique nucleic acid sequence contained in the genome of one or more of Akkermansia muciniphila, Bifidobacterium bifidum, Bifidobacterium longum, Blautia sphaeroides, Campylobacter conciseness, Campylobacter curviformis, Campylobacter jejuni, Campylobacter rectum, Clostridium difficile, Escherichia coli, Eubacterium rectum, Fusobacterium nucleatum, Helicobacter cholerae, Helicobacter hepatica, Helicobacter pylori, Klebsiella pneumoniae, Lactobacillus delbrueckii, Parabacteroides distichous, Proteus mirabilis, Ruminococcus brunneri, and Ruminococcus vigorus (referred to herein as "Group D" microorganisms; see Table 2D), which are species believed to be associated with gastrointestinal disorders, including, for example, irritable bowel syndrome, inflammatory bowel disease, and celiac disease. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises a set of nucleic acid primer pairs, wherein each different nucleic acid primer pair specifically amplifies a different unique nucleic acid sequence contained in a different genome of each genome of different microorganisms in Group D. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of nucleic acids and / or nucleic acid primer pairs that bind, hybridize and / or amplify nucleic acids (e.g., nucleic acids from microorganisms (e.g., bacteria)), or specifically bind, hybridize and / or specifically amplify nucleic acids containing a sequence selected from the following: a sequence in Table 17, or a sequence in Table 17 that does not contain SEQ ID NOs: 1807, 1808 and 1971, or a sequence in Table 17A and Table 17B, or a sequence in Table 17B and Table 17A that does not contain SEQ ID NOs: 1807 and 1808, and / or a substantially identical or similar sequence that corresponds to a microorganism in Group D.In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of primers and / or primer pairs capable of amplifying or specifically amplifying a nucleic acid, such as a nucleic acid from a microorganism (e.g., bacteria), wherein the nucleic acid comprises a sequence selected from the group consisting of a sequence in Table 17, or a sequence in Table 17 that does not include SEQ ID NOs: 1807, 1808, and 1971, or a sequence in Table 17A and Table 17B, or a sequence in Table 17B and Table 17A that does not include SEQ ID NOs: NO:1807 and 1808, and / or a substantially identical or similar sequence corresponding to a microorganism of Group D (or a sequence substantially identical or similar to any of the foregoing sequences) to produce an amplicon sequence of less than about 500, less than about 475, less than about 450, less than about 400, less than about 375, less than about 350, less than about 300, less than about 275, less than about 250, less than about 200, less than about 175, less than about 150, or less than about 100 nucleotides in length, or an amplicon sequence consisting essentially of a nucleotide sequence selected from the group consisting of the following, and optionally containing a nucleic acid primer sequence at the 5' end and the 3' end of the sequence: a sequence in Table 17, or a sequence in Table 17 not comprising SEQ ID NO:1807, 1808 and 1971, or a sequence in Table 17A and Table 17B, or a sequence in Table 17B and Table 17A not comprising SEQ ID NO: NO:1807 and 1808, and / or substantially identical or similar sequences, which correspond to Group D microorganisms, or sequences substantially identical or similar to any of the above sequences. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of primers and / or primer pairs capable of amplifying or specifically amplifying nucleic acids (such as nucleic acids from microorganisms (e.g., bacteria)), wherein the nucleic acids contain a sequence selected from the group consisting of: a sequence in Table 17, or a sequence in Table 17 that does not contain SEQ ID NO:1807, 1808, and 1971, or a sequence in Table 17A and Table 17B, or a sequence in Table 17B and Table 17A that does not contain SEQ ID NO:1807, 1808, and 1971. NO:1807 and 1808 (or a sequence substantially identical or similar to any of the foregoing sequences), which corresponds to Group D microorganisms, to produce an amplicon sequence of less than about 500, less than about 475, less than about 450, less than about 400, less than about 375, less than about 350, less than about 300, less than about 275, less than about 250, less than about 200, less than about 175, less than about 150 or less than about 100 nucleotides in length, or an amplicon sequence consisting essentially of a nucleotide sequence selected from a sequence in Table 17 or Table 17A and Table 17B (which corresponds to Group D microorganisms) or a sequence substantially identical or similar to any of the foregoing sequences and optionally containing nucleic acid primer sequences at the 5' and 3' ends of the sequence.In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of the following nucleic acids and / or nucleic acid primer pairs, which contain (in the case of primer pairs) or consist essentially of one or more nucleotide sequences selected from the following: sequences corresponding to Group D microorganisms in Table 16; or sequences not comprising SEQ ID NOs: 453-456, 809, 810, 1231-1234, and 1587-1588 in Table 16; or SEQ ID NOs: 49-520 of Table 16; or SEQ ID NOs: 49-452 and 457-520 of Table 16; or SEQ ID NOs: 49-492 of Table 16; or SEQ ID NOs: 49-452 and 457-492 of Table 16; or SEQ ID NOs: 49-480 of Table 16A; or SEQ ID NOs: 49-521 of Table 16; or SEQ ID NOs: 49-453 and 457-453 ... or SEQ ID NOs:1299-1598 of Table 16F; or substantially the same or similar sequences, or the nucleotide sequences of any of the above nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases.In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of the following nucleic acids and / or nucleic acid primer pairs, which contain (in the case of primer pairs) or consist essentially of one or more nucleotide sequences selected from the following: sequences corresponding to Group D microorganisms in Table 16; or sequences not comprising SEQ ID NOs: 453-456, 809, 810, 1231-1234, and 1587-1588 in Table 16; or SEQ ID NOs: 49-520 of Table 16; or SEQ ID NOs: 49-452 and 457-520 of Table 16; or SEQ ID NOs: 49-492 of Table 16; or SEQ ID NOs: 49-452 and 457-492 of Table 16; or SEQ ID NOs: 49-480 of Table 16A; or SEQ ID NOs: 49-521 of Table 16; or SEQ ID NOs: 49-453 and 457-453 ... or SEQ ID NOs:1299-1598 of Table 16F; or substantially the same or similar sequences, or the nucleotide sequences of any of the above nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases.In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of the following nucleic acids and / or nucleic acid primer pairs, which contain (in the case of primer pairs) or consist essentially of one or more nucleotide sequences selected from the following: sequences corresponding to Group D microorganisms in Table 16; or sequences not comprising SEQ ID NOs: 453-456, 809, 810, 1231-1234, and 1587-1588 in Table 16; or SEQ ID NOs: 49-520 of Table 16; or SEQ ID NOs: 49-452 and 457-520 of Table 16; or SEQ ID NOs: 49-492 of Table 16; or SEQ ID NOs: 49-452 and 457-492 of Table 16; or SEQ ID NOs: 49-480 of Table 16A; or SEQ ID NOs: 49-521 of Table 16; or SEQ ID NOs: 49-453 and 457-453 of Table 16; or SEQ ID NOs: 49-492 ...80 of Table 16A; or SEQ ID NOs: 49-481 of Table 16A; or SEQ ID NOs: 49-493 of Table 16A; or SEQ ID NOs: 49-494 of Table 16A; or SEQ ID NOs: 49-495 of Table 16A; or SEQ ID NOs: 49-496 of Table 16A; or SEQ ID NOs: or SEQ ID NOs:1299-1598 of Table 16F; or substantially the same or similar sequences, or the nucleotide sequences of any of the above nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases.In some embodiments, the combination comprises or consists essentially of different nucleic acids or primers, each of which contains or consists essentially of the following different sequences: a sequence corresponding to a microorganism in Group D in Table 16; or a sequence not comprising SEQ ID NOs: 453-456, 809, 810, 1231-1234, and 1587-1588 in Table 16; or SEQ ID NOs: 49-520 of Table 16; or SEQ ID NOs: 49-452 and 457-520 of Table 16; or SEQ ID NOs: 49-492 of Table 16; or SEQ ID NOs: 49-452 and 457-492 of Table 16; or SEQ ID NOs: 49-480 of Table 16A; or SEQ ID NOs: 49-452 and 457-480 of Table 16A; or SEQ ID NOs: 49-452 and 457-480 of Table 16C. or SEQ ID NOs:1299-1598 of Table 16F; or substantially the same or similar sequences, or the nucleotide sequences of any of the above nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases.
[0097] In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises two or more nucleic acids and / or nucleic acid primer pairs that specifically bind, hybridize and / or specifically amplify a unique nucleic acid sequence contained in the genome of one or more of: Akkermansia muciniphila, Bacteroides fragilis, Bacteroides vulgaris, Bifidobacterium adolescentis, Campylobacter conciseness, Campylobacter jejuni, Citrobacter rodentium, Clostridium difficile, Enterococcus gallinarum, Escherichia coli, Helicobacter cholerae, Lactobacillus delbrueckii, Lactobacillus murinus, Lactobacillus reuteri, Lactobacillus rhamnosus, Lactobacillus lactis and Prevotella hominis (referred to herein as "Group E" microorganisms; see Table 2E), which are species believed to be associated with autoimmune disorders, including but not limited to lupus and rheumatoid arthritis. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises a set of nucleic acid primer pairs, wherein each different nucleic acid primer pair specifically amplifies a different unique nucleic acid sequence contained in a different genome of each genome of a different microorganism in Group E. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of nucleic acids and / or nucleic acid primer pairs that bind to, hybridize to, and / or amplify nucleic acids, such as nucleic acids from microorganisms (e.g., bacteria), or specifically bind to, hybridize to, and / or specifically amplify nucleic acids containing a sequence selected from the group consisting of sequences in Table 17 or Table 17A and Table 17B, and / or substantially identical or similar sequences, corresponding to Group E microorganisms. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of primers and / or primer pairs capable of amplifying or specifically amplifying a nucleic acid, such as a nucleic acid from a microorganism (e.g., bacteria), comprising a sequence selected from the group consisting of: a sequence in Table 17 or Table 17A and Table 17B, and / or a substantially identical or similar sequence corresponding to a Group E microorganism (or a sequence substantially identical or similar to any of the foregoing sequences), to produce a sequence less than about 500, less than about 475, less than about 450, less than about 4 00, less than about 375, less than about 350, less than about 300, less than about 275, less than about 250, less than about 200, less than about 175, less than about 150, or less than about 100 nucleotides, or an amplicon sequence consisting essentially of a nucleotide sequence selected from a sequence in Table 17 or Table 17A and Table 17B and / or a substantially identical or similar sequence (which corresponds to Group E microorganisms) or a sequence substantially identical or similar to any of the above sequences and optionally containing nucleic acid primer sequences at the 5' and 3' ends of the sequence.In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of primers and / or primer pairs capable of amplifying or specifically amplifying a nucleic acid, such as a nucleic acid from a microorganism (e.g., bacteria), wherein the nucleic acid comprises a sequence selected from the group consisting of a sequence in Table 17 or Table 17A and Table 17B (or a sequence substantially identical or similar to any of the foregoing sequences) corresponding to a Group E microorganism to produce a sequence less than about 500, less than about 475, less than about 450, less than about 4 00, less than about 375, less than about 350, less than about 300, less than about 275, less than about 250, less than about 200, less than about 175, less than about 150, or less than about 100 nucleotides, or an amplicon sequence consisting essentially of a nucleotide sequence selected from a sequence in Table 17 or Table 17A and Table 17B (which corresponds to Group E microorganisms) or a sequence substantially identical or similar to any of the above sequences and optionally containing nucleic acid primer sequences at the 5' and 3' ends of the sequence. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of the following nucleic acids and / or nucleic acid primer pairs, which contain or consist essentially of one or more nucleotide sequences selected from the following (in the case of primer pairs): sequences corresponding to Group E microorganisms in Table 16; SEQ ID NOs: 49-520 of Table 16; SEQ ID NOs: 49-492 of Table 16; SEQ ID NOs: 49-480 of Table 16A; SEQ ID NOs: 521-826 of Table 16C; SEQ ID NOs: 521-820 of Table 16C; SEQ ID NOs: 827-1298 of Table 16; SEQ ID NOs: 827-1258 of Table 16D; SEQ ID NOs: 1299-1604 of Table 16F; or SEQ ID NOs: 16F NO:1299-1598; or a substantially identical or similar sequence, or the nucleotide sequence of any of the above nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases.In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of the following nucleic acids and / or nucleic acid primer pairs, which contain or consist essentially of one or more nucleotide sequences selected from the following (in the case of primer pairs): sequences corresponding to Group E microorganisms in Table 16; SEQ ID NOs: 49-520 of Table 16; SEQ ID NOs: 49-492 of Table 16; SEQ ID NOs: 49-480 of Table 16A; SEQ ID NOs: 521-826 of Table 16C; SEQ ID NOs: 521-820 of Table 16C; SEQ ID NOs: 827-1298 of Table 16; SEQ ID NOs: 827-1258 of Table 16D; SEQ ID NOs: 1299-1604 of Table 16F; or SEQ ID NOs: 16F NO:1299-1598; or a substantially identical or similar sequence, or the nucleotide sequence of any of the above nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises or consists essentially of the following nucleic acids and / or nucleic acid primer pairs, which contain or consist essentially of one or more nucleotide sequences selected from the following (in the case of primer pairs): sequences corresponding to Group E microorganisms in Table 16; SEQ ID NOs: 49-520 of Table 16; SEQ ID NOs: 49-492 of Table 16; SEQ ID NOs: 49-480 of Table 16A; SEQ ID NOs: 521-826 of Table 16C; SEQ ID NOs: 521-820 of Table 16C; SEQ ID NOs: 827-1298 of Table 16; SEQ ID NOs: 827-1258 of Table 16D; SEQ ID NOs: 1299-1604 of Table 16F; or SEQ ID NOs: 16F NO:1299-1598; or a substantially identical or similar sequence, or the nucleotide sequence of any of the above nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases.In some embodiments, the combination comprises or consists essentially of different nucleic acids or primers, each of which contains or consists essentially of the following different sequences: sequences corresponding to group E microorganisms in Table 16; SEQ ID NOs: 49-520 of Table 16; SEQ ID NOs: 49-492 of Table 16; SEQ ID NOs: 49-480 of Table 16A; SEQ ID NOs: 521-826 of Table 16C; SEQ ID NOs: 521-820 of Table 16C; SEQ ID NOs: 827-1298 of Table 16; SEQ ID NOs: 827-1258 of Table 16D; SEQ ID NOs: 1299-1604 of Table 16F; or SEQ ID NOs: 1299-1598 of Table 16F; or substantially the same or similar sequences, or the nucleotide sequence of any of the above nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases.
[0098] Nucleic acid combination
[0099] In order to accurately assess, analyze and characterize microbial populations, so as to establish meaningful correlations between the microbiome of an animal or environment and a healthy state or homeostasis and imbalance or disease, and then assess and characterize microbial sample to detect and / or diagnose imbalance, susceptibility, illness and / or disease, it is necessary to be able to comprehensively, specifically and proportionally assess the microbial composition in the microbial population. Accurate analysis of microbial populations relies on comprehensively detecting and identifying all microorganisms (e.g., bacteria) present in the population at least at the genus level, as well as detecting and identifying some microorganisms, such as microorganisms that are particularly important in health and disease, most, most or substantially all microbial species present in the population, to achieve sufficient depth of microbial composition in the population. Compositions and methods are provided herein, as well as combinations, kits and systems comprising the compositions and methods, for accurately, comprehensively, information-rich, sensitive, specific, fast, high-throughput and cost-effective assessment, analysis or characterization of a mixture or population of microorganisms (e.g., bacteria). In certain embodiments, a mixture or population of microorganisms is present in a sample (e.g., a biological sample), e.g., a sample of the digestive tract contents of an organism (e.g., an animal). In some embodiments, compositions provided herein for such evaluation, analysis or characterization of a mixture or population of microorganisms (e.g., bacteria) comprise a combination of: (1) one or more kingdom-covered nucleic acid primer pairs that are capable of amplifying sequences in homologous genes or genomic regions that are common to multiple, most, most, substantially all or all microorganisms (e.g., bacteria) in a kingdom but differ between microorganisms of different kingdoms; and / or (2) microorganism-specific nucleic acids and / or nucleic acid primer pairs that are capable of amplifying or specifically or selectively amplifying specific nucleic acid sequences that are unique to a particular microorganism (e.g., a species, subspecies or strain of a microorganism such as bacteria). Many embodiments of kingdom-covered nucleic acid primer pairs and microorganism-specific nucleic acid primer pairs that can be used in nucleic acid combinations are provided herein.
[0100] For example, in some embodiments, the nucleic acid of the nucleic acid combination includes one or more primer pairs that amplify two or more regions in the prokaryotic (e.g., bacterial) 16S rRNA gene, such as hypervariable regions. In some embodiments, the nucleotide sequences of any two 16s rRNA gene primers of the nucleic acid that amplify the sequence located in multiple hypervariable regions are slightly overlapped (e.g., less than or equal to 7 nucleotides, or 6 nucleotides, or 5 nucleotides, or 4 nucleotides, or 3 nucleotides, or 2 nucleotides, or 1 nucleotide) or do not overlap. In some aspects, the nucleic acid primer pair of the nucleic acid is less than or equal to about 200 nucleotides in length, such as a 16s rRNA gene sequence between about 125 and 200 nucleotides in length. In some embodiments, the nucleic acid of the nucleic acid combination includes a plurality of nucleic acid primer pairs, which include at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8 or at least 9 separate primer pairs, and optionally degenerate variants thereof, which separately amplify nucleic acids containing sequences located in 2, 3, 4, 5, 6, 7, 8 or 9 different hypervariable regions of prokaryotic 16s rRNA genes in nucleic acid amplification reactions. In some embodiments, the nucleic acid of the nucleic acid combination includes at least 8 separate primer pairs, and optionally degenerate variants thereof, which separately amplify nucleic acids containing sequences located in 8 different hypervariable regions of prokaryotic 16s rRNA genes in nucleic acid amplification reactions. In some embodiments, the nucleic acid of the nucleic acid combination includes a plurality of primer pairs, which separately amplify nucleic acids containing sequences located in 3 or more hypervariable regions of prokaryotic 16S rRNA genes, and one of the 3 or more regions is a V5 region. Degenerate primer variants are included in some compositions, for example, different nucleotides at 1 or 2 positions in the primer sequence to ensure amplification of 16S rRNA genes containing minor changes in conserved regions. Restricted examples of nucleotide sequences of primer pairs that amplify 8 hypervariable regions (V2, V3, V4, V5, V6, V7, V8, and V9) of prokaryotic 16S rRNA genes are listed in Table 15. In some embodiments, the nucleic acids in the nucleic acid combination include at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 24, at least 30, at least 35, at least 40, at least 45 or more, or all primers, or all primer pairs, wherein the primers or primer pairs have a sequence listed in Table 15 or SEQ ID NOs: 1-24 in Table 15 and / or SEQ ID NOs: 25-48 in Table 15 or SEQ ID NOs: 11-16, 23 and 24 in Table 15 and / or SEQ ID NOs: 35-40, 47 and 48 in Table 15 or consist essentially of the sequence.In some embodiments, the kingdom-encompassing nucleic acids in the nucleic acid combination comprise one or more primer pairs that provide at least 85%, or at least 90%, or at least 92%, or at least 95%, or at least 98%, or at least 99%, or 100% coverage of different bacterial 16S rRNA gene sequences in a given database containing bacterial 16S rRNA gene sequences (e.g., GreenGenes bacterial 16S rRNA gene sequences; www.greengenes.lbl.gov; SILVA database (www.arb-silva.de)). In some embodiments, each of the one or more microorganism-specific nucleic acid primer pairs contained in the primer pair combination is capable of amplifying or specifically amplifying a specific nucleic acid (e.g., a nucleic acid sequence from a microorganism such as bacteria) selected from SEQ ID NOs: 1605-1979 of Table 17, or SEQ ID NOs: 1605-1806, 1809-1816, 1821-1970, 1972-1974, and 1977-1979 of Table 17, or SEQ ID NOs: 1605-1826 of Table 17, or SEQ ID NOs: 1605-1806, 1809-1816, and 1821-1826 of Table 17, or SEQ ID NOs: 1605-1820 of Table 17A, or SEQ ID NOs: 1605-1806 and 1809-1816 of Table 17A, or SEQ ID NOs: 1605-1820 of Table 17C. NO:1827-1979 or the nucleotide sequence of SEQ ID NO:1827-1976 in Table 17C or essentially consists of the nucleotide sequence.In some embodiments, each of the one or more microorganism-specific nucleic acid primer pairs contained in the primer pair combination is capable of amplifying or specifically amplifying a specific nucleic acid sequence, the specific nucleic acid sequence containing a nucleotide sequence selected from the following: SEQ ID NO: 1605-1979 in Table 17; or SEQ ID NO: 1605-1806, 1809-1816, 1821-1970, 1972-1974 and 1977-1979 in Table 17; or SEQ ID NO: 1605-1826 in Table 17; or SEQ ID NO: 1605-1806, 1809-1816 and 1821-1826 in Table 17; or SEQ ID NO: 1605-1820 in Table 17A; or SEQ ID NO: 1605-1806 and 1809-1816 in Table 17A; or SEQ ID NO: 1605-1806 and 1809-1816 in Table 17C or SEQ ID NOs: 1827-1979 in Table 17C, or substantially identical or similar sequences to produce an amplicon sequence of less than about 500, less than about 475, less than about 450, less than about 400, less than about 375, less than about 350, less than about 300, less than about 275, less than about 250, less than about 200, less than about 175, less than about 150, or less than about 100 nucleotides in length, or consisting essentially of SEQ ID NOs: 1605-1979 in Table 17, or SEQ ID NOs: 1605-1806, 1809-1816, 1821-1970, 1972-1974, and 1977-1979 in Table 17, or SEQ ID NOs: 1605-1826 in Table 17, or SEQ ID NOs: ID NOs: 1605-1806, 1809-1816 and 1821-1826, or SEQ ID NOs: 1605-1820 in Table 17A, or SEQ ID NOs: 1605-1806 and 1809-1816 in Table 17A, or SEQ ID NOs: 1827-1979 in Table 17C, or SEQ ID NOs: 1827-1976 in Table 17C, or substantially identical or similar sequences and optionally an amplicon sequence containing nucleic acid primer sequences at the 5' and 3' ends of the sequence.In some embodiments, the collection of microorganism-specific nucleic acid primer pairs in the combination is capable of amplifying or specifically amplifying at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 125, at least 150, at least 175, at least 200, at least 225, or at least 230 or more different nucleic acids in a multiplex reaction, wherein the different nucleic acids contain different sequences of: SEQ ID NO: 1605-1979 in Table 17, or SEQ ID NO: 1605-1979 in Table 17. NO: 1605-1806, 1809-1816, 1821-1970, 1972-1974 and 1977-1979, or SEQ ID NO: 1605-1826 in Table 17, or SEQ ID NO: 1605-1806, 1809-1816 and 1821-1826 in Table 17, or SEQ ID NO: 1605-1820 in Table 17A, or SEQ ID NO: 1605-1806 and 1809-1816 in Table 17A, or SEQ ID NO: 1827-1979 in Table 17C, or SEQ ID NO: 1827-1976 in Table 17C.In some such embodiments, the microorganism-specific nucleic acid primer pairs in the combination can amplify different nucleic acids, wherein the different nucleic acids contain the following different sequences: SEQ ID NO: 1605-1979 in Table 17, or SEQ ID NO: 1605-1806, 1809-1816, 1821-1970, 1972-1974 in Table 17, or SEQ ID NO: 1605-1826 in Table 17, or SEQ ID NO: 1605-1806, 1809-1816 and 1821-1826 in Table 17, or SEQ ID NO: 1605-1820 in Table 17A, or SEQ ID NO: 1605-1806 and 1809-1816 in Table 17A, or SEQ ID NO: 1827-1979 in Table 17C, or SEQ ID NO: 1827-1979 in Table 17C. NO: 1827-1976, or a substantially identical or similar sequence to produce an amplicon sequence of less than about 500, less than about 475, less than about 450, less than about 400, less than about 375, less than about 350, less than about 300, less than about 275, less than about 250, less than about 200, less than about 175, less than about 150, or less than about 100 nucleotides in length, or consisting essentially of SEQ ID NO: 1605-1979 in Table 17, or SEQ ID NO: 1605-1806, 1809-1816, 1821-1970, 1972-1974, and 1977-1979 in Table 17, or SEQ ID NO: 1605-1826 in Table 17, or SEQ ID NO: 1605-1979 in Table 17, or SEQ ID NO: 1605-1826 in Table 17 NOs: 1605-1806, 1809-1816 and 1821-1826, or SEQ ID NOs: 1605-1820 in Table 17A, or SEQ ID NOs: 1605-1806 and 1809-1816 in Table 17A, or SEQ ID NOs: 1827-1979 in Table 17C, or SEQ ID NOs: 1827-1976 in Table 17C, or substantially identical or similar sequences, and optionally an amplicon sequence consisting of a nucleotide sequence containing nucleic acid primer sequences at the 5' and 3' ends of the sequence.In some embodiments, the microorganism-specific nucleic acid primer pairs in the combination comprise one or more primer pairs having or consisting essentially of a nucleotide sequence or sequence pair selected from the group consisting of: Table 16; or SEQ ID NOs: 49-520 of Table 16; or SEQ ID NOs: 49-452, 457-472, and 481-520 of Table 16; or SEQ ID NOs: 49-492 of Table 16; or SEQ ID NOs: 49-452, 457-472, and 481-492 of Table 16; or SEQ ID NOs: 49-480 of Table 16A; or SEQ ID NOs: 49-452 and 457-472 of Table 16A; or SEQ ID NOs: 521-826 of Table 16C; or SEQ ID NOs: 521-820 of Table 16C; or SEQ ID NOs: 521-820 of Table 16C; or SEQ ID NOs:1299-1598 of Table 16F; or one or more sequences substantially identical or similar thereto, or the nucleotide sequence of any of the foregoing nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases.In some embodiments, the microorganism-specific nucleic acid primer pairs in the combination include at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 125, at least 150, at least 175, at least 200, at least 225, at least 230, or all nucleic acid primer pairs, wherein the nucleic acid primer pairs have a sequence selected from the following or consist essentially of a sequence selected from the following: Table 16; or SEQ ID NO: 49-520 of Table 16; or SEQ ID NO: 49-520 of Table 16; NO:49-452, 457-472 and 481-520; or SEQ ID NO:49-492 of Table 16; or SEQ ID NO:49-452, 457-472 and 481-492 of Table 16; or SEQ ID NO:49-480 of Table 16A; or SEQ ID NO:49-452 and 457-472 of Table 16A; or SEQ ID NO:521-826 of Table 16C; or SEQ ID NO:521-820 of Table 16C; or SEQ ID NO:827-1298 of Table 16; or SEQ ID NO:827-1230, 1235-1250 and 1259-1298 of Table 16; or SEQ ID NO:827-1270 of Table 16; or SEQ ID NO:827-1270 of Table 16 ID NO: 827-1230, 1235-1250 and 1259-1270; or SEQ ID NO: 827-1258 of Table 16D; or SEQ ID NO: 827-1230 and 1235-1250 of Table 16D; or SEQ ID NO: 1299-1604 of Table 16F; or SEQ ID NO: 1299-1598 of Table 16F; or one or more sequences substantially identical or similar thereto, or the nucleotide sequence of any of the above nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases. In some embodiments, the combination comprises one or more microorganism-specific nucleic acid primer pairs that amplify specific nucleic acid sequences unique to one or more microorganisms (e.g., bacteria) associated with one or more conditions, disorders, and / or diseases.
[0101] In any embodiment of the composition of the combination of one or more or multiple nucleic acids, primers or nucleic acid primer pairs or nucleic acids, primers or nucleic acid primer pairs described herein, one or more nucleic acids, or one or more primers or primer pairs may include modifications. In certain embodiments, modifications are modifications that promote nucleic acid manipulation, amplification, connection and / or amplification product sequencing and / or reduce or eliminate primer dimers. In certain embodiments, modifications are modifications that promote multiple nucleic acid amplification, connection and / or multiple amplification product sequencing. In certain embodiments, at least one primer in a primer pair or two primers in a primer pair contain modifications relative to a nucleic acid sequence to be amplified, which increase the susceptibility of primer pairs to cutting. For example, in certain embodiments, one or more nucleic acids or primers in a primer pair or two primers have at least one cleavable group, which is located at a) 3' end or 5' end and / or at b) approximately the central nucleotide position of a nucleic acid or primer, and wherein nucleic acid, primer or primer pair may be substantially non-complementary with other nucleic acids, primers or primer pairs in the composition. In some embodiments, the composition includes at least 50, 100, 150, 200, 250, 300, 350, 398 or more primer pairs. In some embodiments, the primer pair includes a length of about 15 nucleotides to about 40 nucleotides. In some embodiments, at least one nucleotide of one or more primers is replaced by a cleavable group. In some embodiments, the cleavable group can be a uridine nucleotide. In some embodiments, the template, one or more primers and / or amplification products include nucleotides or nucleobases that can be recognized by a specific enzyme. In some embodiments, the nucleotides or nucleobases can be bound by a specific enzyme. Optionally, the specific enzyme can also cut the template, one or more primers and / or amplification products at one or more sites. In some embodiments, such cutting can occur at a specific nucleotide in the template, one or more primers and / or amplification products. For example, the template, one or more primers and / or amplification products can include one or more nucleotides or nucleobases, including uracil, and the one or more nucleotides or nucleobases can be recognized and / or cut by enzymes such as uracil DNA glycosylase (UDG, also known as UNG) or formamidopyrimidine DNA glycosylase (Fpg). The template, one or more primers and / or the amplified product may comprise one or more nucleotides or nucleobases, including RNA-specific bases, which may be recognized and / or cut by enzymes such as RNAseH. In some embodiments, the template, one or more primers and / or the amplified product may comprise one or more abasic sites, which may be recognized and / or cut by various proofreading polymerases or apyrase treatments. In some embodiments, the template, one or more primers and / or the amplified product may comprise 7,8-dihydro-8-oxoguanine (8-oxoG) nucleobases, which may be recognized or cut by enzymes such as Fpg.In some embodiments, one or more amplified target sequences can be partially digested by FuPa reagents. In some embodiments, the primer contains a sufficient number of modified nucleotides to allow the primer to be functionally completely degraded by the cleavage treatment, but not to interfere with the specificity or functionality of the primer before such cleavage treatment, such as in the amplification reaction. In some embodiments, the primer contains at least one modified nucleotide, but no more than 75% of the nucleotides of the primer are modified. For example, the primer can contain a nucleobase containing uracil, which can be selectively cut using UNG / UDG (optionally heated and / or alkali). In some embodiments, the primer can contain a nucleotide containing uracil, which can be selectively cut using UNG and Fpg. In some embodiments, the cleavage treatment includes exposure to oxidative conditions to selectively cut dithiols, treatment with RNAseH to selectively cut modified nucleotides, including RNA-specific parts (e.g., ribose, etc.), etc. The cleavage treatment can effectively split the original amplification primer and non-specific amplification product into small nucleic acid fragments, which contain relatively few nucleotides. Such fragments are generally not able to promote additional amplification at elevated temperatures. Such fragments can also be relatively easily removed from the reaction pool by various post-amplification cleanup procedures known in the art (e.g., spin columns, NaEtOH precipitation, etc.).
[0102] In some embodiments, the compositions provided herein include samples containing various microorganisms, or nucleic acids from such samples containing various microorganisms, and one or more nucleic acids, primers and / or primer pairs of any embodiment of the compositions described herein, and optional polymerases, such as DNA polymerases. In some embodiments, the sample is a biological sample, such as, for example, an environmental sample or a sample from an animal subject (e.g., a human). The sample includes, but is not limited to, a biological fluid sample, a blood sample, a skin sample, a mucus sample, a saliva sample, a sputum sample, a sample from the oral cavity or nasal cavity of a subject, a respiratory sample, a vaginal sample, a digestive tract sample, and a fecal sample. In some embodiments, the sample is from the digestive tract of an animal, such as, for example, feces or fecal samples. In a specific embodiment, the compositions include one or more kingdoms covering nucleic acid primer pairs, which can amplify sequences in homologous genes or genomic regions that are common to multiple, most, most, substantially all or all microorganisms (e.g., bacteria) in a kingdom, but that differ between microorganisms in different kingdoms, and / or one or more microorganism-specific nucleic acid primer pairs, which amplify specific microorganisms (e.g., species, subspecies or strains of microorganisms such as bacteria) that are unique to specific nucleic acid sequences. Many embodiments of the nucleic acid primer pairs and microorganism-specific nucleic acid primer pairs that can be used in nucleic acid combinations are provided herein. For example, in some embodiments, the nucleic acid primer pairs of the nucleic acid combinations include one or more primer pairs that amplify two or more, three or more, four or more, five or more, six or more, seven or more, or eight or more regions, such as hypervariable regions, in a prokaryotic (e.g., bacterial) 16S rRNA gene alone.In some embodiments, at least one of the one or more microorganism-specific nucleic acid primer pairs is capable of amplifying or specifically amplifying a specific nucleic acid sequence comprising a nucleotide sequence selected from the group consisting of: SEQ ID NOs: 1605-1979 in Table 17; or SEQ ID NOs: 1605-1806, 1809-1816, 1821-1970, 1972-1974, and 1977-1979 in Table 17; or SEQ ID NOs: 1605-1826 in Table 17; or SEQ ID NOs: 1605-1806, 1809-1816, and 1821-1826 in Table 17; or SEQ ID NOs: 1605-1820 in Table 17A; or SEQ ID NOs: 1605-1806 and 1809-1816 in Table 17A; or SEQ ID NOs: 1605-1806 and 1809-1816 in Table 17C or SEQ ID NOs: 1827-1979 in Table 17C, or substantially identical or similar sequences to produce an amplicon sequence of less than about 500, less than about 475, less than about 450, less than about 400, less than about 375, less than about 350, less than about 300, less than about 275, less than about 250, less than about 200, less than about 175, less than about 150, or less than about 100 nucleotides in length, or consisting essentially of SEQ ID NOs: 1605-1979 in Table 17, or SEQ ID NOs: 1605-1806, 1809-1816, 1821-1970, 1972-1974, and 1977-1979 in Table 17, or SEQ ID NOs: 1605-1826 in Table 17, or SEQ ID NOs: ID NOs: 1605-1806, 1809-1816 and 1821-1826, or SEQ ID NOs: 1605-1820 in Table 17A, or SEQ ID NOs: 1605-1806 and 1809-1816 in Table 17A, or SEQ ID NOs: 1827-1979 in Table 17C, or SEQ ID NOs: 1827-1976 in Table 17C, or substantially identical or similar sequences and optionally an amplicon sequence containing nucleic acid primer sequences at the 5' and 3' ends of the sequence.In some embodiments, the one or more microorganism-specific nucleic acid primer pairs are a plurality of such primer pairs, wherein each of the primer pairs is capable of amplifying or specifically amplifying a specific nucleic acid sequence, the specific nucleic acid sequence comprising a nucleotide sequence selected from the group consisting of: SEQ ID NOs: 1605-1979 in Table 17, or SEQ ID NOs: 1605-1806, 1809-1816, 1821-1970, 1972-1974, and 1977-1979 in Table 17, or SEQ ID NOs: 1605-1826 in Table 17, or SEQ ID NOs: 1605-1806, 1809-1816, and 1821-1826 in Table 17, or SEQ ID NOs: 1605-1820 in Table 17A, or SEQ ID NOs: 1605-1806 and 1809-1816 in Table 17A, or SEQ ID NOs: 1605-1820 in Table 17A, or SEQ ID NOs: 1605-1806 and 1809-1816 in Table 17C. NO: 1827-1979 or SEQ ID NO: 1827-1976 in Table 17C.In some embodiments, the one or more microorganism-specific nucleic acid primer pairs are a plurality of such primer pairs, wherein each of the primer pairs is capable of amplifying or specifically amplifying a specific nucleic acid sequence, the specific nucleic acid sequence comprising a nucleotide sequence selected from the group consisting of: SEQ ID NOs: 1605-1979 in Table 17; or SEQ ID NOs: 1605-1806, 1809-1816, 1821-1970, 1972-1974, and 1977-1979 in Table 17; or SEQ ID NOs: 1605-1826 in Table 17; or SEQ ID NOs: 1605-1806, 1809-1816, and 1821-1826 in Table 17; or SEQ ID NOs: 1605-1820 in Table 17A; or SEQ ID NOs: 1605-1806 and 1809-1816 in Table 17A; or SEQ ID NOs: 1605-1806 and 1809-1816 in Table 17C. or SEQ ID NOs: 1827-1979 in Table 17C, or substantially identical or similar sequences to produce an amplicon sequence of less than about 500, less than about 475, less than about 450, less than about 400, less than about 375, less than about 350, less than about 300, less than about 275, less than about 250, less than about 200, less than about 175, less than about 150, or less than about 100 nucleotides in length, or consisting essentially of SEQ ID NOs: 1605-1979 in Table 17, or SEQ ID NOs: 1605-1806, 1809-1816, 1821-1970, 1972-1974, and 1977-1979 in Table 17, or SEQ ID NOs: 1605-1826 in Table 17, or SEQ ID NOs: ID NOs: 1605-1806, 1809-1816 and 1821-1826, or SEQ ID NOs: 1605-1820 in Table 17A, or SEQ ID NOs: 1605-1806 and 1809-1816 in Table 17A, or SEQ ID NOs: 1827-1979 in Table 17C, or SEQ ID NOs: 1827-1976 in Table 17C, or substantially identical or similar sequences and optionally an amplicon sequence consisting of a nucleotide sequence containing nucleic acid primer sequences at the 5' and 3' ends of the sequence.
[0103] Also provided herein are compositions containing a mixture of nucleic acids, wherein most or substantially all of the nucleic acids contain a sequence of a portion of a microorganism (e.g., bacteria) genome. In some embodiments, the mixture of nucleic acids comprises nucleic acids containing a sequence of a portion of at least 2, at least 5, at least 10, at least 20, at least 25, at least 30, at least 35, at least 40, at least 50, at least 75, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400 or at least 500 or more different microorganisms (e.g., different kinds of microorganisms, such as bacteria). In some embodiments, the length of the sequence of the microbial genome portion is less than about 1000 nucleotides, less than about 900 nucleotides, less than about 1000 nucleotides, less than about 900 nucleotides, less than about 800 nucleotides, less than about 700 nucleotides, less than about 600 nucleotides, less than about 500 nucleotides, less than about 450 nucleotides, less than about 400 nucleotides, less than about 350 nucleotides, less than about 300 nucleotides, less than about 250 nucleotides or less than about 200 nucleotides. In some embodiments, the length of the sequence of the microbial genome portion is less than 250 nucleotides or about 250 nucleotides. In some embodiments, the nucleic acid comprises a double-stranded, partially double-stranded and / or single-stranded nucleic acid. In some embodiments, the nucleic acid is included in an amplicon produced in a nucleic acid amplification reaction of a nucleic acid from one or more or more microorganisms (such as a variety of different microorganisms, for example, bacteria). In some embodiments, the nucleic acid comprises a nucleotide containing a uracil nucleobase. In some embodiments, the nucleic acid contains a 5' and / or 3' overhang. In some embodiments, the composition contains one or more or more primers, such as the nucleic acids and / or primer pairs of any embodiment described herein. In some embodiments, the composition comprises a DNA polymerase, a DNA ligase, and / or at least one uracil cleavage or modification enzyme. In some embodiments, the nucleic acid comprises one or more of the following:
[0104] (1) one or more nucleic acids comprising or essentially consisting of a nucleotide sequence of a hypervariable region (e.g., V1, V2, V3, V4, V5, V6, V7, V8 and / or V9 region) of a prokaryotic 16S rRNA gene,
[0105] (2) a plurality of nucleic acids comprising or consisting essentially of a nucleotide sequence of a hypervariable region (e.g., V1, V2, V3, V4, V5, V6, V7, V8 and / or V9 region) of a prokaryotic 16S rRNA gene,
[0106] (3) one or more nucleic acids comprising or consisting essentially of a nucleotide sequence of a hypervariable region (e.g., V1, V2, V3, V4, V5, V6, V7, V8 and / or V9 region) of a prokaryotic 16S rRNA gene, wherein the sequence has sequence from only one hypervariable region,
[0107] (4) one or more nucleic acids comprising or consisting essentially of a nucleotide sequence of a hypervariable region (e.g., V1, V2, V3, V4, V5, V6, V7, V8 and / or V9 region) of a prokaryotic 16S rRNA gene, wherein the sequence comprises one or more sequences selected from the group consisting of SEQ ID NOs: 1-24 in Table 15 and / or SEQ ID NOs: 25-48 in Table 15 or SEQ ID NOs: 11-16, 23 and 24 in Table 15 and / or SEQ ID NOs: 35-40, 47 and 48 in Table 15, and / or
[0108] (5) one or more single-stranded nucleic acids containing SEQ ID NOs: 1605-1979 in Table 17, or SEQ ID NOs: 1605-1806, 1809-1816, 1821-1970, 1972-1974 and 1977-1979 in Table 17, or SEQ ID NOs: 1605-1826 in Table 17, or SEQ ID NOs: 1605-1806, 1809-1816 and 1821-1826 in Table 17, or SEQ ID NOs: 1605-1820 in Table 17A, or SEQ ID NOs: 1605-1806 and 1809-1816 in Table 17A, or SEQ ID NOs: 1827-1979 in Table 17C, or SEQ ID NOs: 1827-1979 in Table 17C NO: 1827-1976 or a substantially identical or similar sequence and optionally containing a nucleotide sequence of one or more primer sequences at the 3' end and / or the 5' end (e.g., selected from Table 16, or SEQ ID NO: 49-520 of Table 16, or SEQ ID NO: 49-452, 457-472 and 481-520 of Table 16, or SEQ ID NO: 49-492 of Table 16, or SEQ ID NO: 49-452, 457-472 and 481-492 of Table 16, or SEQ ID NO: 49-480 of Table 16A, or SEQ ID NO: 49-452 and 457-472 of Table 16A, or SEQ ID NO: 521-826 of Table 16C, or SEQ ID NO: 521-820 of Table 16C, or SEQ ID NO: 827-1298 of Table 16, or SEQ ID NO: 827-1299 of Table 16, or SEQ ID NO: 827-1300 of Table 16 NO: 827-1230, 1235-1250 and 1259-1298 of Table 16, or SEQ ID NO: 827-1270 of Table 16, or SEQ ID NO: 827-1230, 1235-1250 and 1259-1270 of Table 16, or SEQ ID NO: 827-1258 of Table 16D, or SEQ ID NO: 827-1230 and 1235-1250 of Table 16D, or SEQ ID NO: 1299-1604 of Table 16F, or SEQ ID NO: 1299-1598 of Table 16F) or the complement thereof, or consisting essentially of said nucleotide sequence or the complement thereof, and / or one or more double-stranded or partially double-stranded nucleic acids,It contains SEQ ID NOs: 1605-1979 in Table 17, or SEQ ID NOs: 1605-1806, 1809-1816, 1821-1970, 1972-1974 and 1977-1979 in Table 17, or SEQ ID NOs: 1605-1826 in Table 17, or SEQ ID NOs: 1605-1806, 1809-1816 and 1821-1826 in Table 17, or SEQ ID NOs: 1605-1820 in Table 17A, or SEQ ID NOs: 1605-1806 and 1809-1816 in Table 17A, or SEQ ID NOs: 1827-1979 in Table 17C, or SEQ ID NOs: 1827-1979 in Table 17C NO: 1827-1976 or a substantially identical or similar sequence and optionally containing one or more primer sequences at the 3' end and / or the 5' end (e.g., a nucleotide sequence selected from Table 16, or SEQ ID NO: 49-520 of Table 16, or SEQ ID NO: 49-452, 457-472 and 481-520 of Table 16, or SEQ ID NO: 49-492 of Table 16, or SEQ ID NO: 49-452, 457-472 and 481-492 of Table 16, or SEQ ID NO: 49-480 of Table 16A, or SEQ ID NO: 49-452 and 457-472 of Table 16A, or SEQ ID NO: 521-826 of Table 16C, or SEQ ID NO: 521-820 of Table 16C, or SEQ ID NO: 827-1298 of Table 16, or SEQ ID NO: 827-1299 of Table 16, or SEQ ID NO: 827-1300 of Table 16 NO:827-1230, 1235-1250 and 1259-1298 of Table 16, or SEQ ID NO:827-1270 of Table 16, or SEQ ID NO:827-1230, 1235-1250 and 1259-1270 of Table 16, or SEQ ID NO:827-1258 of Table 16D, or SEQ ID NO:827-1230 and 1235-1250 of Table 16D, or SEQ ID NO:1299-1604 of Table 16F, or SEQ ID NO:1299-1598 of Table 16F) and complementary nucleotide sequences that hybridize therewith, or consist essentially of said nucleotide sequences and complementary nucleotide sequences that hybridize therewith.
[0109] In some embodiments, the nucleic acid comprises any combination of the nucleic acid of (1), (2), (3) or (4) above and the nucleic acid of (5) above. In some embodiments of the composition containing a mixture of nucleic acids, wherein most or substantially all of the nucleic acids contain a sequence of a portion of the genome of a microorganism (e.g., bacteria) provided herein, the composition is or contains one or more microorganism (e.g., bacteria) nucleic acid libraries. In some embodiments, the mixture of nucleic acids is produced by amplifying nucleic acids in or from a sample containing a microorganism (e.g., bacteria) using primers and / or primer pairs provided herein. In some embodiments, the mixture of nucleic acids can be produced by amplifying nucleic acids using: (1) one or more kingdom-encompassing nucleic acid primer pairs that are capable of amplifying sequences in homologous genes or genomic regions that are common to multiple, most, most, substantially all or all microorganisms (e.g., bacteria) in a kingdom but differ between microorganisms of different kingdoms; and / or (2) one or more microorganism-specific nucleic acids and / or nucleic acid primer pairs that are capable of amplifying or specifically or selectively amplifying specific nucleic acid sequences that are unique to a particular microorganism (e.g., a species, subspecies or strain of a microorganism such as bacteria). Provided herein are many embodiments of nucleic acid primer pairs and microorganism-specific nucleic acid primer pairs that can be used to produce nucleic acid combinations. In some embodiments, a mixture is produced by using nucleic acid primer pairs and microorganism-specific nucleic acid primers and / or nucleic acid primer pairs to amplify microorganism nucleic acids in a single reaction mixture. In some embodiments, a mixture is produced by, for example, using nucleic acid primer pairs in a single sample to amplify microorganism nucleic acids in a separate amplification reaction and then combining the products of two amplification reactions. In some embodiments, the nucleic acid mixture includes or is substantially composed of: a portion of a prokaryotic 16S rRNA gene, such as a nucleotide sequence of a hypervariable region (e.g., V1, V2, V3, V4, V5, V6, V7, V8, and / or V9 region) of a prokaryotic (e.g., bacterium) 16S rRNA gene from one or more or more microorganisms; and a portion of a microorganism (e.g., bacterium) genome from one or more or more microorganisms that is not included in the prokaryotic 16S rRNA gene.
[0110] Methods for nucleic acid amplification
[0111] The method provided herein includes a method for amplifying and / or detecting nucleic acid. In a specific embodiment, the nucleic acid amplified and / or detected is from a microorganism, including, for example, bacteria and archaea. As further described herein, the method provided herein for amplifying and / or detecting nucleic acids from microorganisms represents a significant improvement relative to previous methods, including but not limited to microbial nucleic acid amplification and / or detection coverage, sensitivity, efficiency, scale, cost-effectiveness and / or improvements in the application or use of other methods. In some embodiments, nucleic acid hybridization and / or amplification are performed on nucleic acid using, for example, any nucleic acid provided herein as a probe and / or amplification primer. In some embodiments, the presence or absence of one or more hybridization and / or nucleic acid amplification products is detected. In some embodiments, the nucleic acid amplification is multiplex amplification. In some embodiments, the amplification is performed using multiple nucleic acid primers and is performed in a single multiplex amplification reaction mixture. In some embodiments, one or more nucleic acids provided herein are used as probes (e.g., detectable or labeled probes) to detect the presence or absence of one or more nucleic acids and / or amplification products. In some embodiments, the presence or absence of one or more nucleic acid amplification products is detected by obtaining the nucleotide sequence information of one or more nucleic acid amplification products.
[0112] Method for amplifying nucleic acid of selected microorganisms
[0113] In some embodiments, the methods provided herein for amplifying target nucleic acids of one or more microorganisms include (a) obtaining nucleic acids of one or more microorganisms, wherein the one or more microorganisms are selected from the microorganisms listed in Table 1 (or Table 1, excluding or excluding Actinomycetes viscosus and / or Blautia sphericalus; or Table 1, excluding or excluding Actinomycetes viscosus, Blautia sphericalus and / or Helicobacter zabbii); and (b) performing nucleic acid amplification on the nucleic acid using at least one primer pair capable of specifically amplifying a target nucleic acid sequence contained in the genome of a microorganism, wherein the microorganism is selected from the microorganisms of Table 1 (or Table 1, excluding or excluding Actinomycetes viscosus and / or Blautia sphericalus; or Table 1, excluding or excluding Actinomycetes viscosus, Blautia sphericalus and / or Helicobacter zabbii), thereby producing an amplified copy of the target nucleic acid. In some embodiments, the target nucleic acid is unique to the microorganism. In some embodiments, the target nucleic acid is not contained in the prokaryotic 16S rRNA gene. In some embodiments, the nucleic acid to be amplified includes nucleic acids from a plurality of different microorganisms listed in Table 1. In some such embodiments, amplified copies of a plurality of different microorganisms in Table 1 (or Table 1, excluding or excluding Actinomyces viscosus and / or Blautia sphaeroides; or Table 1, excluding or excluding Actinomyces viscosus, Blautia sphaeroides and / or Helicobacter zabbii) are generated, for example, in multiple nucleic acid amplifications. In some embodiments, the nucleic acid to be amplified comprises a mixture of nucleic acids of one or more microorganisms selected from the microorganisms listed in Table 1 (or Table 1, excluding or excluding Actinomyces viscosus and / or Blautia sphaeroides; or Table 1, excluding or excluding Actinomyces viscosus, Blautia sphaeroides and / or Helicobacter zabbii) and one or more microorganisms not listed in Table 1 (e.g., bacteria). In some embodiments, nucleic acids of one or more microorganisms selected from the microorganisms listed in Table 1 are obtained from a biological sample, such as, for example, a sample of the contents of an animal's digestive tract. In some embodiments, the sample is a fecal sample.In some embodiments, at least one, or one or more target nucleic acid sequences comprises or consists essentially of a nucleotide sequence selected from the group consisting of SEQ ID NOs: 1605-1979 in Table 17, or SEQ ID NOs: 1605-1806, 1809-1816, 1821-1970, 1972-1974, and 1977-1979 in Table 17, or SEQ ID NOs: 1605-1826 in Table 17, or SEQ ID NOs: 1605-1806, 1809-1816, and 1821-1826 in Table 17, or SEQ ID NOs: 1605-1820 in Table 17A, or SEQ ID NOs: 1605-1806 and 1809-1816 in Table 17A, or SEQ ID NOs: 1827-1979 in Table 17C, or SEQ ID NOs: 1827-1979 in Table 17C. NO:1827-1976, or a substantially identical or similar sequence. In some embodiments, at least one, or one or more nucleic acid amplification products include SEQ ID NO:1605-1979 in Table 17, or SEQ ID NO:1605-1806, 1809-1816, 1821-1970, 1972-1974 and 1977-1979 in Table 17, or SEQ ID NO:1605-1826 in Table 17, or SEQ ID NO:1605-1806, 1809-1816 and 1821-1826 in Table 17, or SEQ ID NO:1605-1820 in Table 17A, or SEQ ID NO:1605-1806 and 1809-1816 in Table 17A, or SEQ ID NO:1827-1979 in Table 17C, or SEQ ID NO:1827-1979 in Table 17C NO:1827-1976 or substantially identical or similar sequence or its complement and optionally at 5 ' and / or 3 ' end of sequence there is the nucleotide sequence of one or more primer sequences (any primer sequence provided as this paper), or is basically made up of described nucleotide sequence.In certain embodiments, at least one primer pair can not detectably increase the nucleotide sequence that is included in any genus except the microorganism genus that contains the target nucleic acid sequence.In certain embodiments, at least one primer pair can not detectably increase the nucleotide sequence that is included in any kind except the microorganism species that contains the target nucleic acid sequence.In some embodiments, at least one primer in the primer pair, or at least one primer pair, contains or consists essentially of one or more sequences of the primers or primer pairs in Table 16; or SEQ ID NOs: 49-520 of Table 16; or SEQ ID NOs: 49-452, 457-472, and 481-520 of Table 16; or SEQ ID NOs: 49-492 of Table 16; or SEQ ID NOs: 49-452, 457-472, and 481-492 of Table 16; or SEQ ID NOs: 49-480 of Table 16A; or SEQ ID NOs: 49-452 and 457-472 of Table 16A; or SEQ ID NOs: 521-826 of Table 16C; or SEQ ID NOs: 521-820 of Table 16C; or SEQ ID NOs: 827-1298 of Table 16; or SEQ ID NOs: 527-1299 of Table 16. or SEQ ID NOs: 1299-1598 of Table 16F, or substantially identical or similar sequences, or the nucleotide sequence of any of the foregoing nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases.In some embodiments, nucleic acid amplification is performed on a nucleic acid using a plurality of primers or primer pairs, each of which contains or consists essentially of one or more sequences of a primer pair in Table 16; or SEQ ID NOs: 49-520 of Table 16; or SEQ ID NOs: 49-452, 457-472, and 481-520 of Table 16; or SEQ ID NOs: 49-492 of Table 16; or SEQ ID NOs: 49-452, 457-472, and 481-492 of Table 16; or SEQ ID NOs: 49-480 of Table 16A; or SEQ ID NOs: 49-452 and 457-472 of Table 16A; or SEQ ID NOs: 521-826 of Table 16C; or SEQ ID NOs: 521-820 of Table 16C; or SEQ ID NOs: 827-1298 of Table 16; or SEQ ID NOs: 827-1299 of Table 16. NO:827-1230, 1235-1250 and 1259-1298; or SEQ ID NO:827-1270 of Table 16; or SEQ ID NO:827-1230, 1235-1250 and 1259-1270 of Table 16; or SEQ ID NO:827-1258 of Table 16D; or SEQ ID NO:827-1230 and 1235-1250 of Table 16D; or SEQ ID NO:1299-1604 of Table 16F; or SEQ ID NO:1299-1598 of Table 16F; or substantially the same or similar sequence, or the nucleotide sequence of any of the above nucleic acids or primer pairs, wherein one or more thymine bases are replaced with uracil bases. In some embodiments where nucleic acid amplification is performed on a nucleic acid using more than one or more primers or primer pairs, the amplification is a multiplex amplification performed in a single reaction mixture. In some embodiments, at least one primer or a primer pair comprises a modification that facilitates nucleic acid manipulation, amplification, ligation and / or sequencing of amplification products and / or reduces or eliminates primer dimers. In a specific embodiment, the modification is a modification that facilitates multiplex nucleic acid amplification, ligation and / or sequencing of multiplex amplification products.
[0114] Methods for multiplex amplification of multiple regions of a gene
[0115] In some embodiments, the multiple amplification method provided herein is used to amplify multiple regions of the gene of one or more microorganisms (e.g., bacteria). In one embodiment, the method comprises (a) obtaining nucleic acid of one or more microorganisms including 16S rRNA gene and (b) using a primer pair combination comprising at least two primer pairs to carry out nucleic acid amplification to nucleic acid, wherein the at least two primer pairs individually amplify nucleic acid containing different hypervariable region sequences of prokaryotic 16S rRNA gene, thereby generating an amplified copy of the nucleic acid sequence containing different hypervariable region sequences of 16S rRNA gene of one or more microorganisms. In some embodiments, microorganism is a bacterium. In some embodiments, prokaryotic 16S rRNA gene is a bacterial gene. In some embodiments, the nucleic acid amplified comprises nucleic acid from a variety of different microorganisms. In some embodiments, nucleic acid of one or more microorganisms including 16SrRNA gene is obtained from a biological sample, such as a sample of, for example, animal digestive tract contents. In some embodiments, the sample is a fecal sample. In some embodiments, the primers in the primer pair combination are directed to the nucleic acid sequence contained in the conserved region of 16SrRNA gene, or are combined or hybridized with the nucleic acid sequence. In some embodiments, each primer in the primer pair combination contains less than 10, less than 9, less than 8, less than 7, less than 6, less than 5, less than 4, less than 3 or less than 2 consecutive nucleotides, and its sequence is the same as the consecutive nucleotide sequence of another primer in the primer pair combination. In some embodiments, the length of the amplified nucleic acid sequence is less than about 300bp, less than about 250bp, less than about 200bp, less than about 175bp, less than about 150bp or less than about 125bp. In some embodiments, the combination of primer pairs individually amplifies nucleic acids containing sequences of 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more or 9 different hypervariable regions of the prokaryotic 16S rRNA gene, thereby generating amplified copies of nucleic acids containing sequences of 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more or 9 different hypervariable regions of the 16S rRNA gene of one or more microorganisms, wherein the amplified copies of the different hypervariable regions are separate amplicons. In some embodiments, the combination of primer pairs amplifies 8 different nucleic acids individually, and the nucleic acids contain sequences of 8 different hypervariable regions of prokaryotic 16SrRNA genes, respectively. In some embodiments, the 8 different hypervariable regions are V2-V9. In some embodiments, the combination of primer pairs amplifies at least 3 different nucleic acids individually, and each nucleic acid contains sequences of different hypervariable regions of prokaryotic 16S rRNA genes, respectively, wherein one of the 3 or more regions is the V5 region, thereby generating amplified copies of nucleic acids containing sequences of 3 or more hypervariable regions of 16S rRNA genes of one or more microorganisms, respectively.In some embodiments, the combination of primer pairs comprises a degenerate sequence of one or more primers in one or more primer pairs. For example, in some embodiments, for at least one hypervariable region amplified by a primer pair combination, at least two different primer pairs in the primer pair combination individually amplify nucleic acid sequences in the same hypervariable region of 2 or more species of the same prokaryotic genus or 2 or more strains of the same prokaryotic genus, and the prokaryotic genus has differences in the nucleic acid sequences in the same hypervariable region. In some such cases, at least two different primer pairs in the primer pair combination individually amplify nucleic acid sequences in the V2 hypervariable region of 2 or more species of the same prokaryotic genus or 2 or more strains of the same prokaryotic species, and there are differences in the nucleic acid sequences in the V2 hypervariable region, and / or at least two different primer pairs in the primer pair combination individually amplify nucleic acid sequences in the V8 hypervariable region of 2 or more species of the same prokaryotic genus or 2 or more strains of the same prokaryotic species, and the prokaryotic genus has differences in the nucleic acid sequences in the V8 hypervariable region. In some embodiments, the primer pair combination for amplifying a nucleic acid containing a hypervariable region sequence of a prokaryotic 16S rRNA gene includes primers and / or primer pairs, which contain or are substantially composed of the following: one or more sequences of primers or primer pairs in Table 15; or SEQ ID NO: 1-24 in Table 15 and / or SEQ ID NO: 25-48 in Table 15; or SEQ ID NO: 11-16, 23 and 24 in Table 15 and / or SEQ ID NO: 35-40, 47 and 48 in Table 15, or substantially the same or similar sequences, and optionally wherein one or more thymine bases are replaced by uracil bases. In some embodiments, amplification is a multiple amplification performed in a single reaction mixture. In some embodiments, at least one primer or primer pair in the combination comprises a modification that promotes nucleic acid manipulation, amplification, connection and / or sequencing of amplified products and / or reduces or eliminates primer dimers. In a specific embodiment, the modification is a modification that promotes multiple nucleic acid amplification, connection and / or sequencing of multiple amplified products.
[0116] Methods for amplifying multiple regions of the genome of a microorganism
[0117] In some embodiments, a method for amplifying multiple regions of the genome of one or more microorganisms is provided. In some embodiments, the method comprises (a) obtaining nucleic acid of one or more microorganisms including a 16S rRNA gene and (b) performing nucleic acid amplification on the nucleic acid using a primer pair combination, wherein the primer pair combination comprises: (i) one or more primer pairs for amplifying nucleic acid containing a hypervariable region sequence of a prokaryotic 16S rRNA gene (referred to as "16S rRNA gene primers and primer pairs"); and (ii) one or more primer pairs for amplifying a target nucleic acid sequence contained in the genome of a microorganism, wherein the target nucleic acid sequence is not contained in the hypervariable region of the prokaryotic 16S rRNA gene, wherein different primer pairs amplify different target nucleic acid sequences contained in different genomes (referred to as "non-16S rRNA gene primers and primer pairs"), thereby generating amplified copies of at least two different regions of the genome of one or more microorganisms. In some embodiments, the microorganism is a bacterium. In some embodiments, the prokaryotic 16S rRNA gene is a bacterial gene and / or the prokaryotic microorganism is a bacterium. In some embodiments, one or more primer pairs for amplifying nucleic acid containing a hypervariable region sequence of a prokaryotic 16S rRNA gene amplify nucleic acid sequences of different hypervariable regions separately. In some embodiments, the primers in the one or more primer pairs of (i) are directed against, bind to, or hybridize to a nucleic acid sequence contained in a conserved region of a prokaryotic 16S rRNA gene. In some embodiments, the amplification is a multiplex amplification performed in a single reaction mixture. In some embodiments, the amplification method for amplifying multiple regions of the genome of one or more microorganisms comprises (a) obtaining nucleic acid of one or more microorganisms including a 16S rRNA gene and (b) performing two or more separate nucleic acid amplification reactions on the nucleic acid using a first primer pair set for one nucleic acid amplification reaction and a second primer pair set for another nucleic acid amplification reaction, wherein (i) the first primer pair set includes one or more primer pairs for amplifying nucleic acid containing a hypervariable region sequence of a prokaryotic 16S rRNA gene (referred to as "16S rRNA gene primers and primer pairs"), and (ii) the second primer pair set includes one or more primer pairs for amplifying a target nucleic acid sequence contained in the genome of the microorganism, the target nucleic acid sequence is not contained in the hypervariable region of the prokaryotic 16S rRNA gene, wherein different primer pairs amplify different target nucleic acid sequences contained in different genomes (referred to as "non-16S rRNA gene primers and primer pairs"), thereby generating amplified copies of at least two different regions of the genome of one or more microorganisms. In some embodiments, the microorganism is a bacterium. In some embodiments, the prokaryotic 16S rRNA gene is a bacterial gene and / or the prokaryotic microorganism is a bacterium. In some embodiments, one or more primer pairs that amplify nucleic acid containing a hypervariable region sequence of a prokaryotic 16S rRNA gene individually amplify nucleic acid sequences of different hypervariable regions.In some embodiments, the primers in the one or more primer pairs of (i) are directed against, bind to, or hybridize to a nucleic acid sequence contained in a conserved region of a prokaryotic 16S rRNA gene. In some embodiments, the amplification is a multiplex amplification performed in a single reaction mixture.
[0118] In some embodiments, in the amplification method for amplifying multiple regions of the genome of one or more microorganisms, the target nucleic acid sequence contained in the genome of a prokaryotic microorganism (e.g., bacteria) is unique to the microorganism. In some embodiments, one or more 16S rRNA gene primers amplify nucleic acid sequences from multiple microorganisms (e.g., bacteria) of different genera. In some embodiments, a nucleic acid mixture of at least two different microorganisms (e.g., bacteria) is obtained and nucleic acid amplification is performed, and only the genome of one microorganism contains a target sequence specifically amplified by a non-16S rRNA gene primer pair. In some such embodiments, the amplified copy produced contains a copy of the target nucleic acid sequence amplified from the nucleic acid of the genome of a microorganism by a non-16S rRNA gene primer pair, but does not contain a copy of the target nucleic acid sequence amplified from the nucleic acid of the genome of any other microorganism subjected to nucleic acid amplification by a non-16S rRNA gene primer pair. Also in some such embodiments, the amplified copy produced contains a copy of the nucleic acid sequence of the hypervariable region amplified from the nucleic acid of the genome of multiple microorganisms by a 16S rRNA gene primer pair. In some embodiments, the nucleic acid for nucleic acid amplification comprises nucleic acids from multiple different microorganisms. In some embodiments, nucleic acid of one or more microorganisms (e.g., bacteria) including 16S rRNA genes is obtained from a biological sample, such as a sample of, for example, an animal digestive tract content. In some embodiments, the sample is a fecal sample. In some embodiments, each primer in one or more 16S rRNA gene primer pairs contains less than 10, less than 9, less than 8, less than 7, less than 6, less than 5, less than 4, less than 3 or less than 2 continuous nucleotides, and its sequence is identical to the continuous nucleotide sequence of another primer in the primer pair combination. In some embodiments, the length of the nucleic acid sequence amplified by one or more 16S rRNA gene primer pairs is less than about 300bp, less than about 250bp, less than about 200bp, less than about 175bp, less than about 150bp or less than about 125bp. In some embodiments, the 16S rRNA gene primer pairs individually amplify nucleic acids containing different hypervariable regions in 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more or 9 different hypervariable regions of prokaryotic 16S rRNA genes, thereby producing amplified copies of nucleic acids containing the sequence of one of 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more or 9 different hypervariable regions of the 16S rRNA genes of one or more microorganisms, wherein the amplified copies of the different hypervariable regions are separate amplicons. In some embodiments, the 16S rRNA gene primer pairs individually amplify nucleic acids containing the sequences of 8 different hypervariable regions of prokaryotic 16S rRNA genes, respectively. In some embodiments, the 8 different hypervariable regions are V2-V9.In some embodiments, the 16S rRNA gene primer pairs individually amplify nucleic acids containing sequences of three or more different hypervariable regions of prokaryotic 16S rRNA genes, wherein one of the three or more regions is the V5 region, thereby generating amplified copies of nucleic acids containing sequences of three or more different hypervariable regions of 16S rRNA genes of one or more microorganisms. In some embodiments, the combination of primer pairs comprises a degenerate sequence of one or more primers in one or more primer pairs. In some embodiments, the 16S rRNA gene primer pair comprises a primer and / or a primer pair containing or consisting essentially of one or more sequences of a primer or primer pair in Table 15; or SEQ ID NOs: 1-24 in Table 15 and / or SEQ ID NOs: 25-48 in Table 15; or SEQ ID NOs: 11-16, 23 and 24 in Table 15 and / or SEQ ID NOs: 35-40, 47 and 48 in Table 15, or substantially identical or similar sequences, and optionally wherein one or more thymine bases are replaced with uracil bases. In some embodiments, the at least one non-16S rRNA gene primer pair specifically amplifies a target nucleic acid sequence contained within the genome of a microorganism selected from Table 1 (or Table 1, excluding or excluding Actinomyces viscosus and / or Blautia sphaeroides; or Table 1, excluding or excluding Actinomyces viscosus, Blautia sphaeroides and / or Helicobacter zascherii) microorganism. In some embodiments, the target nucleic acid is unique to the microorganism. In some embodiments, the nucleic acid to be amplified comprises nucleic acids from a plurality of different microorganisms listed in Table 1. In some such embodiments, amplified copies of a plurality of different microorganisms in Table 1 (or Table 1, excluding or excluding Actinomyces viscosus and / or Blautia sphaeroides; or Table 1, excluding or excluding Actinomyces viscosus, Blautia sphaeroides and / or Helicobacter zabbii) are produced. In some embodiments, the nucleic acid to be amplified comprises a mixture of nucleic acids of one or more microorganisms selected from the microorganisms listed in Table 1 (or Table 1, excluding or excluding Actinomyces viscosus and / or Blautia sphaeroides; or Table 1, excluding or excluding Actinomyces viscosus, Blautia sphaeroides and / or Helicobacter zabbii) and one or more microorganisms not listed in Table 1 (e.g., bacteria).In some embodiments, at least one, or one or more target nucleic acid sequences comprises or consists essentially of a nucleotide sequence selected from the group consisting of SEQ ID NOs: 1605-1979 in Table 17, or SEQ ID NOs: 1605-1806, 1809-1816, 1821-1970, 1972-1974, and 1977-1979 in Table 17, or SEQ ID NOs: 1605-1826 in Table 17, or SEQ ID NOs: 1605-1806, 1809-1816, and 1821-1826 in Table 17, or SEQ ID NOs: 1605-1820 in Table 17A, or SEQ ID NOs: 1605-1806 and 1809-1816 in Table 17A, or SEQ ID NOs: 1827-1979 in Table 17C, or SEQ ID NOs: 1827-1979 in Table 17C. NO:1827-1976, or a substantially identical or similar sequence. In some embodiments, at least one, or one or more nucleic acid amplification products include SEQ ID NO:1605-1979 in Table 17, or SEQ ID NO:1605-1806, 1809-1816, 1821-1970, 1972-1974 and 1977-1979 in Table 17, or SEQ ID NO:1605-1826 in Table 17, or SEQ ID NO:1605-1806, 1809-1816 and 1821-1826 in Table 17, or SEQ ID NO:1605-1820 in Table 17A, or SEQ ID NO:1605-1806 and 1809-1816 in Table 17A, or SEQ ID NO:1827-1979 in Table 17C, or SEQ ID NO:1827-1979 in Table 17C. NO:1827-1976 or substantially identical or similar sequence or its complement and optionally at the 5 ' and / or 3 ' end of sequence there is the nucleotide sequence of one or more primer sequences (any primer sequence provided herein), or is basically composed of said nucleotide sequence.In certain embodiments, at least one, or the length of one or more nucleic acid amplification products is less than about 500, less than about 475, less than about 450, less than about 400, less than about 375, less than about 350, less than about 300, less than about 275, less than about 250, less than about 200, less than about 175, less than about 150 or less than about 100 nucleotides.In certain embodiments, at least one non-16S rRNA gene primer is to can not detectably amplify the nucleotide sequence contained in any genus except the microorganism genus containing the target nucleic acid sequence.In certain embodiments, at least one non-16S rRNA gene primer is to can not detectably amplify the nucleotide sequence contained in any kind except the microorganism species containing the target nucleic acid sequence.In some embodiments, at least one primer in the non-16S rRNA gene primer pair, or at least one non-16S rRNA gene primer pair, contains or consists essentially of one or more sequences of a primer or primer pair in Table 16; or SEQ ID NOs: 49-520 of Table 16; or SEQ ID NOs: 49-452, 457-472, and 481-520 of Table 16; or SEQ ID NOs: 49-492 of Table 16; or SEQ ID NOs: 49-452, 457-472, and 481-492 of Table 16; or SEQ ID NOs: 49-480 of Table 16A; or SEQ ID NOs: 49-452 and 457-472 of Table 16A; or SEQ ID NOs: 521-826 of Table 16C; or SEQ ID NOs: 521-820 of Table 16C; or SEQ ID NOs: 521-820 of Table 16C; or SEQ ID NOs:1299-1598 of Table 16F, or substantially the same or similar sequences, or the nucleotide sequence of any of the foregoing nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases.In some embodiments, nucleic acid amplification is performed on a nucleic acid using a plurality of non-16S rRNA gene primers or primer pairs, each of which contains or consists essentially of one or more sequences of a primer pair in Table 16; or SEQ ID NOs: 49-520 of Table 16; or SEQ ID NOs: 49-452, 457-472, and 481-520 of Table 16; or SEQ ID NOs: 49-492 of Table 16; or SEQ ID NOs: 49-452, 457-472, and 481-492 of Table 16; or SEQ ID NOs: 49-480 of Table 16A; or SEQ ID NOs: 49-452 and 457-472 of Table 16A; or SEQ ID NOs: 521-826 of Table 16C; or SEQ ID NOs: 521-820 of Table 16C; or SEQ ID NOs: 521-820 of Table 16C; or SEQ ID NOs:1299-1598 of Table 16F; or substantially the same or similar sequences, or the nucleotide sequences of any of the above nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases. In some embodiments, at least one primer or one primer pair in the primer pair combination comprises a modification that promotes nucleic acid manipulation, amplification, connection and / or sequencing of amplified products and / or reduces or eliminates primer dimers. In a specific embodiment, the modification is a modification that promotes multiplex nucleic acid amplification, connection and / or sequencing of multiplex amplified products.
[0119] Procedures / techniques used for nucleic acid amplification methods
[0120] Methods for obtaining nucleic acids, for example, from samples are described herein and / or are known to those skilled in the art. Samples containing microorganisms can come from a variety of sources, including, for example, environmental sources, such as water, soil, and organism sources, such as animals, including but not limited to insects, livestock (e.g., cattle, sheep, pigs, horses, dogs, cats, etc.), mammals (e.g., humans). Common animal samples include, but are not limited to, saliva, biopsies, tumors, scrapings, swabs, blood, mucus, urine, plasma, semen, hair, laser capture microdissection, surgical resection, feces, and other clinical or laboratory-obtained samples. Fecal samples are commonly used as a source of animal digestive tract or intestinal microorganisms. Kits, protocols, and instruments for extracting nucleic acids from animal samples are available from commercial public sources, including, for example, MagMAX TM Microbiome Ultra Nucleic Acid Isolation Kit (Thermo Fisher Scientific; Catalog No. A42357 (with plates) or A42358 (with tubes)), which can be used with a Thermo Scientific TM Kingfisher TM The method of the present invention is to use a Flex magnetic particle processor (Thermo Fisher Scientific; Catalog No. 5400630). The amount of nucleic acid material required for a successful multiple amplification reaction as may be performed in an embodiment of the method provided herein may be about 1 ng. In some embodiments, the amount of nucleic acid material may be about 10 ng to about 50 ng, about 10 ng to about 100 ng, or about 1 ng to about 200 ng of nucleic acid material. Higher amounts of input material may be used, however, one aspect of the present disclosure is to selectively amplify multiple target sequences from about low (ng) starting materials.
[0121] The amplification methods provided herein generally include preparing an amplification reaction mixture containing reagents, the reagents being used to react and subjecting the mixture to conditions to achieve repeated cycles of primer annealing with template nucleic acid, primer extension, and dissociation of extended primers and template strands (e.g., denaturation). Various techniques for amplifying nucleic acids can be used for amplification methods, such as techniques based on polymerase chain reaction (PCR), helicase-dependent amplification (HDA), loop-mediated isothermal amplification (LAMP), and strand displacement amplification. In some embodiments, the method includes hybridizing one or more primers in a primer pair with a target template sequence, extending the first primer in the primer pair, denaturing the extended first primer product from a nucleic acid molecule population, hybridizing the extended first primer product with the second primer in the primer pair, extending the second primer to form a double-stranded product, and in some embodiments, digesting the target-specific primer pair away from the double-stranded product to produce a plurality of amplified target sequences. In some embodiments, digestion includes partially digesting one or more target-specific primers from the amplified target sequence. In some embodiments, the method for performing multiple PCR amplification includes: contacting a plurality of primer pairs having forward and reverse primers with a population of template nucleic acid sequences, such as in or from a sample, to form a plurality of template / primer duplexes; adding a mixture of DNA polymerase and dNTP to the plurality of template / primer duplexes, at a sufficient temperature for a sufficient time to extend the forward or reverse primer (or both) in each target-specific primer pair via template-dependent synthesis, thereby producing a plurality of extended primer products / template duplexes; denaturing the extended primer products / template duplexes; annealing complementary primers from the target-specific primer pair to the extended primer products; and extending the annealed primers in the presence of DNA polymerase and dNTPs to form a plurality of target-specific double-stranded nucleic acid molecules. In some embodiments, the steps of the amplification PCR method can be performed in any order. In some cases, the method disclosed herein can be further optimized to remove one or more steps, and still obtain sufficient amplified target sequences for various downstream processes. For example, the number of purification or clearing steps can be modified to include more or less steps than disclosed herein, provided that the amplified target sequence is produced at a sufficient yield. In some embodiments, multiplex PCR includes hybridizing one or more target-specific primer pairs to nucleic acid molecules, extending the primers in the target-specific primer pairs by template-dependent synthesis in the presence of a DNA polymerase and dNTPs; repeating the hybridization and extension steps for a sufficient time and a sufficient temperature to produce a plurality of amplified target sequences. In some embodiments, the steps of the multiplex amplification reaction method can be performed in any order. The multiplex PCR amplification reaction disclosed herein can include multiple "cycles" typically performed on a thermal cycler. Each cycle includes at least one annealing step and at least one extension step.In one embodiment, a multiplex PCR amplification reaction is performed, wherein a target-specific primer pair is hybridized to a target sequence; the hybridized primer is extended to produce an extended primer product / nucleic acid duplex; the extended primer product / nucleic acid duplex is denatured, thereby hybridizing a complementary primer to the extended primer product, wherein the complementary primer is extended to produce a plurality of amplified target sequences. In one embodiment, the method disclosed herein has about 5 to about 18 cycles per preamplification reaction. The annealing temperature and / or annealing duration of each cycle may be the same; may include an incremental increase or decrease, or a combination of both. The extension temperature and / or extension duration of each cycle may be the same; may include an incremental increase or decrease, or a combination of both. For example, the annealing temperature or the extension temperature may be kept constant per cycle. In some embodiments, the annealing temperature may be kept constant in each cycle, and the extension duration may be increased incrementally per cycle. In some embodiments, the increase or decrease in duration may occur in increments of 15 seconds, 30 seconds, 1 minute, 2 minutes, or 4 minutes. In some embodiments, the increase or decrease in temperature may occur as a deviation of 0.5, 1, 2, 3, or 4 degrees Celsius. In some embodiments, hot start PCR techniques can be used to perform the amplification reaction. These techniques include using a heating step (>60°C) before the start of polymerization to reduce the formation of undesirable PCR products. Other techniques, such as reversible inactivation or physical separation of one or more key reagents of the reaction, for example, magnesium or DNA polymerase can be isolated in wax beads, which melt when the wax beads are heated during the denaturation step, thereby releasing the reagents only at higher temperatures. DNA polymerase can also be maintained in an active state by binding to an aptamer or antibody. This binding is destroyed at higher temperatures, thereby releasing a functional DNA polymerase that can perform PCR unhindered.
[0122] In some embodiments, the amplified target sequence can be connected to one or more adapters. In some embodiments, the adapter can include one or more nucleic acid barcodes or marker sequences. In some embodiments, once the amplified target sequence is connected to the adapter, it can undergo nick translation reaction and / or further amplification to produce an amplified target sequence library connected to the adapter. In one embodiment, the amplification method involves multiple primers with cleavable groups to carry out multiple PCR on nucleic acid samples. In some embodiments, the multiple PCR amplification reaction is carried out using a variety of primers provided herein, which have cleavable groups and include DNA polymerase, adapters, dATP, dCTP, dGTP and dTTP. In some embodiments, the cleavable group can be a uracil nucleotide. In some embodiments, the forward and reverse primers contain uracil nucleotides as one or more cleavable groups. In one embodiment, the primer pair can include a uracil nucleotide in each forward and reverse primer of each primer pair. In one embodiment, the forward or reverse primer contains one, two, three or more uracil nucleotides. In some embodiments, the method involves amplifying at least 10, 50, 100, 150, 200, 250, 300, 350, 398 or more target sequences from a nucleic acid population having a plurality of target sequences using target-specific forward and reverse primer pairs containing at least two uracil nucleotides. The reaction may also contain one or more antibodies and / or nucleic acid barcodes. In some embodiments, the method comprises a process for reducing the formation of amplification artifacts in multiplex PCR. In some embodiments, primer dimers or non-specific amplification products are obtained in lower quantities or yields compared to standard multiplex PCR of the prior art. In some embodiments, the reduction of amplification artifacts is controlled in part by the use of specific primer pairs in the multiplex PCR reaction. In one embodiment, the number of specific primer pairs in the multiplex PCR reaction may be greater than 50, 100, 150, 200, 250, 300 or more. In some embodiments, multiplex PCR is performed using primers containing cleavable groups. In one embodiment, the primers containing cleavable groups may contain one or more cleavable portions of each primer in each primer pair. In some embodiments, the primer containing a cleavable group comprises a nucleotide that is neither normally present in the sample nor naturally present in the nucleic acid population undergoing multiplex PCR. For example, the primer can comprise one or more non-natural nucleic acid molecules, such as but not limited to thymine dimers, 8-oxo-2'-deoxyguanosine, inosine, deoxyuridine, bromodeoxyuridine, apurinic nucleotides, etc.
[0123] In some embodiments, the disclosed methods may optionally include destroying amplification artifacts containing one or more primers, such as primer-dimers, dimer-dimers, or superamplicons. In some embodiments, destruction may optionally include treating primers and / or amplification products to cut specific cleavable groups present in primers and / or amplification products. In some embodiments, treatment may include partial or complete digestion of one or more target-specific primers. In one embodiment, treatment may include removing at least 40% of target-specific primers from the amplification product. The cleavable treatment may include enzyme, acid, base, heat, light, or chemical activity. The cleavable treatment may result in the cutting or other destruction of the connection between one or more nucleotides of the primer or between one or more nucleotides of the amplification product. The primer and / or amplification product may optionally include one or more modified nucleotides or nucleobases. In some embodiments, cutting may occur selectively at these sites, or adjacent to modified nucleotides or nucleobases. In some embodiments, the primer includes a sufficient number of modified nucleotides to allow the primer to be functionally completely degraded by the cutting treatment, but not to interfere with the specificity or functionality of the primer before such cutting treatment, such as in the amplification reaction. In some embodiments, the primer comprises at least one modified nucleotide, but no more than 75% of the nucleotides of the primer are modified. In some embodiments, the cleavage or treatment of the amplified target sequence may result in the formation of a phosphorylated amplified target sequence. In some embodiments, the amplified target sequence is phosphorylated at the 5' end.
[0124] In some embodiments, primers can be designed de novo using an algorithm that generates oligonucleotide sequences according to specified design criteria. For example, primers can be selected according to any one or more of the criteria specified herein. In some embodiments, one or more primers are selected or designed to meet any one or more of the following criteria: (1) two or more modified nucleotides are included within the primer sequence, at least one of the nucleotides is included near or at the end of the primer, and at least one of the nucleotides is included at or around the central nucleotide position of the primer sequence; (2) the primer length is about 15 to about 40 bases; (3) the primer length is about 15 to about 40 bases; (4) the primer length is about 15 to about 40 bases; (5) the primer length is about 15 to about 40 bases; (6) the primer length is about 15 to about 40 bases; (7) the primer length is about 15 to about 40 bases; (8) the primer length is about 15 to about 40 bases; (9) the primer length is about 15 to about 40 bases; (10) the primer length is about 15 to about 40 bases; (11) the primer length is about 15 to about 40 bases; (12) the primer length is about 15 to about 40 bases; (13) the primer length is about 15 to about 40 bases; (14) the primer length is about 15 to about 40 bases; (15) the primer length is about 15 to about 40 bases; (16) the primer length is about 15 to about 40 bases; (17) the primer length is about 15 to about 40 bases; (18) the primer length is about 15 to about 40 bases; (19) the primer length is about 15 to about 40 bases; (20) the primer length is about 15 to about 40 bases; (21) the primer length is about mof about 60°C to about 70°C; (4) low cross-reactivity with non-target sequences present in the target genome or target sample; (5) for each primer in a given reaction, the sequence of at least the first four nucleotides (from 3' to 5' direction) is not complementary to any sequence within any other primer present in the same reaction; and (6) no amplicon contains any continuous stretch of at least 5 nucleotides that is complementary to any sequence within any other amplicon. In some embodiments, the primers include one or more primer pairs designed to amplify a target sequence of about 100 base pairs to about 500 base pairs in length from a sample. In some embodiments, the primers include a plurality of primer pairs designed to amplify target sequences, wherein the amplified target sequences are expected to vary in length from each other by no more than 50%, typically by no more than 25%, and even more typically by no more than 10% or 5%. For example, if one primer pair is selected or predicted to amplify a product of 100 nucleotides in length, other primer pairs are selected or predicted to amplify a product of 50-150 nucleotides in length (typically, between 75-125 nucleotides in length, even more typically, between 90-110 nucleotides in length, or 95-105 nucleotides, or 99-101 nucleotides in length). In some embodiments, at least one primer pair in the amplification reaction is not designed from scratch according to any predetermined selection criteria. For example, at least one primer pair can be an oligonucleotide sequence selected or generated at random, or an oligonucleotide sequence previously selected or generated for other applications. In an exemplary embodiment, the amplification reaction can include a sequence selected from At least one primer pair for the probe reagent (Roche Molecular Systems). Reagents include labeled probes, and are particularly useful for measuring the amount of target sequences present in a sample, optionally in real time. Some examples of TaqMan technology are disclosed in U.S. Patent Nos. 5,210,015, 5,487,972, 5,804,375, 6,214,979, 7,141,377, and 7,445,900, which are incorporated herein by reference as a whole. In certain embodiments, at least one primer in an amplification reaction may be labeled, for example, with an optically detectable label, to promote specific target applications. For example, labeling may promote the quantification of a target template and / or an amplified product, the separation of a target template and / or an amplified product, etc. In certain embodiments, a primer does not include a carbon spacer or an end joint. In certain embodiments, a primer or an amplified target sequence does not include an enzyme label, a magnetic label, an optical label, or a fluorescent label.
[0125] In certain embodiments, the primer that is complementary to the discontinuous section of nucleic acid template chain and can hybridize with it comprises: the primer that can hybridize with the 5' region of template, it covers the sequence complementary to forward or reverse amplification primer.In certain embodiments, forward primer, reverse primer or both do not share common nucleic acid sequence, so that they are hybridized with different nucleic acid sequences.For example, target specific forward and reverse primers can be prepared, and it does not compete with other primers in primer pool to increase the same nucleic acid sequence.In this example, the primer that does not compete with other primers in primer pool helps to reduce non-specific or false amplification product.In certain embodiments, the forward and reverse primers of each primer pair are unique, because the nucleotide sequence of each primer is non-complementary and different from other primers in primer pair.In certain embodiments, the difference of primer pair can be at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85% or at least 90% nucleotide identity. In certain embodiments, the forward and reverse primers in each primer pair are non-complementary or different from other primer pairs in a primer pool or multiplex reaction. For example, the primer pair in a primer pool or multiplex reaction can differ from other primer pairs in a primer pool or multiplex reaction by at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60% or at least 70% nucleotide identity. Primers are designed to minimize the formation of primer-dimers, dimer-dimers or other non-specific amplification products. Typically, primers are optimized to reduce GC bias and low melting temperature (T) during amplification reactions. m In some embodiments, the primers are designed to have a T of about 55°C to about 72°C. m In some embodiments, the primers of the primer pool may have a T of about 59°C to about 70°C, 60°C to about 68°C, or 60°C to about 65°C. m In some embodiments, the primer pool may have a T value that deviates by no more than 5°C. m .
[0126] In some embodiments, the primer pairs used to generate the amplicon library can result in amplification of target-specific nucleic acid molecules having one or more of the following metrics: greater than 97% target coverage at 20x if normalized to 100x average coverage depth; greater than 97% of bases at an average greater than 0.2x; greater than 90% of bases without strand bias; greater than 95% of all on-target reads; greater than 99% of bases at an average greater than 0.01x; and per-base accuracy greater than 99.5%.
[0127] In certain embodiments, primers can be provided in a single amplification container as a primer pair set. In certain embodiments, primers can be provided in one or more primer pairs aliquots, and before multiplex PCR reaction is carried out in a single amplification container or reaction chamber, these aliquots can be merged. In one embodiment, primers can be provided as a forward primer pool and a separate reverse primer pool. In another embodiment, primer pairs can be merged into subsets, such as non-overlapping primer pairs. In certain embodiments, primer pair pools can be provided in a single reaction chamber or microwell, for example, on a PCR plate, to use a thermal cycler to carry out multiplex PCR. In certain embodiments, forward and reverse primer pairs can be substantially complementary to the target sequence. In certain embodiments, primer pairs do not include a common extension (tail) at the 3' end or 5' end of the primer. In another embodiment, primers do not include a label or a universal sequence. In certain embodiments, primer pairs are designed to eliminate or reduce the interaction that promotes the formation of non-specific amplification.
[0128] Method for detecting and / or measuring the presence of microorganisms in a sample
[0129] This paper also provides the method based on nucleic acid of whether there is microorganism in detection and / or measurement sample.In some embodiments of detection and / or measurement method, use any nucleic acid provided by this paper as probe and / or amplification primer, nucleic acid hybridization and / or amplification are carried out to the nucleic acid in sample or from sample.In certain embodiments, detect whether there is one or more hybridization and / or nucleic acid amplification products, thereby detect whether there is microorganism.
[0130] In some embodiments, the methods provided herein for detecting, determining the presence or absence of one or more microorganisms in a sample and / or measuring one or more microorganisms in a sample comprise (a) performing nucleic acid amplification on a nucleic acid in or from a sample using one or more primer pairs, wherein the one or more primer pairs specifically amplify a target nucleic acid sequence contained in the genome of a microorganism selected from the microorganisms listed in Table 1 (or Table 1, excluding or excluding Actinomyces viscosus and / or Blautia sphaeroides; or Table 1, excluding or excluding Actinomyces viscosus, Blautia sphaeroides and / or Helicobacter zabbii); and (b) detecting one or more amplification products (or their presence or absence), thereby detecting and / or measuring one or more microorganisms selected from the microorganisms listed in Table 1 (or Table 1, excluding or excluding Actinomycetes viscosus and / or Blautia sphaeroides; or Table 1, excluding or excluding Actinomycetes viscosus, Blautia sphaeroides and / or Helicobacter zadoni), or determining the presence or absence of one or more microorganisms selected from the microorganisms listed in Table 1 (or Table 1, excluding or excluding Actinomycetes viscosus and / or Blautia sphaeroides; or Table 1, excluding or excluding Actinomycetes viscosus, Blautia sphaeroides and / or Helicobacter zadoni). In some embodiments, the target nucleic acid is unique to the microorganism. In some embodiments, the target nucleic acid is not contained in the prokaryotic 16S rRNA gene. Any embodiment provided herein for amplifying nucleic acids using one or more primer pairs can be used in any embodiment of the method to detect the presence or absence of microorganisms in a sample, wherein the one or more primer pairs specifically amplify a target nucleic acid sequence contained in the genome of a microorganism listed in Table 1. In some embodiments, at least one primer pair does not detectably amplify a nucleic acid sequence contained in any genus other than the genus of the microorganism. In some embodiments, at least one primer pair does not detectably amplify a nucleic acid sequence contained in any species other than a microbial species. In some embodiments, the nucleic acid in or from the sample comprises nucleic acids from a plurality of different microorganisms listed in Table 1 (or Table 1, excluding or excluding Actinomycetes viscosus and / or Blautia sphaeroides; or Table 1, excluding or excluding Actinomycetes viscosus, Blautia sphaeroides and / or Helicobacter zabbii) and / or a plurality of different microorganisms not listed in Table 1 (e.g., bacteria). In some embodiments, the sample is a biological sample, such as, for example, a sample of the contents of an animal's digestive tract. In some embodiments, the sample is a fecal sample.In some embodiments, the target nucleic acid sequence comprises or consists essentially of a nucleotide sequence selected from the group consisting of SEQ ID NOs: 1605-1979 in Table 17, or SEQ ID NOs: 1605-1806, 1809-1816, 1821-1970, 1972-1974, and 1977-1979 in Table 17, or SEQ ID NOs: 1605-1826 in Table 17, or SEQ ID NOs: 1605-1806, 1809-1816, and 1821-1826 in Table 17, or SEQ ID NOs: 1605-1820 in Table 17A, or SEQ ID NOs: 1605-1806 and 1809-1816 in Table 17A, or SEQ ID NOs: 1827-1979 in Table 17C, or SEQ ID NOs: 1827-1979 in Table 17C. NO:1827-1976, or substantially the same or similar sequence. In some embodiments, detecting the presence or absence of one or more amplification products comprises detecting the presence or absence of one or more nucleotide sequences selected from the group consisting of SEQ ID NOs: 1605-1979 in Table 17, or SEQ ID NOs: 1605-1806, 1809-1816, 1821-1970, 1972-1974, and 1977-1979 in Table 17, or SEQ ID NOs: 1605-1826 in Table 17, or SEQ ID NOs: 1605-1806, 1809-1816, and 1821-1826 in Table 17, or SEQ ID NOs: 1605-1820 in Table 17A, or SEQ ID NOs: 1605-1806 and 1809-1816 in Table 17A, or SEQ ID NOs: 1827-1979 in Table 17C, or SEQ ID NOs: 1827-1979 in Table 17C. NO:1827-1976, or substantially the same or similar sequence, or its complement.In certain embodiments, at least one primer pair can not detectably increase the nucleotide sequence contained in any genus except the microorganism genus containing the target nucleic acid sequence.In certain embodiments, at least one primer pair can not detectably increase the nucleotide sequence contained in any kind except the microorganism species containing the target nucleic acid sequence.In some embodiments, at least one primer in the primer pair, or at least one primer pair, contains or consists essentially of one or more sequences of the primers or primer pairs in Table 16; or SEQ ID NOs: 49-520 of Table 16; or SEQ ID NOs: 49-452, 457-472, and 481-520 of Table 16; or SEQ ID NOs: 49-492 of Table 16; or SEQ ID NOs: 49-452, 457-472, and 481-492 of Table 16; or SEQ ID NOs: 49-480 of Table 16A; or SEQ ID NOs: 49-452 and 457-472 of Table 16A; or SEQ ID NOs: 521-826 of Table 16C; or SEQ ID NOs: 521-820 of Table 16C; or SEQ ID NOs: 827-1298 of Table 16; or SEQ ID NOs: 527-1299 of Table 16. or SEQ ID NOs: 1299-1598 of Table 16F, or substantially identical or similar sequences, or the nucleotide sequence of any of the foregoing nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases.In some embodiments, nucleic acid amplification is performed on a nucleic acid using a plurality of primers or primer pairs, each of which contains or consists essentially of one or more sequences selected from the group consisting of the sequences of the primers in Table 16; or SEQ ID NOs: 49-520 of Table 16; or SEQ ID NOs: 49-452, 457-472, and 481-520 of Table 16; or SEQ ID NOs: 49-492 of Table 16; or SEQ ID NOs: 49-452, 457-472, and 481-492 of Table 16; or SEQ ID NOs: 49-480 of Table 16A; or SEQ ID NOs: 49-452 and 457-472 of Table 16A; or SEQ ID NOs: 521-826 of Table 16C; or SEQ ID NOs: 521-820 of Table 16C; or SEQ ID NOs: 521-820 of Table 16C; or SEQ ID NOs:1299-1598 of Table 16F; or substantially the same or similar sequences, or the nucleotide sequences of any of the above nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases. In some embodiments where nucleic acid amplification is performed on a nucleic acid using more than one or more primers or primer pairs, the amplification is a multiplex amplification performed in a single reaction mixture. In some embodiments, at least one primer or primer pair comprises a modification that promotes nucleic acid manipulation, amplification, connection, and / or sequencing of amplified products and / or reduces or eliminates primer dimers. In a specific embodiment, the modification is a modification that promotes multiple nucleic acid amplification, connection, and / or sequencing of multiple amplified products.
[0131] Methods for detecting and / or measuring microbial populations
[0132] In some embodiments of the method for detecting, determining the presence or absence of one or more microorganisms in a sample and / or measuring one or more microorganisms in a sample, the method is designed to focus on detecting and / or measuring a certain group of microorganisms. In such embodiments, the one or more primer pairs used in the method are primers and / or combinations of primer pairs that contain a selected group or subset of microorganism-specific nucleic acids, enabling targeted investigation of the sample to identify microbial species that may be important, for example, in certain health and disease states or microbial imbalances (e.g., dysbiosis). In some embodiments, the combination of nucleic acids comprises microorganism-specific nucleic acids and / or primer pairs that specifically amplify nucleic acid sequences contained in the genome of one or more microorganisms (e.g., bacteria) that are associated with one or more pathologies, disorders, and / or diseases (referred to herein as a "conditional group" of microorganisms. In specific embodiments, the combination of nucleic acid primers and / or primer pairs comprises primers that specifically amplify sequences contained in a genome selected from Table 1 (or Table 1, excluding or excluding Actinomyces viscosus and / or Blautia sphaeroides; or Table 1, excluding or excluding Actinomyces viscosus In some embodiments, the combination comprises a plurality of nucleic acids and / or primer pairs, wherein the plurality of nucleic acids and / or primer pairs comprises at least one nucleic acid primer pair that specifically amplifies a target nucleic acid in each microorganism in Table 1 (or Table 1, excluding or excluding Actinomyces viscosus and / or Blautia sphaeroides; or Table 1, excluding or excluding Actinomyces viscosus, Blautia sphaeroides and / or Helicobacter zabbii). In some embodiments, the plurality of primer pairs comprises Primer pairs that specifically amplify genomic target nucleic acids contained in at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, or 70 microorganisms in Table 1 (or Table 1, excluding or excluding Actinomycetes viscosus and / or Blautia sphaeroides; or Table 1, excluding or excluding Actinomycetes viscosus, Blautia sphaeroides, and / or Helicobacter zascheroni). In certain embodiments, the target nucleic acid sequences contained in the genomes of different microorganisms are unique to each microorganism. In some embodiments, the nucleic acids and / or nucleic acids The combination of nucleic acid primer pairs comprises two or more nucleic acids and / or nucleic acid primer pairs that specifically amplify a unique nucleic acid sequence contained in the genome of one or more Group A microorganisms (see Table 2A), wherein the Group A microorganisms are species involved in playing a role in a variety of conditions, diseases and / or disorders, including, for example, neoplastic conditions (including, for example, response to immuno-oncology therapy and cancer), gastrointestinal disorders (including, for example, irritable bowel syndrome, inflammatory bowel disease and celiac disease), and autoimmune diseases (including, for example, lupus and rheumatoid arthritis).In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises a set of nucleic acid primer pairs, wherein each different nucleic acid primer pair specifically amplifies a different unique nucleic acid sequence contained in a different genome of each genome of different microorganisms in Group A. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises nucleic acids and / or nucleic acid primer pairs, which specifically bind, hybridize and / or specifically amplify the sequences listed in Table 2A for exemplary nucleic acids, primers and primer pairs for Group A microorganism genomes. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises nucleic acids and / or nucleic acid primer pairs, which have one or more nucleotide sequences listed in Table 2A for exemplary nucleic acids, primers and primer pairs for Group A microorganism genomes.
[0133] In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs in the method for detecting and / or measuring a certain group of microorganisms comprises two or more nucleic acids and / or nucleic acid primer pairs that specifically amplify unique nucleic acid sequences contained in the genome of one or more Group B microorganisms (see Table 2B), which are species that are considered to be associated with immuno-oncology treatment response. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises a set of nucleic acid primer pairs, wherein each different nucleic acid primer pair specifically amplifies a different unique nucleic acid sequence contained in a different genome of each genome of different microorganisms in Group B. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises nucleic acids and / or nucleic acid primer pairs that specifically bind, hybridize and / or specifically amplify sequences listed in Table 2B for exemplary nucleic acids, primers and primer pairs for Group B microorganism genomes. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises nucleic acids and / or nucleic acid primer pairs that have one or more nucleotide sequences listed in Table 2B for exemplary nucleic acids, primers and primer pairs for Group B microorganism genomes.
[0134] In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs in the method for detecting and / or measuring a certain group of microorganisms comprises two or more nucleic acids and / or nucleic acid primer pairs that specifically amplify a unique nucleic acid sequence contained in the genome of one or more C group microorganisms (see Table 2C) or C group microorganisms that do not contain Helicobacter zavodii (subgroup 1 of C group microorganisms), which are species believed to be associated with cancer. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises a set of nucleic acid primer pairs, wherein each different nucleic acid primer pair specifically amplifies a different unique nucleic acid sequence in a different genome of each genome of different microorganisms in Group C or Group C microorganisms that do not contain Helicobacter zavodii (subgroup 1 of C group microorganisms). In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises nucleic acids and / or nucleic acid primer pairs that specifically bind, hybridize, and / or specifically amplify sequences listed in Table 2C for exemplary nucleic acids, primers, and primer pairs for genomes of Group C microorganisms. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises the following nucleic acids and / or nucleic acid primer pairs, which specifically bind to, hybridize to, and / or specifically amplify the following sequences: SEQ ID NO: 17 ID NO: 1616, 1619, 1620, 1625-1628, 1635-1640, 1699, 1700, 1705-1708, 1752, 1753, 1784-1786, 1817-1820, 1827, 1828, 1840, 1841, 1844, 1845, 1852-1859, 1899, 1900, 1904, 1905, 1932, 1933, 1956-1958, 1975, 1976 and / or a substantially identical or similar sequence, or a SEQ ID NO selected from Table 17 NO: 1616, 1619, 1620, 1625-1628, 1635-1640, 1699, 1700, 1705-1708, 1752, 1753, 1784-1786, 1827, 1828, 1840, 1841, 1844, 1845, 1852-1859, 1899, 1900, 1904, 1905, 1932, 1933, 1956, 1957, 1958 and / or substantially identical or similar sequences. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises nucleic acids and / or nucleic acid primer pairs having one or more nucleotide sequences as listed in Table 2C for exemplary nucleic acids, primers, and primer pairs for Group C microbial genomes.In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises the following nucleic acids and / or nucleic acid primer pairs, wherein the nucleic acids and / or nucleic acid primer pairs have one or more nucleotide sequences selected from the group consisting of: SEQ ID NO: 71, 72, 77-80, 89-96, 109-120, 237-242, 249-256, 343-346, 407-412, 493-496, 511-520, 521-524, 547-550, 555-558, 561-568, 571-586, 665-668, 675-678, 731-734, 779-784 and / or SEQ ID NO: NO:849, 850, 855-858, 867-874, 887-898, 1012-1020, 1025-1034, 1121-1124, 1185-1190, 1271-1276, 1289-1298, 1299-1302, 1325-1328, 1333-1336, 1339-1346, 1349-1364, 1443-1446, 1453-1456, 1509-1512, 1557-1562, or a substantially identical or similar sequence, or the nucleotide sequence of any of the foregoing nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises the following nucleic acids and / or nucleic acid primer pairs, wherein the nucleic acids and / or nucleic acid primer pairs have one or more nucleotide sequences selected from the group consisting of: SEQ ID NO: 71, 72, 77-80, 89-96, 109-120, 237-242, 249-256, 343-346, 407-412, 493-496, 511-520 and / or SEQ ID NO: NO:849, 850, 855-858, 867-874, 887-898, 1012-1020, 1025-1034, 1121-1124, 1185-1190, 1271-1276, 1289-1298, or a substantially identical or similar sequence, or the nucleotide sequence of any of the foregoing nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases.In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises nucleic acids and / or nucleic acid primer pairs having one or more nucleotide sequences selected from the group consisting of SEQ ID NOs: 71, 72, 77-80, 89-96, 109-120, 237-242, 249-256, 343-346, 407-412 and / or SEQ ID NOs: 849, 850, 855-858, 867-874, 887-898, 1012-1020, 1025-1034, 1121-1124, 1185-1190 in Table 16, or a substantially identical or similar sequence, or the nucleotide sequence of any of the foregoing nucleic acids or primer pairs, wherein one or more thymine bases are replaced by uracil bases.
[0135] In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs used in the methods for detecting and / or measuring a group of microorganisms comprises two or more nucleic acids and / or nucleic acid primer pairs that specifically amplify unique nucleic acid sequences contained in the genomes of one or more Group D microorganisms (see Table 2D), which are species believed to be associated with gastrointestinal disorders, including, for example, irritable bowel syndrome, inflammatory bowel disease, and celiac disease. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises a set of nucleic acid primer pairs, wherein each different nucleic acid primer pair specifically amplifies a different unique nucleic acid sequence contained in a different genome of each genome of different microorganisms in Group D. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises nucleic acids and / or nucleic acid primer pairs that specifically bind, hybridize, and / or specifically amplify sequences listed in Table 2D for exemplary nucleic acids, primers, and primer pairs for Group D microorganism genomes. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises nucleic acids and / or nucleic acid primer pairs having one or more nucleotide sequences as listed in Table 2D for exemplary nucleic acids, primers, and primer pairs for Group D microbial genomes.
[0136] In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs used in the method for detecting and / or measuring a certain group of microorganisms comprises two or more nucleic acids and / or nucleic acid primer pairs that specifically amplify unique nucleic acid sequences contained in the genomes of one or more Group E microorganisms (see Table 2E), which are species believed to be associated with autoimmune disorders, including but not limited to lupus and rheumatoid arthritis. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises a set of nucleic acid primer pairs, wherein each different nucleic acid primer pair specifically amplifies a different unique nucleic acid sequence contained in a different genome of each genome of different microorganisms in Group E. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises nucleic acids and / or nucleic acid primer pairs that specifically bind, hybridize, and / or specifically amplify sequences listed in Table 2E for exemplary nucleic acids, primers, and primer pairs for Group E microorganism genomes. In some embodiments, the combination of nucleic acids and / or nucleic acid primer pairs comprises nucleic acids and / or nucleic acid primer pairs having one or more nucleotide sequences as listed in Table 2E for exemplary nucleic acids, primers, and primer pairs for Group E microbial genomes.
[0137] In some embodiments, detection, determination of the presence or absence of one or more nucleic acid amplification products and / or measurement of one or more nucleic acid amplification products as provided herein are detected by a labeled probe having a nucleic acid sequence of a composition provided herein, and the labeled probe specifically identifies a specific nucleic acid product. For example, such methods can use a nucleic acid microarray of an oligonucleotide connected to a substrate (e.g., a chip) to capture amplification products, and then contact the amplification products with a labeled (e.g., fluorescently labeled) specific probe under conditions suitable for hybridization, which can be detected when combined with a complementary product, thereby detecting the presence of microorganisms in a sample. Methods for detecting labels are known in the art and include, for example, optical methods, such as scanning using a confocal laser microscope or a CCD camera. Such methods also allow hybridization to be quantified to assess the abundance of labeled nucleic acid products. In some embodiments, the presence or absence of one or more nucleic acid amplification products is detected by obtaining nucleotide sequence information of one or more nucleic acid amplification products. Methods for sequencing nucleic acids are described herein and / or are known in the art. The sequence of amplification products can also be used to identify microorganisms at various specificity levels, for example, kingdoms, phyla, classes, orders, families, genera, and / or species. If there is a microorganism in Table 1 or 2 to be determined whether there is present in the sample, the sequence of at least one amplified product will be a target sequence specifically amplified by one or more primers from the microorganism genome, so the presence of the amplified product is detected. If there is no microorganism in the sample, an amplified product containing a target sequence specifically amplified by one or more primers will not be produced, so the absence of the amplified product is detected. In some embodiments, detecting whether there is a nucleic acid amplification product includes comparing the sequence of one or more nucleic acid amplification products with the nucleic acid sequence of the genome of one or more microorganisms in Table 1 or 2. The genome sequence of the microorganism in Table 1 or 2 can be obtained in public databases (e.g., NCBI public database; www.ncbi.nlm.nih.gov / genome / microbes / ). In some embodiments, comparing the sequence of the nucleic acid amplification product with the reference genome sequence includes performing a computer-assisted comparison of the sequence and mapping it to the reference genome. An exemplary nucleotide sequence analysis workflow for mapping the sequence reads of the amplified product is provided herein.In certain embodiments, detecting the presence or absence of a nucleic acid amplification product comprises detecting the presence or absence of an amplification product comprising the following sequences: SEQ ID NOs: 1605-1979 in Table 17, or SEQ ID NOs: 1605-1806, 1809-1816, 1821-1970, 1972-1974, and 1977-1979 in Table 17, or SEQ ID NOs: 1605-1826 in Table 17, or SEQ ID NOs: 1605-1806, 1809-1816, and 1821-1826 in Table 17, or SEQ ID NOs: 1605-1820 in Table 17A, or SEQ ID NOs: 1605-1806 and 1809-1816 in Table 17A, or SEQ ID NOs: 1827-1979 in Table 17C, or SEQ ID NOs: 1827-1979 in Table 17C. NO:1827-1976, or substantially the same or similar sequence.
[0138] Method for improving the accuracy, specificity and / or sensitivity of detecting and / or measuring microorganisms in a sample
[0139] In some embodiments, the methods provided herein for detecting, determining whether there are one or more microorganisms in a sample and / or measuring one or more microorganisms in a sample include (a) using one or more primer pairs or a combination of primer pairs to perform nucleic acid amplification on nucleic acids in or from a sample, wherein the one or more primer pairs or a combination of primer pairs can individually amplify nucleic acids containing sequences of one or more different hypervariable regions of a prokaryotic 16S rRNA gene, respectively; and (b) detecting one or more amplification products, thereby detecting one or more microorganisms in a sample or determining whether there are one or more microorganisms in a sample. In some embodiments, the microorganism is a bacterium. In some embodiments, the presence or absence of one or more microorganisms is detected at the level of a microbial genus. In some embodiments, the presence or absence of one or more microorganisms is detected at the level of a microbial species. In some embodiments, the prokaryotic 16S rRNA gene is a bacterial gene. In some embodiments, the one or more primer pairs or a combination of primer pairs individually amplify nucleic acids containing sequences of different variable regions. In some embodiments, the primers in the primer pairs are directed to nucleic acid sequences contained in the conserved regions of the prokaryotic 16S rRNA gene, or are combined with or hybridized to the nucleic acid sequences. In some embodiments, the one or more primer pairs or combinations of primer pairs amplify or individually amplify nucleic acids containing 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more or 9 different hypervariable regions of prokaryotic 16S rRNA genes, thereby generating amplified copies of nucleic acids containing 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more or 9 different hypervariable regions of 16S rRNA genes of one or more microorganisms, wherein in some embodiments, the amplified copies of different hypervariable regions are separate amplicons. In some embodiments, the one or more primer pairs or combinations of primer pairs individually amplify nucleic acids containing 8 different hypervariable regions of prokaryotic 16S rRNA genes, respectively. In some embodiments, 8 different hypervariable regions are V2-V9. In some embodiments, the one or more primer pairs or combinations of primer pairs individually amplify nucleic acids containing sequences of 3 or more different hypervariable regions of prokaryotic 16S rRNA genes, wherein one of the 3 or more different regions is a V5 difference region, thereby generating amplified copies of nucleic acids containing sequences of 3 or more different hypervariable regions of 16S rRNA genes of one or more microorganisms. In some embodiments, the primer pairs comprise the degenerate sequence of one or more primers in the one or more primer pairs. In some embodiments, the one or more primer pairs or combinations of primer pairs are selected from any primer pairs described herein, which individually amplify nucleic acids containing sequences of one or more hypervariable regions of prokaryotic 16S rRNA genes.For example, in some embodiments, the one or more primer pairs or combinations of primer pairs capable of individually amplifying a nucleic acid containing a sequence of one or more hypervariable regions of a prokaryotic 16S rRNA gene include one or more primer pairs containing a sequence selected from the following: one or more sequences of primers or primer pairs in Table 15; or SEQ ID NOs: 1-24 in Table 15 and / or SEQ ID NOs: 25-48 in Table 15; or SEQ ID NOs: 11-16, 23 and 24 in Table 15 and / or SEQ ID NOs: 35-40, 47 and 48 in Table 15, or substantially identical or similar sequences, and optionally in which one or more thymine bases are replaced by uracil bases. In some embodiments, the one or more primer pairs or combinations of primer pairs that individually amplify nucleic acids containing sequences of one or more hypervariable regions of prokaryotic 16S rRNA genes include one or more primer pairs containing sequences selected from the following: SEQ ID NO: 25-48 in Table 15 and / or SEQ ID NO: 35-40, 47 and 48 in Table 15, or substantially identical or similar sequences, wherein one or more thymine bases are replaced by uracil bases. Substantially identical or similar sequences. In some embodiments, the presence or absence of one or more nucleic acid amplification products is detected by obtaining nucleotide sequence information of one or more nucleic acid amplification products. If only one sequence is obtained for each of the one or more hypervariable regions amplified, the sequence information indicates that only one microorganism is present in the sample. If 2 or more different sequences are obtained for each of the one or more hypervariable regions amplified, the sequence information indicates that 2 or more different microorganisms are present in the sample. In addition, the combined results of amplification using 16S rRNA gene primers that individually amplify multiple different hypervariable regions improve the accuracy of detecting and / or measuring one or more microorganisms in a sample by reducing false negatives or false positives that may occur when determining the results based on the use of 16S rRNA primers that amplify only one or only a few (e.g., less than 8 or less than 5) hypervariable regions or amplify a combined region (e.g., V2-V3 or V3-V4, etc.) as a single amplicon (if the amplification primers are not directed to the highly conserved sequences on both sides of the multiple regions). For example, due to possible sequence variations in the conserved regions of the 16S rRNA gene in certain species or strains of microorganisms (e.g., bacteria), primers designed to amplify specific hypervariable regions of all bacteria may fail to amplify certain bacterial nucleic acids present in the sample, thereby producing false negative results. If the primers are designed to amplify the combined region as a single amplicon and do not target the highly conserved sequences located on both sides of the multiple regions, this failure may be more complicated because not only one region may not be amplified, but two or more regions may also not be amplified.However, using the 16S rRNA gene primers that are designed to increase the multiple hypervariable regions in this type of microorganism individually to amplify the same nucleic acid may amplify at least one hypervariable sequence in the microorganism and can detect the presence of microorganism in the sample. Amplifying multiple hypervariable regions individually has increased the coverage of the sequence and provided more useful information and increased the accuracy and resolution of detection. In addition, when using the 16S rRNA gene primers that amplify multiple (for example, 3, 4, 5, 6, 7, 8 or 9) hypervariable regions individually, the quantity (and sequence) of the amplified products using different hypervariable region primers is for filtering and eliminating the hypervariable region sequence of the specific microorganism that is amplified only in one (or less than 3, or less than 4, or less than 5, or less than 6 or another threshold value amount) 16S rRNA gene hypervariable region amplification so that it is not considered to be unreliable and provides the basis, because the 16S rRNA gene nucleic acid from the bacterial microorganism that really exists in the sample should be amplified in most of amplifications using the primer that amplifies multiple hypervariable regions individually. The sequence of the amplified product can also be used to identify microorganisms at various specific levels, for example, kingdoms, phyla, classes, orders, families, genera and / or species. In some embodiments, detecting, determining whether a nucleic acid amplification product exists and / or measuring a nucleic acid amplification product comprises comparing the sequence of one or more nucleic acid amplification products with the nucleic acid sequence of the prokaryotic (e.g., bacterial) 16S rRNA gene of one or more microorganisms. The sequence of the prokaryotic 16S rRNA gene is available in public databases (see, for example, GreenGenes bacterial 16S rRNA gene sequence; for example, www.greengenes.lbl.gov). In some embodiments, comparing the sequence of the nucleic acid amplification product with the reference genome sequence comprises performing a computer-assisted alignment of the sequence and mapping it to the reference genome. An exemplary nucleotide sequence analysis workflow for mapping sequence reads of the amplified product is provided herein. In some embodiments, the relative and / or absolute levels of one or more microorganisms are determined or measured. For example, in some embodiments, the abundance level of one or more nucleic acid amplification products and / or sequence reads can be measured to provide the relative and / or absolute levels of one or more microorganisms. The technology for quantitatively quantifying nucleic acids (e.g., amplified products) and / or sequence reads is known in the art and / or provided herein.
[0140] In some embodiments, the methods provided herein for detecting, determining the presence or absence of one or more microorganisms in a sample and / or measuring one or more microorganisms in a sample comprise (a) performing nucleic acid amplification on a nucleic acid in or from a sample using a combination of primer pairs, the combination of primer pairs comprising: (i) one or more primer pairs capable of amplifying nucleic acids containing one or more hypervariable region sequences of a prokaryotic 16S rRNA gene (referred to as "16S rRNA gene primers and primer pairs"); and (ii) one or more primer pairs capable of amplifying a target nucleic acid sequence contained in the genome of a microorganism, the target nucleic acid sequence not being contained in the hypervariable region of the prokaryotic 16S rRNA gene, wherein different primer pairs amplify different target nucleic acid sequences contained in different genomes (referred to as "non-...
Claims
1. A system comprising: machine readable storage; and a processor configured to execute machine-readable instructions that, when executed by the processor, cause the system to perform a method comprising: receiving a plurality of nucleic acid sequence reads, wherein the sequence reads comprise a plurality of 16S sequence reads; first mapping the plurality of 16S sequence reads to a plurality of compressed 16S reference sequences, wherein each compressed 16S reference sequence comprises a set of hypervariable segments for a corresponding strain of a species, wherein a computer PCR simulation is applied to a full-length 16S rRNA gene sequence contained in a 16S rRNA gene sequence database to generate the compressed 16S reference sequences based on primers in a 16S primer pool; generating a read count matrix containing read counts of 16S sequence reads mapped to each hypervariable segment in the set of hypervariable segments, wherein rows of the read count matrix correspond to strains of species and columns correspond to the hypervariable segments; reducing the read count matrix by applying a threshold to the read counts to form a reduced read count matrix; compressing the database of full-length 16S reference sequences to form a reduced set of full-length 16S reference sequences based on selecting corresponding full-length 16S reference sequences from the database of full-length 16S reference sequences using the reduced read count matrix, the reduced set of full-length 16S reference sequences being stored in a memory; mapping the plurality of 16S sequence reads to the reduced set of full-length 16S reference sequences for a second time; counting 16S sequence reads of each full-length reference in the reduced set that map to the full-length 16S reference sequence to form a second read count set; Normalizing the read counts in the second read count set to form normalized counts; aggregating the normalized counts for a given level to form aggregated counts, wherein the given level is a species level, a genus level, or a family level; and A threshold is applied to the aggregate counts to detect the presence of a given level of microorganism in the sample.
2. The system of claim 1, wherein reducing the read count matrix further comprises eliminating rows of the read count matrix to form a first reduced read count matrix when a sum of read counts within a row is less than a row sum threshold.
3. The system of claim 2, wherein reducing the read count matrix further comprises summing the read counts of the rows of the first reduced read count matrix corresponding to the same expected feature of the corresponding species to form column sums; and The column sums are added to form a combined sum, where an expected feature comprises a binary value corresponding to whether a hypervariable segment in the set of hypervariable segments is expected to be present (=1) or absent (=0) in a strain.
4. The system of claim 3, wherein reducing the read count matrix further comprises eliminating rows of the first reduced read count matrix to form a second reduced read count matrix when the combined sum is less than a combined sum threshold.
5. The system of claim 3, wherein reducing the read count matrix further comprises applying a feature threshold to the column sums to assign binary values to form observed features for each row of a second reduced read count matrix, the observed features and expected features each having a total number of categories.
6. The system of claim 5, wherein the compressing further comprises determining a ratio of classes having matching binary values in the observed features and the expected features to the total number of classes.
7. The system of claim 6, wherein the compressing further comprises selecting a corresponding full-length 16S reference sequence for the first reduced set of full-length 16S reference sequences from a database of full-length 16S reference sequences stored in a memory when the ratio is greater than a ratio threshold.
8. The system of claim 1, wherein the plurality of nucleic acid sequence reads further comprises a plurality of target species sequence reads.
9. The system of claim 8, further comprising mapping the target species sequence reads to segmented reference sequences to form target species mapped reads, wherein each segmented reference sequence includes segments corresponding to expected amplicons of strains of the target species.
10. The system of claim 1, wherein the plurality of 16S sequence reads correspond to amplicons generated by amplifying a nucleic acid sample in the presence of one or more primer pairs targeting one or more hypervariable regions of a prokaryotic 16S rRNA gene.
11. The system of claim 8, wherein the plurality of target species sequence reads correspond to amplicons generated by amplifying a target nucleic acid sequence contained within the genome of a microorganism outside a hypervariable region of a prokaryotic 16S rRNA gene, wherein different primer pairs amplify different target nucleic acid sequences contained within the genomes of different microorganisms in the nucleic acid sample.
12. A system comprising: receiving, at a processor, a plurality of nucleic acid sequence reads, wherein the sequence reads comprise a plurality of 16S sequence reads; first mapping the reads of the plurality of 16S sequence reads to a plurality of compressed 16S reference sequences, wherein each compressed 16S reference sequence comprises a set of hypervariable segments for a corresponding strain of the species; Counting 16S sequence reads mapped to each hypervariable segment in the hypervariable segment set to form a first read count set; compressing a database of full-length 16S reference sequences to form a reduced set of full-length 16S reference sequences based on the first set of read counts of 16S sequence reads mapped to the compressed 16S reference sequences, the reduced set of full-length 16S reference sequences stored in a memory; mapping the plurality of 16S sequence reads to the reduced set of full-length 16S reference sequences for a second time; counting 16S sequence reads that map to each full-length reference sequence in the reduced set of full-length 16S reference sequences to form a second read count set; as well as The presence of microorganisms in the sample at the species level, genus level, or family level is detected based on the second read count set.
13. The system of claim 12, wherein the plurality of nucleic acid sequence reads further comprises a plurality of target species sequence reads.
14. The system of claim 13, further comprising mapping the target species sequence reads to segmented reference sequences to form target species mapped reads, wherein each segmented reference sequence includes segments corresponding to expected amplicons of strains of the target species.
15. The system of claim 14, further comprising aggregating counts of mapped reads of the target species to form aggregated read counts per species.
16. The system of claim 15, further comprising detecting the presence of a target species in the sample based on the aggregated read counts for each species.
17. The system of claim 14, further comprising generating the segmented reference sequence by applying computer PCR based on primers in a seed primer pool.
18. The system of claim 12, wherein the plurality of 16S sequence reads correspond to amplicons generated by amplifying a nucleic acid sample in the presence of one or more primer pairs targeting one or more hypervariable regions of a prokaryotic 16S rRNA gene.
19. The system of claim 13, wherein the plurality of target species sequence reads correspond to amplicons generated by amplifying a target nucleic acid sequence contained within the genome of a microorganism outside a hypervariable region of a prokaryotic 16S rRNA gene, wherein different primer pairs amplify different target nucleic acid sequences contained within the genomes of different microorganisms in the nucleic acid sample.
Citation Information
Patent Citations
Methods and apparatus for measuring analytes using large scale FET arrays
US20090026082A1
Methods and apparatus for measuring analytes using large scale FET arrays
US20090127589A1
Methodology for analysis of sequence variations within the HCV NS5b genomic region
US20090325145A1
Method for multiplexed nucleic acid patch polymerase chain reaction
US20100129874A1
Scaffolded nucleic acid polymer particles and methods of making and using
US20100304982A1