Methods and systems for species identification using sequences common to multiple species

By using nucleic acid sequences common to multiple species to generate a molecular identification code, the method addresses the inefficiencies of current species identification techniques, achieving high-throughput, cost-effective, and accurate species detection.

WO2026041101A1PCT designated stage Publication Date: 2026-02-26HANGZHOU BIOCHIP AI LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/116225
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-23
Filing Date
2025-08-21
Publication Date
2026-02-26

AI Technical Summary

Technical Problem

Current methods for species identification, such as those based on phenotypical differences or single genetic markers, struggle with accuracy and efficiency, particularly in distinguishing between species from different genera, and require extensive resources and time, while high-throughput sequencing is costly and complex.

Method used

A method utilizing nucleic acid sequences common to multiple species to generate a molecular identification code, enabling efficient detection through techniques like PCR, FISH, or DNA microarrays, reducing the number of assays needed and enhancing sensitivity and specificity.

Benefits of technology

This approach allows for high-throughput, cost-effective, and rapid species identification with increased accuracy, capable of detecting a larger number of species with fewer sequences and less sample material, suitable for miniaturized point-of-care testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025116225_26022026_PF_FP_ABST
    Figure CN2025116225_26022026_PF_FP_ABST
Patent Text Reader

Abstract

Methods and systems are disclosed for determining a set of target sequences for detecting polynucleic acid originating from a plurality of candidate species, including obtaining a genomic sequence of each candidate species; generating a set of k-mers for each candidate species; and incorporating iteratively k-mers into the set of target sequences, and calculating a molecular identification code for each of the plurality of candidate species, until all candidate species have a unique molecular identification code. Methods and system are also disclosed for analyzing in a sample comprising polynucleic acid originating from a plurality of candidate species, including detecting if each one of a set of target sequences is present in the polynucleic acid; generating a molecular identification code comprising binary digits, wherein each binary digit encodes whether a corresponding target sequence is present in the polynucleic acid; and comparing the molecular identification code with known molecular identification codes.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS AND SYSTEMS FOR SPECIES IDENTIFICATION USING SEQUENCES COMMON TO MULTIPLE SPECIESBACKGROUND

[0001] The present application relates to the field of biological detection technology. Specifically, present application discloses methods of identifying biological species based on features of nucleic acid sequences. The methods disclosed herein can efficiently identify microorganisms of different genera and species, and they may be widely applied in the fields of microbial detection and pathogen screening.

[0002] The methods disclosed in the present application generate a unique molecular identification code for each candidate species by utilizing nucleic acid sequences that are shared between or among the genomes of the candidate species.

[0003] More specifically, the methods disclosed herein generate designs of molecular probes; detect the presence of nuclear acid sequences via molecular biology experiments (such as Quantitative PCR (QPCR) , Fluorescence In Situ Hybridization (FISH) , DNA microarray, microfluidic chips, mass spectrometry, or Next Generation Sequencing (NGS) ) , and in particular, by using nuclear acid sequences that are common to at least two candidate species; generate a molecular identification code that is specific to each candidate species; and use the molecular identification code to determine if the candidate species, and specifically genetic materials originating from such species, are present in a sample.

[0004] The methods disclosed may be used for metagenomic survey and analysis of human, animals, plants, microorganisms, or viruses. Because the methods detect candidate species via the presence or absence of molecular features (i.e., nucleic acid sequences) in a combinatorial manner, the methods are capable of detecting a far larger number of candidate species than the number of molecular features, thus giving the methods great throughput and efficiency. These methods may be applied in a variety of molecular diagnostic platforms.

[0005] Accurately distinguishing organisms from different genera and species is essential to all biological sciences from taxonomy to public health. Current methods usually rely on phenotypical differences (in morphology or biochemistry, for example) or genetic differences. Detection methods that rely on phenotypic differences (such as Gram staining and biochemical reactions) cannot effectively distinguish between species from different genera with similar metabolic pathways (such as Pseudomonas and Burkholderia) , and each detection cycle can last as long as 48–72 hours.

[0006] High-throughput whole-genome sequencing can be used for species identification, but the high cost of equipment and the complexity of data analysis make it difficult to meet the needs of rapid on-site testing and large-scale screening.

[0007] There are also methods that rely on a single genetic marker, such as sequences in 16S or Internal Transcribed Spacer (ITS) rRNA. These methods, however, suffer from interferences due to the presence of homologous sequences in species of different genera. For example, homology between Klebsiella pneumoniae and Enterobacter aerogenes in the 16S rRNA is as high as 97%, and current methods based on primers binding to this region can only identify the species with an accuracy rate of about 72%.

[0008] There are also methods that rely on multiple genetic markers. They rely on specific sequences that are unique to each candidate species. For M candidate species each having N specific sequences, M × N assays need to be performed on a sample to test the presence of each sequence.

[0009] Overall, current methods suffer from several deficiencies. First, the primers employed in current methods are focused solely on conserved sequences within the target genus, and there is no mechanism for excluding sequences from non-target genera. For example, primers designed for Escherichia coli often produce false-positive amplification in Shigella (which has 95%homology with E. Coli) . Second, there is insufficient combinatorial possibilities. PCR tests 3–5 pairs of primer can only generate 8–32possible combinations of outcomes, which cannot meet the requirements for cross-genus identification of diverse species. Three, existing bioinformatics tools (such as ClustalW) primarily optimize intra-genus sequence alignments and are inefficient in discovering cross-genus differences, and therefore likely miss potential signature sequences.

[0010] The methods disclosed herein utilizes in a combinatorial manner the sequences that are shared between or among the candidate species. For M candidate species, the required number of sequences to be tested is only log2 M in theory, which greatly reduces the number of assays that need be performed. For example, for 255 candidate species and a negative result indicating the absence of all 255 candidate species, current methods need to test the presence of at least 255 species-specific sequences, while the methods disclosed herein only require log2 (255 + 1) = 8 sequences to be tested in an ideal situation. Practically, more than 8 sequences may be required, but the number is nonetheless significantly less than 255.

[0011] The methods disclosed herein can bring high-throughput capabilities to existing tools and techniques of molecular biology. For example, current methods of detecting pathogenic microorganisms by PCR are incapable of processing a large number of samples and are limited to detecting about a dozen species at most. The methods disclosed herein may be combined with techniques such as PCR, FISH, or DNA microarray to give them high-throughput capabilities.

[0012] In addition, the methods disclosed herein have many other advantages. For example, the methods disclosed herein have simpler PCR protocols because there are significantly fewer sequences to be tested. In contrast, methods such as targeted next generation sequencing (tNGS) require multiple rounds of PCR amplification that are difficult to design and implement.

[0013] The methods disclosed herein detect a single species via a combination of events involving the detection of multiple sequences, these methods therefore enjoy increased detection sensitivity and specificity.

[0014] The methods disclosed herein require less data analysis for metagenomic species identification, which translates to cost savings.

[0015] The methods disclosed herein require fewer sequences to be tested per sample, the minimal amount of sample material required is also less.

[0016] The methods disclosed herein, when combined with PCR or DNA microarray, can detect the same number of pathogenic microorganisms in a manner that is simpler, cheaper, and faster when compared with tNGS methods.

[0017] The methods disclosed herein may be implemented in microfluidics as miniaturized point-of-care testing products with high-throughput capabilities.SUMMARY

[0018] Some implementations herein relate to a method of analyzing in a sample comprising polynucleic acid originating from a plurality of candidate species. For example, method may include detecting if each one of a set of target sequences is present in the polynucleic acid. Method may also include generating a molecular identification code having a plurality of binary digits, where each binary digit in the molecular identification code encodes whether a corresponding target sequence in the set of target sequences is present in the polynucleic acid. Method may furthermore include comparing the molecular identification code with a list of known molecular identification codes, where each known molecular identification uniquely corresponds to one of the plurality of the candidate species. Method may in addition include determining whether any one of the plurality of the candidate species is present in the sample. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.

[0019] The described implementations may also include one or more of the following features. Method where the polynucleic acid is DNA. Method where amplifying the polynucleic acid in the sample may include: amplifying the polynucleic acid in the sample in a polymerase chain reaction. Method where the polynucleic acid is RNA. Method may include amplifying the polynucleic acid in the sample. Method where detecting if each one of the set of target sequences is present in the polynucleic acid may include: detecting if each one of the set of target sequences is present in the polynucleic acid in a molecular biology experiment. Method where the molecular biology experiment is quantitative PCR (QPCR) , fluorescence in situ hybridization (FISH) , mass spectrometry, or Next Generation Sequencing (NGS) . Method where the molecular biology experiment is conducted with a DNA microarray. Method where the molecular biology experiment is conducted with a microfluidic device. Method where the molecular biology experiment is conducted with a nanofluidic device. Method where each one of the set of target sequences is a polynucleic acid sequence that is present in more than one candidate species’ genome. Method where none of the set of target sequences is a polynucleic acid sequence that is present in all candidate species’ genome. Method may include generating a plurality of binary digits, where each binary digit encodes whether a corresponding target sequence in the set of target sequences is present in the polynucleic acid. Method where each binary digit in the molecular identification code is either one or zero, where one encodes that the corresponding target sequence is present in the polynucleic acid, zero encodes that the corresponding target sequence is absent from the polynucleic acid. Method where generating the molecular identification code having one or more of binary digits having: concatenating the one or more of binary digits together as a single string. Method where each of the known molecular identification codes having one or more of binary digits, and where each binary digit encodes whether the corresponding target sequence in the set of target sequences is present in a genome of the candidate species to which the known molecular identification code uniquely corresponds. Method where all known molecular identification codes may include an identical number of binary digits. Method where the molecular identification code may include an identical number of binary digits as a total number of target sequences in the set of target sequences. Method where a total number of the plurality of candidate species is less than or equal to two to a power of a total number of the plurality of binary digits in the molecular identification code. Method where a total number of the plurality of candidate species is less than or equal to two to a power of a total number of target sequences in the set of target sequences. Method may include determining which one of the plurality of the candidate species is present in the sample. Method where the sample originates from a human in a form of a nasal swab, a nasopharyngeal swab, a rectal swab, a vaginal swab, a penile swab, an ear swab, or a wound stab, and where each of the plurality of candidate species is a pathogenic microorganism. Method where the sample is a human bodily fluid. Method where the sample is a ruminal sample. Method where the sample is a soil sample. Method where the sample is a rhizosphere sample. Method where the sample is a marine or freshwater sample. Method where the sample is collected from goods, cargo, or passengers at a port of entry, and where each of the plurality of candidate species is an invasive species. Method where each of the plurality of candidate species is a prokaryote, eukaryote, bacteria, eubacteria, archaea, archaebacteria, fungus, plant, animal, protist, or virus species. Method where each of the plurality of candidate species is an insect, fish, bird, or mammal species. Implementations of the described techniques may include hardware, a method or process, or a computer tangible medium.

[0020] Some implementations herein relate to a method of determining a set of target sequences for detecting polynucleic acid originating from a plurality of candidate species. For example, method may include obtaining a genomic sequence of each of the plurality of candidate species. Method may also include generating a set of k-mers for each of the plurality of candidate species, where each k-mer is a continuous fragment of the genomic sequence. Method may furthermore include incorporating iteratively k-mers as target sequences into the set of target sequences, and calculating a molecular identification code for each of the plurality of candidate species each time the set of target sequences is changed, until each of the plurality of candidate species has a unique molecular identification code. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.

[0021] The described implementations may also include one or more of the following features. Method may include: obtaining a genomic sequence of a background species; and generating a set of k-mers for the background species, where the set of k-mers is generated in an identical manner as generating the set of k-mers for each of the plurality of candidate species. Method may include: excluding from the set of target sequences any k-mer that is identical to any k-mer in the set of k-mers for the background species. Method may include: calculating a molecular identification code for the background species each time the set of target sequences is changed; and determining if the molecular identification code for the background species is identical to a molecular identification code for any of the plurality of candidate species. Method may include: incorporating one or more additional k-mers as target sequences into the set of target sequences, until the molecular identification code for the background species is different from all molecular identification codes for the plurality of candidate species. Method where the background species and one of the plurality of candidate species are in a same biological taxonomic unit. Method where the same biological taxonomic unit is a kingdom, a phylum, a class, an order, a family, or a genus. Method where the polynucleic acid is DNA, RNA, or a combination thereof. Method where each of the plurality of candidate species is a prokaryote, eukaryote, bacteria, eubacteria, archaea, archaebacteria, fungus, plant, animal, protist, or virus species. Method where each of the plurality of candidate species is an insect, fish, bird, or mammal species. Method where obtaining the genomic sequence of each of the plurality of candidate species may include: downloading the genomic sequence of each of the plurality of candidate species from a database. Method where the database is an Internet database of genomic sequences. Method where the database is GenBank. Method where generating the set of k-mers for each of the plurality of candidate species may include: generating a k-mer which is a continuous fragment of the genomic sequence of a candidate species, where the continuous fragment is of a predetermined size, and the continuous fragment starts from a locus in the genomic sequence; and generating one or more further k-mers, where each of the one or more further k-mers starts from a new locus that is shifted by a predetermined step size from a previous locus in a same direction. Method where the predetermined step size is one base. Method may include: removing duplicative k-mers within the set of k-mers for each of the plurality of candidate species and retaining only one copy of such duplicative k-mers, where duplicative k-mers are k-mers having an identical sequence. Method may include: removing any k-mer whose sequence is only present in a genomic sequence of a single candidate species of the plurality of candidate species. Method may include: removing any k-mer whose sequence is present in a genomic sequence of every one of the plurality of candidate species. Method where incorporating iteratively k-mers into the set of target sequences may include: incorporating iteratively k-mers, one at a time, into the set of target sequences, in an order based on a preference of each k-mer, where a more preferred k-mer is incorporated into the set of target sequences before a less preferred k-mer. Method where the preference of each k-mer is determined based on how closely the k-mer can exactly bifurcate the plurality of candidate species; where a k-mer can exactly bifurcate the plurality of candidate species if the k-mer can divide the plurality of candidate species into two sets of equal cardinalities, where one set may include the candidate species whose genome sequence may include a sequence of the k-mer, and the other set may include the candidate species whose genome sequence does not may include the sequence of the k-mer; and where a k-mer that can exactly bifurcate the plurality of candidate species has the highest preference. Method wherein the preference of each k-mer is an integer p, a smaller value of p is preferred over a larger value of p, and a k-mer has a preference of p when the k-mer can divide the plurality of candidate species into two sets, wherein one set comprises the candidate species whose genome sequence comprises a sequence of the k-mer, and the other set comprises the candidate species whose genome sequence does not comprise the sequence of the k-mer; wherein the two sets consisting of and candidate species respectively when n is even, or and species respectively when n is odd; and wherein n is a total number of the plurality of candidate species. The method where the preference of each k-mer is a ratio r, a smaller value of r is preferred over a larger value of r, and and wherein a is a number of candidate species whose genome sequence comprises a sequence of the k-mer, and n is the number of the plurality of candidate species. The method where calculating the molecular identification code for each of the plurality of candidate species may include: for each k-mer in the set of target sequences, determining if the k-mer is present in the genomic sequence of the candidate species, and generating a binary digit of one or zero, where one encodes that the k-mer is present, and zero encodes that the k-mer is absent; and concatenating all binary digits together as a single. Method may include optimizing the set of target sequences to reduce a total number of k-mers in the set.

[0022] Some implementations herein relate to a system for analyzing in a sample having polynucleic acid originating from a plurality of candidate species. For example, system may include a thermal cycler configured for amplifying the polynucleic acid in the sample in a polymerase chain reaction. System may furthermore include a DNA microarray configured for detecting if each one of a set of target sequences is present in an amplified product. System may in addition include an optical scanner configured for reading results from the DNA microarray. System may moreover include a computer having at least one processor, and memory for storing instructions which, when executed by the at least one processor, result in operations comprising: generating a molecular identification code having a plurality of binary digits, where each binary digit in the molecular identification code encodes whether a corresponding target sequence in the set of target sequences is present in the polynucleic acid; comparing the molecular identification code with a list of known molecular identification codes, where each known molecular identification uniquely corresponds to one of the plurality of the candidate species; and determining whether any one of the plurality of the candidate species is present in the sample. System may include an interface configured for receiving the results from the optical scanner.

[0023] Some implementations herein relate to a system for designing a set of target sequences for detecting polynucleic acid originating from a plurality of candidate species. System may at least one processor. System may furthermore include memory for storing instructions which, when executed by the at least one processor, result in operations comprising: generating a set of k-mers for each of the plurality of candidate species, where each k mer is a continuous fragment of a genomic sequence of each of the plurality of candidate species; and incorporating iteratively k-mers as target sequences into the set of target sequences, and calculating a molecular identification code for each of the plurality of candidate species each time the set of target sequences is changed, until each of the plurality of candidate species has a unique molecular identification code.BRIEF DESCRIPTION OF THE DRAWINGS

[0024] FIG. 1 shows a flowchart of a method of species identification using sequences that are common to two or more candidate species.

[0025] FIGS. 2A, 2B, and 2C depict example systems for implementing certain steps in the methods of species identification disclosed herein.DETAILED DESCRIPTION

[0026] 1. Introduction

[0027] This disclosure describes methods of detecting and identifying candidate species that may be present in a sample using nucleic acid sequences that are common to two or more candidate species. This section describes certain general terminology and specific terms that are referred to in later sections of the disclosure.

[0028] 1.1 General Terminology

[0029] Throughout this specification and claims, the following definitions, general statements, and illustrations are applicable.

[0030] The patents, published applications, and scientific literature referred to herein establish the knowledge of those with skill in the art and are hereby incorporated by reference in their entireties to the same extent as if each were specifically and individually indicated to be incorporated by reference. Any conflict between any reference cited herein and the specific teachings of this specification shall be resolved in favor of the latter. Likewise, any conflict between an art-understood definition of a word or phrase and a definition of the word or phrase as specifically taught in this specification shall be resolved in favor of the latter.

[0031] As used herein, whether in a transitional phrase or in the body of a claim, the terms “comprise (s) ” and “comprising” are to be interpreted as having an open-ended meaning. That is, the terms are to be interpreted synonymously with the phrases “having at least” or “including at least. ” When used in the context of a process, the term “comprising” means that the process includes at least the recited steps, but may include additional steps. When used in the context of a composition, the term “comprising” means that the composition includes at least the recited features or components, but may also include additional features or components.

[0032] The terms “consists essentially of” or “consisting essentially of” have a partially closed meaning, that is, they do not permit inclusion of steps or features or components which would substantially change the essential characteristics of a process or composition; for example, steps or features or components which would significantly interfere with the desired properties of the compounds or compositions described herein, i.e., the process or composition is limited to the specified steps or materials and those which do not materially affect the basic and novel characteristics of the process or composition.

[0033] The terms “consists of” and “consists” are closed terminology and allow only for the inclusion of the recited steps or features or components.

[0034] Also, as used herein, the terms “has, ” “have, ” “having, ” or the like are intended to be open-ended terms. Further, the phrase “based on” is intended to mean “based, at least in part, on” unless explicitly stated otherwise.

[0035] As used herein, the singular forms “a” and “an” specifically also encompass the plural forms of the terms to which they refer, unless the content clearly dictates otherwise. Conversely, a term in its plural form may also encompass the singular form of the term, unless the content clearly dictates otherwise. Where only one item is intended, the phrase “only one” or similar language is used.

[0036] Further, as used herein, the article “the” is intended to include one or more items referenced in connection with the article “the” and may be used interchangeably with “the one or more. ”

[0037] Furthermore, as used herein, the term “set” is intended to include one or more items (e.g., related items, unrelated items, a combination of related and unrelated items, and / or the like) , and may be used interchangeably with “one or more. ”

[0038] The term “about” is used herein to mean approximately, in the region of, roughly, or around. When the term “about” is used in conjunction with a numerical range, it modifies that range by extending the boundaries above and below the numerical values set forth. In general, the term “about” or “approximately” is used herein to modify a numerical value above and below the stated value by a variance of 20%.

[0039] As used herein, the recitation of a numerical range for a variable is intended to convey that the variable can be equal to any values within that range. Thus, for a variable which is inherently discrete, the variable can be equal to any integer value of the numerical range, including the end-points of the range. Similarly, for a variable which is inherently continuous, the variable can be equal to any value of the numerical range, including the end-points of the range. As an example, a variable which is described as having values between 0 and 2, can be 0, 1 or 2 for variables which are inherently discrete, and can be 0. 0, 0.1, 0. 01, 0. 001, or any other value for variables which are inherently continuous.

[0040] Also, as used herein, the term “or” is intended to be inclusive when used in a series and may be used interchangeably with “and / or, ” unless explicitly stated otherwise (e.g., if used in combination with “either” or “only one of” ) .

[0041] 1.2 Genome

[0042] A “genome” generally refers to the genomic information or material from an organism, which can be, for example, at least a portion of, or the entirety of, the organism’s hereditary information. A genome can include coding sequences (e.g., that code for proteins) as well as non-coding sequences. A genome can include the sequences of some or all of the organism’s chromosomes. For example, the human genome ordinarily has a total of 46 chromosomes. The sequences of some or all of these chromosomes can constitute the genome.

[0043] 1.3 Nucleic Acid and Nucleotide

[0044] The terms “nucleic acid” and “nucleotide” are intended to be consistent with their use in the art and to include naturally-occurring species or functional analogs thereof. Particularly useful functional analogs of nucleic acids are capable of hybridizing to a nucleic acid in a sequence-specific fashion (e.g., capable of hybridizing to two nucleic acids such that ligation can occur between the two hybridized nucleic acids) or are capable of being used as a template for replication of a particular nucleotide sequence. Naturally-occurring nucleic acids generally have a backbone containing phosphodiester bonds. An analog structure can have an alternate backbone linkage including any of a variety of those known in the art. Naturally-occurring nucleic acids generally have a deoxyribose sugar (e.g., found in deoxyribonucleic acid (DNA) ) or a ribose sugar (e.g., found in ribonucleic acid (RNA) ) .

[0045] A nucleic acid can contain nucleotides having any of a variety of analogs of these sugar moieties that are known in the art. A nucleic acid can include native or non-native nucleotides. In this regard, a native deoxyribonucleic acid can have one or more bases selected from the group consisting of adenine (A) , thymine (T) , cytosine (C) , or guanine (G) , and a ribonucleic acid can have one or more bases selected from the group consisting of uracil (U) , adenine (A) , cytosine (C) , or guanine (G) . Useful non-native bases that can be included in a nucleic acid or nucleotide are known in the art.

[0046] 1.4 Polynucleotide

[0047] The terms “polynucleotide” is used to refer to a single-stranded or double-stranded multimer of nucleotides of any length. Polynucleotides may be natural or synthetic. Polynucleotides can include ribonucleotide monomers (i.e., can be polyribonucleotides) and / or deoxyribonucleotide monomers (i.e., polydeoxyribonucleotides) . In some embodiments, polynucleotides can include a combination of both deoxyribonucleotide monomers and ribonucleotide monomers (e.g., random or ordered combination of deoxyribonucleotide monomers and ribonucleotide monomers) . A polynucleotide can be of any length that is more than 1, for example, it can be of 2 to 10, 11 to 20, 21 to 30, 31 to 40, 41 to 50, 51 to 60, 61 to 70, 71 to 80, 81 to 90, 91 to 100, 100 to 150, 150 to 200, 200 to 250, 250 to 300, 300 to 350, 350 to 400, or 400to 500 nucleotides in length. Polynucleotides can include one or more functional moieties that are attached (e.g., covalently or non-covalently) to the multimer structure. For example, a polynucleotide can include one or more detectable labels (e.g., a radioisotope or fluorophore) .

[0048] 1.5 Sequence and k-mer

[0049] The term “sequence” refers to a fragment of a polynucleotide, and especially the polynucleotide’s information content. The sequence of a polynucleotide is the information regarding the identities of nucleotides, one by one, and in order, in the polynucleotide. A sequence of k nucleotides is referred to as a “k-mer. ” In some embodiments, k is 2, and the k-mers are dimers. In some embodiments, k is 3, and the k-mers are trimers. In some embodiments, k is 4, and the k-mers are tetramers. In some embodiments, k is 5, and the k-mers are pentamers. In some embodiments, k is 6, and the k-mers are hexamers. In some embodiments, k is 7, and the k-mers are heptamers. In some embodiments, k is 8, and the k-mers are octamers. In some embodiments, k is 9, and the k-mers are nonamers. In some embodiments, k is 10, and the k-mers are decamers.

[0050] 1.6 Species

[0051] As used herein, a “species” may be a species of an animal (e.g., human or a non-human animals) , a plant, a bacterium, a fungus, or any other living organism, or a virus. Examples of species include, but are not limited to, a mammal such as a rodent, mouse, rat, bat, rabbit, guinea pig, ungulate, horse, sheep, pig, goat, cow, cat, dog, primate (i.e. human or non-human primate) ; a plant such as Arabidopsis thaliana, corn, sorghum, oat, wheat, rice, canola, or soybean; an algae such as Chlamydomonas reinhardtii; a nematode such as Caenorhabditis elegans; an insect such as Drosophila melanogaster, mosquito, fruit fly, or honey bee; an arachnid such as a spider; a fish such as zebrafish; a reptile; an amphibian such as a frog or Xenopus laevis; an amoeba such as Dictyostelium discoideum; a fungi such as Pneumocystis carinii, Takifugu rubripes, yeast, Saccharamoyces cerevisiae, or Schizosaccharomyces pombe; a protozoa such as Plasmodium falciparum; or a virus such as SARS‐CoV‐2. In some embodiments, a species may refer to a subspecies, a variety, a subvariety, a cultivar, a form, a subform, a breed, a race, or a strain at a taxonomic rank below a species.

[0052] 1.7 Candidate Species and Background Species

[0053] Certain aspects of the methods disclosed herein pertain to “candidate species” and “background species. ” As used herein, “candidate species” are those species whose presence or absence (and more specifically, the presence of genetic materials originating from such candidate species) is of interest, and is to be determined by the methods disclosed herein. Candidate species may be one species or multiple species. In some embodiments, the number of candidate species is approximately 3, 4, 5, 6, 7, 8, 9, 10, 12, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 500, 1,000, 2,000, 5,000, or 10,000. In some embodiments, the number of candidate species is between a minimum number and a maximum number; wherein the minimum number is 1, 2, 3, or 4; and the maximum number is approximately 24, 25, 26, 27, 28, 29, 210, 212, 216, or 220. Candidate species may consist of one or more of the species describe in Section 1. 6 above. The methods disclosed herein, when executed on a sample, will determine the presence or absence of the candidate species. In some embodiments, the methods disclosed herein will determine the presence or absence of any one of the candidate species. In some embodiments, the methods disclosed herein will determine the presence or absence of each one of the candidate species.

[0054] “Background species” are the species whose presence or absence in a sample is not relevant to the purpose of the test, but whose presence may lead to “false positive” results, meaning the presence of a background species is incorrectly detected and interpreted as the presence of a candidate species. In some embodiments, the background species consists of the non-candidate species that are in the same domain of life as the candidate species. In some embodiments, the background species consists of all of the prokaryote, eukaryote, bacteria, eubacteria, archaea, archaebacteria, fungus, plant, animal, protist, or virus species, and excluding all candidate species. In some embodiments, the background species consists of the non-candidate species that are in a same taxonomic unit as the candidate species. In some embodiments, the background species consists of the non-candidate species that are in a same kingdom, phylum, class, order, family, or genus as the candidate species. In some embodiments, the background species consists of all species of insects, human parasites, cattle parasites, human pathogenic organisms, or animal pathogenic organisms, and excluding all candidate species.

[0055] The methods disclosed herein, when executed on a sample, will determine the presence or absence of any or all of the candidate species. In some embodiments according to the methods disclosed herein, the possibility of a “false positive” detection of a candidate species is minimized, given that some or all of the background species may be present in the sample.

[0056] 1.8 Molecular Identification Code and Target Sequence

[0057] Certain aspects of the methods disclosed herein pertain to molecular identification code. As used herein, the term “molecular identification code” is an identifier assigned to a candidate species. In some embodiments, the molecular identifier code may be represented as a string of binary digits, wherein each binary digit represents whether the species’ genome contains a certain specific sequence. In some embodiments, “1” represents the presence of the specific sequence in the genome of a candidate species, and “0” represents the absence of the specific sequence in the genome.

[0058] In some embodiments, a molecular identification code of a certain number of digits encodes the presence of absence of the same number of specific sequences in the genome of a particular species. Each one of the specific sequences whose presence or absence is encoded in the molecular identification code is herein termed a “target sequence” or “target k-mer. ”

[0059] In some embodiments, some or all of the target sequences are sequences that are common to two or more candidate species.

[0060] 1.9 Sample

[0061] As used herein, a “sample” is a specimen in which the presence or absence of the candidate species (and more specifically, the presence of genetic materials originating from such candidate species) is of interest, and is to be determined by the methods disclosed herein. The samples for use with the methods disclosed herein can be of any type as long as it contains some genetic material (such as polynuclear acid) that is capable of being analyzed by techniques in molecular biology (such as such as QPCR, FISH, DNA microarray, microfluidic chips, mass spectrometry, or NGS) . In some embodiments, the methods disclosed herein may also employ “blank” samples that do not contain any polynuclear acid or other analytes, and such “blank” samples may be used as negative controls, and to reduce the likelihood of “false positive” results.

[0062] The samples for use with the method disclosed herein can be derived from a homogeneous culture or population of the organisms mentioned herein or alternatively from a collection of several different organisms, for example, in a community or ecosystem.

[0063] In some embodiments, the samples for use with the methods disclosed provided herein encompass without limitation, an animal sample (e.g., mammal, reptile, bird) , soil (e.g., rhizosphere) , air, water (e.g., marine, freshwater, wastewater sludge) , sediment, oil, agricultural product, plant, and extreme environmental sample (e.g., acid mine drainage, hydrothermal systems) . In the case of marine or freshwater samples, the sample can be from the surface of the body of water, or any depth of the body water, e.g., a deep-sea sample. The water sample, in one embodiment, is an ocean, river or lake sample.

[0064] In some embodiments, the samples for use with the methods disclosed herein can be of any type that includes a microbial community.

[0065] In some embodiments, the sample is an animal sample in the form of a body fluid. In another embodiment, the animal sample is a tissue sample. Non-limiting animal samples include tooth, perspiration, fingernail, skin, hair, feces, urine, semen, mucus, saliva, gastrointestinal tract. The animal sample can be, for example, a human, primate, bovine, porcine, canine, feline, rodent (e.g., mouse or rat) , or bird sample. In one embodiment, the bird sample comprises a sample from one or more chickens. In another embodiment, the sample is a human sample. In some embodiments, the human sample may be collected in a healthcare context, such as from a patient. The human sample may be in the form of a nasal swab, a nasopharyngeal swab, a rectal swab, a vaginal swab, a penile swab, an ear swab, or a wound stab. The human microbiome comprises the collection of microorganisms found on the surface and deep layers of skin, in mammary glands, saliva, oral mucosa, conjunctiva and gastrointestinal tract. In some embodiments, the human sample may comprise pathogenic microorganisms. The microorganisms found in the microbiome include bacteria, fungi, protozoa, viruses, and archaea. Different parts of the body exhibit varying diversity of microorganisms. The quantity and type of microorganisms may signal a healthy or diseased state for an individual. The number of bacteria taxa are in the thousands, and viruses may be as abundant. The bacterial composition for a given site on a body varies from person to person, not only in type, but also in abundance or quantity.

[0066] In another embodiment, the sample is a ruminal sample. Ruminants such as cattle rely upon diverse microbial communities to digest their feed. These animals have evolved to use feed with poor nutritive value by having a modified upper digestive tract (reticulorumen or rumen) where feed is held while it is fermented by a community of anaerobic microbes. The rumen microbial community is very dense, with about 3 × 1010 microbial cells per milliliter. Anaerobic fermenting microbes dominate in the rumen. The rumen microbial community includes members of three domains of life: bacteria, archaea, and eukarya. Ruminal fermentation products are required by their respective hosts for body maintenance and growth, as well as milk production. Moreover, milk yield and composition has been reported to be associated with ruminal microbial communities. Ruminal samples, in one embodiment, are collected via the process described in Jewell et al., Appl. Environ. Microbiol. 81 (14) : 4697–4710 (2015) .

[0067] In another embodiment, the sample is a soil sample (e.g., bulk soil or rhizosphere sample) . It has been estimated that 1 gram of soil contains tens of thousands of bacterial taxa, and up to 1 billion bacteria cells as well as about 200 million fungal hyphae. Bacteria, actinomycetes, fungi, algae, protozoa, and viruses are all found in soil. Soil microorganism community diversity has been implicated in the structure and fertility of the soil microenvironment, nutrient acquisition by plants, plant diversity and growth, as well as the cycling of resources between above-and below-ground communities. Accordingly, assessing the microbial contents of a soil sample over time and the co-occurrence of active microorganisms (as well as the number of the active microorganisms) provides insight into microorganisms associated with an environmental metadata parameter such as nutrient acquisition and / or plant diversity.

[0068] The soil sample in one embodiment is a rhizosphere sample, i.e., the narrow region of soil that is directly influenced by root secretions and associated soil microorganisms. The rhizosphere is a densely populated area in which elevated microbial activities have been observed and plant roots interact with soil microorganisms through the exchange of nutrients and growth factors. As plants secrete many compounds into the rhizosphere, analysis of the organism types in the rhizosphere may be useful in determining features of the plants which grow therein.

[0069] In another embodiment, the sample is a marine or freshwater sample. Ocean water contains up to one million microorganisms per milliliter and several thousand microbial types. These numbers may be an order of magnitude higher in coastal waters with their higher productivity and higher load of organic matter and nutrients. Marine microorganisms are crucial for the functioning of marine ecosystems; maintaining the balance between produced and fixed carbon dioxide; production of more than 50%of the oxygen on Earth through marine phototrophic microorganisms such as Cyanobacteria, diatoms and pico-and nanophytoplankton; providing novel bioactive compounds and metabolic pathways; ensuring a sustainable supply of seafood products by occupying the critical bottom trophic level in marine food webs. Organisms found in the marine environment include viruses, bacteria, archaea and some eukarya. Marine viruses may play a significant role in controlling populations of marine bacteria through viral lysis. Marine bacteria are important as a food source for other small microorganisms as well as being producers of organic matter. Archaea found throughout the water column in the ocean are pelagic Archaea and their abundance rivals that of marine bacteria.

[0070] In another embodiment, the sample comprises a sample from an extreme environment, i.e., an environment that harbors conditions that are detrimental to most life on Earth. Organisms that thrive in extreme environments are called extremophiles. Though the domain Archaea contains well-known examples of extremophiles, the domain bacteria can also have representatives of these microorganisms. Extremophiles include: acidophiles which grow at pH levels of 3 or below; alkaliphiles which grow at pH levels of 9 or above; anaerobes such as Spinoloricus Cinzia which does not require oxygen for growth; cryptoendoliths which live in microscopic spaces within rocks, fissures, aquifers and faults filled with groundwater in the deep subsurface; halophiles which grow in about at least 0. 2M concentration of salt; hyperthermophiles which thrive at high temperatures (about 80–122 ℃. ) such as found in hydrothermal systems; hypoliths which live underneath rocks in cold deserts; lithoautotrophs such as Nitrosomonas europaea which derive energy from reduced mineral compounds like pyrites and are active in geochemical cycling; metallotolerant organisms which tolerate high levels of dissolved heavy metals such as copper, cadmium, arsenic and zinc; oligotrophs which grow in nutritionally limited environments; osmophiles which grow in environments with a high sugar concentration; piezophiles (or barophiles) which thrive at high pressures such as found deep in the ocean or underground; psychrophiles / cryophiles which survive, grow and / or reproduce at temperatures of about -15 ℃ or lower; radioresistant organisms which are resistant to high levels of ionizing radiation; thermophiles which thrive at temperatures between 45–122 ℃; xerophiles which can grow in extremely dry conditions. Polyextremophiles are organisms that qualify as extremophiles under more than one category and include thermoacidophiles (prefer temperatures of 70–80° C. and pH between 2 and 3) . The Crenarchaeota group of Archaea includes the thermoacidophiles.

[0071] The sample can include microorganisms from one or more domains. For example, in one embodiment, the sample comprises a heterogeneous population of bacteria and / or fungi.

[0072] In some embodiments, the sample may be collected for detecting and identifying invasive species. Such samples, for example, may be collected from goods, cargo, or passengers at a port of entry by customs personnel. As used herein, the term “invasive species” is used here to refer to any animal or plants that is classified as invasive or considered undesirable, unnecessary, or harmful to the environment in which it is located or to other vegetation, animal, or persons in proximity.

[0073] 1.10 PCR

[0074] A “PCR amplification” or “PCR” refers to the use of a polymerase chain reaction to generate copies of genetic material, including DNA and RNA sequences. Suitable reagents and conditions for implementing PCR are described, for example, in U.S. Patent Nos. 4,683,202, 4,683,195, 4,800,159, 4,965,188, and 5,512,462. In a typical PCR amplification, the reaction mixture includes the genetic material to be amplified, an enzyme, one or more primers that are employed in a primer extension reaction, and reagents for the reaction. The oligonucleotide primers are of sufficient length to provide for hybridization to complementary genetic material under annealing conditions. The length of the primers generally depends on the length of the amplification domains, but will typically be at least 4 bases and can be as long as 40 bases or longer, where the length of the primers will generally range from 18 to 50bases. The genetic material can be contacted with a single primer or a set of two primers (forward and reverse primers) , depending upon whether primer extension, linear or exponential amplification of the genetic material is desired.

[0075] In some embodiments, the PCR amplification process uses a DNA polymerase enzyme. The DNA polymerase activity can be provided by one or more distinct DNA polymerase enzymes. In certain embodiments, the DNA polymerase enzyme is from a bacterium, e.g., the DNA polymerase enzyme is a bacterial DNA polymerase enzyme. For instance, the DNA polymerase can be from a bacterium of the genus Escherichia, Bacillus, Thermus, or Pyrococcus.

[0076] Suitable examples of DNA polymerases that can be used include, but are not limited to: E. coli DNA polymerase I, Bsu DNA polymerase, Bst DNA polymerase, Taq DNA polymerase, VENTTM DNA polymerase, DEEPVENTTM DNA polymerase,  Taq DNA polymerase,  Hot Start Taq DNA polymerase, Crimson Taq DNA polymerase, Crimson Taq DNA polymerase,  DNA polymerase,  DNA polymerase, Hemo DNA polymerase,  DNA polymerase,  DNA polymerase,  High-Fidelity DNA polymerase, Platinum Pfx DNA polymerase, AccuPrime Pfx DNA polymerase, Phi29 DNA polymerase, Klenow fragment, Pwo DNA polymerase, Pfu DNA polymerase, T4 DNA polymerase and T7 DNA polymerase enzymes.

[0077] The term “DNA polymerase” includes not only naturally-occurring enzymes but also all modified derivatives thereof, including derivatives of naturally-occurring DNA polymerase enzymes. For instance, in some embodiments, the DNA polymerase is modified to remove 5’ -3’ exonuclease activity. Sequence-modified derivatives or mutants of DNA polymerase enzymes that can be used include, but are not limited to, mutants that retain at least some of the functional, e.g., DNA polymerase activity of the wild-type sequence.

[0078] Mutations can affect the activity profile of the enzymes, e.g., enhance or reduce the rate of polymerization, under different reaction conditions, e.g., temperature, template concentration, primer concentration, et cetera. Mutations or sequence-modifications can also affect the exonuclease activity and / or thermostability of the enzyme.

[0079] In some embodiments, PCR amplification can include reactions such as, but not limited to, a strand-displacement amplification reaction, a rolling circle amplification reaction, a ligase chain reaction, a transcription-mediated amplification reaction, an isothermal amplification reaction, and / or a loop-mediated amplification reaction.

[0080] In some embodiments, PCR amplification uses a single primer that is complementary to the 3’ tag of target DNA fragments. In some embodiments, PCR amplification uses a first and a second primer, where at least a 3’ end portion of the first primer is complementary to at least a portion of the 3’ tag of the target nucleic acid fragments, and where at least a 3’ end portion of the second primer exhibits the sequence of at least a portion of the 5’ tag of the target nucleic acid fragments. In some embodiments, a 5’ end portion of the first primer is non-complementary to the 3’ tag of the target nucleic acid fragments, and a 5’ end portion of the second primer does not exhibit the sequence of at least a portion of the 5’ tag of the target nucleic acid fragments. In some embodiments, the first primer includes a first universal sequence and / or the second primer includes a second universal sequence.

[0081] In some embodiments (e.g., when the PCR amplification amplifies captured DNA) , the PCR amplification products can be ligated to additional sequences using a DNA ligase enzyme. The DNA ligase activity can be provided by one or more distinct DNA ligase enzymes. In some embodiments, the DNA ligase enzyme is from a bacterium, e.g., the DNA ligase enzyme is a bacterial DNA ligase enzyme. In some embodiments, the DNA ligase enzyme is from a virus (e.g., a bacteriophage) . For instance, the DNA ligase can be T4 DNA ligase. Other enzymes appropriate for the ligation step include, but are not limited to, Tth DNA ligase, Taq DNA ligase, Thermococcus sp. (strain 9oN) DNA ligase (9oNTM DNA ligase, available from New England Biolabs, Ipswich, MA) , and  (available from Lucigen, Middleton, WI) . Derivatives, e.g., sequence-modified derivatives, and / or mutants thereof, can also be used.

[0082] In some embodiments, genetic material is amplified by reverse transcription polymerase chain reaction (RT-PCR) . The desired reverse transcriptase activity can be provided by one or more distinct reverse transcriptase enzymes (i.e., RNA dependent DNA polymerases) , suitable examples of which include, but are not limited to: M-MLV, MuLV, AMV, HIV, ArrayScriptTM, MultiScribeTM, ThermoScriptTM, and I, II, III, and IV enzymes. “Reverse transcriptase” includes not only naturally occurring enzymes, but all such modified derivatives thereof, including also derivatives of naturally-occurring reverse transcriptase enzymes.

[0083] In addition, reverse transcription can be performed using sequence-modified derivatives or mutants of M-MLV, MuLV, AMV, and HIV reverse transcriptase enzymes, including mutants that retain at least some of the functional, e.g., reverse transcriptase, activity of the wild-type sequence. The reverse transcriptase enzyme can be provided as part of a composition that includes other components, e.g., stabilizing components that enhance or improve the activity of the reverse transcriptase enzyme, such as RNase inhibitor (s) , inhibitors of DNA-dependent DNA synthesis, e.g., actinomycin D. Many sequence-modified derivative or mutants of reverse transcriptase enzymes, e.g., M-MLV, and compositions including unmodified and modified enzymes are commercially available, e.g., ArrayScriptTM, MultiScribeTM, ThermoScriptTM, and I, II, III, and IV enzymes.

[0084] Certain reverse transcriptase enzymes (e.g., Avian Myeloblastosis Virus (AMV) Reverse Transcriptase and Moloney Murine Leukemia Virus (M-MuLV, MMLV) Reverse Transcriptase) can synthesize a complementary DNA strand using both RNA (cDNA synthesis) and single-stranded DNA (ssDNA) as a template. Thus, in some embodiments, the reverse transcription reaction can use an enzyme (reverse transcriptase) that is capable of using both RNA and ssDNA as the template for an extension reaction, e.g., an AMV or MMLV reverse transcriptase.

[0085] In some embodiments, the quantification of RNA and / or DNA is carried out by real-time PCR (also known as quantitative PCR or qPCR) , using techniques well known in the art, such as but not limited to“TAQMANTM, ” or dyes such as or on capillaries ( “ Capillaries” ) . In some embodiments, the quantification of genetic material is determined by optical absorbance and with real-time PCR. In some embodiments, the quantification of genetic material is determined by digital PCR. In some embodiments, the genes analyzed can be compared to a reference nucleic acid extract (DNA and RNA) corresponding to the expression (mRNA) and quantity (DNA) in order to compare expression levels of the target nucleic acids.

[0086] 2. Methods of Species Identification

[0087] 2.1 Overview

[0088] In one embodiment, disclosed herein is a method that comprises the following steps:

[0089] constructing a genome database: collecting into a database the whole-genome sequences of the target species and other species in the same taxonomy unit (e.g., a genus, order or family) ;

[0090] selecting characteristic sequences: selecting sequences that are common to multiple target species, and optionally, selecting sequences that are highly specific to individual target species;

[0091] designing primers;

[0092] amplifying samples by PCR;

[0093] detecting fluorescence signals, and assigning each signal as “1” or “0” depending on whether it exceeds or is below a baseline level; and

[0094] matching the fluorescence signals with species information known to be associated with the signals.

[0095] In some embodiments of the method, the primers are designed using Primer3.

[0096] In some embodiments of the method, the primers’s equences are designed or optimized such that they have a melting temperature Tm in the range of 55–65 ℃. In some embodiments of the method, the primers’s equences are designed or optimized such that the Tm difference between forward and reverse primers is within 2 ℃. In some embodiments of the method, the primers’s equences are designed or optimized such that the amplification products are 100–500 bp in length. In some embodiments, the primers are optimized for formation of dimer or hairpin structures (e.g., ΔG ≥ -5 kcal / mol) with OligoAnalyzer 3. 1.

[0097] In another embodiment, disclosed herein is a method that comprises the following steps:

[0098] (1) downloading genome sequences of candidate species and extract common sequences;

[0099] (2) downloading genome databases of bacteria, fungi, animals, plants, or viruses; and perform k-mer segmentation (i.e., splitting genomic sequences into k-mers) ;

[0100] (3) constructing an index of the genome databases to facilitate searching;

[0101] (4) selecting target sequences from common sequences, and arranging the target sequences until there is a unique molecular identification code for each candidate species;

[0102] (5) using the index of the genome databases to check if the molecular identification codes of the candidate species coincide with any background species; and removing any target sequences that cause code collisions;

[0103] (6) repeating steps 4) and 5) until each candidate species are assigned a unique molecular identification code;

[0104] (7) selecting appropriate molecular biology detection method based on the molecular identification codes;

[0105] (8) making DNA microarray chips according to the target sequences on which the molecular identification codes are based;

[0106] (9) performing whole-genome amplification of the genetic materials (DNA  / RNA) in a sample;

[0107] (10) hybridizing the amplified products with the DNA microarray chips;

[0108] (11) detecting fluorescent signals from the DNA microarray chips by a scanner, and constructing a molecular identification code sequence according to the fluorescent signals, where in the fluorescent signals indicate the presence or absence of the target sequences in the amplified products; and

[0109] (12) a candidate species is detected when the molecular identification code constructed from the fluorescent signals matches a known molecular identification code of the candidate species.

[0110] In some embodiments, for example, steps (9) to (11) may be performed according to the experimental protocol in Example 2.

[0111] FIG. 1 illustrates an exemplary embodiment of the method according to the present disclosure. In some embodiments, the method of the present application may include steps of experiment design, experiment execution, or both.

[0112] 2.2 Experiment Design Steps

[0113] In some embodiments, the experiment design steps may include a step of obtaining the genomic sequences 101 of the candidate species to be tested for presence or absence in a sample. In some embodiments, the genomic sequences are downloaded from an online database, such as GenBank. In some embodiments, the genomic sequences are downloaded as binary or text files. In some embodiments, the genomic sequences are downloaded as files in FASTA format.

[0114] In some embodiments, the experiment design steps may include a sequence splitting step 103, in which a computer processes each genomic sequence by splitting it into multiple fragments, each of which is termed a “k-mer, ” where k is the length of the fragment measured in bases. In some embodiments, k is 2, 3, 4, 5, 6, 7, 8, 9, or 10. In some embodiments, all of the k-mers are of the same lengths, or in other words, all of the k-mers contain the same number of bases or nucleotides. In some embodiments, the k-mers are 1000-mers, 900-mers, 800-mers, 600-mers, 500-mers, 400-mers, 300-mers, 200-mers, 100-mers, 90-mers, 80-mers, 70-mers, 60-mers, 50-mers, 40-mers, 30-mers, 25-mers, 20-mers, 15-mers, or 10-mers. In some embodiments, k is between about 5000 and about 4000, between about 4000 and about 3000, between about 3000 and about 2000, between about 2000 and about 1000, between about 1000 and about 500, between about 900 and about 400, between about 800 and about 300, between about 700 and about 200, between about 600 and about 100, between about 500 and about 50, between about 400 and about 40, between about 300 and about 30, between about 200 and about 20, between about 100 and about 10, between about 100 and about 50, or between about 50 and about 5.

[0115] In some embodiments, the computer processes each genomic sequence in by splitting it into multiple k-mers that do not overlap in sequence. In some embodiments, the computer processes each genomic sequence by splitting it into multiple k-mers that overlap in sequence. In some embodiments, the computer processes each genomic sequence by splitting it into multiple k-mers that overlap in a staggered fashion. In some embodiments, the computer processes each genomic sequence by splitting it into multiple k-mers at a step size of s, meaning that every pair of consecutive k-mers overlaps except for the first s bases of one k-mer and the last s bases of the other k-mer. In other words, every pair of consecutive k-mers originates from the same locus in the genome sequence, except that one k-mer is shifted by s bases relative to the other k-mer. In some embodiments, the step size s is less than k, the length of the k-mer. In some embodiments, the step size s is more than k. In some embodiments, the step size s is equal to k. In some embodiments, the step size in the number of bases is 1000, 900, 800, 700, 600, 500, 400, 300, 200, 100, 90, 80, 70, 60, 50, 40, 30, 25, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1. In some embodiments, the step size is between about 1000 and about 500, between about 900 and about 400, between about 800 and about 300, between about 700 and about 200, between about 600 and about 100, between about 500 and about 50, between about 400 and about 40, between about 300 and about 30, between about 200 and about 20, between about 100 and about 10, between about 100 and about 50, between about 90 and about 40, between about 80 and about 30, between about 70 and about 20, between about 60 and about 10, between about 50and about 1, between about 50 and about 40, between about 40 and about 30, between about 30 and about 20, between about 20 and about 10, between about 15 and about 5, between about 10 and about 5, or between about 5 and about 1. In some embodiments, the step size is 1.

[0116] In some embodiments, the computer processes each genomic sequence at step 103 by splitting it into multiple k-mers at a step size of 1, resulting in a collection of k-mers. In some embodiments, these k-mers are stored in one or more computer files 105 in FASTA format. In some embodiments, these k-mers are stored together in a single computer file. In some embodiments, these k-mers are stored individually in multiple computer files, each of which corresponds to a single k-mer. In some embodiments, the computer files are in binary or text files. In some embodiments, the computer files are in FASTA format.

[0117] In some embodiments, the experiment design steps may include a k-mer selection step 107, which compares the k-mers obtained from step 103 across the candidate species, and retains only the k-mers that are common to at least two of the candidate species whose presence or absence is to be determined in the sample to be tested. For example, a k-mer is common to two species when each of the two species contains the same k-mer sequence in its genome.

[0118] In some embodiments, the k-mer selection step may include eliminating any k-mer that is common to all of the candidate species, since detection of k-mers does not yield any information that can be used to distinguish the candidate species from one another.

[0119] In some embodiments, the k-mer selection step may include preferentially retaining any k-mer that is common to exactly half of the number of the candidate species. For example, if there are n candidate species whose presence or absence need to be determined from a sample, and n is an even number, a preferred k-mer is present in exactly in the genomes of species, and is absent in the genomes of the remaining species. If, for example, there are n candidate species and n is an odd number, a preferred k-mer is present in exactly in the genomes of species, and is absent in the genomes of the remaining species, or alternatively, a preferred k-mer is present in exactly in the genomes of species, and is absent in the genomes of the remaining species

[0120] In some embodiments, the k-mer selection step may include selecting k-mers according to their ranked preference. The most highly ranked k-mers can exactly bifurcate the set of candidate species, meaning each of such k-mers can divide the candidate species into two sets of equal numbers, wherein one set contains the species whose genome contains the sequence of the k-mer, and the other set contains the species whose genome does not contain the sequence of the k-mer. Other k-mers are ranked according to how close they are to, or how far they deviate from, the ideal target of exact bifurcation. In other words, k-mers are ranked higher in preference if they are closer to the target of exact bifurcation, and are ranked lower in preference if they are further away from the target of exact bifurcation.

[0121] In some embodiments, the k-mer selection step may include selecting k-mers according to their ranked preference, which is expressed as an integer. The most preferred k-mers have a preference of 0, and they each can divide n candidate species into two sets, wherein the two sets both consisting of species when n is even, or wherein the two sets consisting of and species respectively when n is odd, and one set contains the candidate species whose genome contains the sequence of the k-mer, and the other set contains the candidate species whose genome does not contain the sequence of the k-mer. The immediate lower level of preferred k-mers have a preference of 1 (which is less preferred than 0) , and they each can divide n candidate species into two sets, wherein the two sets consisting of and species respectively when n is even, or and species respectively when n is odd. Similarly, the next lower level of preferred k-mers have a preference of 2 (which is less preferred than 1) , and they each can divide n candidate species into two sets, wherein the two sets consisting of and species respectively when n is even, or and species respectively when n is odd. In general, each k-mer has a preference of p, when it can divide n candidate species into two sets, wherein the two sets consisting of and species respectively when n is even, or and species respectively when n is odd, and a k-mer having a smaller value of p is preferred over another k-mer having a larger value of p.

[0122] In some embodiments, the k-mer selection step 107 may include verifying that a k-mer is not present in the genome of any background species. In some embodiments, verifying that a k-mer is not present in the genome of any background species may involve steps that are analogous to the processing of the genomic sequences of the candidate species. For example, in some embodiments, the genomic sequences 119 of all background species are downloaded from an online database, such as GenBank. In some embodiments, the genomic sequences are downloaded as binary or text files. In some embodiments, the genomic sequences are downloaded as files in FASTA format. The genomic sequences 119 of all background species are then processed in a sequence splitting sequence step 121, in which a computer processes each genomic sequence by splitting it into multiple k-mers. In some embodiments, the sequence splitting step 121 is the same as the sequence splitting step 103, except that it is the genome sequences of the background species being processed in step 121. In some embodiments, the sequence splitting steps 121 and 103 use the same parameters, including the length of the fragments k and the step size s.

[0123] In some embodiments, verifying that a k-mer is not present in the genome of any background species may involve comparing k-mers 105 obtained from the candidate species, and k-mers 123 obtained from the background species, and eliminating k-mers that are common to a candidate species and a background species.

[0124] In some embodiments, n target k-mers are selected for detecting a maximum of 2n candidate species. In some embodiments, a minimum of log2 N target k-mers are used for detecting N candidate species. In some embodiments, the set of selected target k-mers comprises the minimum number of k-mers, and approximately 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 90%, 100%, 120%, 150%, 200%, 300%, 400%, 500%, or 1000%extra k-mers above the minimum number.

[0125] In some embodiments, the experiment design steps may include a step 109 of generating a molecular identifier code for each candidate species from a set of selected target k-mers. In some embodiments, generating a molecular identifier code for a candidate species from a set of selected target k-mers comprises determining if each target k-mer is present or absence in the genomic sequence of the candidate species; generating a binary digit “0” if the target k-mer is not present in the genomic sequence; generating a binary digit “1” if the target k-mer is not present in the genomic sequence; repeating these steps until the presence or absence of all target k-mers have been determined; and stringing together all binary digits to give the molecular identifier code.

[0126] In some embodiments, the binary digits in a molecular identification code are ordered based on the preference of the corresponding target sequence. In some embodiments, the first digit in a molecular identification code corresponds to the most preferred k-mer, and the subsequent digits corresponds to k-mers that are successively lower in preference. In some embodiments, the first digit in a molecular identification code corresponds to the least preferred k-mer, and the subsequent digits corresponds to k-mers that are successively higher in preference.

[0127] In some embodiments, n digits in a molecular identification code are used to code for a maximum of 2n candidate species. In some embodiments, a minimum of log2 N digits is used in a molecular identifier to code for N candidate species. In some embodiments, the molecular identification code comprises the minimum number of digits, and approximately 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 90%, 100%, 120%, 150%, 200%, 300%, 400%, 500%, or 1000%extra digits above the minimum number.

[0128] In some embodiments, the experiment design steps may include a step 111 of verifying that each candidate species has a unique molecular identification code, or in other words, each candidate species has a different molecular identification code. In some embodiments, if two or more candidate species have the same molecular identification code ( “code collision” ) , the experiment design steps may include repeating the k-mer selection step 107 to select a different set of target k-mers. In some embodiments, in case of code collision, the experiment design steps may include adding more target k-mers (i.e., adding more digits to the molecular identification code) . In some embodiments, in case of code collision, the experiment design steps may include swapping one target k-mer with another of high, equal, or lower preference. In some embodiments, in case of code collision, the experiment design steps may include eliminating from the set of target k-mers one or more k-mers, starting from the ones having the lowest preference; and then adding to the set of target k-mers an equal number of other k-mers produced in step 103, starting from the ones with the highest preference.

[0129] In some embodiments, step 111 may include an optimization process, in which a computer program seeks to shorten the length of the molecular identification code (i.e., decreasing the number of digits in the molecular identification code) . When two target k-mers are always present or absence at the same time in all of the candidate species, one of them can be eliminated. In other words, if two digits in all molecular identification codes are always “1” or “0” at the same time across all candidate species, only one digit it necessary, and the other digit may be eliminated in step 111 to shorten the length of the molecular identification code by one. From the optimization step 111, a final set of molecular identification codes 133 and the associated target k-mers are generated.

[0130] In certain embodiments, the experiment design steps may include a step 113 of selecting the appropriate method of detecting the target k-mers. In some embodiments, the presence of the target k-mer sequences in a sample may be detected via Quantitative PCR (QPCR) , Fluorescence In Situ Hybridization (FISH) , DNA microarray, microfluidic chips, mass spectrometry, or Next Generation Sequencing (NGS) ) .

[0131] 2.3 Experiment Execution Steps

[0132] In some embodiments, the methods of the present application may include steps of experiment execution, including sample preparation steps, sample testing steps, and data analysis steps. In some embodiments, the method of the present application may include sample preparation steps. In some embodiments, the sample preparation steps may include obtaining a sample 125 to be analyzed for the presence of genetic materials originating from certain candidate species. The sample 125 is then processed in step 127 to isolate, purify, and / or extract the nucleic acids (e.g., DNA or RNA) in the sample, which are amplified at the full-genome level in step 129, and result in the amplified products 131 to be tested according to the method disclosed herein.

[0133] In some embodiments, the methods of the present application may include sample testing steps. In some embodiments, the sample testing step 115 includes testing the products 131 for the presence of certain k-mer sequences. The testing, for example, may be conducted on the amplified products 131 via DNA hybridization probes in a microarray. The testing results-more specifically, the presence or absence of each k-mer sequence-may be indicated by presence or absence of optical signals (such as fluorescence) . Binary digits of “1” and “0, ” where they respectively represent the presence or absence of a certain k-mer sequence, may be used to record the results of the hybridization experiments. These digits of “1” sand “0”s, in combination, constitute a molecular identification code that is representative of the amplified products 131, and of the sample 125.

[0134] In some embodiments, the methods of the present application may include data analysis steps. In some embodiments, the data analysis step 117 includes matching the molecular identification code obtained from step 115 with known unique identification codes for the candidate species. When the molecular identification code read out from the sample is the same as the molecular identification codes for a particular species, it is determined that said species is present in sample 125. Under ideal conditions, the number of candidate species that may be identified in one round of experiments is two to the power of n (i.e., 2n) , where n is number of target k-mer sequences that underlie the molecular identification code.

[0135] 3. Computer-Implemented Methods

[0136] FIGS. 2A, 2B, and 2C depict example systems for implementing certain experiment design steps described in the preceding section 2. For example, FIG. 2A depicts an exemplary system 200 that includes a standalone computer architecture where a processing system 202 (e.g., one or more computer processors located in a given computer or in multiple computers that may be separate and distinct from one another) includes a computer-implemented sequence processing engine 204 being executed on the processing system 202. The processing system 202 has access to a computer-readable memory 207 in addition to one or more data stores 208. The one or more data stores 208 may include a first data structure 234 as well as a second data structure 236. In some embodiments, the data stores 208 may be a database system. In some embodiments, the first data structure comprises the genomic sequences of the candidate species, and the second data structure comprises the genomic sequences of the background species. In some embodiments, the first data structure comprises the genomic sequences of the candidate species and the background species; and the second data structure comprises the k-mer sequences that are selected and optimized for identifying the candidate species. In some embodiments, the first data structure comprises the k-mer sequences of the candidate species and the background species; and the second data structure comprises the molecular identification codes.

[0137] The processing system 202 may be a distributed parallel computing environment, which may be used to handle very large-scale data sets. In some embodiments, for example, the segmentation of genomic sequences into k-mers may be distributed to multiple processors on a per-species basis, and carried out in parallel.

[0138] FIG. 2B depicts a system 220 that includes a client-server architecture. One or more user client computers or devices 222 access one or more servers 224 running a data processing engine 237 on a processing system 227 via one or more networks 228. The one or more servers 224 may access a computer-readable memory 230 as well as one or more data stores 232. The one or more data stores 232may include a first data structure 234 as well as a second data structure 236. In some embodiments, the client 222 may be an instrument computer that transmits experimental data (e.g., data from DNA microarray experiments) to the server 224 for analysis and identification of species. In some embodiments, the client 222 may be an instrument computer that transmits experimental data (e.g., data from DNA microarray experiments) to the server 224 for analysis and identification of species. In some embodiments, the client 222 may be a personal computer or a mobile device on which a user specifies certain parameters (such as the candidate species) , and the server 224 performs the steps of experiment design.

[0139] FIG. 2C shows a block diagram of exemplary hardware for a standalone computer architecture 250, such as the architecture depicted in FIG. 2A that may be used to include and / or implement the program instructions of system embodiments of the present disclosure. A bus 252 may serve as the information highway interconnecting the other illustrated components of the hardware. A processing system 254labeled CPU (central processing unit) (e.g., one or more computer processors at a given computer or at multiple computers) , may perform calculations and logic operations required to execute a program. A non-transitory processor-readable storage medium, such as read only memory (ROM) 258 and random access memory (RAM) 259, may be in communication with the processing system 254 and may include one or more programming instructions for performing certain steps in the methods disclosed herein, such as the steps of experiment design and data analysis. Optionally, program instructions may be stored on a non-transitory computer-readable storage medium such as a magnetic disk, optical disk, recordable memory device, flash memory, or other physical storage medium.

[0140] In FIGS. 2A, 2B, and 2C, computer readable memories 207, 230, 258, 259 or data stores 208, 232, or disks 283, 284, 285 may include one or more data structures for storing various data used in the methods disclosed herein. For example, a data structure stored in any of the aforementioned locations may be used to store sequence data in FASTA format. A controller 290 interfaces one or more optional disk drives to the system bus 252. These disk drives may be external or internal disk drives such as 283, external or internal optical disc drives such as 284, or external or internal solid state drives 285. As indicated previously, these various disk drives and disk controllers are optional devices.

[0141] The system 250 may include a software application stored in one or more of the disk drives connected to the disk controller 290, the ROM 25 and / or the RAM 259. The processor 254 may access the software application stored on one or more of these components as required.

[0142] A graphical processing unit (GPU) 287 may permit information from the bus 252 to be displayed on a display 280 in audio, graphic, or alphanumeric format. In some embodiments, the computational power of the GPU 287 may be leveraged, especially for tasks that are easily divided and executed in parallel.

[0143] Communication with external devices may optionally occur using various communication ports 282. In certain embodiments, various laboratory instruments may communicate with the system 250 via the ports 282.

[0144] In addition to these computer-type components, the hardware may also include data input devices, such as a keyboard 279, or other input device, such as a laboratory plate reader 281, microphone, remote control, pointer, mouse, or joystick.

[0145] Additionally, the methods described herein may be implemented on many different types of processing devices by program code comprising program instructions that are executable by the device processing subsystem. The software program instructions may include source code, object code, machine code, or any other stored data that is operable to cause a processing system to perform the methods and operations described herein and may be provided in any suitable language such as C, C++, JAVA, Python, for example, or any other suitable programming language. Other implementations may also be used, however, such as firmware or even appropriately designed hardware configured to carry out the methods and systems described herein.

[0146] The data used in the methods disclosed herein (e.g., genomic sequences, k-mers, molecular identification codes, fluorescent signals, etc. ) may be stored and implemented in one or more different types of computer-implemented data stores, such as different types of storage devices and programming constructs (e.g., RAM, ROM, Flash memory, flat files, databases, programming data structures, programming variables, etc. ) . It is noted that data structures describe formats for use in organizing and storing data in databases, programs, memory, or other computer-readable media for use by a computer program.

[0147] The computer components, software modules, functions, data stores and data structures described herein may be connected directly or indirectly to each other in order to allow the flow of data needed for their operations. It is also noted that a module or processor includes but is not limited to a unit of code that performs a software operation, and can be implemented for example as a subroutine unit of code, or as a software function unit of code, or as an object (as in an object-oriented paradigm) , a package, a module, or a class, or in a computer script language, or as another type of computer code. The software components and / or functionality may be located on a single computer or distributed across multiple computers depending upon the situation at hand.

[0148] One or more aspects or features of the subject matter described herein can be realized in digital electronic circuitry, integrated circuitry, specially designed application specific integrated circuits (ASICs) , field programmable gate arrays (FPGAs) computer hardware, firmware, software, and / or combinations thereof. These various aspects or features can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device. The programmable system or computing system may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0149] These computer programs, which can also be referred to as programs, software, software applications, applications, components, or code, include machine instructions for a programmable processor, and can be implemented in a high-level procedural language, an object-oriented programming language, a functional programming language, a logical programming language, and / or in assembly / machine language. As used herein, the term “machine-readable medium” refers to any computer program product, apparatus and / or device, such as for example magnetic discs, optical disks, memory, and Programmable Logic Devices (PLDs) , used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor. The machine-readable medium can store such machine instructions non-transitorily, such as for example as would a non-transient solid-state memory or a magnetic hard drive or any equivalent storage medium. The machine-readable medium can alternatively or additionally store such machine instructions in a transient manner, such as for example as would a processor cache or other random access memory associated with one or more physical processor cores.

[0150] 4. Microfluidics-Implemented Methods

[0151] In some embodiments, the experiment execution steps in the methods disclosed here may be implemented in microfluidics or nanofluidics. As used herein, the terms “microfluidic” and “nanofluidic” in reference to devices are used interchangeably herein, each means an integrated system for capturing, moving, mixing, dispensing or analyzing small volumes of fluid, including samples (which, in turn, may contain or comprise cellular or molecular analytes of interest) , reagents, dilutants, buffers, or the like. Generally, reference to “microfluidics” and “nanofluidics” denotes different scales in the size of devices and volumes of fluids handled. In some embodiments, features of a microfluidic device have cross-sectional dimensions of less than a few hundred square micrometers and have passages, or channels, with capillary dimensions, for example, having maximal cross-sectional dimensions of from about 500 μm to about 0. 1 μm. In some embodiments, microfluidics devices have volume capacities in the range of from 1 μL to a few nL, e.g., 10–100 nL. Dimensions of corresponding features, or structures, in nanofluidics devices are typically from 1 to 3 orders of magnitude less than those for microfluidics devices. One skilled in the art would know from the circumstances of a particular application which dimensionality would be pertinent. In some embodiments, microfluidic or nanofluidic devices have one or more chambers, ports, and channels that are interconnected and in fluid communication and that are designed for carrying out one or more analytical reactions or processes, either alone or in cooperation with an appliance or instrument that provides support functions, such as sample introduction, fluid and / or reagent driving means, such as positive or negative pressure, acoustical energy, or the like, temperature control, detection systems, data collection and / or integration systems, and the like. In some embodiments, microfluidics and nanofluidics devices may further include valves, pumps, filters, and specialized functional coatings on interior walls, e.g., to prevent adsorption of sample components or reactants, facilitate reagent movement by electroosmosis, or the like. Such devices may be fabricated as an integrated device in a solid substrate, which may be glass, plastic, or other solid polymeric materials, and may have a planar format for ease of detecting and monitoring sample and reagent movement, especially via optical or electrochemical methods. In some embodiments, such devices are disposable after a single use. The fabrication and operation of microfluidics and nanofluidics devices are exemplified by the following references: Ramsey, U.S. Pat. Nos. 6,001,229; 5,858,195; 6,010,607; and U.S. Pat. No. 6,033,546; Soane et al, U.S. Pat. Nos. 5,126,022 and 6,054,034; Nelson et al, U.S. Pat. No. 6,613,525; Maher et al, U.S. Pat. No. 6,399,952; Ricco et al, International Patent Publication WO 02 / 24322; Bjornson et al, International Patent Publication WO 99 / 19717; Wilding et al, U.S. Pat. Nos. 5,587,128; 5,498,392; Sia et al, Electrophoresis, 24: 3563–3576 (2003) ; Unger et al, Science, 288: 113–116 (2000) ; Enzelberger et al, U.S. Pat. No. 6,960,437; Cao, “Nanostructures &Nanomaterials: Synthesis, Properties &Applications, ” (Imperial College Press, London, 2004) ; Haeberle et al, LabChip, 7: 1094–1110 (2007) ; Cheng et al, Biochip Technology (CRC Press, 2001) ; and the like.

[0152] 5. Examples

[0153] The following examples further illustrate the methods described herein. These are illustrative and are not to be considered limiting in any way whatsoever, as many modifications in methods will be apparent to those skilled in the art.

[0154] 5.1 Example 1 -Method of Species Identification

[0155] The following steps 1–9 describe an exemplary method for species identification using sequences that are common to candidate species.

[0156] 1. Download genome sequences of the candidate species to be detected.

[0157] Download the genomic sequences, in FASTA format, of the species to be detected from GenBank (https:  / / ftp. ncbi. nlm. nih. gov / genomes / genbank) according to the species names or their TaxIDs. Assign a sequential numerical identifier to each of the candidate species to be detected, and generate a set of associative relationships between the species names and the corresponding numerical identifier, SS = {species name: species i} , wherein 0 <i ≤ N, i is a natural number, and N is the total number of samples to be detected.

[0158] 2. Analyze and extract common sequences for candidate species.

[0159] The downloaded FASTA genome sequence file of each species is split at a fixed length k and a step length s. Each split sequence obtained is called a k-mer, where k and s are both natural numbers, k and s are both greater than 0.

[0160] For each k-mer of a candidate species, first construct a dictionary associative relationship {kmer: species name x} . The k-mers of all candidate species are then deduplicated to obtain a set of unique non-duplicative k-mers {uniq-kmer} of all candidate species to be detected. Finally, using {uniq-kmer} and {k-mer: species name x} , we can obtain information about the presence of a uniq-kmer in each candidate species to be tested, that is, {uniq-kmer: species 1, species 2, ... } . For example, if k is 8 and a uniq-kmer sequence is “ACTGTACG, ” { “ACTGTACG” : species 1, species 7, species 41} means that the uniq-kmer appears in species 1, species 7, and species 41, and this information will be used in the later steps.

[0161] 3. Download the whole-genome databases of bacteria, fungi, plants, and / or viruses.

[0162] Download the genomic sequences, in FASTA format, of all species in each domain of life from GenBank.

[0163] 4. Analyze and extract common sequences from the whole-genome databases.

[0164] Process the genome sequences from the whole-genome databases using the same method and parameters as in step 2 to obtain a set AS= {allSpecies-uniq-kmer: species 1, species 2, ... } .

[0165] 5. Design unique molecular identification code (MID) for species.

[0166] Using the set {uniq-kmer: species 1, species 2, ... } from step 2, for each “uniq-kmer: species 1, species 2, ..., ” generate a set of species {species 1, species 2, ... } . Let N be the number of species in the set, and M be the number of candidate species. In order to minimize the number of uniq-kmers need to cover all candidate species, uniq-kmers with N  / M = 0. 5 are preferentially selected for use in constructing the unique molecular identification codes (MID) for species, followed by selection of uniq-kmers with N =(M × 0. 5 ± k) , wherein k is 1, 2, 3, ... M × 0. 5 in turn.

[0167] Each time a uniq-kmer is included (as a target k-mer) , a species set S = {species 1, species 2, ... } is generated, and a set TS is expanded in the manner that TS = TS ∪ S (where TS is initially defined as an empty set) . The selection priority of the uniq-kmer is A = S ∩ TS, |A|  / |S|= 0. 5 ± k, where k is 0, 0.01, …0. 5. The selection of uniq-kmer (as target k-mers) is completed if and only if |TS| is equal to the number of candidate species, and the target k-mer set K = {uniq-kmer1, uniq-kmer2, …} is generated.

[0168] For each element ki in the set K, ki is either 0 or 1. If ki appears in a candidate species, it is 1, and otherwise it is 0. For all candidate species, the species MID keyX = “k1k2k3…kn” and the target k-mer set kmerX= {uniq-kmer} can then be generated, where X is the number of candidate species.

[0169] 6. Verify the uniqueness of MIDs across all species.

[0170] In order to prevent the target k-mer set kmerX obtained in step 5 from appearing in other species (i.e., background species) , it is necessary to use the set AS from step 4 to verify each member of kmerX one by one. For each element uniq-kmer in the set kmerX, extract a species information set Sn = {species 1, species 2, ... } associated with the AS set, n = |kmerX|. Define the set u = S1 ∩ S2 ∩…∩ Sn. If |u| > 0, it means that the species MID keyX (and the target k-mer set kmerX) overlap with some background species, which will cause interference in the actual detection process. If this happens, the following measures can be taken: (1) delete the species MID from the design, repeat steps 4 and 5 until all candidate species are assigned a unique MID; (2) investigate the characteristics of the background species that causes the collision with the candidate species, and if the background species has little effect on the test, or if the likelihood of false positive is very low, the MID may be retained in the experiment design; (3) remove the species as a candidate species for detection and delete the corresponding MID.

[0171] 7. Finalize the species MIDs and the set of target k-mer sequences.

[0172] From steps 1 to 5, a target k-mer set K = {uniq-kmer1, uniq-kmer2, …} is obtained. For each candidate species, the species MID keyX = “k1k2k3…kn” is obtained, wherein kn is either 0 or 1. If kn appears in a candidate species, it is 1, and otherwise it is 0. For example, for Species 1, key1 = “01100110. ”

[0173] In this design, if the number of candidate species is M, the perfectly minimized number of target k-mers is in theory N = log2 M. For example, if 63 candidate species (and a case of negative result) are being detected , 6 k-mer sequences need to be used at a minimum. However, this ideal situation may not exist in practice. For example, if the common sequence |A|  / |S| among the candidate species is much less than or greater than 0. 5, the number N will increase.

[0174] 8. Select appropriate molecular biology methods based on the target k-mer sequences, and use DNA microarrays for detection, which is further detailed in Example 2.

[0175] 9. Determine if a candidate species is detected.

[0176] Generate a molecular identification code S according to negative and positive signals corresponding to each target k-mer sequences, where S = “k1k2k3…kn, ” and each kn is either 0 or 1. For a positive experimental detection, kn is 1, and otherwise it is 0. If S is the same as the molecular identification code KeyX of a candidate species, it is judged that that the candidate species is detected.

[0177] 5.2 Example 2 -Experimental Procedure Using Gene Chips

[0178] The following steps 1–6 describe an exemplary experimental procedure of species identification using gene chips.

[0179] 1. Nucleic Acid Extraction

[0180] Use pathogen nucleic acid extraction reagents, such as the VAMNE Magnetic Pathogen DNA / RNA Kit (Catalog No.: RM601) produced by Nanjing Vazyme Biotech Co., Ltd. ( “Vazyme” ) , and extract all nucleic acids from the sample to be tested according to the instructions.

[0181] 2. Synthesis and Purification of Double-Stranded cDNA

[0182] Synthesize double-stranded cDNA using Vazyme’s ds-cDNA Synthesis Module reagents kit (Catalog No. NR21) .

[0183] Prepare the first-strand synthesis reaction system according to Table 1.

[0184] Table 1. First Chain Synthesis Reaction System.

[0185] Synthesize the first strand using PCR with the following reaction program settings: 25 ℃, 10 min; 42 ℃, 15 min; 70 ℃, 15 min; 4℃, hold (heated lid 105 ℃) .

[0186] Prepare the second chain synthesis reaction system according to Table 2.

[0187] Table 2. Second Chain Synthesis Reaction System.

[0188] Add the prepared second strand synthesis reaction system to the products from the first strand synthesis. Synthesize the second strand using PCR with the following reaction program settings: 16 ℃, 60 min; 4 ℃, hold (heated lid 30 ℃) .

[0189] Purify the double-stranded cDNA products using magnetic beads (VAHTS DNA Clean Beads) .

[0190] 3. Fragmentation of Double-Stranded cDNA

[0191] Fragment the double-stranded cDNA using Covaris M220 Ultrasonicator. Set the fragmentation parameters as shown in Table 3.

[0192] Table 3. Parameters for Fragmentation of Double-Stranded cDNA.

[0193] 4. Preparation of Whole-Genome Library

[0194] Prepare whole-genome library using Vazyme’s VAHTS Universal Pro DNA Library Prep Kit (Catalog No. ND608) .

[0195] (1) Ends Preparation. Prepare an ends preparation reaction system according to Table 4.

[0196] Table 4. Ends Preparation Reaction System.

[0197] Perform ends preparation using PCR, with the following reaction program settings: 30 ℃, 20 min; 65 ℃, 15 min; 4 ℃, hold (heated lid 105 ℃) .

[0198] (2) Adaptor Ligation. Prepare an adaptor ligation reaction system according to Table 5.

[0199] Table 5. Adaptor Ligation Reaction System.

[0200] Transfer the above ligation reaction reagents to a centrifuge tube after the completion of the end preparation. Ligate the adaptors using PCR with the following reaction program settings: 20 ℃, 15 min; 4 ℃, hold (heated lid 105 ℃) .

[0201] (3) Purification of Ligation Products. Purify the ligation products magnetic beads (VAHTS DNA Clean Beads) according to the product instructions to obtain 20 μL of purified DNA.

[0202] (4) First Round of Library Amplification. Prepare a reaction system for a first round of library amplification according to Table 6.

[0203] Table 6. Reaction system for the First Round of Library Amplification.

[0204] Note: Primers amplify the adapter sequences.

[0205] Transfer all the reagents for the first round of amplification reaction to a centrifuge tube containing the purified ligation products. Perform the first round of library amplification using PCR with the following reaction program settings: 98 ℃, 45 s; (98 ℃, 15 s; 65 ℃, 30 s; 72 ℃, 30 s) 12 cycles; 72 ℃, 1 min; 4 ℃, hold (heated lid 105 ℃) .

[0206] (5) Purification of the First Round of Library Amplification Products. Use magnetic beads (VAHTS DNA Clean Beads) to purify the first round of library amplification products according to the product instructions.

[0207] (6) Product Quantification. Measure the concentration of purified DNA with a fluorescence spectrometer.

[0208] (7) Second Round of Library Amplification. Prepare a reaction system for a second round of library amplification according to Table 7.

[0209] Table 7. Reaction System for the Second Round of Library Amplification.

[0210] Use PCR for the second round of library amplification with the following reaction program settings: 98 ℃, 45 s; (98 ℃, 15 s; 65 ℃, 30 s; 72 ℃, 30 s) 35–40 cycles; 72℃, 1 min; 4 ℃, hold (heated lid 105 ℃) .

[0211] The products of the second round of library amplification do not need purification and can be used directly for on-chip detection.

[0212] 5. On-Chip Hybridization and Data Collection.

[0213] Perform hybridization on gene chips and collect the results.

[0214] 6. Data Processing and Analysis

[0215] Process and analyze the results from the experiments on gene chips.

[0216] 5.3 Example 3 -Method Testing with Randomly-Selected Sequences of Multiple Species and Strains

[0217] 1. Preparation of Data of Target Species

[0218] Genomic data of 817 strains were downloaded from NCBI, which included randomly-selected fungal and bacterial genera and strains.

[0219] 2. Selection of Sequences and Primer Design

[0220] From the genomic data of the 817 strains, 50 shared sequences (Table 8) were selected, and one or more of which were present in 795 strains. It was ascertained that the 50 sequences could constitute molecular identification codes in binary that uniquely identified 93 strains. These 50 sequences and 93microbial strains were used for subsequent experimentation.

[0221] Table 8. List of 50 Shared Oligonucleotide Sequences.

[0222] The 93 microbial strains are listed in Table 9. 1.

[0223] Table 9. 1. List of 93 Microbial Strains.

[0224] Tables 9. 2, 9. 3, and 9. 4 detail the presence or absence of each of the 50 share oligonucleotide sequences in the genome of each of the 93 microbial strains.

[0225] Table 9. 2. Presence of SEQ. NOS. 1–17 in Each of the 93 Microbial Strains.

[0226] Table 9. 5. Presence of SEQ. NOS. 18–34 in Each of the 93 Microbial Strains.

[0227] Table 9. 4. Presence of SEQ. NOS. 35–50 in Each of the 93 Microbial Strains.

[0228] For each of the 93 microbial strains, the presence or absence of the 50 share oligonucleotide sequences is represented as a molecular identification code in 40 binary digits. The molecular identification codes are shown in Table 9. 5. It was checked that each one of the molecular identification codes was unique.

[0229] Table 9. 5. Molecular Identification Codes for the 93 Microbial Strains.

[0230] 3. Gene Chips, Sample Sequences, and Standards

[0231] Samples of DNA sequences that were specific to the 93 strains were ordered. They were used for subsequent amplification with primers. A gene chip was ordered that contained the 50 share oligonucleotide sequences as probes. The probes were located at spots at x from 10 to 59 and y equals 235. Four standards were also purchased: Aspergillus oryzae (AO-J1) , Aspergillus japonicus (AJ-J1) , Aspergillus aculeatus (AA-J1) , and Agaricus bisporus (AB-J1) . These were stains used in industrial manufacturing.

[0232] 4. PCR Amplification and Exaction of Fluorescence Signals

[0233] The labels for the 93 ordered samples of DNA and the 4 standards were covered, and they were randomly assigned with serial numbers T1 to T93 for the samples and A1 to A4 for the standards. They were each amplified in a PCR thermocycler with fluorescent dNTP such that the amplified products were fluorescently labeled.

[0234] The fluorescently labeled PCR products from each sample were diluted, mixed with buffer, and then added to the chip to cover the array of probes. The chip was covered with a glass slide to prevent evaporation, and then placed in an oven at 55 ℃ for 2 hours for hybridization.

[0235] After hybridization, the chip was washed 2 × SSC / 0. 1%SDS, 1 × SSC / 0. 1%SDS, and 0. 1 × SSC, with each wash lasting 5 minutes. The chip was then washed with deionized water, dried, and then scanned with a fluorescent scanner to images of fluorescence signals.

[0236] The images were analyzed to determine the signal strength at each probe spot on the chip. Tables 10.1 to 10. 8 list the fluorescence signals at each (x, y) spot, where y is always 235. Each cell in the tables corresponds to a hybridization assay that detects the presence of one of the 60 oligonucleotide sequences in one of the 97 samples (including T1 to T93 and A1 to A4) .

[0237] Fluorescence signals above and below a reference level were judged to be positive and negative, respectively.

[0238] Table 10. 1. Fluorescence Signals from Samples T1 to T12.

[0239] Table 10. 2. Fluorescence Signals from Samples T13 to T25.

[0240] Table 10. 3. Fluorescence Signals from Samples T26 to T38.

[0241] Table 10. 4. Fluorescence Signals from Samples T39 to T51.

[0242] Table 10. 5. Fluorescence Signals from Samples T52 to T64.

[0243] Table 10. 6. Fluorescence Signals from Samples T65 to T77.

[0244] Table 10. 7. Fluorescence Signals from Samples T78 to T90.

[0245] Table 10. 8. Fluorescence Signals from Samples T91 to T93 and Standards A1 to A4.

[0246] 5. Extraction of Results and Verification

[0247] According to the positive / negative detection results from the 60 probes, the fluorescence signals for each sample were converted into a corresponding molecular identification code in binary, which was then compared against the predetermined database of known molecular identification codes and associated species (Table 9. 5) . Having obtained a species identification via the molecular identification codes, the covers of the sample labels were then removed to reveal the original strain ID. The results of this experiments are shown in Table 11.

[0248] Table 11. Species Identified by Experiment.

[0249] Comparing the experimental identifications against the original sample labels, it can be seem from Table 11 that the experiment correctly identified all 93 samples as well as the four standards.

[0250] The methods as described above are characterized by several particular advantages. The methods are generally applicable to identification of species both within a genus as well as across multiple genera. The binary code scheme of the methods can distinguish up to 2n species from n pairs of primers. The theoretical upper limit of 2n may be as high as 1015. The methods disclosed herein are inexpensive. Its costs is only about 5%of whole-genome sequencing. The methods disclosed herein also facilitates the rapid development of new assays as need may arise.

[0251] The foregoing disclosure provides illustration and description but is not intended to be exhaustive or to limit the implementations to the precise form disclosed. Modifications may be made in light of the above disclosure or may be acquired from practice of the implementations. Although particular combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of various implementations. In fact, many of these features may be combined in ways not specifically recited in the claims and / or disclosed in the specification. No element, act, or instruction used herein should be construed as critical or essential unless explicitly described as such.

[0252] Although each dependent claim listed below may directly depend on only one claim, the disclosure of various implementations includes each dependent claim in combination with every other claim in the claim set.

Claims

A method of analyzing in a sample comprising polynucleic acid originating from a plurality of candidate species, the method comprising the following steps:detecting if each one of a set of target sequences is present in the polynucleic acid;generating a molecular identification code comprising a plurality of binary digits, wherein each binary digit in the molecular identification code encodes whether a corresponding target sequence in the set of target sequences is present in the polynucleic acid;comparing the molecular identification code with a list of known molecular identification codes, wherein each known molecular identification uniquely corresponds to one of the plurality of the candidate species; anddetermining whether any one of the plurality of the candidate species is present in the sample.The method of claim 1, wherein the polynucleic acid is DNA.The method of claim 1, wherein the polynucleic acid is RNA.The method of claim 1, further comprising amplifying the polynucleic acid in the sample.The method of claim 2, wherein amplifying the polynucleic acid in the sample comprises:amplifying the polynucleic acid in the sample in a polymerase chain reaction.The method of claim 1, wherein detecting if each one of the set of target sequences is present in the polynucleic acid comprises:detecting if each one of the set of target sequences is present in the polynucleic acid in a molecular biology experiment.The method of claim 6, wherein the molecular biology experiment is quantitative PCR (QPCR) , fluorescence in situ hybridization (FISH) , mass spectrometry, or Next Generation Sequencing (NGS) .The method of claim 6, wherein the molecular biology experiment is conducted with a DNA microarray.The method of claim 6, wherein the molecular biology experiment is conducted with a microfluidic device.The method of claim 6, wherein the molecular biology experiment is conducted with a nanofluidic device.The method of claim 1, wherein each one of the set of target sequences is a polynucleic acid sequence that is present in more than one candidate species’ genome.The method of claim 11, wherein none of the set of target sequences is a polynucleic acid sequence that is present in all candidate species’ genome.The method of claim 1, further comprising generating a plurality of binary digits, wherein each binary digit encodes whether a corresponding target sequence in the set of target sequences is present in the polynucleic acid.The method of claim 1, wherein each binary digit in the molecular identification code is either one or zero, wherein one encodes that the corresponding target sequence is present in the polynucleic acid, zero encodes that the corresponding target sequence is absent from the polynucleic acid.The method of claim 1, wherein generating the molecular identification code comprising one or more of binary digits comprising:concatenating the one or more of binary digits together as a single string.The method of claim 1, wherein each of the known molecular identification codes comprising one or more of binary digits, and wherein each binary digit encodes whether the corresponding target sequence in the set of target sequences is present in a genome of the candidate species to which the known molecular identification code uniquely corresponds.The method of claim 1, wherein the molecular identification code comprises an identical number of binary digits as a total number of target sequences in the set of target sequences.The method of claim 16, wherein all known molecular identification codes comprise an identical number of binary digits.The method of claim 1, wherein a total number of the plurality of candidate species is less than or equal to two to a power of a total number of the plurality of binary digits in the molecular identification code.The method of claim 1, wherein a total number of the plurality of candidate species is less than or equal to two to a power of a total number of target sequences in the set of target sequences.The method of claim 1, further comprising which one of the plurality of the candidate species is present in the sample.The method of claim 1, wherein the sample originates from a human in a form of a nasal swab, a nasopharyngeal swab, a rectal swab, a vaginal swab, a penile swab, an ear swab, or a wound stab, and wherein each of the plurality of candidate species is a pathogenic microorganism.The method of claim 1, wherein the sample is a human bodily fluid.The method of claim 1, wherein the sample is a ruminal sample.The method of claim 1, wherein the sample is a soil sample.The method of claim 25, wherein the sample is a rhizosphere sample.The method of claim 1, wherein the sample is a marine or freshwater sample.The method of claim 1, wherein the sample is collected from goods, cargo, or passengers at a port of entry, and wherein each of the plurality of candidate species is an invasive species.The method of claim 1, wherein each of the plurality of candidate species is a prokaryote, eukaryote, bacteria, eubacteria, archaea, archaebacteria, fungus, plant, animal, protist, or virus species.The method of claim 1, wherein each of the plurality of candidate species is an insect, fish, bird, or mammal species.A method of determining a set of target sequences for detecting polynucleic acid originating from a plurality of candidate species, the method comprising the following steps:obtaining a genomic sequence of each of the plurality of candidate species;generating a set of k-mers for each of the plurality of candidate species, wherein each k-mer is a continuous fragment of the genomic sequence; andincorporating iteratively k-mers as target sequences into the set of target sequences, and calculating a molecular identification code for each of the plurality of candidate species each time the set of target sequences is changed, until each of the plurality of candidate species has a unique molecular identification code.The method of claim 31, further comprising:obtaining a genomic sequence of a background species; andgenerating a set of k-mers for the background species, wherein the set of k-mers is generated in an identical manner as generating the set of k-mers for each of the plurality of candidate species.The method of claim 32, further comprising:excluding from the set of target sequences any k-mer that is identical to any k-mer in the set of k-mers for the background species.The method of claim 32, further comprising:calculating a molecular identification code for the background species each time the set of target sequences is changed; anddetermining if the molecular identification code for the background species is identical to a molecular identification code for any of the plurality of candidate species.The method of claim 34, further comprising:incorporating one or more additional k-mers as target sequences into the set of target sequences, until the molecular identification code for the background species is different from all molecular identification codes for the plurality of candidate species.The method of claim 31, wherein the polynucleic acid is DNA, RNA, or a combination thereof.The method of claim 31, wherein each of the plurality of candidate species is a prokaryote, eukaryote, bacteria, eubacteria, archaea, archaebacteria, fungus, plant, animal, protist, or virus species.The method of claim 31, wherein each of the plurality of candidate species is an insect, fish, bird, or mammal species.The method of claim 32, wherein the background species and one of the plurality of candidate species are in a same biological taxonomic unit.The method of claim 39, wherein the same biological taxonomic unit is a kingdom, a phylum, a class, an order, a family, or a genus.The method of claim 31, wherein obtaining the genomic sequence of each of the plurality of candidate species comprises:downloading the genomic sequence of each of the plurality of candidate species from a database.The method of claim 41, wherein the database is an Internet database of genomic sequences.The method of claim 41, wherein the database is GenBank.The method of claim 31, wherein generating the set of k-mers for each of the plurality of candidate species comprises:generating a k-mer which is a continuous fragment of the genomic sequence of a candidate species, wherein the continuous fragment is of a predetermined size, and the continuous fragment starts from a locus in the genomic sequence; andgenerating one or more further k-mers, wherein each of the one or more further k-mers starts from a new locus that is shifted by a predetermined step size from a previous locus in a same direction.The method of claim 44, wherein the predetermined step size is one base.The method of claim 31, further comprising:removing duplicative k-mers within the set of k-mers for each of the plurality of candidate species and retaining only one copy of such duplicative k-mers, wherein duplicative k-mers are k-mers having an identical sequence.The method of claim 31, further comprising:removing any k-mer whose sequence is only present in a genomic sequence of a single candidate species of the plurality of candidate species.The method of claim 31, further comprising:removing any k-mer whose sequence is present in a genomic sequence of every one of the plurality of candidate species.The method of claim 31, wherein incorporating iteratively k-mers into the set of target sequences comprises:incorporating iteratively k-mers, one at a time, into the set of target sequences, in an order based on a preference of each k-mer, wherein a more preferred k-mer is incorporated into the set of target sequences before a less preferred k-mer.The method of claim 49, wherein the preference of each k-mer is determined based on how closely the k-mer can exactly bifurcate the plurality of candidate species; wherein a k-mer can exactly bifurcate the plurality of candidate species if the k-mer can divide the plurality of candidate species into two sets of equal cardinalities, wherein one set comprises the candidate species whose genome sequence comprises a sequence of the k-mer, and the other set comprises the candidate species whose genome sequence does not comprise the sequence of the k-mer; and wherein a k-mer that can exactly bifurcate the plurality of candidate species has the highest preference.The method of claim 49, wherein the preference of each k-mer is an integer p, a smaller value of p is preferred over a larger value of p, and a k-mer has a preference of p when the k-mer can divide the plurality of candidate species into two sets, wherein one set comprises the candidate species whose genome sequence comprises a sequence of the k-mer, and the other set comprises the candidate species whose genome sequence does not comprise the sequence of the k-mer; wherein the two sets consisting ofandcandidate species respectively when n is even, orandspecies respectively when n is odd; and wherein n is a total number of the plurality of candidate species.The method of claim 49, wherein the preference of each k-mer is a ratio r, a smaller value of r is preferred over a larger value of r, andand wherein a is a number of candidate species whose genome sequence comprises a sequence of the k-mer, and n is the number of the plurality of candidate species.The method of claim 31, wherein calculating the molecular identification code for each of the plurality of candidate species comprises:for each k-mer in the set of target sequences, determining if the k-mer is present in the genomic sequence of the candidate species, and generating a binary digit of one or zero, wherein one encodes that the k-mer is present, and zero encodes that the k-mer is absent; andconcatenating all binary digits together as a single string.The method of claim 31, further comprising optimizing the set of target sequences to reduce a total number of k-mers in the set.A system for analyzing in a sample comprising polynucleic acid originating from a plurality of candidate species, the system comprising:a thermal cycler configured for amplifying the polynucleic acid in the sample in a polymerase chain reaction;a DNA microarray configured for detecting if each one of a set of target sequences is present in an amplified product;an optical scanner configured for reading results from the DNA microarray; anda computer comprising at least one processor, and memory for storing instructions which, when executed by the at least one processor, result in operations comprising:generating a molecular identification code comprising a plurality of binary digits, wherein each binary digit in the molecular identification code encodes whether a corresponding target sequence in the set of target sequences is present in the polynucleic acid;comparing the molecular identification code with a list of known molecular identification codes, wherein each known molecular identification uniquely corresponds to one of the plurality of the candidate species; anddetermining whether any one of the plurality of the candidate species is present in the sample.The system of claim 55, further comprising an interface configured for receiving the results from the optical scanner.A system for designing a set of target sequences for detecting polynucleic acid originating from a plurality of candidate species, the system comprising:at least one processor;memory for storing instructions which, when executed by the at least one processor, result in operations comprising:generating a set of k-mers for each of the plurality of candidate species, wherein each k mer is a continuous fragment of a genomic sequence of each of the plurality of candidate species; andincorporating iteratively k-mers as target sequences into the set of target sequences, and calculating a molecular identification code for each of the plurality of candidate species each time the set of target sequences is changed, until each of the plurality of candidate species has a unique molecular identification code.A non-transitory computer-readable medium encoded with instructions for commanding one or more processors to execute the method of claim 31.

Citation Information

Patent Citations

  • High Throughput Testing for Presence of Microorganisms in a Biological Sample

    US20090306230A1

  • Biological sample target classification, detection and selection methods, and related arrays and oligonucleotide probes

    US20130267429A1

  • Method for identifying and classifying sample microorganisms

    US20210202040A1

  • Methods for comparative metagenomic analysis

    US20210249102A1

  • Method and a system for profiling of metagenome

    US20220392565A1