Compositions and methods for enriching populations of nucleic acids
The method enriches non-host nucleic acids in host samples using diverse oligonucleotides to sequence and identify pathogens, addressing misdiagnosis and antibiotic resistance in infectious diseases by enhancing detection and diagnosis.
Patent Information
- Application Number
- JP2021071794
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2016-05-10
- Filing Date
- 2021-04-21
- Publication Date
- 2025-11-13
- Estimated Expiration
- 2036-05-17
AI Technical Summary
Inadequate detection of infectious diseases leads to misdiagnosis and underdiagnosis, contributing to the overuse or misuse of antibiotics and the emergence of antibiotic-resistant bacteria, due to slow diagnostic tests and pathogen-specific testing that requires prior knowledge of etiological factors, and the challenge of secondary and co-infections masking true infection sources.
A method for enriching non-host nucleic acids in a host sample, such as blood or plasma, by using a collection of oligonucleotides with diverse sequences to capture and sequence non-host nucleic acids, followed by sequencing assays like next-generation sequencing to identify pathogens.
Enhances the detection and identification of pathogens, reducing misdiagnosis and antibiotic resistance by providing a comprehensive and affordable method for identifying infectious organisms, enabling early and accurate diagnosis, prognosis, and monitoring of infectious diseases.
Smart Images

Figure 0007769480000001 
Figure 0007769480000002 
Figure 0007769480000003
Abstract
Description
[Technical Field]
[0001] Related Applications This application was filed on May 18, 2015, which application is incorporated herein by reference. No. 62 / 163,273 filed on May 10, 2016, and U.S. Provisional Application No. 62 / 163,273 filed on May 10, 2016. The benefit of US Provisional Application No. 62 / 334348 is claimed. [Background technology]
[0002] Infectious diseases and disorders present challenges for primary care providers and patients alike. Inadequate detection of infectious diseases often leads to poor detection rates, especially compared to other diseases. This is due to several factors, including the lack of meaningful tests that can quickly produce accurate answers. Given the slow speed of some diagnostic tests, many doctors may choose to Treat patients based on suspected etiology before receiving test results, rather than risking worsening symptoms Tests for infectious diseases are also generally pathogen-specific, so physicians Before ordering a test, you must have some idea of the etiological factors of the patient's symptoms. Another confounding factor for some infectious diseases is the sudden change in the infectious agent during the course of infection. As a result, the initial diagnosis may not accurately reflect the nature of the patient's condition at a later time. Secondary and co-infections can also mask other sources of infection or escape detection entirely. This can confound diagnosis and treatment. Summary of the Invention [Problem to be solved by the invention]
[0003] Misdiagnosis and underdiagnosis of pathogen infections can have dire consequences for patients and the community at large. For example, the overuse or misuse of antibiotics can lead to an increase in antibiotic-resistant bacteria. This may be dangerous not only to the patient but also to others who come into contact with the patient. Therefore, it is a reliable, comprehensive and affordable method for identifying pathogens in samples. There is a need in the art for testing. [Means for solving the problem]
[0004] The present disclosure generally relates to the detection of non-host nucleic acids in a sample obtained from a host in which host nucleic acids are present. Such methods include, for example, using a cell-free sample taken from a host. It has a variety of uses, including the identification of infectious or pathogenic organisms within a host through analysis of the virus. Generally, the methods described herein involve a host sample, such as a cell-free plasma sample from a host. The enrichment may include selective enrichment of non-host nucleic acids relative to host nucleic acids derived from the enriched nucleic acids to identify the presence of non-host nucleic acids and the presence of pathogens or infectious organisms within the host. Identification of the presence of non-host nucleic acids, pathogens or infectious organisms can be used to identify if the host is may enable the detection, diagnosis, prognosis, monitoring or staging of an infectious disease or disorder experienced .
[0005] In one example, the disclosure provides a method for identifying a pathogen in a host, comprising: collecting acellular blood from the host; The present invention provides a method that begins by providing a blood or plasma sample. enriching the plasma sample for non-host-derived nucleic acids relative to host-derived nucleic acids, and then The sample is analyzed for non-host derived nucleic acids and then pathogens in the host are identified from the non-host nucleic acids. It can be determined.
[0006] In one aspect, the present disclosure provides a method for priming or detecting non-host sequences in a nucleic acid sample from a host. A method of capturing, comprising the steps of: (a) providing a nucleic acid sample from a host; (b) a nucleic acid sample from the host, the nucleic acid sample comprising host nucleic acid and non-host nucleic acid; A nucleic acid sample is mixed with a collection of oligonucleotides, thereby obtaining a mixture. wherein the collection of oligonucleotides has different nucleotide sequences. The different nucleotide sequences comprise at least 1000 oligonucleotides each having a different nucleotide sequence. a step of specifically selecting the non-host nucleic acid sequence to contain at least 10 nucleotides in length (c) contacting the collection of oligonucleotides with the nucleic acid sample in a mixture. contacting the non-host nucleic acids in the mixture to at least 10 nucleotides; binds to long non-host nucleic acid sequences, thereby priming or capturing non-host nucleic acids. contacting up to 10% of the host nucleic acid with a non-host nucleic acid sequence of at least 10 nucleotides in length; and binding to the
[0007] In some embodiments, the method further comprises: In some embodiments, the method further comprises preferentially amplifying the primary nucleic acid. Next-generation sequencing assays, high-throughput sequencing assays, massively parallel sequencing -Sequencing assay, nanopore sequencing assay or Sanger sequencing By performing a sequencing assay such as the In some embodiments, the method further comprises sequencing the non-host nucleic acid. further comprising preferentially isolating the primed or captured non-host nucleic acid. In some embodiments, the preferential isolation step comprises performing a pull-down assay. In some embodiments, the method comprises: In some embodiments, the method further comprises performing a primer extension reaction on an acid. At least 1000 oligonucleotides with different nucleotide sequences are used to label the nucleic acid. In some embodiments, the primed or captured non-host nucleic acid contains RNase A. A is a non-host nucleic acid. In some embodiments, the method In some embodiments, the method further comprises performing a polymerization reaction on the RNA non-host nucleic acid. In the , the polymerization reaction is carried out by reverse transcriptase.
[0008] In another aspect, the disclosure provides a method for sequencing a non-host sequence in a nucleic acid sample from a host. (a) providing a nucleic acid sample from a host, (b) a nucleic acid sample from the host, the sample comprising host nucleic acid and non-host nucleic acid; with a collection of oligonucleotides, thereby obtaining a mixture. The collection of oligonucleotides is at least 1000 having different nucleotide sequences. The oligonucleotides may contain at least 1000 different nucleotide sequences. (c) a non-host nucleic acid sequence of 100 nucleotides in length; contacting the collection of oligonucleotides with the nucleic acid sample in a mixture; and wherein the contacting converts the non-host nucleic acids in the mixture into non-host nucleic acids of at least 10 nucleotides in length. The nucleic acid sequence binds to the nucleic acid sequence and, upon contact, binds up to 10% of the host nucleic acid to a nucleic acid sequence of at least 10 nucleotides in length. (d) performing a sequencing assay. Thus, non-host nucleic acids linked to non-host nucleic acid sequences of at least 10 nucleotides in length are sequenced. and determining the
[0009] In some embodiments, the method includes a step of preferentially amplifying non-host nucleic acid in the reaction. In some embodiments, the sequencing assay further comprises next generation sequencing. Sequencing assay, high-throughput sequencing assay, massively parallel sequencing assay assay, nanopore sequencing assay, or Sanger sequencing assay. In some embodiments, the method further comprises preferentially isolating non-host nucleic acid. In some embodiments, the isolating step comprises performing a pull-down assay. In some embodiments, the method comprises performing a primer extension reaction on a non-host nucleic acid. In some embodiments, the method further comprises the step of: At least 1000 of the oligonucleotides contain a nucleic acid label. In some embodiments, the non-host nucleic acid is an RNA non-host nucleic acid. In some embodiments, the polymerization reaction further comprises performing a polymerization reaction on the non-host nucleic acid. The ligation reaction is carried out by reverse transcriptase.
[0010] In some embodiments of the methods provided herein, the method further comprises the step of: At least 1000 oligonucleotides having different nucleotide sequences At least 10,000 oligonucleotides. In this embodiment, at least 1000 oligonucleotides having different nucleotide sequences are The oligonucleotides are at least 100,000 oligonucleotides having different nucleotide sequences. In some embodiments of the methods provided herein, different nucleotides At least 1000 oligonucleotides having the sequence At least 1,000,000 oligonucleotides having the sequence provided herein. In some embodiments of the method, at least 1000 nucleotide sequences having different nucleotide sequences are The oligonucleotides provided herein are not conjugated to a solid support. In some embodiments of the method, at least 1000 nucleotide sequences having different nucleotide sequences are The oligonucleotides provided herein have a length of up to 200 nucleotides. In some embodiments of the method, at least 10 nucleotide sequences having different nucleotide sequences are used. Each of the 00 oligonucleotides has a domain of nucleotides 10-20 nucleotides in length. Each domain of nucleotides is 10 to 20 nucleotides long and contains a different nucleotide. In some embodiments of the methods provided herein, the nucleic acid sequence comprises a 10-20 nucleotide sequence. Each domain of nucleotides of length 12 to 15 nucleotides is 12 to 15 nucleotides in length. In some embodiments of the provided methods, each of the nucleotides 10 to 20 nucleotides in length In some embodiments of the methods provided herein, the domain is 13 to 15 nucleotides in length. In an embodiment, each domain of nucleotides of 10 to 20 nucleotides in length is a mammalian nucleic acid sequence In some embodiments of the methods provided herein, the host is not a mammalian host. Thus, a nucleic acid sample from a mammalian host contains mammalian host nucleic acid and non-mammalian nucleic acid. In some embodiments of the methods provided herein, the host is a human host, The nucleic acid sample from the method comprises human host nucleic acid and non-human nucleic acid. In some embodiments, the non-human nucleic acid comprises microbial nucleic acid. In some embodiments, the non-human nucleic acid comprises bacterial nucleic acid. In some embodiments, the nucleic acid sample from the host contains at least five non-host nucleic acid sequences. and the method further comprises detecting at least five non-host nucleic acid sequences. In some embodiments of the methods provided herein, the nucleic acid sample from the host is selected from the group consisting of blood, plasma, and the like. , serum, saliva, cerebrospinal fluid, synovial fluid, lavage fluid, urine and stool, e.g. from blood, plasma and serum In some embodiments of the methods provided herein, the sample is selected from the group consisting of: is selected from the group consisting of blood, plasma, and serum. In some embodiments, the nucleic acid sample from the host is a circulating nucleic acid sample. In some embodiments of the method, the nucleic acid sample from the host is a circulating cell-free nucleic acid sample. In some embodiments of the methods provided herein, a nucleic acid sample from a host is In some embodiments of the methods provided herein, the library comprises a sequenceable nucleic acid library. In embodiments, the nucleic acid sample from the host comprises single-stranded DNA or cDNA. In some embodiments of the provided methods, the nucleic acid is DNA. In some embodiments of the methods, the nucleic acid is RNA. In some embodiments, the nucleic acid sample from the host is free of artificially fragmented nucleic acids. In some embodiments of the methods provided herein, a collection of oligonucleotides The nucleic acid may be DNA, RNA, PNA, LNA, BNA, or any combination thereof. In some embodiments of the methods provided herein, a collection of oligonucleotides In some embodiments of the methods provided herein, the oligonucleotide comprises a DNA oligonucleotide. In, the collection of oligonucleotides comprises RNA oligonucleotides.
[0011] In some embodiments of the methods provided herein, a collection of oligonucleotides In some embodiments of the methods provided herein, In this embodiment, the collection of oligonucleotides are RNA oligonucleotides. In some embodiments of the methods provided herein, the collection of oligonucleotides comprises a nucleic acid sequence. In some embodiments of the methods provided herein, the nucleic acid is labeled with an acid or chemical label. In some embodiments of the methods provided herein, the chemical label is biotin. The collection of oligonucleotides does not include artificially fragmented nucleic acids. In some embodiments of the provided methods, the non-host nucleic acid is a pathogen nucleic acid, a microbial nucleic acid, a bacterial nucleic acid, or a nucleic acids, viral nucleic acids, fungal nucleic acids, parasitic nucleic acids and any combination thereof In some embodiments of the methods provided herein, the non-host nucleic acid is selected from the group In some embodiments of the methods provided herein, the non-host nucleic acid is a microbial nucleic acid. In some embodiments of the methods provided herein, the non-host nucleic acid is bacterial nucleic acid. It is a viral nucleic acid.
[0012] In yet another aspect, the present disclosure provides a method for enriching non-host sequences in a nucleic acid sample from a host. (a) providing a nucleic acid sample from a host, The nucleic acid sample is a single-stranded nucleic acid sample from a host, containing both host and non-host nucleic acids. (b) regenerating at least a portion of the single-stranded nucleic acid from the host, thereby (c) generating a population of double-stranded nucleic acids in the sample using a nuclease. and removing at least a portion of the double-stranded nucleic acid in the nucleic acid sample from the host. and enriching for non-host sequences.
[0013] In some embodiments, the method further comprises performing a sequencing assay. In some embodiments, the nucleic acid sample from the host comprises at least five non-host nucleic acids. and wherein the method further comprises detecting at least five non-host nucleic acid sequences. In some embodiments, the host is a human. In some embodiments, the host is a human. In some embodiments, the nucleic acid sample from the host is a circulating nucleic acid sample. In some embodiments, the nucleic acid sample from the host is a circulating cell-free nucleic acid sample. blood, plasma, serum, saliva, cerebrospinal fluid, synovial fluid, lavage fluid, urine and faeces, e.g. and serum. In some embodiments, the nucleic acid is DNA. In some embodiments, the nucleic acid is RNA. The method further comprises adding a single-stranded host sequence to the single-stranded nucleic acid sample from the host. In one embodiment, the method uses heat to denature at least a portion of the nucleic acids in the nucleic acid sample. In some embodiments, the method further comprises generating a single-stranded nucleic acid sample by: Renaturation of at least a portion of the nucleic acid occurs within a set time frame. In some embodiments, renaturation of at least a portion of the nucleic acid occurs within 96 hours. In some embodiments, the method comprises regeneration in the presence of trimethylammonium chloride. The nuclease may be a double-strand specific nuclease, BAL-31, double-strand specific DNase, or or a combination thereof. In some embodiments, the nuclease is double-strand specific. In some embodiments, the nuclease is BAL-31. In some embodiments, the nuclease is active on double-stranded nucleic acids. In some embodiments, the nuclease is not active against single-stranded nucleic acids. The acid has not been artificially fragmented.
[0014] In yet another aspect, the present disclosure provides a method for enriching non-host sequences in a nucleic acid sample from a host. (a) providing a nucleic acid sample from a host, the acid sample contains host nucleic acid and non-host nucleic acid associated with nucleosomes; b) removing at least a portion of the host nucleic acid associated with nucleosomes, thereby and concentrating non-host nucleic acids in the nucleic acid sample from the host.
[0015] In some embodiments, the method further comprises performing a sequencing assay. In some embodiments, the nucleic acid sample from the host comprises at least five non-host nucleic acids. and wherein the method further comprises detecting at least five non-host nucleic acid sequences. In some embodiments, the host is a human. In some embodiments, the host is a human. In some embodiments, the nucleic acid sample from the host is a circulating nucleic acid sample. In some embodiments, the nucleic acid sample from the host is a circulating cell-free nucleic acid sample. blood, plasma, serum, saliva, cerebrospinal fluid, synovial fluid, lavage fluid, urine and faeces, e.g. and serum. In some embodiments, the antibody in step (b) is selected from the group consisting of: In some embodiments, the removal in step (b) comprises performing electrophoresis. In some embodiments, the removal comprises isotachophoresis. In some embodiments, the removing step comprises using a porous filter. In some embodiments, the removal of the stearate comprises using an ion exchange column. The removal in step (b) is performed using one or more antibodies specific for one or more histones. In some embodiments, the one or more histones are histone H2A N-terminus, solvent-exposed epitope of histone H2A, motif on Lys9 in histone H3 methylation on Lys9 in histone H3, dimethylation on Lys56 in histone H3 trimethylation of histone H2B, phosphorylation on Ser14 in histone H2B, and histone H2A.X In some embodiments, the phosphorylation of Ser139 in 1 In some embodiments, the method comprises immobilizing one or more antibodies on a column. The method further comprises removing the antibody or antibodies.
[0016] In another aspect, the disclosure is a method for enriching non-host sequences in a nucleic acid sample from a host. (a) providing a nucleic acid sample from a host, (b) a sequence of nucleic acids comprising one or more length intervals; Remove or isolate DNA, thereby enriching non-host nucleic acids in a nucleic acid sample from a host. and shrinking the
[0017] In some embodiments, step (b) removes one or more length intervals of DNA. In some embodiments, step (b) comprises measuring one or more length intervals. In some embodiments, the method comprises isolating DNA of one or more length intervals About 180 base pairs, about 360 base pairs, about 540 base pairs, about 720 base pairs, and about 900 base pairs In some embodiments, one or more of the length intervals is selected from the group consisting of: 150 base pairs, about 300 base pairs, about 450 base pairs, about 600 base pairs, and about 750 base pairs In some embodiments, one or more length intervals are selected from the group consisting of: 60 base pairs, approximately 320 base pairs, approximately 480 base pairs, approximately 640 base pairs, and approximately 800 base pairs In some embodiments, one or more length intervals are selected from the group consisting of: From 0 base pairs, about 340 base pairs, about 510 base pairs, about 680 base pairs, and about 850 base pairs In some embodiments, one or more length intervals are selected from the group consisting of: base pairs, about 380 base pairs, about 570 base pairs, about 760 base pairs, and about 950 base pairs In some embodiments, one or more length intervals are selected from the group consisting of: pairs or multiples thereof, 160 base pairs or multiples thereof, 170 base pairs or multiples thereof, 190 It is selected from the group consisting of a base pair or a multiple thereof, and any combination thereof. In some embodiments, step (b) is about 100, 120, 150, 175, 200, This includes removing DNA that is greater than 250, 300, 400, or 500 bases in length. In some embodiments, step (b) is at most about 100, 120, 150, 175, 200 In some embodiments, the method further comprises isolating DNA that is 250 or 300 bases in length. In step (b), the length is about 10 to about 100 bases, about 10 to about 120 bases, Approximately 10 bases to approximately 150 bases in length, approximately 10 bases to approximately 175 bases in length, approximately 10 bases to approximately 200 Base length: about 10 bases to about 250 bases, about 10 bases to about 300 bases, about 30 bases ~ about 100 bases long, about 30 bases long to about 120 bases long, about 30 bases long to about 150 bases long, about 30 bases to approximately 175 bases, approximately 30 bases to approximately 200 bases, approximately 30 bases to approximately 250 bases The method includes isolating DNA having a length of about 30 bases or a length of about 30 bases to about 300 bases.
[0018] In some embodiments, the method further comprises performing a sequencing assay. In some embodiments, the nucleic acid sample from the host comprises at least five non-host nucleic acids. and wherein the method further comprises detecting at least five non-host nucleic acid sequences. In some embodiments, the host is a human. In some embodiments, the host is a human. In some embodiments, the nucleic acid sample from the host is a circulating nucleic acid sample. The sample is a circulating cell-free nucleic acid sample.
[0019] In yet another aspect, the present disclosure provides a method for enriching non-host sequences in a nucleic acid sample from a host. (a) providing a nucleic acid sample from a host, (b) the acid sample contains host nucleic acids, non-host nucleic acids, and exosomes; and removing or isolating at least a portion of the nucleic acid fragment from the host. and enriching the non-host sequences in the vector.
[0020] In some embodiments, step (b) removes at least a portion of the exosomes. In some embodiments, step (b) comprises extracting at least a portion of the exosomes. In some embodiments, the method comprises isolating the In some embodiments, the nucleic acid sample from the host further comprises: each of the non-host nucleic acid sequences comprises five non-host nucleic acid sequences, and the method detects at least five non-host nucleic acid sequences. In some embodiments, the host is a human. In this state, nucleic acid samples from the host may be collected from blood, plasma, serum, saliva, cerebrospinal fluid, synovial fluid, lavage fluid, urine and feces, e.g., selected from the group consisting of blood, plasma, and serum. In some embodiments, the nucleic acid sample from the host is a circulating nucleic acid sample. The nucleic acid sample from the host is a circulating cell-free nucleic acid sample. The method further comprises removing host nucleic acids within the exosomes. In some embodiments, the method further comprises isolating the non-host nucleic acid. wherein the removal or isolation in step (b) removes or isolates leukocyte-derived exosomes. In some embodiments, the leukocytes are macrophages. In embodiments, the removal or isolation in step (b) removes leukocyte-derived exosomes. or using immunoprecipitation to isolate.
[0021] In some embodiments of the methods provided herein, the method comprises: Further comprising adding an acid barcode to one or more samples. In some embodiments of the provided methods, the method comprises detecting one or more pathogenicity loci. ;Antimicrobial resistance marker;Antibiotic resistance marker;Antiviral resistance marker;Antiparasitic Biologic resistance markers; informative genotyping regions; two or more microorganisms, pathogens, bacteria, viruses sequences common to bacteria, fungi and / or parasites; non-host genomes integrated into the host genome sequence;masking non-host sequence;non-host mimic sequence;masking host sequence;host mimic sequence;if and one or more microorganisms, pathogens, bacteria, viruses, fungi and / or parasites The method further includes adding one or more nucleic acids specific to the specific sequence to the sample. nothing.
[0022] In yet another aspect, the present disclosure provides a method for priming or detecting sequences in a nucleic acid sample from a host. A method of capturing, comprising: (a) providing a nucleic acid sample from a host; (b) Nucleic acid samples from the host are analyzed for one or more virulence loci; antimicrobial resistance markers; Antibiotic resistance markers; Antiviral resistance markers; Antiparasitic resistance markers; Informative Genotyping region; two or more microorganisms, pathogens, bacteria, viruses, fungi and / or parasites Sequences common to living organisms; non-host sequences integrated into the host genome; masking non-host sequences; a non-host mimicking sequence; a masking host sequence; a host mimicking sequence; and one or more microorganisms , pathogen, bacterial, viral, fungal and / or parasitic specific sequences and mixing the nucleic acid of one or more regions of the nucleic acid sequence with the nucleic acid of one or more regions of the nucleic acid sequence, thereby obtaining a mixture; (c) contacting the nucleic acid of one or more regions of interest with the nucleic acid sample in a mixture; contacting one or more regions of interest with nucleic acids in a nucleic acid sample; and binding to nucleic acids in the region, thereby priming or capturing the nucleic acids. Provide a method for
[0023] In some embodiments, the method further comprises: identifying one or more virulence loci; antimicrobial resistance; Sex markers;Antibiotic resistance markers;Antiviral resistance markers;Antiparasitic resistance markers Car; Informative genotyping region; Two or more microorganisms, pathogens, bacteria, viruses, fungi and and / or parasite-shared sequences; non-host sequences integrated into the host genome; masking non-host sequences; non-host mimicking sequences; masking host sequences; host mimicking sequences; and one or Multiple microorganism, pathogen, bacterial, viral, fungal and / or parasitic specific sequences The method further includes a step of performing a nucleic acid amplification reaction using the specific nucleic acid or nucleic acids. In some embodiments, the nucleic acid amplification reaction is a polymerase chain reaction, reverse transcription, transcription-mediated amplification, or In some embodiments, the method comprises the step of: virulence locus;antimicrobial resistance marker;antibiotic resistance marker;antiviral resistance marker Car; antiparasitic resistance marker; informative genotyping region; two or more microorganisms, pathogens sequences common to bacteria, viruses, fungi, and / or parasites; integrated into the host genome Masking non-host sequences; Masking non-host sequences; Non-host mimic sequences; Masking host sequences; Host mimic sequences and one or more microorganisms, pathogens, bacteria, viruses, fungi and / or The method further comprises isolating nucleic acids specific to the parasite-specific sequences. In embodiments, the isolating step comprises performing a pull-down assay. In embodiments, the method further comprises performing a sequencing assay. In some embodiments, the host is a human. In some embodiments, nucleic acid samples from the host are In some embodiments, the nucleic acid sample from the host is a circulating nucleic acid sample. A cell-free nucleic acid sample.
[0024] In another aspect, the present disclosure provides a method for detecting a nucleotide sequence comprising: A collection of oligonucleotides comprising: (a) at least one of: Each of the 1000 oligonucleotides contains a domain of nucleotides; (b) a nucleotide (c) a domain of 10-20 nucleotides in length; (d) each domain of nucleotides having a length of 100 μg has a different nucleotide sequence; Each domain of nucleotides, 10-20 nucleotides in length, is composed of one or more At least 1000 oligonucleotides specifically selected to be absent from the genome The present invention provides a collection of oligonucleotides comprising:
[0025] In some embodiments, at least 1000 oligonucleotides are 200 nucleotides long. In some embodiments, the one or more genomes are one or more fragments of a single nucleotide or fragment length. In some embodiments, one or more mammalian genomes. The human genome, dog genome, cat genome, rodent genome, pig genome, cow genome, sheep genome, goat genome, rabbit genome, horse genome, and any combination thereof In some embodiments, the one or more mammalian genomes are selected from the group consisting of human genomes. In some embodiments, the one or more genomes is one genome. In some embodiments, each domain of nucleotides has a length of 10 to 20 nucleotides. In some embodiments, the length of the fragment is 12 to 15 nucleotides. Each domain has a length of 13 to 15 nucleotides. In some embodiments, the oligonucleotide is a DNA oligonucleotide. In embodiments, the oligonucleotide is an RNA oligonucleotide. In some embodiments, the oligonucleotide is a synthetic oligonucleotide. In some embodiments, the oligonucleotide does not include artificially fragmented nucleic acids. In some cases, the oligonucleotide is labeled with a nucleic acid label, a chemical label, or an optical label. In an embodiment, the at least 1000 oligonucleotides are at least 5000 oligonucleotides. In some embodiments, at least 1000 oligonucleotides In some embodiments, the oligonucleotide is at least 10,000 oligonucleotides. At least 1,000 oligonucleotides are In some embodiments, the oligonucleotide sequence is at least 1000 oligonucleotides. In some embodiments, the number of oligonucleotides is at least 1,000,000. At least 1000 oligonucleotides are identified as one or more microorganisms, pathogens, cells, or the like. 10-20 nucleotides containing different sequences present in bacterial, viral, fungal, or parasitic genomes It contains a domain of nucleotides of length 10 nucleotides.
[0026] In yet another aspect, the disclosure provides a method for generating a collection of oligonucleotides. (a) providing at least 1000 oligonucleotides; At least 1,000 oligonucleotides have different sequences of 10-20 nucleotides. (b) using a nucleic acid sample from a host, the nucleic acid sample comprising a domain of nucleotides of length 100 nucleotides; (c) extracting at least 1000 oligonucleotides from the nucleic acid sequence of the host. (d) mixing at least 100 nucleic acids that do not hybridize with nucleic acids from the host; and isolating at least a subset of the 0 oligonucleotides. and generating a collection of the sequences.
[0027] In some embodiments, at least 1000 oligonucleotides are In some embodiments, the domain of nucleotides has a length of 12 to 16 nucleotides. In some embodiments, the domain of nucleotides is 13 to 15 nucleotides in length. In some embodiments, the oligonucleotide is DNA or In some embodiments, the oligonucleotide is single-stranded. In some embodiments, the oligonucleotide is a synthetic oligonucleotide. In some embodiments, the nucleic acid from the host is labeled with a nucleic acid label or a chemical label. In some embodiments, the chemical label is biotin. In some embodiments, the isolating step (d) comprises performing pull-down. In some embodiments, the method comprises performing a denaturing assay. The method further comprises the step of denaturing at least a portion of the nucleic acid.
[0028] In yet another aspect, the disclosure provides a method for generating a collection of oligonucleotides. (a) a sequence selected from the group consisting of a genome, an exome, and a transcriptome; Domains of nucleotides 10–20 nucleotides long present in the background population (b) determining a sequence having a different sequence not present in the background population; At least 1000 nucleotides containing domains of nucleotides 10-20 nucleotides in length and generating a collection of oligonucleotides comprising the oligonucleotides. A method is provided.
[0029] In some embodiments, at least 1000 oligonucleotides are In some embodiments, the domain of nucleotides has a length of 12 to 16 nucleotides. In some embodiments, the domain of nucleotides is 13 to 15 nucleotides in length. In some embodiments, the host is a human. In some embodiments, the background population is genomic. In some embodiments, the background population is the exome. In some embodiments, the determination is performed computationally.
[0030] In another aspect, the disclosure provides a method for identifying a pathogen in a host, comprising: providing a sample enriched for non-host-derived nucleic acids relative to host-derived nucleic acids; and enriching the sample by preferentially extracting nucleic acids greater than about 300 bases in length. analyzing the non-host derived nucleic acid; and identifying the pathogen in the host from the pathogen.
[0031] In some embodiments of the methods described herein, the concentrating step comprises concentrating the soluble fraction of ... , preferentially removing nucleic acids greater than about 150, about 200, or about 250 bases in length from the sample. In some embodiments of the methods described herein, the concentrating step However, the length of about 10 bases to about 60 bases, about 10 bases to about 120 bases, and about 1 0 base length to approximately 150 base length, approximately 10 base length to approximately 300 base length, approximately 30 base length to approximately 60 base length Length, about 30 bases to about 120 bases, about 30 bases to about 150 bases, about 30 bases to about and preferentially concentrating nucleic acids having a length of 200 bases, or between about 30 bases and about 300 bases. In some embodiments of the methods described herein, the enriching step comprises enriching a host-derived In some embodiments of the methods described herein, the method comprises preferentially digesting nucleic acids. In some embodiments, the enrichment step comprises preferentially replicating non-host-derived nucleic acids. In some embodiments of the described methods, the non-host nucleic acid is a nucleic acid that is not present in the host-derived nucleic acid. one or more priming oligos complementary to one or more domains of the nucleotide The nucleic acid sequence is preferentially replicated using a nucleotide or capture oligonucleotide. In some embodiments of the described methods, the domain of nucleotides is 10-20 nucleotides. In some embodiments of the methods described herein, the length of the nucleotide In some embodiments of the methods described herein, the domains are 12 to 15 nucleotides in length. In embodiments, the nucleotide domain is a 13-mer or a 14-mer. In some embodiments of the methods described, the host is a eukaryotic host. In some embodiments of the methods described herein, the host is a vertebrate host. In some embodiments of the methods, the host is a mammalian host. In some embodiments of the method, the non-host DNA comprises DNA from a pathogenic organism. In some embodiments of the method described in In some embodiments of the methods described herein, the sample is an acellular sample. In some embodiments of the methods described herein, the enriching step comprises enriching host-derived The ratio of non-host-derived nucleic acid to native nucleic acid is at least 2-fold, at least 3-fold, at least 4-fold, fold, at least 5 times, at least 6 times, at least 7 times, at least 8 times, at least 9 times times, at least 10 times, at least 11 times, at least 12 times, at least 13 times, at least at least 14 times, at least 15 times, at least 16 times, at least 17 times, at least 1 8-fold, at least 19-fold, at least 20-fold, at least 30-fold, at least 40-fold, at least 50 times, at least 60 times, at least 70 times, at least 80 times, at least 90-fold, at least 100-fold, at least 1000-fold, at least 5000-fold or less Increase it by at least 10,000 times.
[0032] In another aspect, the present disclosure provides a method for identifying a nucleic acid sequence comprising 10 oligonucleotide sequences in a single oligonucleotide. an ultramer oligonucleotide comprising: (a) 10 Each of the oligonucleotide sequences is bounded by a uracil residue or an apurinic / apyrimidinic site. (b) each of the ten oligonucleotide sequences is separated by a domain of nucleotides; (c) the nucleotide domain has a length of 10 to 20 nucleotides; (d) Each domain of nucleotides with a length of 10 to 20 nucleotides is composed of different nucleotides. In some embodiments, ultramer oligonucleotides having a nucleotide sequence are provided. The ten oligonucleotide sequences are 200, 150, 100, 50, 40, 30 or In some embodiments, the length of the oligonucleotide is 10 or less nucleotides. In some embodiments, the sequences have a length of 10 to 20 nucleotides. Each domain of nucleotides present in one or more genomes is absent. In embodiments, the one or more genomes are one or more mammalian genomes. In some embodiments, the one or more mammalian genomes are a human genome, a dog genome, a cat genome, or a mammalian genome. Genome, rodent genome, pig genome, cow genome, sheep genome, goat genome, rabbit genome, horse genome, and any combination thereof. In some embodiments, the one or more mammalian genomes are human genomes. In some embodiments, the one or more genomes is a genome. Each domain of nucleotides having a length of 12 to 15 nucleotides is 12 to 15 nucleotides. In some embodiments, each of the nucleotides having a length of 10 to 20 nucleotides In some embodiments, the domain is 13-15 nucleotides in length. Each domain of nucleotides having a length of nucleotides contains mixed bases. In this embodiment, each domain of nucleotides having a length of 10 to 20 nucleotides is one or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, Contains 10 or more, 11 or more, 12 or more, or 13 or more mixed bases. In embodiments, the mixed bases are N(A,C,G,T), D(A,G,T), V(A,C,G), B(C, G, T), H(A, C, T), W(A, T), S(C, G), K(G, T), M (A, C), Y(C, T), R(A, G), and any combination thereof In some embodiments, the 10 oligonucleotide sequences are selected from the following: In some embodiments, the 10 oligonucleotide sequences have 2, 3, 4, 6, 8, 9, 12, 16, 18, 24, 27, 32, 36, 48, 54, 64, 72, 8 1, 96, 108, 128, 144, 162, 192, 216, 243 or 256 In some embodiments, the 10 oligonucleotide sequences are 2 and 3 In some embodiments, the degeneracy is a prime factor selected from 10 oligonucleotides. In some embodiments, the oligonucleotide sequences are DNA. In some embodiments, 10 oligonucleotide sequences are synthesized. In some embodiments, the ultramer oligonucleotide is about or less All are approximately 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 30 In some embodiments, the oligonucleotide sequence comprises 0, 400, or 500 oligonucleotides. Ultramer oligonucleotides up to approximately 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400 or 500 oligonucleotide sequences In some embodiments, the 10 oligonucleotide sequences comprise one or more Different sequences present in the genome of a microorganism, pathogen, bacterium, virus, fungus or parasite It contains a domain of nucleotides 10 to 20 nucleotides long that has
[0033] In another aspect, the disclosure provides a method of generating a collection of oligonucleotides, comprising: (a) providing an Ultramer oligonucleotide as disclosed herein; (b) hydrolyzing the Ultramer oligonucleotide, thereby and generating a collection of the codes. , that step (b) hydrolyzes the apurinic / apyrimidinic moiety in the ultramer. In some embodiments, step (b) comprises endonuclease IV or endonuclease IV. In some embodiments, the method is carried out using nuclease VII. - Hydrolysis of uracil residues in oligonucleotides to form apurinic / apyrimidinic moieties In some embodiments, the step of hydrolyzing the uracil residues further comprises forming In some embodiments, the step is performed using uracil DNA glycosylase. The method comprises biotinylating oligonucleotides in a collection of oligonucleotides. In some embodiments, the biotinylation is enzymatic (e.g., T 4 using polynucleotide kinase) or chemically. In the present invention, the method comprises: In some embodiments, the probe further comprises attaching a digoxigenin to the target protein. The probes are fluorescent probes or fluorescent probes.
[0034] Incorporation by Reference All publications, patents, and patent applications mentioned herein are to be construed as if each individual publication were incorporated by reference. The product, patent, or patent application is specifically and individually indicated to be incorporated by reference. No. 6,399,423, filed on Oct. 1, 2004, and which is incorporated herein by reference in its entirety to the same extent as if fully set forth herein.
[0035] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments in which the principles of the present invention are utilized. and by reference to the accompanying drawings. [Brief explanation of the drawings]
[0036] [Figure 1] FIG. 1 illustrates a priming strategy that uses a collection of oligonucleotides to enrich a population of nucleic acids. [Figure 2] FIG. 1 shows a pull-down strategy that uses a collection of oligonucleotides to enrich a population of nucleic acids. [Figure 3] FIG. 1 shows how a collection of oligonucleotides can be generated using a hybridization-based method. [Figure 4] Figure 1 shows nucleotide domain lengths and sequence coverage. Approximately 96.5% of nucleotide domains with a length of 13 nucleotides are human. [Figure 5] 1 shows the sequence coverage of a population of interest estimated using a collection of oligonucleotides, showing coverage of the E. coli genome using a collection of oligonucleotides containing non-human nucleotide domains 13 nucleotides in length. [Figure 6] FIG. 1 shows the periodicity of human cell-free DNA length. [Figure 7]FIG. 1 shows nucleosome depletion with nucleosome-specific antibodies. [Figure 8A] FIG. 1 shows a plot of cell-free DNA amount versus DNA fragment length for both host and non-host cell-free DNA. [Figure 8B] FIG. 1 shows the results of the selection and enrichment process for short fragments of cell-free DNA. [Figure 9A] FIG. 1 shows representative Ultramer oligonucleotide designs. [Figure 9B] FIG. 1 shows Ultramer oligonucleotides containing oligonucleotide sequences with N-mixed bases. [Figure 9C] FIG. 1 shows an Ultramer oligonucleotide containing an oligonucleotide sequence with two N-mixed bases. [Figure 9D] FIG. 1 shows an Ultramer oligonucleotide containing an oligonucleotide sequence with two mixed bases. [Figure 10A] FIG. 1 shows a reaction scheme for generating a collection of oligonucleotides from Ultramer oligonucleotides. [Figure 10B] FIG. 1 shows examples of digestion reaction products after digestion of Ultramer oligonucleotides. [Figure 11A] FIG. 1 shows representative oligonucleotides with mixed base sites. [Figure 11B] FIG. 1 shows the analysis of non-human 13-mer probe pools grouped according to degeneracy. [Figure 12A] FIG. 1 is a schematic diagram of a process for enzymatic biotinylation of oligonucleotides. [Figure 12B] FIG. 1 is a schematic diagram of a process for chemical biotinylation of oligonucleotides. DETAILED DESCRIPTION OF THE INVENTION
[0037] overview The present disclosure provides a method for analyzing patient samples in which a predominantly high percentage of the sample is made up of the patient's own nucleic acids. The present disclosure provides a novel and rapid method for concentrating pathogen nucleic acids in samples. This allows for the detection of multiple pathogen nucleic acids in the patient's blood, allowing the patient's caregiver to determine which pathogens are infecting the patient. hypothetical situations, such as situations where you have no clear idea or suspicion about what could have happened. Generally, the compositions provided herein can detect pathogens even in conditions where the pathogen is not readily detectable. The methods and methods involve the use of nucleic acids to increase the representation of pathogen nucleic acids relative to nucleic acids from the host. The present invention is designed to concentrate biological samples containing pathogen nucleic acids. reduces the time and costs associated with analyzing samples, especially if the analysis involves sequencing reactions. can be reduced.
[0038] The present disclosure provides several methods for performing concentration. In some examples, the present disclosure provides methods for performing concentration on samples. A novel collection of oligonucleotides that preferentially bind to a set of non-host nucleic acids in a cell The present invention provides a method for producing a collection of oligonucleotides of multiple different types. It can be used to preferentially detect non-host nucleic acids in molecular biology assays. For example, a collection of oligonucleotides can be used in a primer extension reaction, PCR reaction, or reverse transcription reaction. It can be used as a primer in a transcription reaction. Collection for hybridization and / or pull-down assays can be used to preferentially bind and isolate pathogen nucleic acids. The collection of oligonucleotides is used not only to enrich for pathogen nucleic acids but also to identify such nucleic acids. It can also be used to add tags to acids.
[0039] Additional enrichment techniques are also provided herein. Such techniques may be used alone or in combination with oligonucleotides. This can be used in conjunction with collection or another enrichment method for the leukocyte. Examples of such additional enrichment techniques include: (a) determining whether a major population in a nucleic acid sample is present in the sample; Self-hybridization (self) (f) hybridization technique, (b) nucleosome assembly from free DNA. (c) removal and / or isolation of specific length intervals of DNA; (d) exosomal DNA depletion and (e) strategic capture of targeted regions.
[0040] The enrichment techniques provided herein are directed to detecting circulating HIV-1 antibodies present in blood samples from infected patients. These are particularly suitable for detecting microbial nucleic acids, such as circulating cell-free microbial nucleic acids. The technique also applies to any other method involving the detection of a minority population of nucleic acids in a mixture dominated by a majority population. It is also useful in situations.
[0041] FIG. 1 illustrates a method for detecting pathogens or other non-human nucleic acids in human patient samples, as described herein. In some instances, a blood sample (120) or The plasma sample (130) is free of pathogens (e.g., microorganisms, bacteria, viruses, fungi, or parasites). obtained from a human patient (110) infected (or suspected of being infected) with a bacterial or bacterial infection. The fluid sample is contacted with a collection of oligonucleotides (150) provided herein. The nucleic acid may contain nucleic acids, such as circulating cell-free nucleic acids (140), which may be exposed to oligonucleotides. A collection of otides can preferentially bind to pathogen or non-human nucleic acid sequences (170) The collection of oligonucleotides includes nucleic acid labels, barcodes, and sample-specific barcodes. The code, universal primer sequence, sequencing primer binding site, and barcode amplification primer binding sites, sequencer compatible sequences and / or adapters (e.g. For example, the label may be linked to a label (160), which may include a sequencing adapter. If the nucleic acid sample contains RNA, the collection of oligonucleotides may be To preferentially prime A, cDNA synthesis from an RNA template (170) was performed. In some cases, a collection of oligonucleotides can be used to identify used in a primer extension reaction (170) to preferentially identify pathogen DNA sequences in the sample. Primer extension reactions can also prime overhanging sequences. For example, the label (160) can be attached to the nucleic acid via a primer extension reaction. A sequence (e.g., an adapter, a barcode, etc.) can be added to the pathogen nucleic acid. A CR reaction can be performed to prepare the final library (180), which is then used to identify the Sequencing assays (19) to aid in the detection and identification of pathogen species (195) 0).
[0042] FIG. 2 shows the hybridization and pull-down steps provided herein. 1 shows steps in another method for preparing a sample. The sample may be prepared as in FIG. 1 (210, 220, 230). Samples can be used to prepare sequenceable libraries. The nucleic acids (240) in the sample can be tagged with a label (250). Each of the molecules can be conjugated to a label, such as a biotin tag (270). The oligonucleotides can be contacted with a collection of oligonucleotides (260). The nucleotides hybridize to non-human sequences in the sample (280, 285) and then, e.g., For example, pull-down assays using avidin or streptavidin bound to a solid support. The antibodies can then be preferentially pulled down by the ELISA (285). Libraries were sequenced to identify non-host species (e.g., pathogens) within the host (295). It can be determined (290).
[0043] The methods and compositions provided herein allow for the identification of minor populations of interest in complex mixtures. It offers many advantages over current methods for detecting nucleic acids. One advantage is the library size. Reduce sequencing costs by reducing the size or number of sequence reads analyzed In addition, target oligonucleotides such as the collection of oligonucleotides provided herein By using oligonucleotides, it is possible to specifically target background population nucleic acids (e.g. Compared to other methods, such as those that rely on depleting human nucleic acids from the sample, This can speed up the process of priming, capturing, or amplifying minority populations. These can be used to detect specific populations (e.g., pathogens or microorganisms) in a sample. Furthermore, the sensitivity and specificity of the oligonucleotides can be dramatically increased. The collection can be used to generate DNA or RNA sequencing libraries. In some cases, these can be used to generate DNA sequencing libraries from the same sample. As a result, both RNA sequencing and RNA sequencing libraries can be generated. The provided methods and compositions are useful for treating pneumonia, tuberculosis, HIV infection, hepatitis infection (e.g., A, Hepatitis B or C), sepsis, human papillomavirus (HPV) infection, chlamydia Infectious diseases, syphilis infection, Ebola infection, multidrug-resistant infection, Staphylococcus aureus infection, Enterococcus infection This provides a new and efficient method for detecting infectious diseases, including influenza and flu. In this example, the methods and compositions may be used to treat a variety of conditions, including those that may or may not be experiencing symptoms associated with an infection. The presence of one or more microorganisms in a host, e.g., a microbiome can be monitored.
[0044] Collection of oligonucleotides for sequence enrichment The present disclosure provides a method for concentrating and capturing specific nucleic acid populations ("populations of interest") within a complex mixture of nucleic acids. The present invention provides a collection of oligonucleotides useful for capturing or priming. A target population is a population containing both host (e.g., human) and non-host nucleic acids. The nucleic acid may be a non-human, microbial, or pathogen nucleic acid. is a separate population of nucleic acids ("background") that constitutes a larger portion of a complex mixture of nucleic acids. " population, e.g., host nucleic acids). In some cases, the methods and compositions provided herein may be used to determine whether a population of interest is entirely in a sample. It is particularly useful if it constitutes up to 1% or 0.1% of the nucleic acid.
[0045] Generally, the methods provided herein involve priming, capturing or This involves using oligonucleotides to enrich for specific populations. The oligonucleotides have specific attributes or characteristics and are used to prime and capture the target population. The target nucleic acid may be captured or concentrated in a separate population (e.g., "background") that may constitute a large proportion of the total population of nucleic acids. They are designed not to prime, capture, or enrich the "wound" population. A target population is a non-host (e.g., non-mammalian) population containing both host and non-host nucleic acids. a subset of nucleic acids (human, non-human, pathogen, microorganism, virus, bacteria, fungus, or parasite) In some cases, the oligonucleotide may be present in a background population (e.g., human nucleic acids) to thereby remove, separate, isolate or deplete a population of interest (e.g. , non-human nucleic acids). In some cases, the oligonucleotides Identify (completely or substantially) sequences of a given population of interest (e.g., non-human nucleic acid sequences) of nucleotides having a sequence that may be identical or (fully or substantially) complementary It may include a domain.
[0046] The collection of oligonucleotides provided herein are designed to bind to populations of interest. The oligonucleotide may comprise a nucleic acid sequence capable of identifying or recognizing a population of interest. A collection of samples can be used to identify background populations (e.g., hosts or or human) and samples containing nucleic acids from a population of interest (e.g., non-host or non-human). The oligonucleotides can be used to specifically target and detect populations of interest in a pool. The collection of nucleotides is analyzed by analyzing the non-host nucleic acids in a sample containing both host and non-host nucleic acids. Used as a primer to specifically prime, capture, amplify, replicate, or detect More specifically, in some cases, they can be used to It is possible to prime cDNA synthesis from an existing RNA template. These are used in primer extension reactions to add nucleic acid tags or sequences to target sequences. In some cases, these can be used to generate DNA, cDNA, or RNA libraries. This can be used as bait to capture non-host sequences from the ribosome. The amplified or captured non-host nucleic acids can be subjected to sequencing assays, particularly high-throughput Sequencing assays, next-generation sequencing platforms, massively parallel sequencing sequencing platform, nanopore sequencing assay or another sequencing method known in the art. Identification may be performed by methods known in the art, such as by performing a sequencing assay. It is possible.
[0047] The collection of oligonucleotides may include one or more oligonucleotides. In some cases, the collection of oligonucleotides is about 100; 200; 300; 400;500;600;700;800;900;1000;2000;3000;4 000;5000;6000;7000;8000;9000;10000;20000 ;30000;40000;50000;60000;70000;80000;900 00;100000;200000;300000;400000;500000;60 0000;700000;800000;900000;10 6 ;1.1×10 6 ;1. 2×10 6 ;1.3×10 6 ;1.4×10 6 ;1.5×10 6 ;1.6×10 6 ;1. 7×10 6 ;1.8×10 6 ;1.9×10 6 ;2×10 6 ;2.1×10 6 ;2.2× 10 6;2.3×10 6 ;2.4×10 6 ;2.5×10 6 ;2.6×10 6 ;2.7× 10 6 ;2.8×10 6 ;2.9×10 6 ;3×10 6 ;4×10 6 ;5×10 6 ;6× 10 6 ;7×10 6 ;8×10 6 ;9×10 6 ;10 7 ;5×10 7 ;10 8 ;5×10 8 ; or 10 9 In some cases, the oligonucleotides Chido's collection is up to 100;200;300;400;500;600;700; 800;900;1000;2000;3000;4000;5000;6000;70 00;8000;9000;10000;20000;30000;40000;500 00;60000;70000;80000;90000;100000;200000 ;300000;400000;500000;600000;700000;8000 00;900000;10 6 ;1.1×10 6 ;1.2×10 6 ;1.3×10 6 ;1. 4×10 6 ;1.5×10 6 ;1.6×10 6 ;1.7×10 6 ;1.8×10 6 ;1. 9×10 6 ;2×10 6 ;2.1×10 6 ;2.2×10 6 ;2.3×106 ;2.4× 10 6 ;2.5×10 6 ;2.6×10 6 ;2.7×10 6 ;2.8×10 6 ;2.9× 10 6 ;3×10 6 ;4×10 6 ;5×10 6 ;6×10 6 ;7×10 6 ;8×10 6 ; 9×10 6 ;10 7 ;5×10 7 ;10 8 ;5×10 8 ; or 10 9 Oligonucleotides In some cases, the collection of oligonucleotides may include at least one 00;200;300;400;500;600;700;800;900;1000; 2000;3000;4000;5000;6000;7000;8000;9000; 10000;20000;30000;40000;50000;60000;7000 0;80000;90000;100000;200000;300000;40000 0;500000;600000;700000;800000;900000;10 6 ;1.1×10 6 ;1.2×10 6 ;1.3×10 6 ;1.4×10 6 ;1.5×10 6 ;1.6×10 6 ;1.7×10 6 ;1.8×10 6 ;1.9×10 6 ;2×10 6 ;2 .1×10 6 ;2.2×10 6;2.3×10 6 ;2.4×10 6 ;2.5×10 6 ;2 .6×10 6 ;2.7×10 6 ;2.8×10 6 ;2.9×10 6 ;3×10 6 ;3.5 x10 6 ;4×10 6 ;4.5×10 6 ;5×10 6 ;5.5×10 6 ;6×10 6 ;6 .5×10 6 ;7×10 6 ;7.5×10 6 ;8×10 6 ;8.5×10 6 ;9×10 6 ;9.5×10 6 ;10 7 ;5×10 7 ;10 8 ;5×10 8 ; or 10 9 Oligonucleotides It may contain nucleotides.
[0048] The oligonucleotides in a collection of oligonucleotides may have the same length but may be different. In some cases, two or more of the oligonucleotides in the collection may have a length of The oligonucleotides above have the same length. All oligonucleotides in a collection have the same length. The oligonucleotides in the collection of nucleotides have different lengths. If the oligonucleotides in the collection of oligonucleotides are about 1, 2, 3, 4, Available in 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or 20 different lengths In some cases, the oligonucleotide in the collection of oligonucleotides is the most Large about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or In some cases, the length of the oligonucleotides in the collection of oligonucleotides is Nucleotides of at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 1 They have lengths of 3, 14, 15 or 20.
[0049] In some cases, the oligonucleotide is about 10, 11, 12, 13, 14, 15, 16 , 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170 , 180, 190, 200, 210, 220, 230, 240, 250, 260, 270 , 280, 290, 300, 400, 500, 600, 700, 800, 900 or 1 In some cases, the oligonucleotide may be up to 10, 100, 1000 nucleotides in length. 1, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 , 25, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130 , 140, 150, 160, 170, 180, 190, 200, 210, 220, 230 , 240, 250, 260, 270, 280, 290, 300, 400, 500, 600 , 700, 800, 900 or 1000 nucleotides in length. The oligonucleotides must be at least 10, 11, 12, 13, 14, 15, 16, 17, 1 8, 19, 20, 21, 22, 23, 24, 25, 30, 40, 50, 60, 70, 80 , 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 400, 500, 600, 700, 800, 900 or 1000 Nu In some cases, the oligonucleotides can be 10 to 1000, 10 to 1000, 500, 10-400, 10-300, 10-200, 10-100, 10-90, 10 ~80, 10~70, 10~60, 10~50, 10~40, 10~30, 10~20, 10~15, 11~1000, 11~500, 11~400, 11~300, 11~20 0, 11~100, 11~90, 11~80, 11~70, 11~60, 11~50, 1 1~40, 11~30, 11~20, 11~15, 12~1000, 12~500, 12 ~400, 12~300, 12~200, 12~100, 12~90, 12~80, 12 ~70, 12~60, 12~50, 12~40, 12~30, 12~20, 12~15, 13~1000, 13~500, 13~400, 13~300, 13~200, 13~1 00, 13~90, 13~80, 13~70, 13~60, 13~50, 13~40, 1 3~30, 13~20, 13~15, 13~14, 14~1000, 14~500, 14 ~400, 14~300, 14~200, 14~100, 14~90, 14~80, 14 ~70, 14~60, 14~50, 14~40, 14~30, 14~20 or 14~1 It may be 5 nucleotides in length.
[0050] The oligonucleotides in the collection of oligonucleotides have different nucleotide sequences In some cases, the collection of oligonucleotides may have different nucleotides. having a sequence of approximately 100;200;300;400;500;600;700;800;9 00;1000;2000;3000;4000;5000;6000;7000;80 00;9000;10000;20000;30000;40000;50000;60 000;70000;80000;90000;100000;200000;3000 00;400000;500000;600000;700000;800000;90 0000;10 6 ;1.1×10 6 ;1.2×10 6 ;1.3×10 6 ;1.4×10 6 ;1.5×10 6 ;1.6×10 6 ;1.7×10 6 ;1.8×10 6 ;1.9×10 6 ;2×10 6 ;2.1×10 6 ;2.2×10 6 ;2.3×10 6 ;2.4×10 6 ;2 .5×10 6 ;2.6×10 6 ;2.7×10 6 ;2.8×10 6 ;2.9×10 6 ;3 x10 6 ;4×10 6 ;5×10 6 ;6×10 6 ;7×10 6 ;8×10 6 ;9×10 6 ; or 10 7 In some cases, the oligonucleotides The collection of genes is up to 100; 200; 300; 40 with different nucleotide sequences. 0;500;600;700;800;900;1000;2000;3000;400 0;5000;6000;7000;8000;9000;10000;20000;3 0000;40000;50000;60000;70000;80000;90000 ;100000;200000;300000;400000;500000;6000 00;700000;800000;900000;10 6 ;1.1×10 6 ;1.2× 10 6 ;1.3×10 6 ;1.4×10 6 ;1.5×10 6 ;1.6×10 6 ;1.7× 10 6 ;1.8×10 6 ;1.9×10 6 ;2×10 6 ;2.1×10 6 ;2.2×10 6 ;2.3×10 6 ;2.4×10 6 ;2.5×10 6 ;2.6×10 6 ;2.7×10 6 ;2.8×10 6 ;2.9×10 6 ;3×10 6 ;4×10 6 ;5×10 6 ;6×10 6 ;7×10 6 ;8×10 6 ;9×10 6 ; or 10 7 Contains oligonucleotides In some cases, the collection of oligonucleotides contains different nucleotide sequences. Have at least 100; 200; 300; 400; 500; 600; 700; 800; 900;1000;2000;3000;4000;5000;6000;7000;8 000;9000;10000;20000;30000;40000;50000;6 0000;70000;80000;90000;100000;200000;300 000;400000;500000;600000;700000;800000;9 00000;10 6 ;1.1×10 6 ;1.2×10 6 ;1.3×10 6 ;1.4×10 6 ;1.5×10 6 ;1.6×10 6 ;1.7×10 6 ;1.8×10 6 ;1.9×10 6 ;2×10 6 ;2.1×10 6 ;2.2×10 6 ;2.3×10 6 ;2.4×10 6 ; 2.5×10 6 ;2.6×10 6 ;2.7×10 6 ;2.8×10 6 ;2.9×10 6 ; 3×10 6 ;4×10 6 ;5×10 6 ;6×10 6 ;7×10 6 ;8×10 6 ;9×10 6 ; or 10 7 The oligonucleotides may comprise:
[0051] Oligonucleotides with different sequences are multiplexed in a collection of oligonucleotides. A collection of oligonucleotides can be present in multiple copies. In some cases, the oligonucleotide may contain a sequence of Rection is approximately 1;2;3;4;5;6;7;8;9;10;11;12;13;14; 15;16;17;18;19;20;25;30;35;40;45;50;55;6 0;65;70;75;80;85;90;95;100;150;200;250;3 00;350;400;450;500;600;700;800;900;or 10 In some cases, the oligonucleotide may contain 00 copies of the same nucleotide sequence. Collection up to 1;2;3;4;5;6;7;8;9;10;11;12;13;1 4;15;16;17;18;19;20;25;30;35;40;45;50;55 ;60;65;70;75;80;85;90;95;100;150;200;250 ;300;350;400;450;500;600;700;800;900;or In some cases, the oligonucleotide may contain 1000 copies of the same nucleotide sequence. Collection of at least 1;2;3;4;5;6;7;8;9;10;11;12 ;13;14;15;16;17;18;19;20;25;30;35;40;45; 50;55;60;65;70;75;80;85;90;95;100;150;20 0;250;300;350;400;450;500;600;700;800;90 or 1000 copies of the same nucleotide sequence.
[0052] In some cases, one or more oligonucleotides in the collection of oligonucleotides The oligonucleotide is unlabeled; in some cases, one in a collection of oligonucleotides Alternatively, multiple oligonucleotides may be labeled. In some cases, one or more of the oligonucleotides may be labeled. The oligonucleotides may be labeled, for example, with a nucleic acid label, a chemical label, or an optical label. In this case, one or more oligonucleotides are conjugated to a solid support. In some cases, one or more oligonucleotides are conjugated to a solid support. In some cases, the label is attached to the 5' or 3' end of the oligonucleotide, or In some cases, the oligonucleotide may be attached to the interior of the oligonucleotide. One or more oligonucleotides in the collection of oligonucleotides are labeled with two or more labels. do.
[0053] The oligonucleotide may comprise a nucleic acid label. In some cases, the nucleic acid label is one of the following: or multiple: barcode (e.g., sample barcode), universal platform primer sequences, primer binding sites (for example, but not limited to, DNA sequencing primers) primer binding sites, sample barcode sequencing primer binding sites and various sequencing Sequencing or barcode sequencing kits containing amplification primer binding sites compatible with platform requirements. (for read-through), sequencer-compatible sequences, and sequencing platform attachment sequence, sequencing adapter sequence or adapter. Nucleic acid labels can be used (e.g., Several of these can be attached to oligonucleotides (by polymerase or synthetic design). In some cases, the length of the nucleic acid label is included in the length of the oligonucleotide. The length of the primer is not included in the length of the oligonucleotide.
[0054] The oligonucleotide may include a chemical label. Some non-limiting examples of chemical labels include: Biotin, avidin, streptavidin, radiolabels, polypeptides and polymers The oligonucleotide may include an optical label. Some non-limiting examples of optical labels include: Typical examples include fluorophores, fluorescent proteins, dyes and quantum dots. The oligonucleotide can be conjugated to a solid support. The oligonucleotides are not conjugated to the solid support. Non-limiting examples include beads, magnetic beads, polymers, slides, chips, surfaces, plates, etc. These include gates, channels, cartridges, microfluidic devices and microarrays. One or more oligonucleotides are attached to the immobilized substrate for affinity chromatography. In some cases, each oligonucleotide may be conjugated to a solid support. In some cases, each oligonucleotide has the same label. In some cases, each copy of the oligonucleotide has the same label. Each oligonucleotide having a different sequence has a different label.
[0055] The oligonucleotides provided herein generally comprise one or more nucleotides. In some cases, the nucleotides include deoxyribonucleotides (e.g., A, C, G or T), ribonucleotides (e.g., A, C, G, or U), modified nucleotides or may be synthetic nucleotides. In some cases, the oligonucleotide may be DNA, RNA, or , cDNA, dsDNA, ssDNA, mRNA or cRNA. In some cases, the oligonucleotide may comprise DNA or RNA. Nucleotides can include DNA and RNA. In some cases, oligonucleotides are In some cases, the oligonucleotide may comprise RNA. In some cases, the oligonucleotide may contain modified or synthetic nucleotides. In some cases, oligonucleotides are peptide nucleic acids (PNAs), locked nucleic acids (LNAs), or or bridged nucleic acids (BNAs). In some cases, the oligonucleotide may be DNA, RNA, PNA, LNA, BNA, or any of these. In some cases, the oligonucleotide may be artificially fragmented. In some cases, the collection of oligonucleotides does not include artificially synthesized nucleic acids. In some cases, the oligonucleotide does not contain fragmented nucleic acids. In some cases, a collection of oligonucleotides. Nucleic acids include synthetic nucleic acids (e.g., by DNA synthesis). The collection of oligonucleotides can be lyophilized or dried. The collection of nucleotides can include water or a buffer solution (eg, an aqueous buffer solution).
[0056] In some cases, the oligonucleotide is about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 , 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 1 30, 140, 150, 160, 170, 180, 190, 200, 210, 220, 2 30, 240, 250, 260, 270, 280, 290, 300, 400, 500, 6 00, 700, 800, 900, 1000 or more PNAs, LNAs and / or In some cases, the oligonucleotide may contain about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 5 5, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96 , 97, 98, 99 or 100% PNA, LNA and / or BNA binding. It is possible.
[0057] Nucleotide domains An oligonucleotide may contain one or more domains of nucleotides. In some cases, the oligonucleotide contains one domain of nucleotides. In some cases, the domain of nucleotides is an oligonucleotide. The portion of the oligonucleotide having the nucleotide sequence appears within the domain of nucleotides. In some cases, portions of the oligonucleotide having different nucleotide sequences are In some cases, oligonucleotides with different nucleotide sequences are A portion of a nucleotide is an oligonucleotide, such as a domain of nucleotides within an oligonucleotide. In some cases, a collection of oligonucleotides The oligonucleotides may comprise nucleic acids having different sequences. The oligonucleotide may be present in multiple copies. In some cases, the domain of nucleotides does not contain artificially fragmented nucleic acids. In some cases, the domain of nucleotides is synthesized (eg, by DNA synthesis).
[0058] In some cases, the collection of oligonucleotides comprises oligonucleotides Each oligonucleotide contains a domain of nucleotides with a different nucleotide sequence. In some cases, the collection of oligonucleotides is about 100; 200; 3 00;400;500;600;700;800;900;1000;2000;300 0;4000;5000;6000;7000;8000;9000;10000;20 000;30000;40000;50000;60000;70000;80000; 90000;100000;200000;300000;400000;500000 ;600000;700000;800000;900000;10 6 ;1.1×10 6 ;1.2×10 6 ;1.3×10 6 ;1.4×10 6 ;1.5×10 6 ;1.6×10 6 ;1.7×10 6 ;1.8×10 6 ;1.9×10 6 ;2×10 6 ;2.1×10 6 ;2 .2×10 6 ;2.3×10 6 ;2.4×10 6 ;2.5×10 6 ;2.6×10 6 ;2 .7×10 6 ;2.8×10 6 ;2.9×10 6 ;3×10 6 ;4×10 6 ;5×10 6 ;6×106 ;7×10 6 ;8×10 6 ;9×10 6 ; or 10 7 Oligonucleotides Each oligonucleotide may contain a nucleotide having a different nucleotide sequence. In some cases, the collection of oligonucleotides may contain up to 100;200;300;400;500;600;700;800;900;1000 ;2000;3000;4000;5000;6000;7000;8000;9000 ;10000;20000;30000;40000;50000;60000;700 00;80000;90000;100000;200000;300000;4000 00;500000;600000;700000;800000;900000;10 6 ;1.1×10 6 ;1.2×10 6 ;1.3×10 6 ;1.4×10 6 ;1.5×10 6 ;1.6×10 6 ;1.7×10 6 ;1.8×10 6 ;1.9×10 6 ;2×10 6 ; 2.1×10 6 ;2.2×10 6 ;2.3×10 6 ;2.4×10 6 ;2.5×10 6 ; 2.6×10 6 ;2.7×10 6 ;2.8×10 6 ;2.9×10 6 ;3×10 6 ;4× 10 6 ;5×10 6 ;6×10 6 ;7×106 ;8×10 6 ;9×10 6 ; or 10 7 Each oligonucleotide may contain a different nucleotide. In some cases, the oligonucleotide comprises a domain of nucleotides having a sequence of The collection is at least 100;200;300;400;500;600;700; 800;900;1000;2000;3000;4000;5000;6000;70 00;8000;9000;10000;20000;30000;40000;500 00;60000;70000;80000;90000;100000;200000 ;300000;400000;500000;600000;700000;8000 00;900000;10 6 ;1.1×10 6 ;1.2×10 6 ;1.3×10 6 ;1. 4×10 6 ;1.5×10 6 ;1.6×10 6 ;1.7×10 6 ;1.8×10 6 ;1. 9×10 6 ;2×10 6 ;2.1×10 6 ;2.2×10 6 ;2.3×10 6 ;2.4× 10 6 ;2.5×10 6 ;2.6×10 6 ;2.7×10 6 ;2.8×10 6 ;2.9× 10 6 ;3×10 6 ;4×10 6 ;5×10 6 ;6×10 6 ;7×10 6 ;8×106 ; 9×10 6 ; or 10 7 Each oligonucleotide may contain Otides contain domains of nucleotides with different nucleotide sequences.
[0059] The oligonucleotides in the collection of oligonucleotides have the same nucleotide sequence. In some cases, the oligonucleotide may comprise a domain of nucleotides having a sequence of nucleotides. The collection contains approximately 100 domains of nucleotides that have identical nucleotide sequences. 1;2;3;4;5;6;7;8;9;10;11;12;13;14;15;16;1 7;18;19;20;25;30;35;40;45;50;55;60;65;70 ;75;80;85;90;95;100;150;200;250;300;350; 400; 450; 500; 600; 700; 800; 900; or 1000 oligos In some cases, the collection of oligonucleotides may contain identical containing a domain of nucleotides having a nucleotide sequence of up to 1;2;3;4;5 ;6;7;8;9;10;11;12;13;14;15;16;17;18;19;2 0;25;30;35;40;45;50;55;60;65;70;75;80;85 ;90;95;100;150;200;250;300;350;400;450;5 Contains 00; 600; 700; 800; 900; or 1000 oligonucleotides In some cases, the collection of oligonucleotides may contain identical nucleotide sequences. containing domains of nucleotides having at least 1;2;3;4;5;6;7;8 ;9;10;11;12;13;14;15;16;17;18;19;20;25;3 0;35;40;45;50;55;60;65;70;75;80;85;90;95 ;100;150;200;250;300;350;400;450;500;600 ; 700; 800; 900; or 1000 oligonucleotides.
[0060] In some cases, each oligonucleotide in the collection of oligonucleotides is may contain domains of nucleotides that are not present in one or more background populations In some cases, each oligonucleotide in the collection of oligonucleotides is Domains of nucleotides not present in the primary genome, exome, or transcriptome In some cases, each oligonucleotide in the collection of oligonucleotides may contain Otides are expressed in one or more vertebrate genomes, exomes or transcriptomes. In some cases, the oligonucleotide may contain domains of nucleotides that are not present in the Each oligonucleotide in the collection is a sequence encoding one or more mammalian genomes, exons, It may contain domains of nucleotides that are not present in the genome or transcriptome. In some cases, each oligonucleotide in a collection of oligonucleotides may be or multiple nucleosides not present in the human genome, exome, or transcriptome In some cases, each of the oligonucleotides in the collection may comprise a domain of The oligonucleotides are derived from one or more human and one or more bacterial genomes, may contain domains of nucleotides that are not present in the chromatome or transcriptome In some cases, each oligonucleotide in the collection of oligonucleotides is one or more human and one or more viral genomes, exomes or transcripts It may contain domains of nucleotides that are not present in the scriptome.
[0061] In some cases, the domain of nucleotides is identical to one or more populations of interest. , may contain substantially identical, complementary or nearly complementary nucleotide sequences. A domain of nucleotides may be identical, nearly identical, complementary, or identical to a subset of the population of interest. may contain nearly complementary nucleotide sequences. In some cases, a domain of nucleotides is approximately 0.0001, 0.0005, 0.001, 0.005 of the nucleic acid in the population of interest. ,0.01,0.05,0.1,0.2,0.3,0.4,0.5,0.6,0.7,0 0.8, 0.9, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 6, 7, 8, 9 , 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 and can contain 100% identical, nearly identical, complementary or nearly complementary nucleotide sequences. In some cases, the domain of nucleotides is up to 0.000 of the nucleic acids in the population of interest. 1, 0.0005, 0.001, 0.005, 0.01, 0.05, 0.1, 0.2, 0 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, 2, 2.5, 3 , 3.5, 4, 4.5, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 4 0, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93 , 94, 95, 96, 97, 98, 99 or 100% identical, nearly identical, complementary or In some cases, the domain of nucleotides may be , at least 0.0001, 0.0005, 0.001, 0. 005, 0.01, 0.05, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0. 7, 0.8, 0.9, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 9 9 or 100% identical, nearly identical, complementary or nearly complementary nucleotide sequences In some cases, the domain of nucleotides is about 0.0001 of the nucleotide sequence. ,0.0005,0.001,0.005,0.01,0.05,0.1,0.2,0. 3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40 , 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100% identical to the nucleic acid in the population of interest; It may contain nucleotide sequences that are substantially identical, complementary, or nearly complementary. ,The nucleotide domain is a nucleotide sequence with a maximum of 0.0001, 0.0005, 0 0.001, 0.005, 0.01, 0.05, 0.1, 0.2, 0.3, 0.4, 0.5 ,0.6,0.7,0.8,0.9,1,1.5,2,2.5,3,3.5,4,4.5 , 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55 , 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100% identical, nearly identical, complementary or non-identical to the nucleic acids in the population of interest In some cases, the dots of nucleotides may be identical or nearly complementary. The main sequence is at least 0.0001, 0.0005, 0.001, 0 0.005, 0.01, 0.05, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0 0.7, 0.8, 0.9, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 6, 7 , 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65 , 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100% identical, nearly identical, complementary or nearly complementary to the nucleic acids in the population of interest In some cases, the domain of nucleotides may comprise a nucleotide sequence that is complementary to the Identical, nearly identical, complementary, or nearly complementary to one or more target populations. In some cases, the domain of nucleotides may include a nucleotide sequence not shown in Figure 4 and As shown in Figure 5, the inclusiveness of the target population can be maintained.
[0062] In some cases, the domain of nucleotides is a subset of one or more background populations. It may contain nucleotide sequences that are not identical, nearly identical, complementary, or nearly complementary to In some cases, the domain of nucleotides is a subset of the background population. It may contain nucleotide sequences that are not identical, nearly identical, complementary, or nearly complementary. In some cases, the domain of nucleotides is approximately 0.01% of the nucleic acids in the background population. 0001, 0.0005, 0.001, 0.005, 0.01, 0.05, 0.1, 0. 2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, 2, 2. 5, 3, 3.5, 4, 4.5, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 3 5, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92 , 93, 94, 95, 96, 97, 98, 99 or 100% identical or nearly identical In some cases, the nucleotide sequence may be complementary or non-complementary. The domain of the nucleic acid in the background population is up to 0.0001, 0.000 5, 0.001, 0.005, 0.01, 0.05, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50 , 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100% identical, nearly identical, complementary, or nearly identical In some cases, the domain of nucleotides may contain nucleotide sequences that are not complementary. , at least 0.0001, 0.0005, 0.001 of nucleic acids in the background population ,0.005,0.01,0.05,0.1,0.2,0.3,0.4,0.5,0.6 ,0.7,0.8,0.9,1,1.5,2,2.5,3,3.5,4,4.5,5,6 ,7,8,9,10,15,20,25,30,35,40,45,50,55,60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 9 8, 99 or 100% identical, nearly identical, complementary or nearly complementary In some cases, the domain of nucleotides may comprise a nucleotide sequence. Array of approximately 0.0001, 0.0005, 0.001, 0.005, 0.01, 0.05, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1. 5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 6, 7, 8, 9, 10, 15, 20, 2 5, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90 , 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100% is background Nucleic acids that are not identical, nearly identical, complementary, or nearly complementary to the nucleic acids in the group In some cases, the domain of nucleotides may comprise a nucleotide sequence Maximum of 0.0001, 0.0005, 0.001, 0.005, 0.01, 0.05, 0 .1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5 , 2, 2.5, 3, 3.5, 4, 4.5, 5, 6, 7, 8, 9, 10, 15, 20, 25 , 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100% is background Nucleotides that are not identical, nearly identical, complementary, or nearly complementary to the nucleic acids in the strand population. In some cases, the domain of nucleotides may comprise a sequence of nucleotides. At least 0.0001, 0.0005, 0.001, 0.005, 0.01, 0.05 ,0.1,0.2,0.3,0.4,0.5,0.6,0.7,0.8,0.9,1,1 .5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 9 0, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100% is background Nucleic acids that are not identical, nearly identical, complementary, or nearly complementary to the nucleic acids in the round population In some cases, the domain of nucleotides may comprise one or more Nucleotide sequences identical, nearly identical, complementary or nearly complementary to a background population of may include:
[0063] In some cases, the domain of nucleotides is a subset of one or more background populations. It may contain nucleotide sequences that bind with mismatches to the nucleic acid. The domains of the tid are 1, 2, 3, 4, 5 with one or more background population nucleic acids. , 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 2 nucleotide sequences that bind with 0, 21, 22, 23, 24, or 25 mismatches. In some cases, the domain of nucleotides may be present in one or more background regions. Do population nucleic acid and up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 1 4, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 mistakes In some cases, the domain of nucleotides may be included in the nucleotide sequence that binds the match. The nucleic acid may be at least 1, 2, 3, 4, 5, or 6 with one or more background population nucleic acids. 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 , 21, 22, 23, 24 or 25 mismatched nucleotide sequences obtain.
[0064] The domain of nucleotides can be one or more in length. The nucleotide domains are of a single length. In some cases, the nucleotide domains are different. In some cases, the domains are 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 nuclei In some cases, the domain of nucleotides may be 13 or 14 nucleotides long. In some cases, the nucleotide domain is 13 nucleotides long. In some cases, the domain of nucleotides is 14 nucleotides long. In this case, the domain of nucleotides is up to 10, 11, 12, 13, 14, 15, 16, 17 The length of the fragment may be 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides. In some cases, the nucleotide domains are up to 13 or 14 nucleotides long. In some cases, the domain of nucleotides is up to 13 nucleotides long. In some cases, the nucleotide domain is up to 14 nucleotides long. The domains of the otides are at least 10, 11, 12, 13, 14, 15, 16, 17, 1 It can be 8, 19, 20, 21, 22, 23, 24 or 25 nucleotides in length. In some cases, the domain of nucleotides is at least 13 or 14 nucleotides in length. In some cases, the domain of nucleotides is at least 13 nucleotides in length. In some cases, the nucleotide domain is at least 14 nucleotides long. In the case of , the nucleotide domains are 10–20, 11–20, 12–20, 13–20, 14-20, 10-19, 10-18, 10-17, 10-16, 10-15, 10-1 4, 10-13, 11-19, 12-19, 13-19, 14-19, 11-18, 11 ~17, 11~16, 11~15, 11~14, 11~13, 12~18, 12~17, 12-16, 12-15, 12-14, 12-13, 13-18, 13-17, 13-1 6, 13-15, 13-14, 14-18, 14-17, 14-16 or 14-15 In some cases, the domain of nucleotides may be 12-15 nucleotides in length. In some cases, the nucleotide domain is 13-15 nucleotides long. In some cases, the nucleotide domain is 13-14 nucleotides in length. In some cases, the nucleotide domains are 10-mers, 11-mers, 12-mers, 13-mers, mer, 14mer, 15mer, 16mer, 17mer, 18mer, 19mer, 20mer, 21mer, 22mer, 23mer, 24mer, or 25mer k-mers. In some cases, the domain of nucleotides is a 13-mer or a 14-mer. In some cases, the nucleotide domain is a 13-mer. In this case, the nucleotide domain is a 14-mer.
[0065] Background population As used herein, a background population generally refers to a population of interest. A background population of nucleic acids is used to generate a collection of oligonucleotides. For example, a background population of nucleic acids can be used to generate a population of interest. 2. Creating a collection of oligonucleotides capable of hybridizing to the population of (e.g., background population nucleic acids can be used to identify starting nucleic acids in oligonucleotides. Use as "bait" to retrieve oligonucleotides from a species collection In some cases, the background population is the population of nucleic acids in the sample. The methods and compositions provided herein can be used to detect background levels of nucleic acids in a sample. In some cases, the methods and compositions can be used to preferentially remove or isolate populations of The target population is isolated, removed, or detected using a method that does not affect background populations. It is not always possible to specifically target a target population in a preferential manner. can.
[0066] In some cases, the background population is the host organism or the host's genome, exome, or may be derived from the sequence of the transcriptome (e.g., the sequences present in this sequence, and contains sequences that are identical, nearly identical, complementary or nearly complementary to this sequence). In some cases, the host may be a mammal, a human, a non-human mammal, or a domestic animal (e.g., a laboratory animal). The animal may be a domestic animal (e.g., a domestic pet or livestock), or a non-domestic animal (e.g., a wild animal). In some cases, the hosts are dogs, cats, rodents, mice, hamsters, cows, birds, and chickens. Birds, pigs, horses, goats, sheep, rabbits, microorganisms, pathogens, bacteria, viruses, fungi or may be a parasite. In some cases, the host is a mammal. In some cases, the host The host may be a human. The host may be a patient. In some cases, the host may be administered an antimicrobial, antibacterial, or In some cases, the host may be treated with an antimicrobial, antiviral, or antiparasitic agent. may be or have been treated with an antibacterial, antiviral or antiparasitic agent In some cases, the host may be (e.g., one or more microorganisms, pathogens, bacteria, In some cases, the host is infected (e.g., by a virus, fungus, or parasite). Not infected with one or more microorganisms, pathogens, bacteria, viruses, fungi or parasites In some cases, the host is healthy. In some cases, the host is susceptible or infected. There is a risk of contamination.
[0067] In some cases, background populations include dogs, cats, rodents, mice, and hamsters. -, cattle, birds, chickens, pigs, horses, goats, sheep, rabbits, microorganisms, pathogens, bacteria , viral, fungal or parasitic origin. In some cases, the background population In some cases, the background population is human. In some cases, the background population may be derived from the host. The analysis involves multiple populations, such as human and microbial populations of nucleic acids (e.g., human only, human and viral, human Humans and bacteria, humans and fungi, humans and parasites, etc.
[0068] In some cases, the background population may include one or more genomes, exomes, and In some cases, the background population may be Approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90 or 100 genomes, exomes and / or can be the transcriptome. In some cases, the background population can be up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90 or 100 genomes, exomes and / or transcripts In some cases, the background population may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90 or 100 genomes, exomes and / or transcripts It can be a transcriptome.
[0069] In some cases, the background population consists of mammalian populations from multiple individual mammals of the same species. It may be an animal genome, exome, or transcriptome. A background population is a set of mammalian genomes from multiple individual mammals of one or more species. In some cases, the background A population is a collection of one or more male and one or more female mammals of the same species. The genome may be a mammalian genome, exome, or transcriptome. The population may be mammalian genomic DNA and mammalian RNA. The background population is a mammalian genome, exome or transcriptome and one Or multiple microbial genomes, exomes, or transcriptomes. In some cases, the background population is a mammalian genome, exome, or transcriptome. in the genome and one or more pathogen genomes, exomes or transcriptomes In some cases, the background population may be a mammalian genome, exome, or The transcriptome and one or more bacterial genomes, exomes, or transmems In some cases, the background population may be a mammalian genome, Exome or transcriptome and one or more viral genomes, exo In some cases, the background population may be , a mammalian genome, exome or transcriptome and one or more In some cases, the viral genome, exome, or transcriptome may be , background populations can be mammalian genomes, exomes or transcriptomes and and one or more viral and one or more bacterial genomes, exomes or transcripts In some cases, the background population may be a mammalian genome. the genome, exome or transcriptome and one or more parasite genomes, may be an exome or a transcriptome. In some cases, one or more Approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, It can be 45, 50, 60, 70, 80, 90 or 100. In some cases, it can be 1 or or multiple up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, It can be 35, 40, 45, 50, 60, 70, 80, 90 or 100. If one or more is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90 or 100 In some cases, mammalian genomes, exomes, or transcriptomes may be non- The human genome, exome, or transcriptome.
[0070] In some cases, the mammalian genome, exome, or transcriptome is the human genome. In some cases, the background The population is the human genome, exome, or transcriptome. Background populations are human genomes, exomes, or transgenes from multiple individual humans. In some cases, the background population may be one or more Human genomes, exomes, or transcripts from a male and one or more females The background population may be human genomic DNA and human RNA. In some cases, the background population is a human genome, exome, or transcriptome. The scriptome and one or more microbial genomes, exomes or transcriptomes In some cases, the background population may be the human genome, exome, or transcriptome and one or more pathogen genomes, exomes or transcriptomes In some cases, the background population may be the human genome. , exome or transcriptome and one or more bacterial genomes, exome In some cases, the background population may be a genome or a transcriptome. Human genome, exome or transcriptome and one or more viral genomes It can be the genome, exome, or transcriptome. The round population is the human genome, exome, or transcriptome and one or more It can be several retroviral genomes, exomes or transcriptomes. In some cases, the background population is the human genome, exome, or transcriptome. and one or more viral and one or more bacterial genomes, exomes or In some cases, the background population may be a human genome. the genome, exome or transcriptome and one or more parasite genomes, may be an exome or a transcriptome. In some cases, one or more Approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, It can be 45, 50, 60, 70, 80, 90 or 100. In some cases, it can be 1 or or multiple up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, It can be 35, 40, 45, 50, 60, 70, 80, 90 or 100. If one or more is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90 or 100 It is possible.
[0071] In some cases, the background population is the genome, exome, or transcriptome. Approximately 0.000001, 0.000005, 0.00001, 0.00005, 0. 0001, 0.0005, 0.001, 0.005, 0.01, 0.05, 0.1, 0. 2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60 , 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, In some cases, the background population may comprise 98, 99, or 100% of the genome. Maximum 0.000001, 0.000005 for genome, exome or transcriptome , 0.00001, 0.00005, 0.0001, 0.0005, 0.001, 0.0 0.5, 0.01, 0.05, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7 ,0.8,0.9,1,2,3,4,5,6,7,8,9,10,15,20,25,3 0, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 91 , 92, 93, 94, 95, 96, 97, 98, 99 or 100%. In some cases, the background population is a genome, exome, or transcriptome. At least 0.000001, 0.000005, 0.00001, 0.00005, 0 .0001, 0.0005, 0.001, 0.005, 0.01, 0.05, 0.1, 0 .2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 2, 3, 4, 5 , 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 6 0, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97 , 98, 99 or 100%.
[0072] In some cases, the background population is about 0.01 of the total population of nucleic acids in the sample; 0.05, 0.1, 0.5, 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35 , 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 99.1, 99.2, 99.3, 99.4 , 99.5, 99.6, 99.7, 99.8 or 99.9%. In this case, the background population is at most 0.01, 0.05, or 0.1, 0.5, 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 4 5, 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94 , 95, 96, 97, 98, 99, 99.1, 99.2, 99.3, 99.4, 99.5 , 99.6, 99.7, 99.8 or 99.9%. The background population is at least 0.01, 0.05, 0.06, or 0.1% of the total population of nucleic acids in the sample. 1, 0.5, 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 9 5, 96, 97, 98, 99, 99.1, 99.2, 99.3, 99.4, 99.5, 9 It may constitute 9.6, 99.7, 99.8 or 99.9%.
[0073] In some cases, the background population is DNA, RNA, cDNA, mRNA, or cR NA, dsDNA, ssDNA, miRNA, circular nucleic acid, circular DNA, circular RNA, microorganism cell nucleic acid, cell-free DNA, cell-free RNA, circulating cell-free DNA, circulating cell-free RNA or genome The background population may contain genomic DNA. The background population may be a mixture of DNA and RNA. The background population can be genomic DNA. In some cases, the background The population includes artificially fragmented nucleic acids. In some cases, the background population is Does not contain artificially fragmented nucleic acids. In some cases, background populations (e.g. In some cases, background population nucleic acids may be present. is synthesized (e.g., by DNA synthesis). In some cases, the background population can be a genome, exome or transcriptome.
[0074] In some cases, one or more nucleic acids in the background population are not labeled; In some cases, one or more nucleic acids in the background population are labeled. In this case, one or more nucleic acids in the background population may be, for example, a nucleic acid label, a chemical label, or the like. In some cases, one or more of the Multiple nucleic acids are conjugated to a solid support. In some cases, the label is removed from the background. at the 5' or 3' end of the nucleic acid in the background population or within the nucleic acid in the background population In some cases, one or more of the target proteins in the background population can be attached. The nucleic acid is labeled with two or more labels.
[0075] Nucleic acids in the background population may contain nucleic acid labels. Acid labels may include one or more of the following: barcodes (e.g., sample barcodes) ), universal primer sequences, primer binding sites (e.g., but not limited to, , DNA sequencing primer binding site, sample barcode sequencing primer binding site and amplification primer binding sites that are compatible with various sequencing platform requirements. sequence determination or barcode reading), sequencer compatibility sequences, sequencing program A sequence to which the nucleic acid is attached, a sequencing adapter sequence or an adapter. Attached to nucleic acids in the background population (e.g., by ligation or synthetic design) In some cases, the length of the nucleic acid label can be adjusted to match the length of the nucleic acid label in the background population. In some cases, the length of the nucleic acid label is included in the length of the nucleic acid in the background population. Not included in the length.
[0076] Nucleic acids in the background population may contain chemical labels. Some non-limiting examples of chemical labels include: Examples include biotin, avidin, streptavidin, radiolabels, polypeptides and Nucleic acids in the background population may contain optical labels. Some non-limiting examples of labels include fluorophores, fluorescent proteins, dyes, and quantum dots. Nucleic acids in the background population are conjugated to a solid support. Some non-limiting examples of solid supports include beads, magnetic beads, polymer beads, and polymerizable polymers. Microfluidic devices, slides, chips, surfaces, plates, channels, cartridges, These include devices and microarrays. It can be conjugated to a solid support for antibody chromatography. In some cases, each nucleic acid in the background population has a different label. Each nucleic acid in the background population has the same label. Each copy of the nucleic acid in the population has the same label. In some cases, Each nucleic acid in the ground population has a different label.
[0077] Target population Generally, a population of interest has some characteristic that distinguishes it from other members of the larger population. The population of interest exists in a complex mixture of host and non-host nucleic acids. The nucleic acid may be a non-host nucleic acid (e.g., a bacterial nucleic acid) present in the host.
[0078] In some cases, the population of interest may be DNA, RNA, cDNA, mRNA, cRNA, dsDNA, ssDNA, miRNA, circulating nucleic acid, circulating DNA, circulating RNA, cell-free nucleic acid , cell-free DNA, cell-free RNA, circulating cell-free DNA, circulating cell-free RNA or genomic DNA A. The population of interest may be a mixture of DNA and RNA. The population can be genomic DNA. In some cases, the population of interest is artificially fragmented. In some cases, the population of interest does not contain artificially fragmented nucleic acids. In some cases, the population of interest may contain synthetic nucleic acids (e.g., by DNA synthesis). In some cases, the population of interest is synthetic (e.g., by DNA synthesis). In some cases, the population of interest is the genome, exome, or transcriptome. obtain.
[0079] In some cases, the population of interest is a non-host. In some cases, a non-host is a In some cases, non-host refers to a species other than the host. In some cases, non-host A non-host may refer to other organisms of the same species as the host. In some cases, a non-host may refer to a non-mammalian organism, Non-human, non-dog, non-cat, non-rodent, non-mouse, non-hamster, non-bovine, non-avian, non- It may be a chicken, a non-pig, a non-horse, a non-goat, a non-sheep or a non-rabbit. In this case, the non-host may be a microorganism, pathogen, bacterium, virus, fungus, parasite, or a combination thereof. In some cases, the non-host is a non-mammal. In some cases, the non-host is a In some cases, the non-host is a non-patient.
[0080] In some cases, the population of interest may be non-mammalian or non-human. In this case, the population of interest may be a microorganism, bacterium, virus, fungus, retrovirus, pathogen or In some cases, the population of interest may be non-microbial, non-bacterial, non-viral, In some cases, the pathogen may be non-mycotic, non-retroviral, non-pathogenic, or non-parasitic. The population of interest may be derived from a microorganism, pathogen, bacterium, virus, fungus, or parasite. In some cases, the population of interest is non-mammalian. The population is non-human. In some cases, the population of interest is a microorganism that infects a host or In some cases, the population of interest may be derived from non-mammalian DNA or RNA. In some cases, the population of interest may be non-human DNA or RNA. In some cases, the population of interest may be a microorganism, bacteria, virus, fungus, retrovirus, It may be the DNA or RNA of a pathogen or parasite. The group is non-microbial, non-bacterial, non-viral, non-fungal, non-retroviral, non-pathogen or non-parasitic. It can be DNA or RNA from a living organism.
[0081] In some cases, the population of interest may contain one or more non-host genomes, exomes, or may comprise a transcriptome. In some cases, the population of interest may comprise a transcriptome of about 1;2;3 ;4;5;6;7;8;9;10;15;20;25;30;35;40;45;50; 60;70;80;90;100;500;1000;5000;10000;5000 or 100,000 non-host genomes, exomes, or transcriptomes In some cases, the target population may be up to 1;2;3;4;5;6;7;8;9 ;10;15;20;25;30;35;40;45;50;60;70;80;90; 100; 500; 1000; 5000; 10000; 50000; or 100000 pieces In some cases, the non-host genome, exome, or transcriptome may be included. The target population is at least 1;2;3;4;5;6;7;8;9;10;15;20 ;25;30;35;40;45;50;60;70;80;90;100;500;1 000; 5000; 10000; 50000; or 100000 non-host genomes, In some cases, the population of interest may include a genome or transcriptome. Non-mammalian genomes, exomes, or transcripts from multiple individuals of the same non-mammalian species In some cases, the population of interest may include one or more non-mammalian species. Contains non-mammalian genomes, exomes, or transcriptomes from multiple individuals of a species In some cases, the population of interest contains non-mammalian genomic DNA and non-mammalian R In some cases, the population of interest may include one or more microbial genomes, In some cases, the population of interest may include the exome or transcriptome. , may include one or more pathogen genomes, exomes, or transcriptomes. In some cases, the population of interest comprises one or more bacterial genomes, exomes, or transcripts. In some cases, the population of interest may comprise one or more transcriptomes. It may include the viral genome, exome, or transcriptome. The population of interest may contain one or more retroviral genomes, exomes, or transgenes. In some cases, the population of interest may include one or more viral and one or more bacterial genomes, exomes, or transcriptomes In some cases, the population of interest may contain one or more parasite genomes, exosomal genomes, or In some cases, one or more may comprise about 1, 2, or 3 genomes or transcriptomes. , 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 5 It may be 0, 60, 70, 80, 90 or 100. In some cases, one or more is up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 4 It can be 0, 45, 50, 60, 70, 80, 90 or 100. In some cases, 1 one or more is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 2 It can be 5, 30, 35, 40, 45, 50, 60, 70, 80, 90 or 100. In some cases, the non-mammalian genome, exome, or transcriptome is a non-human genome. genome, exome, or transcriptome. In some cases, non-mammalian genomes The genome, exome or transcriptome can be analyzed for microorganisms, bacteria, viruses, fungi, retroviruses, and viruses. The genome, exome, or transcriptome of a virus, pathogen, or parasite do.
[0082] The population of interest may include a portion of the genome, exome, or transcriptome. In some cases, the population of interest is approximately 100% of the genome, exome, or transcriptome. 0.000001, 0.000005, 0.00001, 0.00005, 0.0001 ,0.0005,0.001,0.005,0.01,0.05,0.1,0.2,0. 3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 9 In some cases, the population of interest may comprise 9 or 100% of the genome, exome, or or transcriptome up to 0.000001, 0.000005, 0.0000 1, 0.00005, 0.0001, 0.0005, 0.001, 0.005, 0.01 ,0.05,0.1,0.2,0.3,0.4,0.5,0.6,0.7,0.8,0. 9, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40 , 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, It may comprise 94, 95, 96, 97, 98, 99 or 100%. Elephant populations must contain at least 0.000 of the genome, exome, or transcriptome. 001, 0.000005, 0.00001, 0.00005, 0.0001, 0.00 0.5, 0.001, 0.005, 0.01, 0.05, 0.1, 0.2, 0.3, 0.4 ,0.5,0.6,0.7,0.8,0.9,1,2,3,4,5,6,7,8,9,1 0, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75 , 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 1 It may constitute 00%.
[0083] The population of interest may represent a portion of the total population of nucleic acids in the sample. The population of potential candidates is approximately 0.000001, 0.000005, or 0.000006 of the total population of nucleic acids in the sample. 0.00001, 0.00005, 0.0001, 0.0005, 0.001, 0.00 5, 0.01, 0.05, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 6, 7, 8, It may comprise 9, 10, 15, 20, 25, 30, 35, 40, 45 or 50%. In some cases, the population of interest is up to 0.000001, 0, or 1 of the total population of nucleic acids in the sample. .000005, 0.00001, 0.00005, 0.0001, 0.0005, 0. 001, 0.005, 0.01, 0.05, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45 or 50% In some cases, the population of interest may comprise at least a portion of the total population of nucleic acids in the sample. Also 0.000001, 0.000005, 0.00001, 0.00005, 0.000 1, 0.0005, 0.001, 0.005, 0.01, 0.05, 0.1, 0.2, 0 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, 2, 2.5, 3 , 3.5, 4, 4.5, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 4 It may constitute 0, 45 or 50%.
[0084] Background population nucleic acids are present in excess relative to the nucleic acids of the population of interest in the sample. The ratio of the background population to the nucleic acids of the population of interest may be about 1:2:3:4. ;5;6;7;8;9;10;11;12;13;14;15;16;17;18;19 ;20;25;30;35;40;45;50;55;60;65;70;75;80; 85;90;95;100;200;300;400;500;600;700;800 ;900;1000;2000;3000;4000;5000;6000;7000; 8000;9000;10000;20000;30000;40000;50000; 60000;70000;80000;90000;100000;200000;30 0000;400000;500000;600000;700000;800000; 900000;1000000;2000000;3000000;4000000;5 000000;6000000;7000000;8000000;9000000;1 The ratio of the background population to the nucleic acids of the population of interest can be Max 1;2;3;4;5;6;7;8;9;10;11;12;13;14;15;16 ;17;18;19;20;25;30;35;40;45;50;55;60;65; 70;75;80;85;90;95;100;200;300;400;500;60 0;700;800;900;1000;2000;3000;4000;5000;6 000;7000;8000;9000;10000;20000;30000;400 00;50000;60000;70000;80000;90000;100000; 200000;300000;400000;500000;600000;70000 0;800000;900000;1000000;2000000;3000000; 4,000,000;5,000,000;6,000,000;7,000,000;8,000,000; 9,000,000; 10,000,000. The background for the nucleic acid of the population of interest The ratio of round populations is at least 1;2;3;4;5;6;7;8;9;10;11;12 ;13;14;15;16;17;18;19;20;25;30;35;40;45; 50;55;60;65;70;75;80;85;90;95;100;200;30 0;400;500;600;700;800;900;1000;2000;3000 ;4000;5000;6000;7000;8000;9000;10000;200 00;30000;40000;50000;60000;70000;80000;9 0000;100000;200000;300000;400000;500000; 600000;700000;800000;900000;1000000;2000 000;3000000;4000000;5000000;6000000;7000 The ratio can be: 000; 8,000,000; 9,000,000; 10,000,000. It can be calculated in terms of moles or mass.
[0085] The population of interest may include nucleic acids from one or more species. The target population is approximately 1;2;3;4;5;6;7;8;9;10;15;20;25; 30;35;40;45;50;60;70;80;90;100;200;300;4 00;500;1000;5000;10000;50000; or 100000 seeds In some cases, the population of interest may include nucleic acids derived from up to 1; 2; 3; 4; 5;6;7;8;9;10;15;20;25;30;35;40;45;50;60; 70;80;90;100;200;300;400;500;1000;5000;1 It may contain nucleic acids from species 0000; 50000; or 100000. In this case, the target population is at least 1;2;3;4;5;6;7;8;9;10;15 ;20;25;30;35;40;45;50;60;70;80;90;100;20 0;300;400;500;1000;5000;10000;50000; or 1 00000. In some cases, the species is a non-mammalian species. In some cases, the species is a non-human species. In some cases, the species is a microorganism, bacterium, virus, etc. In some cases, the target organism is a bacterium, fungus, retrovirus, pathogen, or parasite. The population is approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35 , 40, 45, 50, 60, 70, 80, 90, 100, 200, 300, 400, 50 It may contain nucleic acids from 0 or 1000 bacterial or viral species. The target population is up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25 , 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 200, 300, It may contain nucleic acids from 400, 500 or 1000 bacterial or viral species. In some cases, the target population is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 , 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100 , 200, 300, 400, 500 or 1000 nuclei from bacterial or viral species In some cases, the target population may comprise about 1, 2, 3, 4, 5, 6, 7, 8 , 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 9 0, 100, 200, 300, 400, 500 or 1000 bacterial and viral species In some cases, the population of interest may contain at least one nucleic acid derived from at most one , 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45 , 50, 60, 70, 80, 90, 100, 200, 300, 400, 500 or 10 The nucleic acid may contain at least one nucleic acid derived from 100 bacterial and viral species. In this case, the target population is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 200 , 300, 400, 500 or 1000 bacterial and viral species In some cases, the population of interest may contain about one bacterial species and one viral 2 bacterial and 2 viral species, 3 bacterial and 3 viral species, 4 bacterial and 4 viral species 5 bacterial and 5 viral species, 6 bacterial and 6 viral species, 7 bacterial and 7 viral species, 8 bacterial species and 8 viral species, 9 bacterial species and 9 viral species, 10 bacterial species and 10 virus species, 15 bacterial species and 15 virus species, 20 bacterial species and 20 virus species 25 bacterial species and 25 viral species, 30 bacterial species and 30 viral species, 35 bacterial species and 35 viral species, 40 bacterial species and 40 viral species, 45 bacterial species and 45 viral species 50 bacterial species and 50 viral species, 60 bacterial species and 60 viral species, 70 bacterial species and 70 virus species, 80 bacterial species and 80 virus species, 90 bacterial species and 90 virus species 100 bacterial and 100 viral species, 200 bacterial and 200 viral species, 3 00 bacterial species and 300 viral species, 400 bacterial species and 400 viral species, 500 bacterial species and 500 viral species, or 1000 bacterial species and 1000 viral species In some cases, the population of interest may contain at most one bacterial species and and 1 virus species, 2 bacteria species and 2 virus species, 3 bacteria species and 3 virus species, 4 bacteria species and 4 viral species, 5 bacterial species and 5 viral species, 6 bacterial species and 6 viral species, 7 Bacterial species and 7 viral species, 8 bacterial species and 8 viral species, 9 bacterial species and 9 viral species , 10 bacterial species and 10 viral species, 15 bacterial species and 15 viral species, 20 bacterial species and and 20 viral species, 25 bacterial species and 25 viral species, 30 bacterial species and 30 viral species , 35 bacterial species and 35 viral species, 40 bacterial species and 40 viral species, 45 bacterial species and and 45 viral species, 50 bacterial species and 50 viral species, 60 bacterial species and 60 viral species , 70 bacterial species and 70 viral species, 80 bacterial species and 80 viral species, 90 bacterial species and and 90 virus species, 100 bacterial species and 100 virus species, 200 bacterial species and 200 virus species 300 bacterial and 300 viral species, 400 bacterial and 400 viral species , 500 bacterial species and 500 viral species, or 1000 bacterial species and 1000 viruses. In some cases, the population of interest may comprise at least one nucleic acid from a species. At least one bacterial species and one viral species, two bacterial species and two viral species, three bacterial species and three Virus species, 4 bacterial species and 4 virus species, 5 bacterial species and 5 virus species, 6 bacterial species and and 6 viral species, 7 bacterial species and 7 viral species, 8 bacterial species and 8 viral species, and 9 bacterial species and 9 viral species, 10 bacterial species and 10 viral species, 15 bacterial species and 15 viruses. species, 20 bacterial species and 20 viral species, 25 bacterial species and 25 viral species, 30 bacterial species and and 30 viral species, 35 bacterial species and 35 viral species, 40 bacterial species and 40 viral species species, 45 bacterial species and 45 viral species, 50 bacterial species and 50 viral species, 60 bacterial species and and 60 viral species, 70 bacterial species and 70 viral species, 80 bacterial species and 80 viral species species, 90 bacterial species and 90 viral species, 100 bacterial species and 100 viral species, 200 cellular 200 bacterial species and 200 viral species, 300 bacterial species and 300 viral species, 400 bacterial species and and 400 virus species, 500 bacterial species and 500 virus species, or 1000 bacterial species and and at least one nucleic acid from 1000 viral species.
[0086] In some cases, one or more nucleic acids in the population of interest are unlabeled; In some cases, one or more nucleic acids in the population of interest are labeled. One or more nucleic acids in the resulting population may be labeled, for example, with a nucleic acid label, a chemical label, or an optical label. In some cases, one or more nucleic acids in a population of interest are attached to a solid support. In some cases, the label is conjugated to the 5' or 5' end of the nucleic acids in the population of interest. They can be attached to the 3' end or internally to the nucleic acids in the population of interest. In this case, one or more nucleic acids in the population of interest are labeled with two or more labels.
[0087] The nucleic acids in the population of interest can include a nucleic acid label. In some cases, the nucleic acid label , may include one or more of the following: a barcode (e.g., a sample barcode), a unique general primer sequence, primer binding site (e.g., but not limited to, DNA Sequencing primer binding site, sample barcode sequencing primer binding site and species sequencing or amplification including amplification primer binding sites compatible with the requirements of various sequencing platforms. or barcode reading), sequencer compatibility sequences, sequencing platform A sequence to which the nucleic acid label is attached, a sequencing adapter sequence or adapter. can be attached to nucleic acids in a population of interest (by ligation or synthetic design). In some cases, the length of the nucleic acid label is included in the length of the nucleic acids in the population of interest. In some cases, the length of the nucleic acid label is not included in the length of the nucleic acids in the population of interest.
[0088] The nucleic acids in the population of interest can include chemical labels. Some non-limiting examples of chemical labels include: Typical examples include biotin, avidin, streptavidin, radioactive labels, polypeptides, etc. and polymers. The nucleic acids in the population of interest can include an optical label. Some non-limiting examples of optical labels include fluorophores, fluorescent proteins, dyes, and quantitators. Nucleic acids in a population of interest can be conjugated to a solid support. Some non-limiting examples of solid supports include beads, magnetic beads, poly Microfluidic devices, slides, chips, surfaces, plates, channels, cartridges, These include devices and microarrays. It can be conjugated to a solid support for chromatography. Each nucleic acid in the population of interest has a different label. Each nucleic acid has the same label. In some cases, each copy of a nucleic acid in a population of interest has the same label. In some cases, each nucleic acid in the population of interest has a different sequence. Has a sign.
[0089] Microorganisms that can be detected by the methods provided herein Some non-limiting examples of bacteria (or microbes) include bacteria, archaea, protozoa, and These include insects, protozoa, fungi, algae, viruses, retroviruses, pathogens or parasites. In some cases, microorganisms are prokaryotes. In some cases, microorganisms are eukaryotes. Some non-limiting examples of bacteria include Bacillus, Bacillus subtilis, Bacillus spp., Bacillus subtilis ... Bordetella, Borrelia, Brucella rucella, Campylobacter, Chlamydia (Chlamydia), Chlamydophila, Clostridium Clostridium, Corynebacterium erium), Enterococcus, Escherichia scherichia, Francisella, Haemophilus (Haemophilus), Helicobacter, Legiobacter Legionella, Leptospira, Listeria Listeria, Mycobacterium, Mycoplasma, Neisseria, Pseudomonas, Rickettsia, Salmonella, Shigella, Staphylococcus Staphylococcus, Staphylococcus aureus Aures, Streptococcus, Treponema Treponema, Vibrio and Yersinia Some non-limiting examples of fungi include Candida. dida), Aspergillus, Cryptococcus yptococcus, Histoplasma, Pneumocystis Pneumocystis and Stachybotrys In some examples, the microorganisms detected by the methods provided herein include The organism is a drug-resistant microorganism or a multidrug-resistant pathogen. Non-limiting examples include: In some cases, Clostridium diff Drug resistance of Clostridium difficile (C. difficile) resistant strains, carbapenem-resistant Enterobacteriaceae (CRE), drug-resistant Neisseria gonorrhoeae (Neisseria gonorrhoeae) norrhoeae) (cephalosporin-resistant), multidrug-resistant Acinetobacter etobacter, drug-resistant Campylobacter, fluconazole-resistant Candida (fungal ), extended-spectrum β-lactamase-producing Enterobacteriaceae (ESBL), vancomycin-resistant Enterobacteriaceae Enterococcus (VRE), multidrug-resistant Pseudomonas aeruginosa (Pseudomona aeruginosa), drug-resistant non-typhoidal Salmonella , drug-resistant Salmonella Typhi, drug-resistant Shigella ella), methicillin-resistant Staphylococcus aureus (MRSA), drug-resistant Streptococcus pneumoniae (Stre pneumonia), drug-resistant tuberculosis (MDR and XDR), Drug-resistant Staphylococcus aureus, vancomycin-resistant Staphylococcus aureus (VRSA), Erythromycin Clindamycin-resistant Streptococcus group A or clindamycin-resistant Streptococcus group B.
[0090] In some cases, the methods and compositions provided herein can be used to detect retroviruses. or viruses such as lentiviruses. The virus classification system is divided into groups I, II, III, IV, and V according to the Baltimore virus classification system. In some cases, the virus is a member of Group V, Group VI, or Group VII. Adenoviridae, Anelloviridae dae), Arenaviridae, Astroviridae troviridae), Bunyaviridae, Kalisiwi Caliciviridae, Coronaviridae e), Filoviridae, Flaviviridae iridae), Hepadnaviridae, Hepevirus Family (Hepeviridae), Family (Herpesviridae), Orthomyxoviridae, Papillomaviridae (Papillomaviridae), Papovavirida e), Paramyxoviridae, Parvoviridae ( Parvoviridae), Picornaviridae, Polyomaviridae, Poxviridae viridae), Reoviridae, Retroviridae troviridae), Rhabdoviridae or Toga A member of the virus family (Togaviridae). In some cases, the virus Adenovirus, Amur virus, Andes virus, animal virus, astrovirus Avian nephritis virus, avian orthoreovirus, avian reovirus, Banna virus, Bass-Congo virus, bat-borne virus, BK virus, blueberry shock virus virus, chicken anemia virus, bovine adenovirus, bovine coronavirus, bovine herpes virus Rus 4, bovine parvovirus, brown-eared bulbul coronavirus HKU11, carrisal virus, Catacamas virus, Chandipura virus, Channel catfish virus, Chokrowy Rus, Coltivirus, Coxsackievirus, Cricket paralysis virus, Crimean corn Hemorrhagic fever virus, cytomegalovirus, dengue virus, Dobrava-Bergredovirus Rus, Ebola virus, El Moro Canyon virus, endotheliotropic elephantiasis Rupes virus, Epstein-Barr virus, feline leukemia virus, foot-and-mouth disease virus, Gou virus, Guanarito virus, Hantaan virus, Ha Antiviral drugs, HCoV-EMC / 2012, Hendra virus, Henipavirus, Type A Hepatitis virus, Hepatitis B virus, Hepatitis C virus, Hepatitis D virus, Hepatitis E virus, Pure herpes type 1, herpes simplex type 2, herpes simplex virus type 1, herpes simplex virus 2, HIV, human astrovirus, human bocavirus, human cytomegalovirus, human Human herpesvirus type 8, human herpesvirus type 8, human immunodeficiency virus (HIV) , human metapneumovirus, human papillomavirus, immunization virus, influenza Manza virus, Isla Vista virus, JC virus, Junin virus, Khabarovsk virus Virus, carp herpesvirus, Kunjin virus, Lassa virus, Limestone virus Canion virus, Lloviu cueva virus, Lloviu virus, Lujo Viruses, Machupo virus, Magboi virus, Marburg virus Marburg virus, measles virus, malaria virus, measles virus Nangle virus, Middle East respiratory syndrome coronavirus, Minioptera bat coronavirus virus 1, Minioptera bat coronavirus HKU8, monkeypox virus, Monongah Hira virus, Muju virus, Mumps virus, Nipah virus, Norwalk virus Orbivirus, parainfluenza virus, parvovirus B19, phytoreovirus Viruses, Pigeon coronavirus HKU5, poliovirus, porcine adenovirus Prospect Hill virus, Kalube virus, Rabies virus, Ravn virus Respiratory syncytial virus, Reston virus, reticuloendotheliosis virus, horseshoe bat Coronavirus HKU2, rhinovirus, roseolovirus, Ross River virus, Tavirus, Rouset fruit bat coronavirus HKU9, rubella virus, Saarema -virus, Sabia virus, Sangassou virus, and bat coronavirus S512, Ceran virus, Severe acute respiratory syndrome virus, Shope papilloma virus, Simian foamy virus, Sin Nombre virus, smallpox, Suchon virus, Sudanese virus Bolavirus, Sudan virus, Thai Forest Ebola virus, Thai Forest virus Rus, Tanzania virus, Thottapalayam virus irus, Topografov virus, Tremowi Rus, Tula virus, Turkey coronavirus, Turkey pox virus, Bamboo bat virus Varicella virus HKU4, varicella zoster virus, varicella zoster virus, West Nile virus , woodchuck hepatitis virus, yellow fever virus, Zika virus or Zaire Ebola virus It's Ilse.
[0091] Some non-limiting examples of pathogens include viruses, bacteria, prions, fungi, and parasites. Some non-limiting examples of pathogens include bacteria, protozoa, and microorganisms. Acanthamoeba, Acari, Acinetobacter Acinetobacter baumannii, Actinomycetes Actinomyces israelii, Actinomyces Actinomyces gerencseriae, Propionibacterium Propionibacterium propionicus us), Actinomycetoma, Eumycetoma umycetoma), Adenoviridae, Alphaviridae Alphavirus, Anaplasma, Anaplasma Anaplasma phagocytophilum, Ancylostoma braziliense , Ancylostoma duodenale, Necator americanus, Angiostro Angiostrongylus costaricensis sis), Anisakis, Arachnida ixodidae ida Ixodidae), Argasidae, hemolytic arkanobacter Terrier (Arcanobacterium haemolyticum), primitive thornyhead (Archiacanthocephala), Moniliformis moniliformis ( Moniliformis moniliformis, Arenaviridae (Aren aviridae), roundworms (Ascaris lumbricoides), ascaris Ascaris genus species, roundworms (Ascaris lumbricoides), Ascaris Aspergillus, Astroviridae ae), Babesia polymorphic Babesia (B. divergens), Babesia B. bigemina, Babesia equi, Babesia B.microfti, B.duncani, Babesia, Bacillus anthracis, Bacillus cereus, Bacteroides es), Balamuthia mandrillari s), Balantidium coli, Henselae tonella henselae, Baylisascari s), raccoon roundworm (Baylisascaris procyonis), human tapeworm (Bertiella mucronata), Bertiella studerii (Bertie Illa studeri), BK virus, Blastocystis s), human blastocystis (Blastocystis hominis), blast Blastomyces dermatitidis, 100 days Bordetella pertussis, Borrelia burgdorferi (B Borrelia burgdorferi, Borrelia species, Borrelia, Brucella, and Br ugia malayi), Brugia timori, Bunya Family: Bunyaviridae, Burkholderia c epacia, Burkholderia species, Burkholderia mallei rkholderia mallei), Burkholderia pse udomallei), Caliciviridae, Campylobacter Campylobacter, Candida albicans albicans, Candida species, Cestoda , Taenia multiceps, Chlamydia trachomatis ydia trachomatis, Chlamydia trachomatis rachomatis), Neisseria gonorrhoeae, pneumonia Chlamydia (Chlamydophila pneumoniae), Chlamydia psittacosis Chlamydophila psittaci, family Cimicid ae), bed bugs (Cimex lectularius), liver fluke (Clonorc his sinensis); Clonorchis viverini (Clonorchis viv errini), Clostridium botulinum, Clostridium Clostridium difficile, well Clostridium perfringens, Clos Clostridium perfringens, Clostridium um) species, Clostridium tetani, Coccidioides Coccidioides immitis, Coccidioides posadasii (Coccidioides posadasii), Cochliomia hominivorax ( Cochliomyia hominivorax, Colorado tick fever virus (CTF V), Coronaviridae, Corynebacterium diphtheriae bacterium diphtheriae, Q fever Coxiella burnetii), Crimean-Congo hemorrhagic fever virus, Cryptococcus neoformans Cryptococcus neoformans, Cryptosporidium Cryptosporidium, Cryptosporidium ium), Cyclospora cayetane nsis), cytomegalovirus, Demodex Demodex folliculorum / brevis is) / canis, dengue virus (DEN-1, DEN-2, DEN-3 and DEN-4), Flavivirus, Dermatozoa matobia homini), spear-shaped fluke (Dicrocoelium dendri) ticum), Dientamoeba fragilis, Geoct Dioctophyme renale, genus Diph yllobothrium), Diphyllobothrium la tum), Dracunculus medinensis, Ebola worm EBOV, Echinococcus, Echinococcus granulosus nococcus granulosus), Echinococcus m ultilocularis), Echinococcus vogeli (E.vogeli), Japanese dormouse E. oligarthus, Ehrlichia chaffeensis hia chaffeensis, Ehrlichia ewingii (Ehrlichia e wingii, Ehrlichia, Entamoeba histolytica ba histolytica), Entamoeba histolytica tica), pinworms (Enterobius vermicularis), enterobius Enterobius gregorii, Enterococcus spp. Enterococcus, Enterovirus, Enterovirus Viruses, Coxsackie A virus, Enterovirus 71 (EV71), Epidermophyton floccosum , Trichophyton rubrum, Trichophyton rubrum yton mentagrophytes), Epstein-Barr virus (EBV), Escherichia coli O157:H7, O111, and O104 :H4, Fasciola hepatica, Giant liver fluke (Fasciola) gigantica), Fasciolopsis buski, filari Filarioidea superfamily, Filoviridae loviridae), Flaviviridae, Fonseca Fonsecaea pedrosoi, Francisel tularensis la tularensis), Fusobacterium, Geotrichum candidum, intestinal flagellates ( Giardia intestinalis, Giardia lamblia mblia), Gnathostoma spinigerum, Gnathostoma spinigerum (Gnathostoma hispidum), Group A Streptococcus (Group AS streptococcus, staphylococcus, guanari Tovirus, Haemophilus ducreyi, Haemophilus influenzae, Halicifolia Halicephalobacter gingivalis (Halicephalobus gingivalis), Hartley Endovirus, Helicobacter pylori, Hepadnaviridae, Hepatitis A virus, Hepatitis B virus Hepatitis C virus, Hepatitis D virus, Hepatitis E virus, Hepeviridae (H epeviridae), herpes simplex virus 1 and 2 (HSV-1 and HSV- 2), Herpesviridae, Histoplasma capsula Histoplasma capsulatum, HIV (human immunodeficiency virus) Rus), Hortaea werneckii, Human Human herpesvirus (HBoV), human herpesvirus 6 (HHV-6), human herpesvirus 7 (HHV-7), human metapneumovirus (hMPV), human papillomavirus (H PV), Human Parainfluenza Virus (HPIV), Hymenolep is nana), Hymenolepis diminuta, Isospora Isospora belli, JC virus, Junin virus, Kingae virus Kingella kingae, Klebsiella granulomatis ( Klebsiella granulomatis, Lassa virus, Legionella pneumophila (Legionella pneumophila), Leishmania (L eishmania, Leptospira, Ling uatula serrata), Listeria monocyto genes), Loa loa filariasis, lymphocytic choriomeningitis virus LCMV, Machupo virus, Malassezia, Mansone Mansonella streptocerca, Marlb Lug virus, measles virus, Metagonimus yokagawai ), Microsporidia phylum, Middle East respiratory syndrome coronavirus Viruses, molluscum contagiosum virus (MCV), monkeypox virus, Mucorales a les order) (mucormycosis), Entomophthorales order) (Entomophthora infection), mumps virus, Mycobacterium leprae (Mycobacterium leprae) Mycobacterium leprae, Mycobacterium lepromatosis rium lepromatosis), Mycobacterium tuberculosis (Mycobacterium tuberculosis) erculosis), Mycobacterium ulcerans M ulcerans, Mycoplasma pneumoniae iae), Naegleria fowleri, Neisseria gonorrhoeae (N eisseria gonorrhoeae), meningococcus (Neisseria men ingitidis), Nocardia asteroides ides), Nocardia species, Oestroi dea), Calliphoridae, Sarcopha gidae), Onchocerca volvulus, and the liver fluke ( Opisthorchis viverrini), cat liver fluke (Opisthorch is felineus), liver fluke (Clonorchis sinensis), Orthomyxoviridae, Papillomaviridae apillomaviridae, South American Paracoccidioides es brasiliensis), Paragonimus africanus (Paragonim us africanus); Paragonimus caliensis (Paragonimus caliensis; Paragonimus kell Paragonimus skrj abini); Paragonimus uterobilateralis (Paragonimus uter obilateralis), Westerman's lung fluke (Paragonimus wes) termani), Paragonimus species, paramyxovirus Paramyxoviridae, parasitic dipteran fly larvae, Parvoviridae (P arvoviridae, Parvovirus B19, Pasteurella spp. la), human body louse (Pediculus humanus), human head louse (Pe diculus humanus capitis), body lice (Pediculu) s humanus corporis), pubic lice (Phthirus pubis) , Picornaviridae, Piedraia hortae (P iedraia hortae, Plasmodium falciparum ciparum, Plasmodium vivax, ovale Plasmodium ovale curtisi, oval marrow Plasmodium ovale wallikeri, quartan fever Plasmodium malariae, Plasmodium vivax (Pl asmodium knowlesi), Plasmodium, Pneumocystis jirovecii, Poly Polyomaviruses, Polyomaviridae, Poxviruses Poxviridae, Prevotella, PRNP, Ke Lice (Pthirus pubis), human flea (Pulex irritans), Rabies virus, Reoviridae, Respiratory syncytial virus (RSV), Retroviridae, Rhabdoviridae ridae), Rhinosporidium seeb eri), Rhinovirus, Rhinovirus, Coronavirus Rickettsia akari, Rickettsia spp. ttsia), Rickettsia prowazekii , Rickettsia rickettsii, typhus Kettsia (Rickettsia typhi), Rift Valley fever virus, rotavirus rubella virus, Sabia, Salmonella terica subspecies enterica, serovar typhi, Salmonella, Sarcocystis bovicanis tis bovihominis, Sarcocys suihominis tis suihominis), Sarcoptes scabiei (Sarcoptes s cabiei, SARS coronavirus, Schistosoma , Schistosoma haematobium, Japanese Schistosoma Trematode (Schistosoma japonicum), Schistosoma mansoni (Sc Histosoma mansoni and Schistosoma intercalatum osoma intercalatum Mekong schistosoma (Schistosoma m ekongi, Schistosoma species, Shigella species ella), Sin Nombre virus, Mansoni's diphyllobothrium tapeworm (Spirometra eri naceieuropaei), Sporothrix schenckii (Sporothrix schenckii), Staphylococcus, Streptococcus agalactiae ), Streptococcus pneumoniae, Streptococcus pyogenes (Streptococcus pyogenes), Strongyloides es stercoralis), tapeworm (Taenia), tapeworm (Taenia) saginata), Taenia solium, Bacteriaceae, Enterobacteriaceae (Enterobacteriaceae), California Thelazia a californiensis, Oriental eyeworm (Thelazia callipae) da), Togaviridae, Toxocara canis anis), Toxocara cati, Toxoplasma gondii plasma gondii, Treponema pallidum um), Trichinella spiralis, Trichinella brittlebi (Trichinella britovi), Trichinella nelsoni (Trichin ella nelsoni), Trichinella nati va), Trichobilharzia regen ti), Schistosomatidae, Trichomonas vaginalis homonas vaginalis, Trichophyton ), Trichophyton rubrum, Trichophyton tonsuran Trichophyton tonsurans, Trichophyton tonsurans osporon beigelii), whipworm (Trichuris trichiura) ), whipworm (Trichuris trichiura), whipworm (Trichuris trichiura), whipworm (Trichuris trichiura), whipworm (Trichuris trichiura) vulpis, Trypanosoma brucei ), Trypanosoma cruzi, sand fleas (T unga penetrans, Ureaplasma urealyticum (Ureaplas ma urealyticum), varicella zoster virus (VZV), variola major (Variola major) Variola major, Variola minor, Venezuelan equine encephalitis Vibrio cholerae, West Nile virus, Bancroft virus Wuchereria bancrofti, Wuc hereria bancrofti, Malaysian heartworm (Brugia malayi) , yellow fever virus, Yersinia enterocolitica olitica), Yersinia pestis and Yersinia pseudotuberculosis Examples include Yersinia pseudotuberculosis.
[0092] Methods for Producing Collections of Oligonucleotides The present disclosure provides methods for targeting or detecting populations of interest (e.g., microbial or pathogen nucleic acids). To identify the target gene, a collection of oligonucleotides for use in the methods provided herein is provided. In some cases, the oligonucleotides The methods for generating the collection are bioinformatics, computational, hybridization-based, or may be a digestion-based method.
[0093] In some cases, background population nucleic acid, such as host (e.g., human) genomic DNA oligonucleotides containing host and non-host (e.g., non-human, pathogen) sequences using This method can be used to remove host sequences from a population of host (e.g., human) and an oligonucleotide containing an oligonucleotide capable of binding to non-host nucleic acid. A heterogeneous collection of nucleotides (e.g., 5'-NNNNNNNNNNNNN-3' (where N is A, providing a collection of oligonucleotide randomers, such as a randomized set of randomized oligonucleotides (e.g., a randomized set of randomized oligonucleotides, each of which is C, G, or T); In some instances, the host genomic DNA may be transfected with a heterologous sequence of oligonucleotides. The oligonucleotides are then extracted under conditions that promote binding of the host genomic DNA to the complementary sequences in the collection. This method allows the introduction of oligonucleotides into a heterogeneous collection of nucleotides. Depleting oligonucleotides bound to host genomic DNA from a heterogeneous collection of DNA Typically, such methods involve transforming host nucleic acid (e.g., host genomic nucleic acid) into The heterogeneous collection of oligonucleotides is provided in excess. Hybridized or conjugated oligonucleotides from a heterogeneous collection of nucleotides removing the hybridized or bound oligonucleotides; or by preferentially isolating unbound oligonucleotides, etc. This can be completed by any method.
[0094] FIG. 3 shows a pictorial example of creating a collection of oligonucleotides (370). The method includes providing a heterogeneous collection of oligonucleotides (310). In some cases, the heterogeneous collection of oligonucleotides may be nucleic acid labeled, chemically labeled, or Alternatively, the method may include an oligonucleotide labeled with an optical label (320). Background population nucleic acids (e.g., host nucleic acids, Nucleotides having a sequence identical, nearly identical, complementary, or nearly complementary to a nucleic acid (e.g., human nucleic acid) In some cases, the step of depleting oligonucleotides containing the domain of , a heterogeneous collection of oligonucleotides, a background population of nucleic acids, Complementary or nearly complementary nucleotides of oligonucleotides within a heterogeneous collection of oligonucleotides The background population is then hybridized or specifically bound to the complementary domain. Depletion by contacting (360) with nucleic acid (340) (e.g., human genomic DNA). The hybridization reaction may include a denaturation step and / or a renaturation step. For example, the nucleic acid may be heated or denatured in order to denature the nucleic acid in the reaction. or subjected to stepwise heating (e.g., 95°C for 10 seconds and 65°C for 3 minutes). To regenerate the nucleic acid, the nucleic acid is incubated at 36°C for a period of time, e.g., several hours or weeks. In some cases, background population nucleic acids can be isolated from the host (33 0). In some cases, background population nucleic acids are derived or isolated from nucleic acid targets. In some cases, the method further comprises labeling the antibody with a biomarker, a chemical label, or an optical label (350). and further comprising denaturing at least a portion of the background population nucleic acids, e.g., by heat. include.
[0095] In some cases, one or more blocker oligonucleotides may be hybridized It may be present during the annealing step (360) or before the hybridization step. In some instances, the oligonucleotide may be DNA, RNA, or P Blocker oligonucleotides include NA, LNA, BNA, or any combination thereof. The nucleotides may be complementary to a sequence of the oligonucleotide outside the domain of nucleotides. Therefore, they do not hybridize to domains of nucleotides, but rather to other nucleotides present on the same strand. The blocker oligonucleotides can be designed to "block" the nucleotides. The presence of sequences outside the domain of nucleotides in the background of the oligonucleotide This can help reduce the likelihood of binding to populations (e.g., human genomic DNA). The blocker oligonucleotides provide a signal to the background population of oligonucleotides. The binding specificity of a collection of oligonucleotides can be enhanced. In a specific example, the heterogeneous collection of oligonucleotides comprises blocker oligonucleotides. (e.g., 0.5X PBS, blocker oligonucleotides, RNase inhibitors) It can hybridize to genomic DNA in a buffer solution containing
[0096] Generally, the starting point for generating the collection of oligonucleotides provided herein is The heterogeneous collection of oligonucleotides used as a population was randomly generated. A collection of nucleic acid sequences (e.g., 5'NNNNNNNNNNNNN-3' (where N is A, C , G, or T). In some cases, a heterogeneous collection of oligonucleotides is prepared (e.g., by DNA synthesis). In some cases, a heterogeneous collection of oligonucleotides can be synthesized. In some cases, the heterologous components of the oligonucleotides may be present. The selection is based on approximately 5, 10, 20, 30, or 40 of the possible sequences for a domain of nucleotides. 40, 50, 60, 70, 80, 90, 91, 92, 93, 94, 95, 96, 97, 9 8, 99, 99.1, 99.2, 99.3, 99.4, 99.5, 99.6, 99.7, Oligonucleotides containing domains of nucleotides present in 99.8, 99.9 or 100% of the In some cases, the heterogeneous collection of oligonucleotides comprises nucleotides. Up to 5, 10, 20, 30, 40, 50, 60 possible sequences for the domain of the otide 70, 80, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 99.1 , 99.2, 99.3, 99.4, 99.5, 99.6, 99.7, 99.8, 99.9 or 100% are present, including oligonucleotides containing domains of nucleotides. In some cases, the heterogeneous collection of oligonucleotides is organized into domains of nucleotides. At least 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 99.1, 99.2, 99.3, 99.4, 99.5, 99.6, 99.7, 99.8, 99.9 or 100 % of the nucleotides present in the oligonucleotides.
[0097] In some cases, the heterogeneous collection of oligonucleotides may be a background population of nuclei. a sequence identical, nearly identical, complementary, or nearly complementary to a nucleic acid from a nucleic acid or population of interest In some cases, the oligonucleotide comprises a domain of nucleotides having the A heterogeneous collection of oligonucleotides is identical to one or more background populations. nucleotides having a substantially identical, complementary, or nearly complementary sequence. In some cases, the heterogeneous collection of oligonucleotides includes: Identical, nearly identical, complementary, or nearly complementary to one or more background populations. It includes oligonucleotides that contain domains of nucleotides with sequences that are not complementary to each other. In some cases, the heterogeneous collection of oligonucleotides may be composed of one or more target A domain of nucleotides having an identical, nearly identical, complementary, or nearly complementary sequence to a population In some cases, heterogeneous collections of oligonucleotides The fusion may be identical, nearly identical, complementary, or nearly identical to one or more target populations. It includes oligonucleotides that contain domains of nucleotides that have sequences that are not complementary.
[0098] In some cases, one or more oligonucleotides in a heterogeneous collection of oligonucleotides Nucleotides are unlabeled; in some cases, heterogeneous collections of oligonucleotides One or more of the oligonucleotides in the The oligonucleotides may be labeled, for example, with a nucleic acid label, a chemical label, or an optical label. In some cases, one or more oligonucleotides are conjugated to a solid support. In some cases, one or more oligonucleotides are conjugated to a solid support. In some cases, the label is attached to the 5' or 3' end of the oligonucleotide. or attached to the interior of the oligonucleotide. One or more oligonucleotides in a heterogeneous collection of nucleotides may be two or more target oligonucleotides. It is marked with a recognition.
[0099] Background population nucleic acids are present in excess relative to the heterogeneous collection of oligonucleotides. The ratio of background population to oligonucleotide may be about 0.1; 0.2;0.3;0.4;0.5;0.6;0.7;0.8;0.9;1.0;1.1; 1.2;1.3;1.4;1.5;1.6;1.7;1.8;1.9;2.0;2.5; 3.0;3.5;4.0;4.5;5.0;5.5;6.0;6.5;7.0;7.5; 8.0;8.5;9.0;9.5;10;11;12;13;14;15;16;17; 18;19;20;25;30;35;40;45;50;55;60;65;70;7 5;80;85;90;95;100;200;300;400;500;600;70 0;800;900;1000;2000;3000;4000;5000;6000; 7000; 8000; 9000; or 10000. The ratio of the background population to the target population is up to 0.1; 0.2; 0.3; 0.4; 0.5; 0. 6;0.7;0.8;0.9;1.0;1.1;1.2;1.3;1.4;1.5;1. 6;1.7;1.8;1.9;2.0;2.5;3.0;3.5;4.0;4.5;5. 0;5.5;6.0;6.5;7.0;7.5;8.0;8.5;9.0;9.5;10 ;11;12;13;14;15;16;17;18;19;20;25;30;35; 40;45;50;55;60;65;70;75;80;85;90;95;100; 200;300;400;500;600;700;800;900;1000;200 0;3000;4000;5000;6000;7000;8000;9000;or The ratio of background population to oligonucleotides can be as low as 10,000. At most 0.1;0.2;0.3;0.4;0.5;0.6;0.7;0.8;0.9;1 .0;1.1;1.2;1.3;1.4;1.5;1.6;1.7;1.8;1.9;2 .0;2.5;3.0;3.5;4.0;4.5;5.0;5.5;6.0;6.5;7 .0;7.5;8.0;8.5;9.0;9.5;10;11;12;13;14;15 ;16;17;18;19;20;25;30;35;40;45;50;55;60; 65;70;75;80;85;90;95;100;200;300;400;500 ;600;700;800;900;1000;2000;3000;4000;500 It can be 0; 6000; 7000; 8000; 9000; or 10000. The ratio of background population to nucleotide may be saturating. The ratio of background population to the target may be non-saturating. The ratio may be expressed in terms of concentration, molar or or can be calculated in terms of mass.
[0100] In some cases, the nucleic acid is single-stranded. In some cases, double-stranded nucleic acid is converted to single-stranded nucleic acid. Nucleic acids can be denatured using heat. In some cases, the nucleic acid is denatured for about 35, 4 0, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93 , 94, 95, 96, 97, 98, or 99°C. In some cases, the nucleic acid is heated to about 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50 or 60 minutes In some cases, nucleic acids are denatured by the addition of chemical denaturants (e.g., acids, bases, solvents, chaotic Nucleic acids are denatured using tropic agents (anticoagulants, salts).
[0101] Single-stranded nucleic acid samples can be renatured or hybridized. , hybridizing a single-stranded nucleic acid sample with a heterogeneous collection of oligonucleotides In some cases, a heterogeneous collection of oligonucleotides is used as a blocker oligonucleotide. In some cases, at least a portion of the single-stranded nucleic acid is renatured or hybridized to the nucleic acid. In some cases, the nucleic acid is hybridized to about 0, 1, 2, 3, 4, 5, 6, 7 , 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, Regenerate or regenerate at 35, 36, 37, 40, 45, 50, 55, 60, 65, 68 or 70°C In some cases, nucleic acids are renatured or hybridized on ice. In some cases, the nucleic acids are renatured or hybridized at room temperature. The nucleic acid is dissolved in about 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 , 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50 or 60 minutes, or 2, 3, 4, 5, 10, 15, 20, 22, 24, 30, 40, 4 Regenerate or hybridize for 6, 48, 50, 60, 70, 72, 80, or 96 hours. Temperature, buffer composition, reaction time, and concentration affect the degree of hybridization. In some cases, nucleic acids are renatured in the presence of trimethylammonium chloride. do.
[0102] Without removing hybridized nucleic acids from a heterogeneous collection of oligonucleotides Depletion can be completed by the methods provided herein, such as by The background population is chemically labeled and removed, thereby identifying the hybridized oligos. Nucleotides can be removed. Background populations can be conjugated to magnetic beads. The hybridized oligonucleotides are then removed by magnetization. The background population can be labeled with a nucleic acid tag. The nucleic acid sequence is conjugated to a solid support or tagged with a chemical tag. The hybridized oligonucleotides can be bound or hybridized to the array. The methods include size separation (e.g., gel electrophoresis, capillary electrophoresis, etc.), affinity separation (e.g., For example, DNA pull-down assays, chromatography, or other methods In some cases, the hybridized oligonucleotides can be removed by gel electrophoresis. using separation methods such as electrophoresis, capillary electrophoresis or chromatography; Unhybridized and hybridized oligonucleotides In some cases, hybridization can be performed to isolate the target protein based on differences in size. This does not remove background population nucleic acids.
[0103] In some cases, oligonucleotides hybridized to background population nucleic acids Depletion can be completed by inactivating the ribonucleotides, e.g., oligonucleotides. The nucleotides are chemically bound to the labeled background population nucleic acids and then unbound. By amplifying with only oligonucleotides, background population nucleic acids are not easily detected. Hybridizing oligonucleotides can be prevented from acting as primers .
[0104] In some cases, collections of oligonucleotides are identified using bioinformatic or computational tools. designed to target populations of interest (e.g., pathogen nucleic acids, non-human nucleic acids) using In some cases, the synthesis method is DNA synthesis. The method involves generating a database of randomly generated oligonucleotides of a given length. The method may include obtaining one or more background populations (e.g., Bioinformatically or computationally determining the regions of nucleotides present in the human genome The resulting collection of oligonucleotides may then be subjected to a step of cleaving the oligonucleotides to a specific length. background length (e.g., 10, 11, 12, 13, 14, or 15 nucleotides, etc.) Such domains of nucleotides have little or no cross-sectional sequence. Sequences related to the target gene are calculated from a randomly generated heterogeneous collection of oligonucleotides. It can be deducted in the calculation.
[0105] Synthesis of a collection of non-host oligonucleotides After identifying the non-host oligonucleotide sequences, a collection of non-host oligonucleotides is prepared. Synthesis can be achieved through a myriad of techniques, including the synthesis of non-host oligonucleotides. , followed by nucleofection to produce oligonucleotides of desired length (e.g., 13-mer). The compounds can be synthesized individually by stepwise addition of the diamines.
[0106] In some instances, ultramer oligonucleotides are used as provided herein. A collection of oligonucleotides is synthesized. Generally, Ultramer oligonucleotides are used. The nucleotides are tethered together, each separated by one or more spacer nucleotides. It is a long oligonucleotide containing multiple oligonucleotide sequence units joined together. To generate a set of oligonucleotides from a single Ultramer oligonucleotide In addition, the spacer nucleotides can be cleaved or separated. The general design of a rutramer oligonucleotide is shown here. The sequence is separated by deoxyuracil nucleotides (U) which act as spacer nucleotides. Ultramer oligonucleotides are separated by uracil-DNA glycosylases. uracil dehydrogenase (UDG) (removes uracil residues from DNA by cleaving N-glycosidic bonds) cleaves at abasic sites (e.g., apurinic / apyrimidinic or AP sites) The enzymes are digested by the dual action of endonucleases with specificity for each individual oligonucleotides (e.g., 13-mer oligonucleotides) can be generated. In some cases, Ultramer oligonucleotides are purified by endonuclease IV or is hydrolyzed using endonuclease VII. A representative reaction is depicted in Figure 10A. Figure 10B shows the cleavage of ultramer oligomers by UDG and endonuclease VII. 1 shows examples of digestion reaction products after digestion of nucleotides.
[0107] Ultramer oligonucleotides are separated into individual oligonucleotides within the same Ultramer strand. Degeneracy can be designed so that all code units share the same degeneracy. This refers to the number of different unique sequences that can be generated from a variable sequence. This can be achieved by inserting one or more mixed bases at certain positions within the A mixed base can be any one of a set of bases (e.g., the four standard bases: A, C, G, or T bases). For example, in standard coding schemes, The mixed base can be any one of the A, C, G, or T bases. Similarly, a "D" mixed salt The "V" mixed base can be A, G, or T; the "B" mixed base can be A, C, or G; "H" mixed bases can be C, G, or T; "W" mixed bases can be A, C, or T; " " mixed bases can be A or T; "S" mixed bases can be C or G; "K" mixed bases "M" mixed base can be C or G; "Y" mixed base can be A or C; The group can be C or T; and the "R" mixed base can be A or G. Under this coding scheme, four possible sequences are generated, as shown in Figures 9B and 11A. Therefore, an oligonucleotide sequence containing one N-mixed base has a degeneracy of 4. Similarly, 16 possible sequences will be generated, as shown in Figures 9C and 11A. Therefore, an oligonucleotide sequence containing two N-mixed bases can be decomposed into 4 × 4 = 16 sequences. As shown in Figure 9D, eight possible sequences can be generated, An oligonucleotide containing one N mixed base and one W, S, K, M, Y, or R mixed base The nucleotide sequence will have a degeneracy of 4 x 2 = 8. Six possible sequences can be generated. Therefore, an oligonucleotide sequence containing one R mixed base and one D mixed base is , would have a degeneracy of 2×3=6.
[0108] If the oligonucleotide sequences within an Ultramer share the same degeneracy, Each oligonucleotide sequence in the sequence is likely to be represented by an equal number after Ultramer digestion. For example, if the oligonucleotide sequences in Ultramer each have a degeneracy of 4, If so, four different unique oligonucleotides per oligonucleotide sequence unit Each of these will be present in approximately equal amounts in the resulting composite pool. For example, Ultramer has one oligonucleotide unit with two degrees of degeneracy and If the oligonucleotide contains nine oligonucleotide units with four degrees of degeneracy, one two-degree unit Two oligonucleotides from the For 36 oligonucleotides from 9 four-degree units, each about 2.5% Therefore, the degeneracy of the oligonucleotide units is about 5%. Lutramer is a collection of oligonucleotides with uniformly distributed unique sequences. This can result.
[0109] Computationally identify and group non-host oligonucleotide sequences by degeneracy. Sequence degeneracy can be achieved by sequences that differ only at a certain number of positions, e.g., at position 1. and identified by computationally sorting all sequences that are identical except for two. Multiple degenerate sequences can be combined into a single variable sequence with one or more mixed bases. The number and type of mixed bases required to classify can determine the degree of degeneracy. Figure 11A shows the computational classification of non-host oligonucleotides by degeneracy. shows a representative oligonucleotide with three "N" mixed base sites. The histogram shows the bucketing of non-human 13mers based on degeneracy. Column 1 of Figure 11B are degenerate by 2 (they represent two possible bases, e.g., single mixed bases (e.g., W Mixed bases (S mixed bases, K mixed bases, M mixed bases, Y mixed bases, or R mixed bases) The second column indicates the number of non-human 13mers (meaning that the sequence is degenerate by 3 degrees, e.g. The remaining columns show the number of non-human 13mers with a base sequence (e.g., D, V, B, or H mixed bases). The numbers of non-human 13mers with 4, 6, 8, 9, 12 or 16 degrees of degeneracy are shown. For example, a non-human 13-mer with a degeneracy of 4 may contain one N mixed base or W, S, K, M. , Y and R.
[0110] Representatives of each degenerate group of non-host (e.g., non-human) oligonucleotides (e.g., 13mers) ) with other representatives of the same degeneracy, as discussed elsewhere in this specification, as shown in FIG. can then be combined with the designed The resulting Ultramer oligonucleotides can be synthesized by conventional nucleic acid synthesis methods and service providers, e.g. For example, it can be synthesized by IDT.
[0111] Ultrasound containing oligonucleotide units (or sequences) with shared degeneracy. Digestion of the nucleotides can be performed in one of several different steps. If so, combine groups of ultramers that all contain units that share the same degeneracy, then For example, to maintain equimolar concentrations of the resulting oligonucleotides, A plurality of ultramer oligonucleotides each having an oligonucleotide sequence with the same degeneracy. Alternatively, the leukocytes can be combined in equimolar concentrations and digested. The oligonucleotide units are first individually digested and then the resulting oligonucleotide units are can be combined in equimolar concentrations.
[0112] Ultramers containing oligonucleotide units (or sequences) with different degeneracy Digestion of the can be carried out in one of several different steps. Groups of ultramers, each containing units of different degeneracy, are combined and then digested. For example, to achieve equimolar concentrations of the resulting oligonucleotides, A plurality of Ultramer oligonucleotides having degenerate oligonucleotide sequences In another approach, Ultramar can be first and then the resulting oligonucleotide units are separated into the appropriate The oligonucleotides can be combined in a ratio to result in equimolar concentrations of the resulting oligonucleotides.
[0113] The Ultramer oligonucleotides provided herein may be any number of oligonucleotides. For example, Ultramer oligonucleotides may contain 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100 or 200 or more oligonucleotides In some cases, the oligonucleotide unit or sequence may be The sequence may be 5 to 100 nucleotides (e.g., 5, 10, 11, 12, 13, 14, 15, a certain length, such as 16, 17, 18, 19, 20 or more nucleotides Preferably, the nucleotide domain has the structure: In some embodiments, each domain of nucleotides is one or more, two or more, or a 3-mer. or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 11 or more, 12 or more, or 13 or more mixed bases. In the α-amino acid form, the mixed bases are N(A,C,G,T), D(A,G,T), V(A,C,G), B( C, G, T), H(A, C, T), W(A, T), S(C, G), K(G, T), M(A , C), Y(C, T), R(A, G), and any combination thereof. In some embodiments, the oligonucleotide sequences have the same degree of degeneracy. In some embodiments, the oligonucleotide sequence is 2, 3, 4, 6, 8, 9, 12, 1 6, 18, 24, 27, 32, 36, 48, 54, 64, 72, 81, 96, 108, 1 It has a degeneracy of 28, 144, 162, 192, 216, 243 or 256.
[0114] Individual oligonucleotides can be purified using, for example, T4 polynucleotide kinase (T4 PNK). The antibody can be biotinylated using the ELISA kit or chemically as shown in Figure 12. This step can be particularly useful when surface immobilization or magnetic bead purification is required. For example, biotinylated oligonucleotides can be combined with the sample to identify non-host oligonucleotides in the sample. The antibody can be coated with streptavidin and hybridized to oligonucleotides. The ligated beads are added to the sample and used to pull down non-host oligonucleotides. or can be isolated.
[0115] The Ultramer approach to producing the oligonucleotides provided herein is In some cases, such methods may result in high levels of oligonucleotide A highly diverse collection (e.g., a highly diverse collection of 13mer oligonucleotides) This method can reduce the number of synthesis runs required to create a compound. In addition, it reduces synthesis costs, reduces the synthesis of excess oligonucleotides, and Minimize manual steps for mixing nucleotides or achieve high yields of oligonucleotides It can enable synthesis.
[0116] Use of a collection of oligonucleotides In some cases, a collection of oligonucleotides is used, as shown in Figures 1 and 2. Use primers for PCR amplification, cDNA synthesis, sequencing, or primer extension reactions. or as a capture agent. In some cases, the collection of oligonucleotides The nucleic acid sample (140) is mixed with the nucleic acid sample (150). The sample is a biological sample, such as a blood sample (120) from a host (110) or derived from plasma (130), often using a collection of oligonucleotides When priming amplification, cDNA synthesis, sequencing, or primer extension, oligonucleotides are used. Nucleotides and sample nucleic acids are purified using DNA polymerase, reverse transcriptase, and RNA polymerase. In some cases, the nucleic acid sample may be combined with an appropriate polymerizing enzyme, such as a When RNA is involved, a collection of oligonucleotides is synthesized from an RNA template (150). In some cases, a nucleic acid sample can be used to prime DNA synthesis. When NA is included, a collection of oligonucleotides is used in a primer extension reaction (170). Primer extension reactions can be used to detect, for example, nucleic acid labels, nucleic acid tags, barcodes (e.g., sample barcode), universal primer sequence, primer binding site (e.g. , including but not limited to, DNA sequencing primer binding sites, sample barcode sequences Determinant primer binding sites and amplification primers compatible with various sequencing platform requirements (including mer-binding sites, for sequencing or barcode reading), sequencer Compatible sequences, sequences that attach to the sequencing platform, sequencing adapter sequences or adapters A primer can be added to add overhang sequences to the nucleic acids of the population of interest. In some cases, the primer extension reaction is performed in a population of interest (e.g., non-mammalian, non-human) Sequencing adapters can be added only to the nucleic acids of the target organism (e.g., a plant, a microorganism, or a pathogen). In some cases, a collection of oligonucleotides is used to amplify nucleic acids in a population of interest. They can be used as primers for PCR reactions to detect the presence of a label. A collection of tailored oligonucleotides can be used to label populations of interest. The nucleic acids of the labeled population of interest can then be transferred to, for example, a hybridization vector. The proteins can be isolated by pull-down or pull-down methods.
[0117] A pool of non-host oligonucleotides of any particular length can be prepared from DNA or RNA. may be different if one wishes to enrich for non-host sequences in the generated library. For example, compared to removing N-mers that are perfectly complementary to the host exome, After removing the N-mers that are perfectly complementary to the genome, the larger of all possible N-mers is A larger fraction of the non-host exome may remain. can be probed, potentially offering greater sensitivity in some cases. Furthermore, the same number of oligonucleotides can be used to probe the non-host genome and exome. When using oligonucleotides, a larger number of oligonucleotides that are not complementary to the host exome are used. When selecting which non-complementary oligonucleotides are included in the pool, It may provide additional versatility and allow for the selection of oligonucleotides with desired sequence characteristics. do.
[0118] In some cases, a collection of oligonucleotides is used as a nucleic acid probe, as shown in FIG. For example, a collection of oligonucleotides (250 ) as bait to capture a population of interest from a nucleic acid sample (240). In some cases, the collection of oligonucleotides can be labeled or Conjugated to a solid support. The nucleic acid is directed to a complementary or nearly complementary population of nucleic acids of interest in the sample (260). In some cases, labeled oligonucleotides can be used to hybridize or specifically bind to the target protein. A collection of oligonucleotides serves as primers for PCR amplification. In some cases, PCR amplification products are pulled down. In some cases, oligonucleotides are used. Labeling of the oligonucleotides in the collection is used to identify the oligonucleotide collection. and the nucleic acids of the hybridized or bound population of interest (e.g., biotin -avidin, biotin-streptavidin or nucleic acid hybridization interactions In some cases, the nucleic acid sample contains circulating nucleic acids, such as circulating cell-free nucleic acids. In some cases, the nucleic acid sample is an amplified sample, such as a sequenceable library of nucleic acids. This includes purified, isolated or isolated nucleic acids.
[0119] In some cases, methods for enriching for non-host (e.g., pathogen) sequences after library preparation As such, a collection of oligonucleotides can be used for nucleic acid pull-down. In some cases, a collection of labeled oligonucleotides (e.g., biotinylated) A collection of chemically labeled oligonucleotides (DNA fragments) is used to synthesize single-stranded DNA or cDNA fragments. In some cases, a polymerase is used to hybridize the library. and hybridizing oligonucleotides along the library fragments, e.g., under high fidelity conditions. The nucleotide can be extended. The extended hybridized oligonucleotide The antibody is then drawn using a binding partner for the label, such as streptavidin for biotin labeling. In some cases, the enriched library can be isolated, e.g., by PCR. The advantages of such a method include the ability to amplify the open end of the host nucleic acid of choice. In some cases, the enrichment method is a DNA or cDNA label enrichment method. In some cases, positive selection of non-host library fragments can be achieved. Selection leads to improved kinetics and thermodynamics. In some cases, standard library preparation Applying the method after preparation allows batch processing of multiple samples that can be individually barcoded It becomes possible to do this.
[0120] In some cases, a population of interest (e.g., non-host) in a nucleic acid sample, such as circulating nucleic acid Arrays, (a) providing a nucleic acid sample, the nucleic acid sample comprising: containing a population (e.g., host) nucleic acid and a population of interest nucleic acid; (b) mixing the nucleic acid sample with a collection of oligonucleotides, thereby obtaining a compound, the collection of oligonucleotides comprising a nucleotide domain containing an oligonucleotide having an isoform; (c) contacting the collection of oligonucleotides with the nucleic acid sample in a mixture; a mixture of nucleic acids of a target population by contacting the nucleic acids into domains of nucleotides; thereby priming or capturing nucleic acids of a population of interest. P and It can be primed or captured by
[0121] The primed or captured nucleic acids can be further analyzed or processed. In some cases, the priming or captured nucleic acid may be used as a primary target in a reaction (e.g., a PCR reaction). In some cases, the primed or captured nucleic acid can be amplified first. Sequencing assays (e.g., next-generation sequencing assays, high-throughput sequencing assays) sequencing assay, massively parallel sequencing assay, nanopore sequencing assay can be sequenced by performing a chromatographic or Sanger sequencing assay. In some cases, the primed or captured nucleic acid can be used (e.g., pull-down assay). In some cases, the primer extension reaction is performed by (e.g., to attach a nucleic acid label to the primed or captured nucleic acid) In some cases, the priming or capture The nucleic acid is RNA, which primes or primes the polymerization reaction (e.g., by reverse transcriptase). This is performed on the captured RNA nucleic acid.
[0122] In some cases, a population of interest (e.g., non-host) in a nucleic acid sample, such as circulating nucleic acid The array of (a) providing a nucleic acid sample, the nucleic acid sample comprising: containing a population (e.g., host) nucleic acid and a population of interest nucleic acid; (b) mixing the nucleic acid sample with a collection of oligonucleotides, thereby obtaining a compound, the collection of oligonucleotides comprising a nucleotide domain containing an oligonucleotide having an isoform; (c) contacting the collection of oligonucleotides with the nucleic acid sample in a mixture; a mixture of nucleic acids of a target population by contacting the nucleic acids into domains of nucleotides; and combining the (d) Sequencing assays (e.g., next-generation sequencing assays, high-throughput sequencing assays) Put sequencing assay, massively parallel sequencing assay, nanopore sequencing by performing a sequencing assay or Sanger sequencing assay sequencing the nucleic acids of the bound population; In some cases, the results of interest can be sequenced prior to sequencing. The nucleic acids of the combined population are preferentially amplified in a reaction (e.g., a PCR reaction). Prior to sequencing, the nucleic acids of the bound population of interest may be extracted (e.g., by performing a pull-down assay). In some cases, sequencing is preceded by a primer extension reaction. to the nucleic acids of the bound population of interest (e.g., to attach a nucleic acid label to the bound nucleic acid). In some cases, the nucleic acid of the binding population of interest is RNA. A polymerization reaction (e.g., by reverse transcriptase) is performed on the nucleic acid of the target binding population, which is RNA. This is done.
[0123] By contacting a collection of oligonucleotides with a nucleic acid sample, such as circulating nucleic acids, Thus, a portion of the background population (e.g., host) nucleic acid is converted into a domain of nucleotides. In some cases, the contacting can result in the removal of background population nucleic acids. Approximately 0.1, 0.5, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 1 4, 15, 20, 25, 30, 35, 40, 45 or 50% of the nucleotide domain In some cases, the contact reduces the background population of nucleic acids by up to 0.1 , 0.5, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 , 20, 25, 30, 35, 40, 45 or 50% of the nucleotides bound to the domain In some cases, the contacting reduces the nucleic acid content of the background population by at least 0.1, 0.5, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45 or 50% of the nucleotides are bound to the domain do.
[0124] The collection of oligonucleotides can be provided in excess relative to the sample nucleic acid. The ratio of oligonucleotide to sample nucleic acid is approximately 0.1; 0.2; 0.3; 0.4; 0 .5;0.6;0.7;0.8;0.9;1.0;1.1;1.2;1.3;1.4;1 .5;1.6;1.7;1.8;1.9;2.0;2.5;3.0;3.5;4.0;4 .5;5.0;5.5;6.0;6.5;7.0;7.5;8.0;8.5;9.0;9 .5;10;11;12;13;14;15;16;17;18;19;20;25;3 0;35;40;45;50;55;60;65;70;75;80;85;90;95 ;100;200;300;400;500;600;700;800;900;100 0;2000;3000;4000;5000;6000;7000;8000;900 or 10,000. The ratio of oligonucleotides to sample nucleic acids can be up to 0.1;0.2;0.3;0.4;0.5;0.6;0.7;0.8;0.9;1.0; 1.1;1.2;1.3;1.4;1.5;1.6;1.7;1.8;1.9;2.0; 2.5;3.0;3.5;4.0;4.5;5.0;5.5;6.0;6.5;7.0; 7.5;8.0;8.5;9.0;9.5;10;11;12;13;14;15;16 ;17;18;19;20;25;30;35;40;45;50;55;60;65; 70;75;80;85;90;95;100;200;300;400;500;60 0;700;800;900;1000;2000;3000;4000;5000;6 000; 7000; 8000; 9000; or 10000. The ratio of oligonucleotide to nucleotide should be at least 0.1; 0.2; 0.3; 0.4; 0.5; 0.6;0.7;0.8;0.9;1.0;1.1;1.2;1.3;1.4;1.5; 1.6;1.7;1.8;1.9;2.0;2.5;3.0;3.5;4.0;4.5; 5.0;5.5;6.0;6.5;7.0;7.5;8.0;8.5;9.0;9.5; 10;11;12;13;14;15;16;17;18;19;20;25;30;3 5;40;45;50;55;60;65;70;75;80;85;90;95;10 0;200;300;400;500;600;700;800;900;1000;2 000;3000;4000;5000;6000;7000;8000;9000; The ratio of oligonucleotide to sample nucleic acid can be saturating or 10,000. The ratio of oligonucleotide to sample nucleic acid may be non-saturating. The ratio can be calculated in terms of concentration, moles or mass.
[0125] The collection of oligonucleotides is provided in excess relative to the nucleic acids of the population of interest. The ratio of oligonucleotides to nucleic acids of the population of interest may be about 0.1; 0.2; 0. .3;0.4;0.5;0.6;0.7;0.8;0.9;1.0;1.1;1.2;1 .3;1.4;1.5;1.6;1.7;1.8;1.9;2.0;2.5;3.0;3 .5;4.0;4.5;5.0;5.5;6.0;6.5;7.0;7.5;8.0;8 .5;9.0;9.5;10;11;12;13;14;15;16;17;18;19 ;20;25;30;35;40;45;50;55;60;65;70;75;80; 85;90;95;100;200;300;400;500;600;700;800 ;900;1000;2000;3000;4000;5000;6000;7000; 8000; 9000; or 10000. The ratio of nucleotides is up to 0.1;0.2;0.3;0.4;0.5;0.6;0.7; 0.8;0.9;1.0;1.1;1.2;1.3;1.4;1.5;1.6;1.7; 1.8;1.9;2.0;2.5;3.0;3.5;4.0;4.5;5.0;5.5; 6.0;6.5;7.0;7.5;8.0;8.5;9.0;9.5;10;11;12 ;13;14;15;16;17;18;19;20;25;30;35;40;45; 50;55;60;65;70;75;80;85;90;95;100;200;30 0;400;500;600;700;800;900;1000;2000;3000 ;4000;5000;6000;7000;8000;9000;or 10000 The ratio of oligonucleotides to nucleic acids in the population of interest is at least 0.1; 0.2;0.3;0.4;0.5;0.6;0.7;0.8;0.9;1.0;1.1; 1.2;1.3;1.4;1.5;1.6;1.7;1.8;1.9;2.0;2.5; 3.0;3.5;4.0;4.5;5.0;5.5;6.0;6.5;7.0;7.5; 8.0;8.5;9.0;9.5;10;11;12;13;14;15;16;17; 18;19;20;25;30;35;40;45;50;55;60;65;70;7 5;80;85;90;95;100;200;300;400;500;600;70 0;800;900;1000;2000;3000;4000;5000;6000; 7000; 8000; 9000; or 10000. The ratio of oligonucleotides to nucleic acids of the population of interest may be saturating. The ratio of oligonucleotides may be non-saturating. The ratio may be in terms of concentration, molar or mass. It can be calculated from the above.
[0126] In some cases, the nucleic acid, such as a circulating nucleic acid or a sequenceable library, is single-stranded. In some cases, double-stranded nucleic acids are denatured to single-stranded nucleic acids. Nucleic acids can be denatured using heat. In some cases, the nucleic acid can be purified by about 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98 and In some cases, the nucleic acid is heated to about 0.1, 0.2, 0.3, 0.4, 0 .5, 0.6, 0.7, 0.8, 0.9, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, Heat for 15, 20, 25, 30, 40, 50, or 60 minutes. , denaturing nucleic acids using chemical denaturants (e.g., acids, bases, solvents, chaotropic agents, salts) do.
[0127] Regenerating single-stranded nucleic acid samples, e.g., single-stranded circulating nucleic acids or sequenceable libraries In some cases, a single-stranded nucleic acid sample can be prepared by cleaving or hybridizing the sample. In some cases, the oligonucleotides are hybridized to a collection of A collection of oligonucleotides is hybridized with a blocker oligonucleotide. In some cases, at least a portion of the single-stranded nucleic acid is renatured or hybridized. In this case, the nucleic acid is about 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 2 7, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 40, 45, 50 , 55, 60, 65, 68 or 70°C. In some cases, nucleic acids are renatured or hybridized on ice. In some cases, nucleic acids are renatured or hybridized at room temperature. In some cases, the nucleic acid is allowed to react with or hybridize to a nucleic acid having a concentration of about 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50 or 60 minutes, or 2, 3, 4, 5 , 10, 15, 20, 22, 24, 30, 40, 46, 48, 50, 60, 70, 72, Regenerate or hybridize for 80 or 96 hours. Temperature, buffer composition, reaction time, and The amount and concentration of the hybridization may affect the degree of hybridization. Nucleic acids are renatured in the presence of ethylammonium chloride.
[0128] Self-hybridization method Further methods for priming, capturing, or enriching populations of interest within nucleic acid samples Methods and compositions are also provided herein. In some cases, the subject population is treated with autologous hybridomas. Hybridization methods, such as those utilizing concentration-based hybridization kinetics Nucleic acid hybridization kinetics is characterized by the hybridization of the strands. Generally, high concentrations of nucleic acid hybridize faster than low concentrations of nucleic acid. In a complex population of nucleic acids, abundant nucleic acids (e.g., background population) may overwhelm rare nucleic acids ( This concentration dependence can be used to determine whether the target population is hybridized faster than the target population. After partial hybridization of the population of nucleic acids, at least a portion of the double-stranded nucleic acids are removed. or reducing the amount of abundant nucleic acid by isolating at least a portion of the single-stranded nucleic acid. As a result, the amount of rare nucleic acids in a sample can be enriched. (e.g., human) nucleic acid amount can be reduced, and the population of interest (e.g., non-human, The amount of nucleic acid (microorganism or pathogen) can be concentrated.
[0129] In some cases, background population DNA (e.g., human DNA) and target DNA are used. A sample of DNA (e.g., a sample of DNA containing population DNA (e.g., non-human or pathogen DNA) Circulating DNA (circulating cell-free DNA) is obtained by denaturing the nucleic acids in the sample to form single-stranded DNA. A sample of nucleic acid can be prepared. In some cases, the nucleic acid in the sample is single-stranded DNA. In some cases, background population DNA is then removed. is present in higher concentrations than the population of DNA of interest, so the background population DNA The sample is placed in a defined buffer, with the expectation that it will hybridize faster than the target population DNA. The sample is then renatured in buffer for a defined period of time. The sample is subjected to conditions that remove double-stranded DNA, such as by Background population DNA can be preferentially removed. In some cases, RNA is also present. The sample is combined with a mixture of single-stranded background population (e.g., human) exome sequences. In some cases, samples are regenerated for a defined period to obtain a more abundant exome. The DNA-RNA duplex is then hybridized to the RNA. thereby preserving background population (e.g., human) RNA in the sample. Remove it first.
[0130] In some cases, nucleic acids, such as circulating nucleic acids, are single-stranded. In some cases, double-stranded nucleic acids are Denaturation into single-stranded nucleic acid. In some cases, about 1, 2, 3, 4, 5, 6, 7, 8, 9 , 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99 or 100% are single-stranded In some cases, up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20 , 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99 or 100% are single-stranded. At least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30 of the nucleic acid , 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99 or 100% are single-stranded. Nucleic acids can be denatured using heat. In some cases, the nucleic acid can be 0, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98 or In some cases, the nucleic acid is heated to 99°C for up to 35, 40, 45, 50, 55, 60 , 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, In some cases, the nucleic acid is heated to 98 or 99°C. 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 9 In some cases, the nucleic acid is heated to about 0.1, 0.5, 96, 97, 98, or 99°C. .2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 2, 3, 4, 5 Cook for 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50 or 60 minutes. In some cases, nucleic acids are added at up to 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0 0.7, 0.8, 0.9, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25 , 30, 40, 50, or 60 minutes. In some cases, the nucleic acid is heated for at least 0. 1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 2, 3, Heat for 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50 or 60 minutes In some cases, nucleic acids are denatured by chemical denaturants (e.g., acids, bases, solvents, chaotropes, etc.). Nucleic acids are denatured using blocking agents (blocking agents, salts).
[0131] Single-stranded nucleic acid samples, e.g., single-stranded circulating nucleic acids, can be renatured or hybridized. In some cases, at least a portion of the single-stranded nucleic acid is renatured or hybridized. In some cases, about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, or 20 of the single-stranded nucleic acid , 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, Regenerate or hybridize 85, 90, 95, 96, 97, 98, 99 or 100% In some cases, up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 , 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99 or 100% regenerated or hybridized In some cases, at least 1, 2, 3, 4, 5, 6, 7, or 8 of the single-stranded nucleic acid is 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, Regenerate 70, 75, 80, 85, 90, 95, 96, 97, 98, 99 or 100% In some cases, the nucleic acid is selected from about 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 2 1, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34 , 35, 36, 37, 40, 45, 50, 55, 60, 65, 68 or 70°C In some cases, the nucleic acid is selected from a group consisting of up to 0, 1, 2, 3, 4, 5, 6 ,7,8,9,10,11,12,13,14,15,16,17,18,19,20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 3 Regenerate at 4, 35, 36, 37, 40, 45, 50, 55, 60, 65, 68 or 70°C In some cases, the nucleic acid is hybridized to at least 0, 1, 2, 3, 4 , 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 , 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 40, 45, 50, 55, 60, 65, 68 or 70 In some cases, nucleic acids are renatured or hybridized on ice. In some cases, nucleic acids are allowed to renature or hybridize at room temperature. In some cases, nucleic acids are diluted to about 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, .8, 0.9, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50 or 60 minutes, or 2, 3, 4, 5, 10, 15, 20, 22, 24, 3 0, 40, 46, 48, 50, 60, 70, 72, 80 or 96 hours playback or high In some cases, the nucleic acid is hybridized at up to 0.1, 0.2, 0.3, 0.4, 0 .5, 0.6, 0.7, 0.8, 0.9, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50 or 60 minutes, or 2, 3, 4, 5, 10, 1 5, 20, 22, 24, 30, 40, 46, 48, 50, 60, 70, 72, 80 or The nucleic acid is allowed to regenerate or hybridize for 96 hours. In some cases, the nucleic acid is allowed to regenerate or hybridize for at least 0. 1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50 or 60 minutes, or 2, 3, 4, 5, 10, 15, 20, 22, 24, 30, 40, 46, 48, 50 Allow to regenerate or hybridize for 60, 70, 72, 80, or 96 hours. The buffer composition, reaction time, and concentration can affect the degree of hybridization. In some cases, the nucleic acid is renatured in the presence of trimethylammonium chloride.
[0132] Upon renaturation, the single-stranded nucleic acid hybridizes or reanneals to a double-stranded nucleic acid. In some cases, the double-stranded nucleic acid can be double-stranded DNA, double-stranded RNA, or DNA- The double-stranded nucleic acid is an RNA duplex. At least a portion of the double-stranded nucleic acid is removed to produce an enriched population of nucleic acids. In some cases, about 1, 2, 3, 4, 5, 6, 7, 8, 9 of the double-stranded nucleic acid can be , 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, Removes 75, 80, 85, 90, 95, 96, 97, 98, 99 or 100%. In some cases, up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 2 double-stranded nucleic acids 0, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85 , 90, 95, 96, 97, 98, 99 or 100%. at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25 strands of double-stranded nucleic acid; 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 9 Remove 5, 96, 97, 98, 99 or 100%.
[0133] In some cases, at least a portion of the double-stranded nucleic acids are removed to produce an enriched population of nucleic acids. In some cases, at least a portion of the single-stranded nucleic acids are isolated to produce an enriched population of nucleic acids. In some cases, at least a portion of the double-stranded nucleic acid is purified by gel electrophoresis or capillary electrophoresis. In some cases, at least a portion of the single-stranded nucleic acid is removed using a separation method such as electrophoresis. are isolated using separation methods such as gel electrophoresis or capillary electrophoresis. In some cases, at least a portion of the double-stranded nucleic acid is removed using one or more nucleases. In some cases, nucleases act on double-stranded nucleic acids. Nucleases act on double-stranded DNA. In some cases, nucleases can break down DNA-RNA duplexes. In some cases, nucleases do not act on single-stranded nucleic acids. In some cases, nucleases do not act on single-stranded DNA. Some non-limiting examples of nucleases that do not act on RNA include double-strand specific nucleases. clease (DSN) (e.g., from Kamchatka crab), thermolabile DSN-TL, BAL-31, double-strand specific DNase (e.g., from the northern shrimp, Pandalus borealis) Alice (Pandalus borealis), exonuclease III and and secretory endonucleases (e.g., Culex quinquefa sciatus). Temperature, buffer composition, reaction time, concentration and nucleic acid The ratio of nuclease to substrate can affect nuclease specificity or activity. The substrate preference of double-strand specific DNase from the prawn, Pseudomonas aeruginosa, is different in the presence of magnesium. Although it has been reported that the nuclease is a single-stranded DNA, in the presence of calcium, the nuclease It becomes active against NA (Nilsen et al., "The Enzyme and d the cDNA Sequence of a Thermolabile an d Double-Strand Specific DNase from Nort hern Shrimps (Pandalus borealis) PLOS On e 2010). Nuclease combinations can be used in series or in parallel. In some cases, nucleic acids are purified after nuclease treatment, changing the buffer composition, etc. and / or remove nucleases, nucleotides and / or short nucleic acid fragments.
[0134] In some cases, the sample containing nucleic acid is subjected to a single-stranded background population (e.g., human ) combined with a mixture of nucleic acids. The sample is regenerated for a period of time to obtain a richer background. The nucleic acids in the ring population are allowed to hybridize to the nucleic acids in the sample. removes double-stranded nucleic acids, thereby reducing background populations (e.g., human) in the sample Preferentially removes nucleic acids. In some cases, the sample contains RNA. In some cases, The double-stranded nucleic acid is a DNA-RNA duplex. In some cases, the background population nucleic acid In some cases, depletion is performed to remove background population nucleic acids from the sample. hybridize or specifically hybridize to complementary or nearly complementary background populations within the cell. by contacting a sample containing nucleic acid with background population nucleic acid so that it binds. In some cases, the method comprises: The method further includes denaturing the portion, for example, by heat. The nucleic acid sequences of the cross-population can be identified bioinformatically or computationally. , the background population nucleic acid is synthesized. In some cases, the background population nucleic acid is DNA.
[0135] The background population nucleic acid may be provided in excess relative to the sample nucleic acid. The ratio of round population nucleic acid to sample nucleic acid is approximately 0.1; 0.2; 0.3; 0.4; 0.5; 0. .6;0.7;0.8;0.9;1.0;1.1;1.2;1.3;1.4;1.5;1 .6;1.7;1.8;1.9;2.0;2.5;3.0;3.5;4.0;4.5;5 .0;5.5;6.0;6.5;7.0;7.5;8.0;8.5;9.0;9.5;1 0;11;12;13;14;15;16;17;18;19;20;25;30;35 ;40;45;50;55;60;65;70;75;80;85;90;95;100 ;200;300;400;500;600;700;800;900;1000;20 00;3000;4000;5000;6000;7000;8000;9000;Also can be 10000. The ratio of background population nucleic acid to sample nucleic acid is up to 0.1; 0.2;0.3;0.4;0.5;0.6;0.7;0.8;0.9;1.0;1.1; 1.2;1.3;1.4;1.5;1.6;1.7;1.8;1.9;2.0;2.5; 3.0;3.5;4.0;4.5;5.0;5.5;6.0;6.5;7.0;7.5; 8.0;8.5;9.0;9.5;10;11;12;13;14;15;16;17; 18;19;20;25;30;35;40;45;50;55;60;65;70;7 5;80;85;90;95;100;200;300;400;500;600;70 0;800;900;1000;2000;3000;4000;5000;6000; It can be 7000; 8000; 9000; or 10000. Background population nuclei The ratio of acid to sample nucleic acid is at least 0.1; 0.2; 0.3; 0.4; 0.5; 0.6; 0.7;0.8;0.9;1.0;1.1;1.2;1.3;1.4;1.5;1.6; 1.7;1.8;1.9;2.0;2.5;3.0;3.5;4.0;4.5;5.0; 5.5;6.0;6.5;7.0;7.5;8.0;8.5;9.0;9.5;10;1 1;12;13;14;15;16;17;18;19;20;25;30;35;40 ;45;50;55;60;65;70;75;80;85;90;95;100;20 0;300;400;500;600;700;800;900;1000;2000; 3000; 4000; 5000; 6000; 7000; 8000; 9000; or 10 The ratio of background population to sample nucleic acid can be saturating or The ratio of background population to sample nucleic acid may be non-saturating. , can be calculated in terms of concentration, moles or mass.
[0136] In some cases, at least a portion of the double-stranded nucleic acids are removed to produce an enriched population of nucleic acids. The background population nucleic acid is chemically labeled and removed, thereby obtaining double-stranded nucleic acid. Background population nucleic acids can be conjugated to magnetic beads. , can be removed with a magnet, thereby removing double-stranded nucleic acids. The population nucleic acids can be labeled with a nucleic acid label. The nucleic acid label can be conjugated to a solid support. that bind or hybridize to nucleic acid sequences that are tagged with a chemical label or Double-stranded nucleic acids can be separated by size separation (e.g., gel electrophoresis, capillary electrophoresis, etc.). electrophoresis, etc.), affinity separation (e.g., pull-down assay, chromatography, etc.), or In some cases, double-stranded nucleic acids can be removed by other methods. , single-stranded nuclei using separation methods such as capillary electrophoresis or chromatography. The nucleic acids can be isolated based on the size difference between the nucleic acid and the double-stranded nucleic acid. It does not remove unhybridized background population nucleic acids.
[0137] Enrichment by nucleosome depletion Another example of an enrichment method provided herein is a nucleosome depletion method. Human cell-free DNA has a length periodicity of approximately 180 base pairs, as shown in Figure 6. This suggests that the majority of cell-free DNA is histones associated with nucleosomes. Bacterial DNA does not exhibit any particular length periodicity, as shown in Figure 6. or enriching a population of interest by depleting nucleosome-associated DNA. It can provide a law.
[0138] In some cases, free DNA (e.g., non-nucleosomal DNA or non-nucleosomal DNA) Separation of nucleosomal DNA or nucleosome-associated DNA from nucleosome-associated DNA Methods include electrophoresis and isotachophoresis, which separate based on mass and / or net charge. , porous filters that separate based on shape and / or size, and filters that separate based on charge. ion exchange column (nucleosomes have histones with positively charged tails, and DNA A has a negatively charged backbone), and immunoglobulins that bind host nucleosomes and associated DNA. The antibody specific for the host histone to be depleted is included. In some cases, a portion of the host nucleic acid is nucleosome-associated, and a portion of the host nucleic acid is not nucleosome-associated.
[0139] In some cases, one or more antibodies were used to identify histones or In some cases, antibodies (750) can immunodeplete nucleosomes or nucleosomes. , specific to background populations (e.g., host) histones or nucleosomes (710) In some cases, the antibody may be specific for mammalian histones or nucleosomes. In some cases, the antibody may be specific for human histones or human nucleosomes. In some cases, the one or more antibodies may be specific for one or more histones. In some cases, histones are histone variants but still retain histone modifications. In some cases, immunodepletion may be achieved by nucleating the population of interest (e.g., non-host). nucleosomal DNA (730), a population of targeted non-nucleosomal DNA (740), and and background population (e.g., host) non-nucleosomal DNA (720). In some cases, immunodepletion can be performed using immunoprecipitation, chromatin immunoprecipitation, bulk antibody binding, or or affinity chromatography with column-immobilized antibodies. In some cases, one or more antibodies are immobilized on a column. Several antibodies (e.g., anti-immunoglobulin conjugated to beads or attached to a column) are used. In some cases, one or more antibodies are monoclonal antibodies. In some cases, the one or more antibodies are directed against one or more histones. In some cases, the one or more antibodies target one or more His-terminus. Targeting the N-terminus of t
[0140] Non-limiting examples of histones, histone variants, and histone modifications include histone H2 AN terminus, solvent-exposed epitope on histone H2A, monomer on Lys9 in histone H3 methylation on Lys9 in histone H3, dimethylation on Lys56 in histone H3 remethylation, phosphorylation on Ser14 in histone H2B, and Se in histone H2A.X Phosphorylation on r139, H2A-H2B acidic patch motif, histone H1, histone H 1.0, histone H1.1, histone H1.2, histone H1.3, histone H1.4, Histone H1.5, histone H1.oo, spermatid-specific linker histone H1-like protein Protein, histone H1t, histone H1t2, histone H1FNT, histone H2A type 1-B / E (e.g., histone H2A.2, histone H2A / a, histone H2A / m) , histone H2A type 2-A (e.g., histone H2A.2, histone H2A / o), Histone H2A type 1-D (e.g., histone H2A.3, histone H2A / g), histone Histone H2A type 1 (e.g., H2A.1, histone H2A / p), histone H2A Type 2-C (e.g., histone H2A-GL101, histone H2A / q), histone H 2A type 1-A (e.g., histone H2A / r), histone H2A type 1-C (e.g., histone H2A / 1), histone H2A type 1-H (e.g., histone H2A / s ), histone H2A type 1-J (e.g., histone H2A / e), histone H2A type Type 2-B, histone H2A type 3, histone H2AX (e.g., H2a / x, histone H2A.X, histone H2A.Z, H2A / z, H2A.Z.1, H2A.Z.2) Histone H2A.V (e.g., H2A.F / Z), histone H2A.J (e.g., H2a / j), mH2A1, mH2A2, histone H2A-Bbd type 1 (e.g., H2A-Bbd -body deficiency, H2A.Bbd), histone H2A-Bbd type 2 / 3 (e.g., H2A Barr body deficiency, H2A.Bbd), core histone macroH2A.1 (e.g., histone macromolecules ChrH2A1, mH2A1, histone H2A.y, H2A / y, medulloblastoma antigen MU-MB- 50.205, macroH2A), core histone macroH2A.2, histone H2A, histone Histone H2A.J, histone H2B, histone H2B type 1-C / E / F / G / I (e.g. For example, histone H2B.1A, histone H2B.a, H2B / a, histone H2B.g, H 2B / g, histone H2B.h, H2B / h, histone H2B.k, H2B / k, histone Histone H2B.l, H2B / I), H2BE, histone H2B type 1-H, histone H2B Type 1-A (e.g., testis-derived histone H2B, TSH2B.1, testis-specific histone histone H2B, TSH2B), histone H2B type 1-B (e.g., histone H2B.1, histone H2B.f, H2B / f), histone H2B type 1-D (e.g., HIRA phase interacting protein 2, histone H2B.1B, histone H2B.b, H2B / b), histone Histone H2B type 1-J (e.g., histone H2B.1, histone H2B.r, H2B / r)), histone H2B type 1-O (e.g., histone H2B.2, histone H2B. n, H2B / n), histone H2B type 2-E (e.g., histone H2B-GL105 , histone H2B.q, H2B / q), histone H2B type 1-H (e.g., histone H2B.j, H2B / j), histone H2B type 1-M (e.g., histone H2B. e, H2B / e), histone H2B type 1-L (e.g., histone H2B.c, H2B / c), histone H2B type 1-N (e.g., histone H2B.d, H2B / d), histone Histone H2B type FS (e.g., histone H2B.s, H2B / s), putative histone Histone H2B type 2-C (e.g., histone H2B.t, H2B / t), histone H2B Type 1-K (e.g., H2B K, HIRA-interacting protein 1), histone H2B type 2-F, histone H2B type 2-E, histone H2B type 3-B (e.g., H 2B type 12), histone H2B type FM (e.g., histone H2B.s, H2B / s), putative histone H2B type 2-D, histone H2B type WT (e.g., H2B histone family member (testis-specific), histone 1H2bn isoform CRA_b, histone H2B type 1-N, H2B histone family member M, His Histone H2B type FM, histone H2B type 1-J, H2BFWT, histone H3, Histone H3.1 (e.g., histone H3 / a, histone H3 / b, histone H3 / c, Histone H3 / d, Histone H3 / f, Histone H3 / h, Histone H3 / i, Histone H3 / j, histone H3 / k, histone H3 / l), histone H3.2 (e.g., histone histone H3 / m, histone H3 / o), histone H3.3C (e.g., histone H3.5), Histone H3.1t (e.g., H3 / t, H3t, H3 / g), histone H3.3, histone TonH3-like centromere protein A (e.g., centromere autoantigen A, centromere Protein A, CENP-A), histone H3.3 (cDNA FLJ57905), Histone H3.4, Histone H3.5, Histone H3.X, Histone H3.Y, Histone H 4. Histone H4-like protein type G, HIST1H4J protein, human (Homo sapiens) H4 histone family member N and histone H5 .
[0141] The method can deplete the background population of nucleosome-associated DNA. Nucleosome-associated DNA in the background population was analyzed by approximately 5, 6, 7, 8, 9, 10, 15, and 20 , 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100% depletion The background population of nucleosome-associated DNA can be reduced by up to 5, 6, 7, or 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 9 9 or 100% depletion of background nucleosome associations. DNA at least 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 9 It can be depleted by 4, 95, 96, 97, 98, 99 or 100%.
[0142] Enrichment by removing or isolating DNA within a specific length interval In some cases, the enrichment methods provided herein may be used to enrich for specific length intervals of cell-free DNA, such as Removal of DNA in a specific length interval and / or cell-free DNA in a specific length interval In one example, cell-free DNA samples are isolated from a specific length interval, such as a chromatin. In this case, the ratio of host to non-host cell-free DNA may be minimal. The range of cell-free DNA lengths that may be present can vary significantly with some predictability. If you are interested in analyzing non-host DNA in cell samples, you should separate the non-host DNA and host DNA. By enriching for DNA lengths with more favorable ratios of NAs, non-host DNA is The main DNA can be enriched. For example, Figure 6 shows the top bar representing host or human DNA. The bottom bar represents non-host DNA. The plot shows the association of human DNA with histones in the approximately 175 base pairs. The fragments of the first size that occur in the first step and the fragments that are about 147 bases long (i.e., a single His fragment) As shown in Figure 6, The work shows the size distribution of human DNA, providing maxima and minima for other fragment lengths as well. .
[0143] A size range in which the host to non-host DNA ratio is at a low point, e.g., as shown, e.g., about Choose a length of less than 120 bp, approximately 240 to 280 bp, or approximately 425 to 475 bp. By doing so, non-host DNA can be relatively concentrated. According to certain embodiments, the cell-free nucleic acid sample is enriched for non-host nucleic acids relative to host nucleic acids. In most cases, the enrichment step is carried out on relatively short fragments of about 10 bases to about 3 bases in length. 00 base length, approximately 10 base length to approximately 200 base length, approximately 10 base length to approximately 175 base length, approximately 10 base length base length to approximately 150 bases, approximately 10 bases to 120 bases, approximately 10 bases to approximately 60 bases, approximately 30 bases to approximately 300 bases, approximately 30 bases to approximately 200 bases, approximately 30 bases to approximately 175 bases Base length, about 30 bases to about 150 bases, about 30 bases to 120 bases, or about 30 bases Enrich fragments with lengths between about 60 bases. As will be appreciated, the upper limit of the selection process and The lower limit selection can target either the lower or upper limits of the size selections above. As will be understood, the size selections above are not intended to list exact sizes. Instead, sizes around the ranges above that encompass the normal size distribution around the stated fragment sizes are used. For example, if a given fragment size range is selected, the enumeration If the enriched sample is sized to a size that is less than the range of the predominant size range in the enriched sample, e.g., At least 50%, at least 60%, at least 70%, at least 80% of the fragments in or in some cases, at least 90% are recognized to reflect the listed size range. In other cases, the reflected fragments in the enriched sample fall within the upper and lower limits of the recited range. Substantially outside by about 30% or less, 20% or less, 10% or less, and several In some cases, less than 5% of the fragments are recognized as being outside the original.
[0144] Such methods involve the analysis of population (e.g., pathogen) DNA and background DNA. a sample of nucleic acid (e.g., a sample of circulating DNA) containing human population (e.g., human) DNA; In some cases, DNA of a particular length interval is compared with background population DNA enriched for and depleted from the sample, thereby identifying the DNA of the population of interest In some cases, DNA of a particular length interval can be preferentially enriched. The DNA of a population of interest is enriched and isolated from the sample, thereby In some cases, DNA in a particular length interval can be preferentially enriched. Methods for removing and / or isolating NAs include electrophoresis (e.g., gel electrophoresis or capillary electrophoresis). column electrophoresis), chromatography (e.g., liquid chromatography), cell-free nuclei Acid purification (e.g., silica membrane columns, buffer optimization) and / or mass, size, and and / or mass spectrometry, which separates based on net charge.
[0145] In some cases, cell-free nucleic acid purification is performed using silica membranes (e.g., QIAamp Mini columns). silica membrane columns) or commercially available kits (e.g., QIAamp Circulat This can be done using the Binding Nucleic Acid Kit. Cellular nucleic acid purification involves four steps: lysis, binding, washing, and elution.
[0146] The lysis step releases nucleic acids from proteins, lipids and / or vesicles, allowing DNase and / or RNase inactivation, under denaturing conditions, at high temperature, and / or protein The lysis step can be carried out in the presence of enzyme K. The lysis step can be carried out in buffer ACL and / or Protease. Proteinase K can be used.
[0147] The binding step allows the nucleic acid to be bound or adsorbed onto the silica membrane. The step is performed by adding buffer ACB or about or at least about 35% by volume, 40% by volume, 45% by volume, vol%, 50 vol%, 55 vol%, 60 vol%, 65 vol%, 66 vol%, 70 vol% or a binding buffer containing 75% by volume of alcohol (e.g., isopropanol), e.g., Binding buffers containing about 40% or about 66% isopropanol can be used. In some cases, additional steps were taken during cell-free nucleic acid purification compared to the manufacturer's recommended protocol. Commercially available buffers (e.g., buffer ACB) are used. For example, in some cases, The volume of additional commercially available buffer used should be approximately equal to or less than the volume specified in the manufacturer's recommended protocol. or at least about 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, 3.0, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 3.9, 4.0, 4.1, 4.2, 4.3, 4.4, 4.5, 4.6, 4.7, 4.8, 4.9, 5.0, 6.0, 7.0, 8.0, 9.0 or 10.0 times. per silica membrane column or from 1 mL of sample (e.g., serum or plasma) per solution of about or at least about 1.5, 1.8, 2.0, 2.5, 3.0, 3. 5, 4.0, 5.0, 5.3, 5.4, 5.5, 6.0, 7.0, 8.0, 9.0, 10 0.0, 11.0, 12.0, 13.0, 14.0, 15.0 or 16.0 mL of binding buffer Use buffer or ACB.
[0148] The cleaning step can remove residual contaminants and can include multiple cleaning steps. The washing steps consist of Buffer ACW1, Buffer ACW2, ethanol and / or or about or at least about 50%, 55%, 56.8%, 60%, 6% by volume 5% by volume, 66% by volume, 69.8% by volume, 70% by volume, 75% by volume, 80% by volume, 85 Volume%, 86 volume%, 87 volume%, 87.4 volume%, 88 volume%, 89 volume%, 90 volume% , 95 vol%, 96 vol%, 97 vol%, 98 vol%, 99 vol% or 100 vol% Wash buffers containing alcohol (e.g., ethanol), e.g., about 56.8%, about 69.8% Use a wash buffer containing about 87.4%, about 89%, or about 100% ethanol. In some cases, ethanol can refer to 96-100% ethanol. In some cases, the ethanol is not denatured. In some cases, the ethanol is methylated. In some cases, the washing step does not contain ethanol or methyl ethyl ketone. First wash step with ACW1 (e.g., 600 μL per silica membrane column), A second wash step with buffer ACW2 (e.g., 750 μL per silica membrane column) and a third washing step with ethanol (e.g., 750 μL per silica membrane column). In some cases, during cell-free nucleic acid purification, the manufacturer's recommended protocol may be followed. In comparison, additional ethanol (e.g., absolute ethanol) is added to commercially available buffer solutions (e.g., Qi Add to the buffer (ACW1 buffer, ACW2 buffer). For example, in some cases, The volume of additional ethanol added is per 600 μL or 750 μL of the commercial buffer. , about or at least about 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.05, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.75, 1.8, 1.9 , 2.0, 2.5, 3.0, 3.5, 4.0, 4.5, 5.0, 6.0, 7.0, 8.0 , 9.0 or 10.0 mL. In some cases, the amount of additional ethanol added is The volume is up to 0.5, 0.6, 0.7 per 600 μL or 750 μL of commercially available buffer solution. ,0.8,0.9,1.0,1.05,1.1,1.2,1.3,1.4,1.5,1. 6, 1.7, 1.75, 1.8, 1.9, 2.0, 2.5, 3.0, 3.5, 4.0, 4 0.5, 5.0, 6.0, 7.0, 8.0, 9.0 or 10.0 mL.
[0149] In some cases, additional steps were taken during cell-free nucleic acid purification compared to the manufacturer's recommended protocol. Guanidinium chloride, which is a soluble form of guanidinium chloride, is added to commercially available buffers (e.g., Qiagen ACW1 buffer). For example, in some cases, the amount of additional guanidinium chloride added may be greater than that of a commercially available About or at least about 0.1, 0.2, 0.3, 0.31, 0 per 600 μL of buffer .32, 0.33, 0.34, 0.35, 0.36, 0.37, 0.38, 0.39, 0 .4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1 In some cases, the weight is 0.4, 1.5, 1.6, 1.75, 1.8, 1.9 or 2.0g. In this case, the amount of additional guanidinium chloride added is 60% of the commercial buffer or the wash buffer. Per 0μL, up to 0.1, 0.2, 0.3, 0.31, 0.32, 0.33, 0.34 ,0.35,0.36,0.37,0.38,0.39,0.4,0.5,0.6,0. 7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1. In some cases, the wash buffer is 75, 1.8, 1.9, or 2.0 g. Per 600 μL of liquid, about or at least about 0.1, 0.2, 0.3, 0.31, 0.3 2, 0.33, 0.34, 0.35, 0.36, 0.37, 0.38, 0.39, 0.4 , 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.4 , 1.5, 1.6, 1.75, 1.8, 1.9 or 2.0 g of guanidinium chloride It can be done.
[0150] In some cases, additional steps were taken during cell-free nucleic acid purification compared to the manufacturer's recommended protocol. Ethanol (e.g., absolute ethanol) and additional guanidinium chloride were added to commercially available buffers. Add to buffer (e.g., Qiagen ACW1 buffer). In some cases, a wash step The wash was performed with a first wash buffer containing guanidinium chloride and approximately 89.0% ethanol. a first wash step with a second wash buffer containing about 87.4% ethanol; In some cases, the method may include a first wash step with ethanol, as well as a third wash step with ethanol. , each wash buffer is at least about 60%, 65%, 66%, 69.8%, 70%, 75%, Contains 80%, 85%, 86%, 87% or 87.4% ethanol.
[0151] The nucleic acids can be released from the silica membrane by an elution step. AVE can be used.
[0152] For example, Qiagen Circulating Nucleic Acid (CNA ) The kit manufacturer's recommended protocol can be modified with the following modifications: a) Use 3 volumes of ACB buffer; (b) Use ACW1 buffer according to the manufacturer's recommendations. Prepare an additional 1.75 mL of absolute ethanol and chloride per 600 µL of ACW1 buffer. (c) ACW2 buffer solution prepared according to the manufacturer's recommendations, supplemented with 0.36 g of guanidinium. Prepare and supplement with 1.05 mL of absolute ethanol per 750 µL of ACW2 buffer.
[0153] In addition to other nucleosome targeted depletion methods described elsewhere herein, nucleic acid subunits may be used. Methods for size selection / isolation are well known in the art. For example, gel electrophoresis ... Chromatographic methods such as exclusion chromatography can be used to separate nucleic acids of desired lengths. Furthermore, bead-based charge separation methods can be used to selectively isolate desired Nucleic acids of a range of sizes can be selectively isolated. The SPRI bead system available from Ter, and e.g., New England The AMPure bead system available from Biolabs can be easily used to perform the above-mentioned Size selection of cell-free DNA allows amplification of non-host DNA relative to host DNA. Purification can be carried out.
[0154] In some cases, size selection enrichment is at least about 1.5-fold, at least about 2-fold, or less. at least about 3 times, at least about 4 times, at least about 5 times, at least about 6 times, at least about 7 times, at least about 8 times, at least about 9 times, at least about 10 times, at least about 20 times , at least about 30 times, at least about 40 times, at least about 50 times, at least about 100 times fold and in some cases, a ratio of non-host DNA to host DNA of about 500-fold or more For illustrative purposes only, if non-host DNA is 1:100 to host DNA, If present in a cell-free sample at a ratio of 1:50, an enrichment resulting in a 2-fold increase would yield a ratio of 1:50. In some cases, the increase in the ratio of non-host DNA to host DNA is approximately 1.5-fold, 2x, 3x, 4x or 5x, and about 10x, 20x, 30x, 40x, 50x, 10 It can be 0x, 500x or more.
[0155] One or more length intervals of DNA can be removed. In some cases, one or more The length interval or intervals may be about 140 base pairs, about 145 base pairs, about 150 base pairs, about 155 base pairs, or base pairs, about 160 base pairs, about 165 base pairs, about 170 base pairs, about 175 base pairs, about 180 base pairs base pairs, about 185 base pairs, about 190 base pairs, about 195 base pairs, about 200 base pairs, about 205 base pairs The number of base pairs may be selected from one or more multiples of about 210 base pairs. In this case, the multiples are from 1, 2, 3, 4, 5, 6, 7, 8, 9 and 10 for each occurrence. For example, about 180 base pairs is a multiple of 1 for about 180 base pairs, About 360 base pairs is a multiple of 2 relative to about 180 base pairs. In some cases, one or The length intervals are about 180 base pairs, about 360 base pairs, about 540 base pairs, and about 720 base pairs. Or it may be about 900 base pairs.
[0156] One or more length intervals of DNA can be isolated. The length interval or intervals may be about 10 base pairs, about 20 base pairs, about 30 base pairs, about 40 base pairs, about 50 base pairs, approximately 60 base pairs, approximately 70 base pairs, approximately 80 base pairs, approximately 90 base pairs, approximately 100 bases pairs, about 110 base pairs, about 120 base pairs, about 130 base pairs, about 140 base pairs, about 150 base pairs pairs, about 160 base pairs, about 170 base pairs, up to about 10 base pairs, up to about 20 base pairs, up to about 3 0 base pairs, up to approximately 40 base pairs, up to approximately 50 base pairs, up to approximately 60 base pairs, up to approximately 70 base pairs , up to about 80 base pairs, up to about 90 base pairs, up to about 100 base pairs, up to about 110 base pairs, Maximum approximately 120 base pairs, maximum approximately 130 base pairs, maximum approximately 140 base pairs, maximum approximately 150 base pairs, Larger than about 160 base pairs, maximum about 170 base pairs, and also about 70 base pairs, about 75 base pairs, and about 80 base pairs, about 85 base pairs, about 90 base pairs, about 95 base pairs, about 100 base pairs, and about 105 base pairs The multiples may be selected from one or more multiples of the base pair. In some cases, the multiples may be selected from one or more multiples of the base pair, for each occurrence: Independently selected from multiples of 1, 3, 5, 7, 9, 11, 13, 15, 17 and 19. For example, about 90 base pairs is a multiple of 1 for about 90 base pairs, and about 270 base pairs is a multiple of 1 for about 90 base pairs. In some cases, one or more length intervals are about 90 base pairs, about 270 base pairs, about 450 base pairs, about 630 base pairs, or about 810 base pairs obtain.
[0157] The method can isolate or remove one or more length intervals of DNA. Approximately 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 9 3, 94, 95, 96, 97, 98, 99 or 100% can be isolated or removed. DNA of specific length intervals up to 5, 6, 7, 8, 9, 10, 15, 20, 25, 3 0, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 91 , 92, 93, 94, 95, 96, 97, 98, 99 or 100% isolation or removal At least 5, 6, 7, 8, 9, 10, 15, or 16 fragments of DNA in a specific length interval can be , 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100% It can be isolated or removed.
[0158] Exosomes for background population depletion and / or enrichment of populations of interest Yet another example of the enrichment methods provided herein is for a population of interest (e.g., non-host) A biological sample obtained from a host is used to deplete or enrich the sample for nucleic acids. This includes targeting exosomal nucleic acids within exosomes and other extracellular microorganisms. Vesicles are secreted by cells and are present in biological fluids. They are produced by direct extrusion at the plasma membrane. It can be released by buds or through the body route of multivesicular bodies.
[0159] In some cases, nucleic acids from a population of interest (e.g., pathogen, non-human) are present in exosomes. They may be distributed asymmetrically internally or externally. In some cases, the background population Nucleic acids (e.g., human) are asymmetrically distributed inside or outside the exosomes. If nucleic acids from a target population (e.g., non-host, pathogen) are present outside of exosomes, this method The method involves extracting exosomes from a sample to enrich for nucleic acids of a population of interest outside of exosomes. Similarly, background populations (e.g., host or If the (human) nucleic acid is present within an exosome, the method may further comprise the step of: In some cases, this may include removing exosomes to enrich for a population of nucleic acids. , exosomes, or nucleic acids, such as circulating nucleic acids, from a biological sample prior to isolation of the nucleic acids from the sample. Isolate or remove from.
[0160] Nucleic acids of a population of interest (e.g., a microorganism or pathogen) are present inside exosomes In this case, the method involves extracting exosomes to enrich for nucleic acids of a population of interest within the exosomes. The nucleic acid of the population of interest may include capturing leukocytes (e.g., macrophages). When present within the blood vessels or processed by white blood cells (e.g., macrophages), This method involves immunoprecipitating leukocyte-derived exosomes with an antibody, for example. Similarly, background populations (e.g., If the (host or human) nucleic acid is present outside of the exosome, the method comprises: Capture of exosomes from samples to enrich for nucleic acids of populations of interest within and
[0161] In some cases, background population nucleic acids may be present in unrelated regions inside or outside of exosomes. For example, human nucleic acids, such as human cell-free nucleic acids, are distributed symmetrically within or Generally, the majority of human cell-free RNA in plasma is Circulating nucleic acids are found in exosomes, likely due to the abundance of RNases in plasma. Direct extraction of cell-free RNA using the kit results in high-quality RNA due to degradation of cell-free RNA. Exosome isolation may not yield intact mRNA, 18S and 28S RNA. It can produce high-quality RNA, including exosomal RNA. The majority of human cell-free mRNA is exosomal. Therefore, in some cases, the methods provided herein may be used to Depleting exosomes from samples to deplete human cell-free mRNA from the samples In contrast, exosomes generally likely package the cytosol. Therefore, the majority of human cell-free DNA in plasma may not be present in exosomes. Therefore, in some cases, the methods provided herein can be used to extract human cell-free DNA from a sample. This may involve enriching the exosomes in the sample to deplete NAs.
[0162] In some cases, the methods involve using plasma, isolated exosomes and / or exosomes. The method may include determining the ratio of nucleic acids of the population of interest to the background population in the nucleic acid-depleted plasma. For example, the method may involve the use of plasma, isolated exosomes and / or exosome-depleted In some cases, the method may include determining the ratio of microbial to human nucleic acids in the plasma. Microorganisms and human D in plasma, isolated exosomes and / or exosome-depleted plasma In some cases, the method may include determining the ratio of NA to NA. To determine the ratio of microbial to human RNA in exosome- and / or exosome-depleted plasma In some cases, the method may include using plasma, isolated exosomes, and / or determining the amount, percentage, or concentration of nucleic acids of a population of interest in exosome-depleted plasma; This may include:
[0163] The nucleic acids of the population of interest are present in white blood cells (e.g., macrophages) or When processed by leukocytes (e.g., macrophages), the method involves the production of leukocyte-derived exosomes. By preferentially isolating exosomes derived from leukocytes, such as by immunoprecipitation of exosomes with antibodies, Microbial or pathogenic nucleic acids can be present in or processed by white blood cells. When the exosomes are isolated, the method may further comprise immunoprecipitation of the exosomes from the leukocytes, such as by immunoprecipitation of the exosomes with an antibody. Exosomes derived from blood cells can be preferentially isolated. In some cases, the method comprises: This involves enriching exosomes derived from mammalian leukocytes (e.g., macrophages). In some cases, the method involves using a cell derived from a human leukocyte (e.g., a macrophage). In some cases, the method may include enriching the exosomes in host leukocytes (e.g., This may include enriching for exosomes derived from cells (e.g., macrophages). This method is useful for preventing deep tissue infections because bacteria re-enter the bloodstream, whereas most macrophages do not. You can access information about the disease.
[0164] The method can isolate or remove exosomes. ,7,8,9,10,15,20,25,30,35,40,45,50,55,60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 9 8, 99 or 100% of the exosomes can be isolated or removed. ,7,8,9,10,15,20,25,30,35,40,45,50,55,60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 9 At least 8, 99 or 100% of the exosomes can be isolated or removed. 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 9 7, 98, 99 or 100% of the exosomes can be isolated or removed from the plasma. As some non-limiting examples of kits and protocols available for the isolation of The exosomes were purified using the Exo-Spin Blood Exosome Purification Kit (Cell Guidance Syst. stems), exoRNeasy serum / plasma kit (Qiagen), whole exosomes Isolation reagent (Life Technologies), ExoQuick (System Biosciences), Exo-Flow Exosome Immunopurification (System B iosciences), ME kit (New England Peptide), V n96 Peptide (New England Peptide), PureExo RNA isolation kit (101 Bio) and plasma / serum circulating and exosomal RNA purification Kit (Norgen Biotek Corp.) is an example.
[0165] Targeted host DNA depletion In some cases, depletion of host nucleic acids in a cell-free sample may be achieved by cleaving the host nucleic acid sequence, either host or non-host. A specific sequence, sequence structure, or nucleic acid property or modification that can be specific to one of the host nucleic acids. Targeting can be used, for example, to one of the host or non-host nucleic acids. The specific sequence is used as a mechanism for binding, extraction, sedimentation, digestion, or removal or selection. Targeting can be achieved by sequence motifs, structures, base modifications, etc. For example, The nucleic acids may contain base modifications that may not be reflected in non-host or pathogen-associated nucleic acids. Modifications include, for example, methylation patterns, such as cytosine methylation. Leu-cytosine can be found in CpG islands in higher organisms. Some eukaryotic pathogens Although present in the human body, its presence is significantly higher in vertebrates compared to other organisms.
[0166] In some cases, methylation motif-specific endonucleases (e.g., McrB C, FspEI, LpnPI, MspJI endonucleases) to These endonucleases can target modifications of fragments that normally undergo digestion. The separation requires two adjacent base modifications. , e.g., at least 50 bp, at least 60 bp, at least 70 bp, at least 8 0bp, at least 90bp, at least 100bp, at least 200bp, at least At least 300bp, at least 400bp, at least 500bp, at least 600bp, At least 700bp, at least 800bp, at least 900bp, at least 1k bp, at least 1.5kbp, at least 2kbp, at least 2.5kbp or less Most are 147-175 bp long, which means that this is a human cDNA. This can result in poor digestion efficiency of fDNA fragments during sequencing library preparation. Partner base modification to one of the two adapters provided in the ligation step This can be improved by introducing a modification into the original human cfDNA fragment. Even if only one methylcytosine is present, the P5 sequence-bearing adapter can be used for ligation. The inclusion of methyl-cytosines in CpG islands stimulates McrBC digestion after transcription. Thus, in some cases, "helper" modified bases can be added to the sequence library. By including methylated bases in the adapter sequence used in preparing the sequence Therefore, host-derived library elements can be artificially introduced into the library. allowing non-host-derived library elements to proceed with amplification and sequencing while digesting To avoid this, a digestion step can be used after adapter ligation to the sequence fragments. It will be possible to do this.
[0167] Alternatively or additionally, methods may be used to selectively or selectively isolate sequences longer than this required separation distance. or may require spacing between two or more such sites for preferential digestion. For example, a pair of recognition sequences separated by 50 bases or more (e.g., 50 McrBC endonuclease capable of recognizing two methylated cytosines (from 100 bp to approximately 3 kbp) In this case, methylated bases can be provided to the adapter sequences attached to either end. As a result, fragments that maintain more than 50 bases between the methylated bases on opposing adapters was digested with McrBC endonuclease and purified as described elsewhere herein. As shown in Figure 1, enrichment of samples for short nucleic acid fragments results in greater enrichment of non-host fragments. I think so.
[0168] Combined enrichment method Although the enrichment schemes are described individually, any or all of the above methods may be used in conjunction with Various combinations can be used to enrich non-host DNA in samples relative to the primary DNA. It will be appreciated that a cell-free sample can be first subjecting the DNA to a size-selection-based DNA purification scheme, for example using the SPRI bead system; Subsequent chromatography size selection schemes, nucleosome immunoprecipitation schemes, etc. Any one or more of the following may be used:
[0169] Again, as noted above, in some cases, single or multiple steps of transfection of non-host DNA Enrichment increases the ratio of non-host DNA to host DNA in the sample by approximately 2-fold to approximately 10,000-fold. In some cases, the ratio can be increased by at least two times, at least three times, or at least At least four times, at least five times, at least six times, at least seven times, at least eight times, at least 9 times, at least 10 times, at least 11 times, at least 12 times, at least 13 times, At least 14 times, at least 15 times, at least 16 times, at least 17 times, at least 18 times, at least 19 times, at least 20 times, at least 30 times, at least 40 times , at least 50 times, at least 60 times, at least 70 times, at least 80 times, less at least 90 times, at least 100 times, at least 1000 times, at least 5000 times, or In some cases, an increase of at least 10,000 times can be achieved.
[0170] Molecular barcoding of samples by nucleic acid spike-in Samples can be barcoded with a nucleic acid barcode spike-in. The nucleic acid barcode may be one or more lengths. In some cases, the nucleic acid barcode may be oligonucleotides, double-stranded longmers, PCR products and / or primers In some cases, the nucleic acid barcode may be DNA, RNA, PNA, LNA, or A, BNA, or any combination thereof. In some cases, nucleic acid barcodes The genome is based on one or more background populations, e.g., the human genome or a pathogen genome. In some cases, a nucleic acid barcode may contain sequences that are not present in one or more pairs. It may contain sequences that are not present in the target population.
[0171] In some cases, the barcode is different from the label on the outside of the sample tube and In some cases, the barcode may be present at any time, including before the addition of the biological sample. It can be added at any stage (e.g., by PCR or sequencing). In some cases, barcodes are used to track samples, Cross-contamination can be detected and / or reagents can be tracked. Nucleic acid barcodes can be used to reduce sample mix-ups. The total nucleic acid concentration can be increased using a chromatograph to increase recovery of low-concentration samples. In some cases, bar codes are used to compare known inputs with measured outputs. or providing a reference standard (e.g., normalizing oligonucleotide) for comparison or normalization. In some cases, the sample input can be estimated by providing a barcode. Using the library complexity, sample loss, sensitivity, and / or size bias By measuring performance, it is possible to improve performance through quality control and development.
[0172] In some cases, a sample (e.g., a biological sample or a nucleic acid sample) is In some cases, the normalization oligonucleotide may be spiked with a Retides can be used to monitor the efficiency of DNA manipulation, purification and / or amplification steps. For example, sequencing reads from a population of interest (e.g., a non-host or pathogen) The absolute number of samples was compared to the absolute number of normalized oligonucleotide reads recovered. It is possible to normalize for differences in molecular engineering efficiency between the two. oligonucleotides for barcoding purposes, or for both normalization and barcoding purposes can be used for
[0173] Strategic Capture of Target Areas The oligonucleotide collections and methods provided herein are directed to targeting a region of interest. An oligonucleotide containing a nucleic acid sequence identical, nearly identical, complementary, or nearly complementary to a region of interest. Some non-limiting examples of regions of interest include pathogenicity genes locus (e.g., Clostridium difficile ile); antimicrobial resistance marker; antibiotic resistance marker; antiviral antimicrobial drug resistance markers; antiparasitic drug resistance markers; informative genotyping regions (e.g., host of human, microbial, pathogenic, bacterial, viral, fungal or parasitic origin); two or more microorganisms Sequences common to organisms, pathogens, bacteria, viruses, fungi and / or parasites; host genomes Non-host sequences integrated into the host; masking non-host sequences; non-host mimicking sequences; masking host host-mimicking sequences; and one or more microorganisms, pathogens, bacteria, viruses, fungi In some cases, the region of interest may be a sequence specific to a non-host organism. They can be present in the genomes of bacteria, viruses, pathogens, fungi or microorganisms. The region of interest may be present in a host, mammalian or human genome. genotypic identity by containing a nucleic acid sequence identical, nearly identical, complementary, or nearly complementary to determination, antimicrobial resistance detection, antibiotic resistance detection, antiviral resistance detection, antiparasitic resistance This may allow for increased sensitivity of detection and / or pathogen detection. Identical, complementary, or nearly complementary nucleic acid sequences can be identified, e.g., bioinformatically or computationally. The nucleic acid sequence can be synthesized chemically by DNA synthesis and / or by design. for priming cDNA synthesis from RNA templates for nucleic acid amplification or detection, for sequencing, For the determination of DNA / RNA hybridization in the primer extension reaction, and / or can be used as primers for DNA / RNA pull-down. Cut.
[0174] An antibiotic resistance marker is a mutation or mutations that confer antibiotic resistance. Antibiotic resistance markers may include, but are not limited to, polymorphisms, genes, or gene products. However, antibiotic efflux, antibiotic inactivation, antibiotic target modification, antibiotic target protection, antibiotic one or more functions, such as substance target substitution and / or reduced permeability to antibiotics In some cases, antibiotic resistance markers can confer antibiotic resistance through cross-linking. C. difficile (Carbapenem-Resistant Enterobacteriaceae) , drug-resistant Neisseria gonorrhoeae (cephalosporin-resistant), multidrug-resistant Acinetobacter, drug-resistant Campylobacter Candida albicans, fluconazole-resistant fungi, extended-spectrum beta-lactamase-producing Enterobacteriaceae (ESBL), vancomycin-resistant enterococci (VRE), multidrug-resistant Pseudomonas aeruginosa, drugs Drug-resistant non-typhoidal Salmonella, drug-resistant Salmonella typhi, drug-resistant Shigella, methicillin-resistant yellow b Staphylococcus aureus (MRSA), drug-resistant pneumococcus, drug-resistant tuberculosis (MDR and XDR), multidrug Resistant Staphylococcus aureus, vancomycin-resistant Staphylococcus aureus (VRSA), Erythromycin Clindamycin-resistant Streptococcus group A or clindamycin-resistant Streptococcus group B In some cases, antibiotic resistance markers are acridine dyes, aminocoumaric acid dyes, Antibiotics, aminoglycosides, aminonucleoside antibiotics, β-lactams, diamino Pyrimidines, elfamycin, fluoroquinolones, glycopeptide antibiotics, lincosamine Lipopeptide antibiotics, macrocyclic antibiotics, macrolides, nucleoside antibiotics, Organic arsenic antibiotics, oxazolidinone antibiotics, peptide antibiotics, phenicol, pre- Lomutilin antibiotics, polyamine antibiotics, rifamycin antibiotics, streptogramin tetracycline derivatives, sulfonamides, sulfones, tetracycline derivatives, or any of these In some cases, the combination of antibiotic resistance markers may confer antibiotic resistance. These include β-lactams, penicillins, aminopenicillins, early generation cephalosporins, β-lactams, Combination of cephalosporins, broad-spectrum cephalosporins, carbapenems, fluoxetine, and fluoxetine. Oroquinolones, aminoglycosides, tetracyclines, glycine, polymyxins, Nicorhizoline, methicillin, erythromycin, gentamicin, ceftazidime, vancomycin levofloxacin, imipenem, linezolid, ceftriaxone, ceftaroline, or or any combination thereof.
[0175] Non-limiting examples of antibiotic resistance markers include aac2ia, aac2ib, aac 2ic, aac2id, aac2i, aac3ia, aac3iia, aac3iib, aac3iii, aac3iv, aac3ix, aac3vi, aac3viii, aa c3vii, aac3x, aac6i, aac6ia, aac6ib, aac6ic, a ac6ie, aac6if, aac6ig, aac6iia, aac6iib, aad9 , aad9ib, aadd, acra, acrb, adea, adeb, adec, am ra, amrb, ant2ia, ant2ib, ant3ia, ant4iia, ant 6ia, aph33ia, aph33ib, aph3ia, aph3ib, aph3ic , aph3iiia, aph3iva, aph3va, aph3vb, aph3via, aph3viia, aph4ib, aph6ia, aph6ib, aph6ic, aph 6id, arna, baca, bcra, bcrc, bl1_acc, bl1_ampc , bl1_asba, bl1_ceps, bl1_cmy2, bl1_ec, bl1_f ox, bl1_mox, bl1_och, bl1_pao, bl1_pse, bl1_s m, bl2a_1, bl2a_exo, bl2a_iii2, bl2a_iii, bl2 a_kcc, bl2a_nps, bl2a_okp, bl2a_pc, bl2be_ct xm, bl2be_oxy1, bl2be_per, bl2be_shv2, bl2b_ rob、bl2b_tem1、bl2b_tem2、bl2b_tem、bl2b_tl e、bl2b_ula、bl2c_bro、bl2c_pse1、bl2c_pse3、 bl2d_lcr1、bl2d_moxa、bl2d_oxa10、bl2d_oxa1 、bl2d_oxa2、bl2d_oxa5、bl2d_oxa9、bl2d_r39、 bl2e_cbla、bl2e_cepa、bl2e_cfxa、bl2e_fpm、b l2e_y56、bl2f_nmca、bl2f_sme1、bl2_ges、bl2_ kpc、bl2_len、bl2_veb、bl3_ccra、bl3_cit、bl3 _cpha、bl3_gim、bl3_imp、bl3_l、bl3_shw、bl3_ sim、bl3_vim、ble、blt、bmr、cara、cata10、cata 11、cata12、cat13、cat14、cat15、cat16、ca ta1、cata2、cata3、cata4、cat5、cat6、cata7、 cata8、cata9、catb1、catb2、catb3、catb4、catb 5、ceoa、ceob、cml_e1、cml_e2、cml_e3、cml_e4、 cml_e5、cml_e6、cml_e7、cml_e8、dfra10、dfra1 2、dfra13、dfra14、dfra15、dfra16、dfra17、dfr a19、dfra1、dfra20、dfra21、dfra22、dfra23、df ra24, dfra25, dfra25, dfra25, dfra26, dfra5, d fra7, dfrb1, dfrb2, dfrb3, dfrb6, emea, emrd, e mre、erea、ereb、erma、ermb、ermc、ermd、erme、e rmf、ermg、ermh、ermn、ermo、ermq、ermr、erms、e rmt、ermu、ermv、ermw、ermx、ermy、fosa、fosb、f osc、fosx、fusb、fush、ksga、lmra、lmrb、lnua、l nub, lsa, maca, macb, mdte, mdtf, mdtg, mdth, md tk, mdtl, mdtm, mdtn, mdto, mdtp, meca, mecr1, m efa、mepa、mexa、mexb、mexc、mexd、mexe、mexf、m exh、mexi、mexw、mexx、mexy、mfpa、mpha、mphb、m phc、msra、norm、oleb、opcm、opra、oprd、oprj、o prm、oprn、otra、otrb、pbp1a、pbp1b、pbp2b、pbp 2、pbp2x、pmra、qac、qaca、qacb、qnra、qnrb、qnr s、rosa、rosb、smea、smeb、smec、smed、smee、sme f、srmb、sta、str、sul1、sul2、sul3、tcma、tcr3、 tet30、tet31、tet32、tet33、tet34、tet36、tet3 7、tet38、tet39、tet40、teta、tetb、tetc、tetd、 tete、tetg、teth、tetj、tetk、tetl、tetm、teto、 tetpa、tetpb、tet、tetq、tets、tett、tetu、tetv 、tetw、tetx、tety、tetz、tlrc、tmrb、tolc、tsnr 、vana、vanb、vanc、vand、vane、vang、vanha、van hb, vanhd, vanra, vanrb, vanrc, vanrd, vanre, v anrg, vansa, vansb, vansc, vansd, vanse, vansg , vant, vante, vang, vanug, vanwb, vanwg, vanx a, vanxb, vanxd, vanxyc, vanxye, vanxyg, vanya ,vanyb,vanyd,vanyg,vanz,vata,vatb,vatc,v atd, vate, vgaa, vgab, vgba, vgbb, vph, ykkc and YKKD is one example.
[0176] concentrated The methods described herein can be used to enrich nucleic acids in a population of interest. In some cases, the nucleic acids of the population of interest are reduced to approximately 5%; 10%; 15%; 20%; 25%; 30% ;35%;40%;45%;50%;55%;60%;65%;70%;75%;80% ;85%;90%;95%;100%;150%;200%;250%;300%;35 0%;400%;450%;500%;550%;600%;650%;700%;75 0%;800%;850%;900%;950%;1000%;2000%;3000% ;4000%;5000%;6000%;7000%;8000%;9000%;100 00%;20000%;30000%;40000%;50000%;60000%;7 0000%;80000%;90000%;100000%;200000%;3000 00%;400000%;500000%;600000%;700000%;8000 00%;900000%;1000000%;2000000%;3000000%;4 000000%;5000000%;6000000%;7000000%;80000 00%;9000000%;10000000%;20000000%;3000000 0%;40000000%;50000000%;60000000%;7000000 0%; 80,000,000%; 90,000,000%; or 100,000,000% concentrated In some cases, nucleic acids in a population of interest can be reduced by up to 5%; 10%; 15%; 2%; 0%;25%;30%;35%;40%;45%;50%;55%;60%;65%;7 0%;75%;80%;85%;90%;95%;100%;150%;200%;25 0%;300%;350%;400%;450%;500%;550%;600%;65 0%;700%;750%;800%;850%;900%;950%;1000%;2 000%;3000%;4000%;5000%;6000%;7000%;8000% ;9000%;10000%;20000%;30000%;40000%;50000 %;60000%;70000%;80000%;90000%;100000%;20 0000%;300000%;400000%;500000%;600000%;70 0000%;800000%;900000%;1000000%;2000000%; 3,000,000%;4,000,000%;5,000,000%;6,000,000%;7,000 000%;8000000%;9000000%;10000000%;2000000 0%;30000000%;40000000%;50000000%;6000000 0%; 70,000,000%; 80,000,000%; 90,000,000%; or 1000 In some cases, the nucleic acids of a population of interest can be enriched by at least 100%. Also 5%;10%;15%;20%;25%;30%;35%;40%;45%;50%; 55%;60%;65%;70%;75%;80%;85%;90%;95%;100% ;150%;200%;250%;300%;350%;400%;450%;500% ;550%;600%;650%;700%;750%;800%;850%;900% ;950%;1000%;2000%;3000%;4000%;5000%;6000 %;7000%;8000%;9000%;10000%;20000%;30000% ;40000%;50000%;60000%;70000%;80000%;9000 0%;100000%;200000%;300000%;400000%;50000 0%;600000%;700000%;800000%;900000%;10000 00%;2000000%;3000000%;4000000%;5000000%; 6,000,000%;7,000,000%;8,000,000%;9,000,000%;1,000 0000%;20000000%;30000000%;40000000%;5000 0000%;60000000%;70000000%;80000000%;9000 For example, a sample may be enriched by: Contains nucleic acids of the target population of 5% of the total population of nucleic acids, and 10% of the total population of nucleic acids % of the nucleic acid of the population of interest, the nucleic acid of the population of interest is 00% enrichment. In some cases, the nucleic acids of the population of interest are enriched by approximately 1.5-fold; 2-fold; 2-fold; .5x;3x;3.5x;4x;4.5x;5x;5.5x;6x;6.5x;7x;7 .5x;8x;8.5x;9x;9.5x;10x;15x;20x;25x;30x; 35x; 40x; 45x; 50x; 55x; 60x; 65x; 70x; 75x; 80x; 85x;90x;95x;100x;150x;200x;250x;300x;350 times; 400 times; 450 times; 500 times; 550 times; 600 times; 650 times; 700 times; 750 times; 800 times; 850 times; 900 times; 950 times; 1000 times; 2000 times; 3000 times; 4000x; 5000x; 6000x; 7000x; 8000x; 9000x; 1000 0x;20000x;30000x;40000x;50000x;60000x;70 000x;80000x;90000x;100000x;200000x;30000 0x;400000x;500000x;600000x;700000x;80000 It can be concentrated 0x; 900,000x; 1,000,000x. In some cases, Increase the nucleic acid content of the population by up to 1.5x; 2x; 2.5x; 3x; 3.5x; 4x; 4.5x; 5x; 5.5x; 6x; 6.5x; 7x; 7.5x; 8x; 8.5x; 9x; 9.5x; 10x;15x;20x;25x;30x;35x;40x;45x;50x;55x; 60x;65x;70x;75x;80x;85x;90x;95x;100x;150 times;200 times;250 times;300 times;350 times;400 times;450 times;500 times;550 times; 600 times; 650 times; 700 times; 750 times; 800 times; 850 times; 900 times; 950 times;1000 times;2000 times;3000 times;4000 times;5000 times;6000 times;70 00x;8000x;9000x;10000x;20000x;30000x;400 00x;50000x;60000x;70000x;80000x;90000x;1 00000x;200000x;300000x;400000x;500000x;6 00000x; 700000x; 800000x; 900000x; 1000000x dark In some cases, the nucleic acids of the population of interest can be reduced by at least 1.5-fold; times;2.5 times;3 times;3.5 times;4 times;4.5 times;5 times;5.5 times;6 times;6.5 times;7 times;7.5 times;8 times;8.5 times;9 times;9.5 times;10 times;15 times;20 times;25 times;3 0x;35x;40x;45x;50x;55x;60x;65x;70x;75x;8 0x; 85x; 90x; 95x; 100x; 150x; 200x; 250x; 300x; 350x; 400x; 450x; 500x; 550x; 600x; 650x; 700x; 750x; 800x; 850x; 900x; 950x; 1000x; 2000x; 300 0x;4000x;5000x;6000x;7000x;8000x;9000x;1 0000x; 20000x; 30000x; 40000x; 50000x; 60000x ;70000x;80000x;90000x;100000x;200000x;30 0000x;400000x;500000x;600000x;700000x;80 It can be concentrated 0,000 times; 900,000 times; 1,000,000 times. For example, The sample contains 5% of the nucleic acids of the target population of nucleic acids, and If the target population is enriched to contain 10% of the target population's nucleic acid, The acid is twice as concentrated.
[0177] The methods described herein allow for the depletion of background populations. In some cases, the background population nucleic acid is about 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 9 0, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100% depletion In some cases, the background population nucleic acid can be up to 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 7 5, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or In some cases, background population nucleic acids can be reduced. Tomo 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 5 5, 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96 For example, if the sample is depleted of nucleic acids, The background population contains 50% of the total population of nucleic acids and 10% of the total population of nucleic acids. % background population is depleted to contain nucleic acids. The population nucleic acid is depleted by 80%. In some cases, the background population nucleic acid is depleted by approximately 1.5 times. ;2x;2.5x;3x;3.5x;4x;4.5x;5x;5.5x;6x;6.5x ;7x;7.5x;8x;8.5x;9x;9.5x;10x;15x;20x;25x ;30x;35x;40x;45x;50x;55x;60x;65x;70x;75x ;80x;85x;90x;95x;100x;150x;200x;250x;300 times; 350 times; 400 times; 450 times; 500 times; 550 times; 600 times; 650 times; 700 times; 750 times; 800 times; 850 times; 900 times; 950 times; 1000 times; 2000 times; 3 000x; 4000x; 5000x; 6000x; 7000x; 8000x; 9000x ;10000x;20000x;30000x;40000x;50000x;6000 0x; 70000x; 80000x; 90000x; 100000x; 200000x; 300000x; 400000x; 500000x; 600000x; 700000x; It can be concentrated 800,000 times; 900,000 times; or 1,000,000 times. In this case, increase the background population nucleic acid by up to 1.5x; 2x; 2.5x; 3x; 3.5x; 4x. times;4.5 times;5 times;5.5 times;6 times;6.5 times;7 times;7.5 times;8 times;8.5 times;9 times;9.5 times;10 times;15 times;20 times;25 times;30 times;35 times;40 times;45 times;5 0x;55x;60x;65x;70x;75x;80x;85x;90x;95x;1 00x;150x;200x;250x;300x;350x;400x;450x;5 00x;550x;600x;650x;700x;750x;800x;850x;9 00x;950x;1000x;2000x;3000x;4000x;5000x;6 000x;7000x;8000x;9000x;10000x;20000x;300 00x;40000x;50000x;60000x;70000x;80000x;9 0000x;100000x;200000x;300000x;400000x;50 0000x;600000x;700000x;800000x;900000x;10 In some cases, background population nucleic acids can be concentrated 100,000 times. At least 1.5 times; 2 times; 2.5 times; 3 times; 3.5 times; 4 times; 4.5 times; 5 times; 5.5 times; 6x; 6.5x; 7x; 7.5x; 8x; 8.5x; 9x; 9.5x; 10x; 15x; 20x; 25x; 30x; 35x; 40x; 45x; 50x; 55x; 60x; 65x; 70x;75x;80x;85x;90x;95x;100x;150x;200x;2 50x;300x;350x;400x;450x;500x;550x;600x;6 50x; 700x; 750x; 800x; 850x; 900x; 950x; 1000x; 2000x; 3000x; 4000x; 5000x; 6000x; 7000x; 8000 times;9000 times;10000 times;20000 times;30000 times;40000 times;5000 0x;60000x;70000x;80000x;90000x;100000x;2 00000x;300000x;400000x;500000x;600000x;7 Can be concentrated 00000 times; 800000 times; 900000 times; 1,000,000 times For example, if a sample contains 50% background population nucleic acids of the total population of nucleic acids, The total nucleic acid population is depleted to contain 10% background nucleic acid. If present, the background population nucleic acid is depleted five-fold.
[0178] sample The methods and compositions provided herein may be used to detect multiple antigens obtained from a subject (e.g., a human host). It is useful for detecting nucleic acids in a variety of samples. Non-limiting examples include blood, plasma, serum, whole blood, mucus, saliva, cerebrospinal fluid, synovial fluid, and lavage fluid. Samples include circulating, urine, tissue biopsies, cell samples, skin samples and stool. Nucleic acids, such as circulating nucleic acids, including cell-free nucleic acids (e.g., circulating cell-free DNA, circulating cell-free RNA). In some cases, the sample obtained from the subject may be further processed. For example, the sample can be processed to extract DNA or RNA, which is The results can be analyzed using the methods provided in the specification.
[0179] In some cases, the nucleic acid sample is a biological sample, an isolated nucleic acid sample, or a circulating nucleic acid sample. The sample may be a purified nucleic acid sample containing nucleic acid, such as a nucleic acid. The nucleic acid sample may contain host and non-host sequences, e.g., human and non-human sequences. In some cases, nucleic acids in a nucleic acid sample, such as a sequenceable library of nucleic acids, are In some cases, the nucleic acids in the nucleic acid sample are amplified (e.g., by a PCR amplification reaction). Artificially fragmented (e.g., by sonication, shearing, enzymatic digestion, or chemical fragmentation) In some cases, nucleic acid fragmentation is not necessary due to the relatively short length of the nucleic acid. In some cases, the nucleic acids in the nucleic acid sample are artificially fragmented. The circulating nucleic acid sample is a biological sample, an isolated nucleic acid sample, or a sample containing circulating nucleic acid, such as circulating cell-free nucleic acid The sample may be a purified nucleic acid sample, or in some cases a circulating cell-free nucleic acid sample. is a biological sample, an isolated nucleic acid sample, or circulating cell-free DNA or circulating cell-free The sample may be a purified nucleic acid sample containing circulating cell-free nucleic acids such as RNA. In some cases, the single-stranded nucleic acid sample may be a biological sample, an isolated nucleic acid sample, or a single-stranded nucleic acid sample. A purified nucleic acid sample containing single-stranded nucleic acids, such as double-stranded circulating nucleic acids or single-stranded circulating cell-free nucleic acids. The sample may be a sample of a
[0180] In some cases, the nucleic acid is DNA, RNA, cDNA, mRNA, cRNA, or dsDNA , ssDNA, miRNA, circulating nucleic acid, circulating DNA, circulating RNA, cell-free nucleic acid, cell-free D NA, cell-free RNA, circulating cell-free DNA, circulating cell-free RNA, or genomic DNA. In some cases, the circulating nucleic acid is circulating DNA, circulating RNA, cell-free nucleic acid, or cell-free circulating nucleic acid. In some cases, cell-free nucleic acids are cell-free DNA, cell-free RNA, circulating cell-free DNA, In some cases, the circulating cell-free nucleic acid may be circulating cell-free nucleic acid or circulating cell-free RNA. It may be DNA or circulating cell-free RNA.
[0181] In some cases, the nucleic acids in the sample may be unlabeled; in some cases, the nucleic acids may be unlabeled. The nucleic acid is labeled, for example, with a nucleic acid label, a chemical label, or an optical label. In some cases, the label is conjugated to a solid support. The nucleic acid may be attached to the end or to the interior of the nucleic acid. In some cases, the nucleic acid may be attached to two or more Label with signs.
[0182] Nucleic acids in a sample may be tagged with a nucleic acid label. In some cases, the nucleic acid label is may include one or more of the following: barcodes (e.g., sample barcodes), universal Ape primer sequences, primer binding sites (e.g., but not limited to, DNA sequences Determination primer binding site, sample barcode sequencing primer binding site and various Sequencing or for barcode reading), sequencer-compatible sequences, sequencing platforms The nucleic acid label may be attached to a sequence, a sequencing adapter sequence, or an adapter. The nucleic acid can be attached to the nucleic acid (by gating or synthetic design).
[0183] Nucleic acid labels can include chemical labels. Some non-limiting examples of chemical labels include: , biotin, avidin, streptavidin, radiolabels, polypeptides and polymers Nucleic acid labels can include optical labels. Some non-limiting examples of optical labels include: Examples include fluorophores, fluorescent proteins, dyes and quantum dots. The antibody can be conjugated to a solid support. Some non-limiting examples of solid supports include: Examples include beads, magnetic beads, polymers, slides, chips, surfaces, plates, chips, etc. These include channels, cartridges, microfluidic devices and microarrays. can be conjugated to a solid support for affinity chromatography. In some cases, each nucleic acid has a different label. In some cases, each nucleic acid has the same label. In some cases, the nucleic acids in the sample must be conjugated to a solid support. stomach.
[0184] Sequencing methods In some cases, the enriched population of nucleic acids is sequenced. The described methods further comprise performing a sequencing assay. The nucleic acids of the captured population of interest are then subjected to sequencing assays, particularly high-throughput sequencing. Sequencing assays, next-generation sequencing platforms, massively parallel sequencing platforms Platform, nanopore sequencing assay, Sanger sequencing or the technology By performing other sequencing assays known in the art, such as by performing other sequencing assays known in the art. Some non-limiting examples of sequencing assays include: Examples include high-throughput sequencing assays, next-generation sequencing platforms, Home, Massively Parallel Sequencing Platforms, Nanopore Sequencing and Sangha Some non-limiting examples of types of sequencing machines include: , Illumina, Roche 454, Ion Torrent and Nanopo Examples include re.
[0185] The process methods described herein also distinguish between, for example, host and non-host sequences. and a method for analyzing the host nuclei, including an informatics data filtering method applied to the obtained sequence data, This can be used in conjunction with other methods to distinguish between host and non-host nucleic acid sequences. Examples of such processes are those disclosed herein, the entire disclosure of which is incorporated herein by reference in its entirety for all purposes. Published U.S. Patent Application No. 2015-0133391, which is incorporated herein by reference. These include those described in the specification.
[0186] Purpose The methods and compositions can be used to treat one or more non-host species (e.g., microorganisms, pathogens, bacteria, for analyzing biological samples from hosts infected with viruses, fungi, or parasites The methods and compositions provided herein are useful in treating diseases or disorders, particularly those caused by microbial infections. or specific to the detection, prediction, diagnosis or monitoring of a disease or disorder caused by a pathogen. The methods and compositions provided herein are also useful for detecting circulating cell-free DNA or Certain types of samples obtained from infected hosts, such as samples containing circulating cell-free RNA, This is useful for detecting populations of interest within a sample.
[0187] The methods and compositions can also be used to enable genotyping of pathogens in a sample. The methods and compositions are also particularly useful for detecting alterations in the genomes of microorganisms or pathogens. For example, they can be used to detect antibiotic-resistant strains of bacteria or to identify pathogens. Changes that affect the pathogenicity of agents, particularly viruses and bacteria, can be tracked.
[0188] In other cases, the methods and compositions can be used to identify one or more of the following in a microbiome: The presence of multiple microorganisms can be monitored. For example, the present methods and compositions can be used to: It is possible to monitor the presence of the microbiome in a healthy or uninfected host. The present methods and compositions can be used to identify microbial cycles in a sample containing multiple microorganisms. You can also monitor the
[0189] The methods and compositions provided herein also include methods for the production of microbial vectors in one or more non-host organisms (e.g., , microorganisms, pathogens, bacteria, viruses, fungi or parasites) It may also enable the detection or hypothesis-free diagnosis or monitoring of a disease, disorder or infection. Thus, the methods provided herein can be used to target specific targets (e.g., one or more screening for specific nucleic acid sequences, proteins, or antibodies) and testing of selected targets In some cases, the methods provided herein may differ from other diagnostic methods that are limited to: and the hypothesis-free properties of the composition may be useful for detecting rare infectious diseases, synchronizing the detection of two or more diseases or disorders, Detecting outbreaks, identifying sources of infection, or identifying multiple diseases with similar or common symptoms or may facilitate identification of the fault.
[0190] In some embodiments, the method comprises one or more of the following steps in any order or combination: In combination, it may include: (a) a subject or patient having or suspected of having a pathogen infection (b) providing a nucleic acid sample from a subject; and (b) comparing the nucleic acid sample with a nucleic acid sample prepared according to any of the methods provided herein. A collection of oligonucleotides, selectively enriched to specifically bind to non-host sequences. (c) contacting the sample and the oligonucleotides with a collection of oligonucleotides; A collection of oligonucleotides and a nucleic acid sample (d) subjecting the oligonucleotide to conditions that promote hybridization with the nucleic acid molecule; Amplification assays, pull-down assays, and hybridization assays are performed on nucleic acids hybridized to a collection of nucleotides. (e) performing an assay, such as a sequencing assay or a sequencing assay; and Analyzing the sequence of the hybridized nucleic acid to detect the specific pathogen in the sample Step 1: Typically, the number of specific pathogens detected ranges from about 10, 20, 30, 40, 50, or more. or more different pathogens, as further described herein, are numerous. Typically, pathogen detection is used to identify infectious diseases, infectious disorders, infections, or other diseases or disorders ( For example, it may enable the detection, prognosis, monitoring, or diagnosis of cancer. In such cases, such detection may facilitate staging of a disease or disorder or may be useful in detecting a pathogenic infection. It can provide an indication of the degree of
[0191] In some cases, the methods described herein involve determining the distribution of at least one subject population. further comprising detecting sequences (e.g., pathogenic or other non-host or non-human sequences) In some cases, the methods described herein involve the analysis of at least five populations of interest. In some cases, the methods described herein further include determining a sequence. In some cases, determining at least five non-mammalian sequences is further included. The described method comprises culturing at least one non-mammalian mammal from each of at least five non-mammalian species. In some cases, the methods described herein further comprise detecting at least one sequence. In some cases, determining at least five non-human sequences may be performed. The method comprises obtaining at least one non-human sequence from each of at least five non-human species. In some cases, the methods described herein further comprise detecting at least Further comprising detecting one bacterial sequence and at least one viral sequence. In some cases, the methods described herein involve assaying at two or more time points, e.g., 2, 3, 4, 5, 6, 7, 8, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 2 5, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90 In some cases, the method further comprises taking 95 or 100 time points. The methods described further include taking time points before and after treatment, e.g., antimicrobial agents.
[0192] In some cases, the methods described herein may be performed in a manner that is about 1; 2; 3; 4; 5; 6; 7; 8; 9;10;15;20;25;30;35;40;45;50;60;70;80;90 ;100;200;300;400;500;1000;5000;10000;500 or 100,000 species. In some cases, the methods described herein may include up to 1; 2; 3; 4; 5; 6; 7; 8;9;10;15;20;25;30;35;40;45;50;60;70;80; 90;100;200;300;400;500;1000;5000;10000;5 or 100,000 species. In some cases, the methods described herein include at least 1; 2; 3; 4; 5 ;6;7;8;9;10;15;20;25;30;35;40;45;50;60;7 0;80;90;100;200;300;400;500;1000;5000;10 Detect at least one nucleic acid from 000; 50,000; or 100,000 species In some cases, the species is a non-host species. In some cases, the species is a non-mammalian species. In some cases, the species is a non-human species. In some cases, the species is a microbial species. , bacteria, viruses, fungi, retroviruses, pathogens or parasites. In this case, the methods described herein may be used to , 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 20 At least 0, 300, 400, 500 or 1000 bacterial or viral species In some cases, the methods described herein further comprise detecting one or more nucleic acids. , up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 4 0, 45, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500 or 1,000 bacterial or viral species. In some cases, the methods described herein include at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500 or 1000 bacteria or In some cases, the method further comprises detecting at least one nucleic acid derived from the viral species. In this case, the methods described herein may be used to , 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 20 At least 0, 300, 400, 500 or 1000 bacterial and viral species In some cases, the methods described herein further comprise detecting one or more nucleic acids. , up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 4 0, 45, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500 or 1,000 bacterial and viral species. In some cases, the methods described herein include at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500 or 1000 bacteria and In some cases, the method further comprises detecting at least one nucleic acid derived from the viral species. In this case, the methods described herein may involve approximately one bacterial species and one viral species, two bacterial species and two viral species. Virus species, 3 bacterial species and 3 viral species, 4 bacterial species and 4 viral species, 5 bacterial species and 5 virus species, 6 bacterial species and 6 virus species, 7 bacterial species and 7 virus species, 8 bacterial species and and 8 viral species, 9 bacterial species and 9 viral species, 10 bacterial species and 10 viral species, 1 5 bacterial species and 15 viral species, 20 bacterial species and 20 viral species, 25 bacterial species and 2 5 viral species, 30 bacterial species and 30 viral species, 35 bacterial species and 35 viral species, 4 0 bacterial species and 40 viral species, 45 bacterial species and 45 viral species, 50 bacterial species and 5 0 viral species, 60 bacterial species and 60 viral species, 70 bacterial species and 70 viral species, 8 0 bacterial species and 80 viral species, 90 bacterial species and 90 viral species, 100 bacterial species and 100 virus species, 200 bacterial species and 200 virus species, 300 bacterial species and 300 400 bacterial and 400 viral species, 500 bacterial and 500 viral species or at least one nucleic acid from 1000 bacterial species and 1000 viral species In some cases, the methods described herein further comprise extracting at most one bacterial species. and 1 virus species, 2 bacteria species and 2 viruses, 3 bacteria species and 3 viruses, 4 cells Bacterial species and 4 viral species, 5 bacterial species and 5 viral species, 6 bacterial species and 6 viral species, 7 bacterial and 7 viral species, 8 bacterial and 8 viral species, 9 bacterial and 9 viral species species, 10 bacterial species and 10 viral species, 15 bacterial species and 15 viral species, 20 bacterial species and and 20 viral species, 25 bacterial species and 25 viral species, 30 bacterial species and 30 viral species species, 35 bacterial species and 35 viral species, 40 bacterial species and 40 viral species, 45 bacterial species and and 45 viral species, 50 bacterial species and 50 viral species, 60 bacterial species and 60 viral species species, 70 bacterial species and 70 viral species, 80 bacterial species and 80 viral species, 90 bacterial species and and 90 viral species, 100 bacterial species and 100 viral species, 200 bacterial species and 200 300 viral species, 300 bacterial species and 300 viral species, 400 bacterial species and 400 viruses species, 500 bacterial species and 500 viral species, or 1000 bacterial species and 1000 viral species In some cases, the method further comprises detecting at least one nucleic acid from the species. The methods described herein involve the use of at least one bacterial species and one viral species, two bacterial species and two viral species. Virus species, 3 bacterial species and 3 viral species, 4 bacterial species and 4 viral species, 5 bacterial species and and 5 viral species, 6 bacterial species and 6 viral species, 7 bacterial species and 7 viral species, and 8 bacterial species and 8 viral species, 9 bacterial species and 9 viral species, 10 bacterial species and 10 viral species, 15 bacterial species and 15 viral species, 20 bacterial species and 20 viral species, 25 bacterial species and 25 viral species, 30 bacterial species and 30 viral species, 35 bacterial species and 35 viral species, 40 bacterial species and 40 viral species, 45 bacterial species and 45 viral species, 50 bacterial species and 50 virus species, 60 bacterial species and 60 virus species, 70 bacterial species and 70 virus species, 80 bacterial species and 80 viral species, 90 bacterial species and 90 viral species, 100 bacterial species and and 100 virus species, 200 bacterial species and 200 virus species, 300 bacterial species and 300 400 viral species, 400 bacterial species and 400 viral species, 500 bacterial species and 500 viruses species, or at least one nucleic acid from 1,000 bacterial species and 1,000 viral species Further comprising detecting.
[0193] In some cases, the methods described herein may be used to determine whether the infection is active or latent. In some cases, gene expression quantification can be used to detect and predict active infection. In some cases, the methods described herein may provide methods for diagnosing or monitoring In some cases, gene expression is used to detect disease in a population of interest. In some cases, gene expression quantification can be performed using a gene expression quantification kit (e.g., a gene expression quantification kit) or a gene expression quantification kit (e.g., a gene expression quantification kit). may provide a method for detecting, predicting, diagnosing or monitoring latent infections. The methods described herein include detecting latent infection.
[0194] Representative diseases and disorders include any disease or disorder associated with an infection, e.g., sepsis. , pneumonia, tuberculosis, HIV infection, hepatitis infection (e.g., hepatitis A, B, or C), human Papillomavirus (HPV) infection, chlamydia infection, syphilis infection, Ebola infection, Staphylococcus aureus infection or flu The methods provided herein include the treatment of drug-resistant microorganisms, including multidrug-resistant microorganisms. Some non-limiting examples of diseases and disorders include: These include Alzheimer's disease, amyotrophic lateral sclerosis, anorexia nervosa, anxiety disorders, asthma, Atherosclerosis, attention deficit hyperactivity disorder, autism, autoimmune diseases, bipolar disorder, cancer, chronic fatigue syndrome, chronic obstructive pulmonary disease, Crohn's disease, coronary heart disease, dementia, depression, type 1 diabetes, Type 2 diabetes, dilated cardiomyopathy, epilepsy, Guillain-Barré syndrome, irritable bowel syndrome, lower back pain, Lupus, metabolic syndrome, multiple sclerosis, myocardial infarction, obesity, obsessive-compulsive disorder, panic attacks Neuropathy, Parkinson's disease, psoriasis, rheumatoid arthritis, sarcoidosis, schizophrenia, stroke , thromboangiitis obliterans, Tourette's syndrome, vasculitis, plague, tuberculosis, anthrax, sleeping sickness, dysentery , Toxoplasmosis, Ringworm, Candidiasis, Histoplasmosis, Ebola, Acinetobacter Infectious diseases, actinomycosis, African sleeping sickness (African trypanosomiasis), AIDS (acquired immune Deficiency syndrome), amebiasis, anaplasmosis, anthrax, Arcanobacterium haemolyticus Arcanobacterium haemolyticum infection, Argentina Chin hemorrhagic fever, ascariasis, aspergillosis, astrovirus infection, babesiosis, cerebrospinal fluid Bacterial infections, bacterial pneumonia, bacterial vaginosis (BV), Bacteroides infections, Balantidium disease, Baylisascaris infection, BK virus infection, black sand hair, human blastocyst infection HIV infection, blastomycosis, Bolivian hemorrhagic fever, borreliosis, botulism (and infant botulism), Brazilian hemorrhagic fever, brucellosis, bubonic plague, Burkholderia infection , Buruli ulcer, Calicivirus infection (Norovirus and Sapovirus), Campylovirus Bacteriosis, candidiasis (moniliasis; thrush), cat scratch disease, cellulitis, Chagas American trypanosomiasis, chancroid, chickenpox, chikungunya fever, chlamydia, thyroid Chlamydophila pneumoniae infection ( Taiwan Acute Respiratory Syndrome (TWAR), cholera, chromoblastomycosis, clonorchiasis, clostridium erythrocytosis, Diazolidinium difficile infection, coccidioidomycosis, Colorado tick fever (CTF), colds ( Acute viral nasopharyngeal inflammation (acute coryza), Creutzfeldt-Jakob disease (CJD), Mia-Congo hemorrhagic fever (CCHF), cryptococcosis, cryptosporidiosis, cutaneous larvae CML, cyclosporosis, cysticercosis, cytomegalovirus infection, dengue fever, Dientamoebiasis, diphtheria, diphyllobothriasis, dracunculiasis, Ebola hemorrhagic fever, echinococcus infection, ehrlichiosis, enterobiasis (pinworm infection), enterococcus infection, enterovirus infection, typhus, erythema infectiosum (fifth disease), exanthema subitum (sixth disease), fascioliasis, liver cirrhosis Echinosis, Fatal Familial Insomnia (FFI), Filariasis, Clostridium perfringens food poisoning, Viable amoeba infection, Fusobacterium infection, gas gangrene (clostridial myonecrosis) ), geotrichosis, Gerstmann-Straussler-Scheinker syndrome (GSS), Giardiasis, glanders, gnathostomiasis, gonorrhea, granuloma venereum (donovanosis), group A streptococcal infection Streptococcus aureus, Group B Streptococcus infection, Haemophilus influenzae infection, Hand, Foot and Mouth Disease (HFMD), Hantawi Helicobacter pylori (HPS), Heartland virus disease, Helicobacter pylori (Heli cobacter pylori) infections, hemolytic uremic syndrome (HUS), and renal symptomatology. Blood fever (HFRS), hepatitis A, hepatitis B, hepatitis C, hepatitis D, hepatitis E, herpes simplex , histoplasmosis, hookworm infection, human bocavirus infection, human ehrlichiosis, human condylar Granulocytic anaplasmosis (HGA), human metapneumovirus infection, human monocytic ehrlichiosis ulcerative colitis, human papillomavirus (HPV) infection, human parainfluenza virus infection, Membranous tapeworm disease, Epstein-Barr virus mononucleosis (Mono), influenza (cold), isosporosis, Kawasaki disease, keratitis, Kingella kingae ngae) infection, kuru, Lassa fever, Legionnaires' disease, Legionnaires' disease (Pontiac fever), leishmaniasis, leprosy, leptospirosis, listeriosis , Lyme disease (Lyme borreliosis), lymphatic filariasis (elephantiasis), lymphocytic choriomeningitis flu, malaria, Marburg hemorrhagic fever (MHF), measles, Middle East respiratory syndrome (MERS), Melioidosis (Whitmore's disease), meningitis, meningococcal disease, trichosis, microsporidiosis, contagious Molluscum contagiosum (MC), monkeypox, mumps, murine typhus (typhus fever), mycoplasma pneumonia, mycetoma , myiasis, neonatal conjunctivitis (ophthalmia neonatorum), (new) variant Creutzfeldt-Jakob disease (v CJD, nvCJD), nocardiosis, onchocerciasis (river blindness), paracoccidioidomycosis oides (South American blastomycosis), paragonimiasis, pasteurellosis, head lice infestation (American Body lice), body lice infestation (body lice), pubic lice infestation (pubic lice, pubic lice) pelvic inflammatory disease (PID), pertussis (whooping cough), plague, pneumococcal infection, Pneumocystis pneumonia (PCP), pneumonia, poliomyelitis, Prevotella infection, primary amebic Meningoencephalitis (PAM), progressive multifocal leukoencephalopathy, psittacosis, Q fever, rabies, respiratory syncytial virus infection Infection, rhinosporidiosis, rhinovirus infection, rickettsial infection, rickettsialpox , Rift Valley fever (RVF), Rocky Mountain spotted fever (RMSF), rotavirus infection, Rubella, Salmonellosis, SARS (Severe Acute Respiratory Syndrome), scabies, schistosomiasis, sepsis, Shigella infection (bacterial dysentery), shingles (herpes zoster), smallpox (variola), sporotrichosis Colic, staphylococcal food poisoning, staphylococcal infection, strongyloidiasis, subacute sclerosing panencephalitis, syphilis, Tapeworm disease, tetanus (openis), tinea barbae (barber's itch), tinea capitis (tinea capitis), body Tinea corporis, tinea cruris (jock itch), tinea manus (tinea manu), tinea nigricans, tinea pedis (athlete's foot) Toefoot), Onychomycosis (Onychomycosis), Tinea versicolor (Black catfish), Toxocariasis (Ocular larva migrans LM), toxocariasis (visceral larva migrans (VLM)), trachoma, atrinocriasis septicemia, trichinosis, trichomoniasis, trichuriasis (whipworm infection), tuberculosis, tularemia, typhoid Ureaplasma urealyticum ) infectious diseases, valley fever, Venezuelan equine encephalitis, Venezuelan hemorrhagic fever, viral pneumonia, West Nile Fever, white sand hair (ringworm), Yersinia pseudotuberculosis infection, yersiniosis, yellow fever, and zygomycosis Examples include:
[0195] As used herein, the term "or" means a non-exclusive or unless otherwise indicated. For example, "A or B" is used to refer to "A but Includes "not B," "but B and not A," and "A and B."
[0196] As used herein, the term "about" when referring to a number or range of numbers , the stated numerical value or numerical range is an approximation within experimental variation (or within statistical experimental error). and a number or range of numbers means, for example, 1% of the stated number or range of numbers. In the examples, the term "about" means ±1% of the stated number or value. Indicates 0%.
[0197] [Example] [Example 1] Preparation of cell-free RNA from patient whole blood samples Whole blood is drawn from patients suspected of having an infectious disease and injected into acid citrate dextrose. Place the blood in an ACD tube. Remove a portion of the blood (1.5 mL) from the ACD tube. Place in a 1.5 mL microcentrifuge tube. Add 10 μL of standardized oligonucleotide to the blood and thoroughly mix. The mixture was centrifuged at 1600 g for 10 minutes at 4°C, and the supernatant ("normalized blood") was collected. Remove 550 μL of the standardized plasma and place it in a new microcentrifuge tube. The supernatant ("standardized cell-free plasma") is removed and centrifuged at 16,000 g for 10 minutes. Transfer to a new tube and store at -80°C.
[0198] Thaw the standardized cell-free plasma at room temperature for 10 minutes. Thaw the standardized cell-free plasma at 4°C. Cell-free RNA was extracted using the manufacturer's instructions and centrifuged at 16000 g for 10 minutes to remove debris. Plasma / Serum Circulation and Exosophy according to instructions mal RNA Purification Kit(Slurry Format)( Cell-free RNA was isolated using a 500-kDa PBS (Norgen Biotek Corp.). Store at °C.
[0199] [Example 2] Preparation of a non-human collection of oligonucleotides by computational design and synthesis Approximately 6.7 x 10 of nucleotides 6 The possible different domains have a length of 13 nucleotides. From this theoretical sequence pool, any 13 nucleotide sequence found in human DNA is selected. If we discard the array, it becomes about 2.3 × 10 6 The unique 13 nucleotide sequences, i.e., 3.5% of the body remains. 2.3 × 10 6 From 13 non-human nucleotide sequences ,The following criteria were met: 1) uniformity of melting temperature, 2) sufficient sequence complexity, 3) known pathogen Based on the abundance of binding sites in the body, 4) and the distribution of binding sites among known pathogens, Approximately 1 x 10 for inclusion in the human sequence pool 6 Select this approximately 1 x 10 6 non-humans A set of 13 nucleotide sequences plus additional non-human sequences of 14-20 nucleotides in length In addition, we aim to improve coverage of targeted regions, such as strategic pathogen sequences, and to The markers were then sorted by 1) melting temperature, 2) sequence complexity, 3) abundance of the binding site in known pathogens, and and 4) designed based on the distribution of binding sites among known pathogens. 13–20 nucleotides in length The entire pool of nucleotide domains is called the complete non-human sequence pool. Each 13-20 nucleotide sequence in the vector contains one or more of the following nucleic acid labels: Add an additional 5' sequence of approximately 15-25 nucleotides in length: 1) DNA sequencing primer 1) the mer binding site; 2) the sample barcode; 3) the sample barcode sequencing primer binding site. Binding sites, and 4) amplification primer binding sites compatible with various sequencing platform requirements. This collection of oligonucleotides is called a non-human primer pool. The human primer pool is chemically synthesized. Alternatively, the completely non-human sequence pool is chemically synthesized. and attaching one or more nucleic acid labels (e.g., by ligation) to the non-human plasmid. Form an immersion pool.
[0200] [Example 3] Non-human collection of oligonucleotides using hybridization-based methods Preparation of Approximately 67 million different 13 nucleotide sequences (e.g., where N is A, C, G, or T) 5'-NNNNNNNNNNNNN-3') with one or more of the following nucleic acid labels: Chemically synthesize with ligated 15-25 nucleotide overhangs containing: 1) DNA 1) sequencing primer binding site; 2) sample barcode; 3) sample barcode sequencing 4) amplification primers compatible with various sequencing platform requirements; This heterogeneous collection of oligonucleotides is then hybridized to the primer binding site. Solution buffer (0.5x PBS, 24 μM blocker oligonucleotide, RNase inhibitor 95% of a 1000-fold mass excess of biotinylated human single-stranded genomic DNA (gDNA) fragments in a 100% PBS solution. Hydrate at 10°C for 10 seconds, 65°C for 3 minutes, and 36°C for a set amount of time, ranging from several hours to several weeks. At the end of the incubation period, the human gDNA fragments are hybridized to the fragments. The antibody is removed along with the purified probe by streptavidin beads. The remaining pool of probes that did not bind to human gDNA under the conditions was subjected to 14-20 nucleotide PCR. Additional non-human sequences of any length can be supplemented to improve coverage of regions of interest, such as strategic pathogen sequences. These primers were then compared to improve their 1) melting temperature, 2) sequence complexity, and 3) known pathogen binding. and 4) the distribution of binding sites among known pathogens. The collection of oligonucleotides is called a non-human primer pool.
[0201] [Example 4] Preparation of non-human collections of highly degenerate oligonucleotides in high yields. Non-human 13mers were generated to generate all possible 13mer sequences and matched to the reference human genome. The non-human 13-mer is then degenerated. The histogram in Figure 11B shows the distribution of non-human 13mers based on degeneracy. Indicates packetization.
[0202] Next, the same degenerate non-human 13-mer variable oligonucleotide sequence units were applied to the ultramafic - Classified as oligonucleotides. Degenerate oligonucleotide sequences contained in Ultramer The number of units can be based on the length of each unit and the ultramer length that can be reliably synthesized. Figure 9 shows the general design of some such Ultramer oligonucleotides. The individual degenerate 13mers can be separated by deoxyuracil nucleotides (U). The designed Ultramer oligonucleotides are then synthesized using conventional nucleic acid synthesis methods and It can be synthesized by a service provider, for example IDT.
[0203] The Ultramer oligonucleotides were then purified using uracil-DNA glycosylase (UDG). G) and specificity at abasic sites (e.g., apurinic / apyrimidinic or AP sites) The individual degenerate non-human 13-mer oligonucleotides are synthesized by the enzyme This digestion will result in all the degenerate 13-mers having the same degeneracy. A typical reaction can be performed in the same tube for multiple Ultramer oligonucleotides. The response is shown in FIG.
[0204] Each degenerate 13-mer is then purified using, for example, T4 polynucleotide kinase (T4 PN K) or chemically biotinylated as shown in FIG. This step is optional, for example, if surface immobilization or magnetic bead purification is required. Other probes than biotin (e.g., digoxigenin, fluorescent probes, etc.) can be used here. can be applied.
[0205] [Example 5] Preparation of enriched sequencing libraries using a non-human collection of oligonucleotides The cell-free RNA from Example 1 was diluted in hybridization buffer (0.5x PBS, 2 4 μM blocker oligonucleotide, RNase inhibitor) at 95°C for 10 seconds, 65 3 min at 36 °C and overnight at 36 °C for 3 min. The non-human primer pool (described in Examples 2 or 3) was Hybridize with first strand cDNA synthesis master mix (prepared to transcriptase, blocked second strand synthesis oligonucleotide, dNTP, RNase inhibitor and appropriate buffer) was added at 36°C, after which the temperature was raised to 42°C for 90 minutes, Allow first strand cDNA synthesis. Then incubate the mixture at 70 °C for 10 min. Inactivate the reverse transcriptase by PBS and keep at 4°C.
[0206] The first strand cDNA product is purified using a commercially available kit. The second strand cDNA is purified using the same procedure as in the previous step. The hybridization step hybridizes to fixed sequences added to the 5' and 3' ends of the first strand cDNA. After second strand synthesis by PCR, One or more of the following nucleic acid labels may be added by further rounds of PCR: It can: 1) DNA sequencing primer binding site, 2) sample barcode, 3) sample 1) barcoded sequencing primer binding sites, and 2) compatibility with various sequencing platforms. Amplification primer binding sites compatible with the conditions. The final cDNA library was prepared according to the manufacturer's instructions. Purify using a commercially available kit according to the method described above.
[0207] [Example 6] DNA or cDNA libraries using a non-human collection of oligonucleotides Preparation of a sequencing library enriched for non-human sequences from Prepare a sequenceable library. The non-human primer pool is prepared as described in Example 2 or 3. Prepared as described in and facilitated 5' biotinylation of individual oligonucleotides in the pool. Then, 1 pmol of the 5'-biotinylated non-human primer pool is subjected to conditions for cleavage. Hybridization / Polymerization Buffer (Buffer, Nucleotides, Blocker Oligonucleotides) Add 500 ng of sequenceable library in a 500 μL aliquot of DNA. A. Adding a molecule to an oligonucleotide that enhances the rate or specificity of the hybridization reaction Preincubation of the nucleotide sequence with non-host nucleic acid fragments such as RecA or MutS This can improve capture. Denaturing DNA at 95°C for 10 minutes. Non-human primers Hybridize to the sequencing library at 50°C for 4 hours. Add a strand-displacing DNA polymerase lacking the enzyme to the mixture. Incubate the mixture at 55°C for 15 minutes. Incubate to separate the non-human primers by at least 25 bases, or in subsequent steps The DNA is elongated to a length sufficient for stable double-stranded DNA hybridization. Vidin beads are added to the mixture to bind the DNA fragments. Captured DNA on the beads The captured library DNA fragments are then washed and the DNA sequencing library fragments are analyzed. The primer specific to the adapter at one end and the DNA fragment captured on the beads are then used. The enriched library is amplified using standard PCR amplification methods. is purified using standard DNA purification methods.
[0208] Libraries were quantified using the KAPA DNA Library Quantification Kit (KAPA Biosy) Quantify using NextSeq 500 (Ill) according to the manufacturer's instructions. The DNA was prepared for sequencing on a lumina (Schwarzschild AG). Sequencing was performed using 150 cycles of single-end reading. It consists of a reading and eight cycles of barcode reading.
[0209] Computationally map sequence reads to the pathogen genome to identify one or more nucleic acids Identify the source or sources.
[0210] [Example 7] Human nucleosome assembly using anti-human histone and anti-immunoglobulin antibodies NA depletion Cell-free DNA was extracted using either centrifugation or albumin and immunoglobulin depletion. In either method, 1 mL of human plasma is first heated at 4°C for 16 minutes. Centrifuge at 1000g for 10 minutes. In the case of centrifugation, collect the supernatant (950µL) and Transfer to a suitable tube and centrifuge again at 16000 g for 10 minutes at 4°C. Collect 900 μL of the supernatant containing albumin and immunoglobulin A and transfer it to a new tube. For the debulinization method, the supernatant (950 μL) from the first centrifugation was collected and analyzed for human albumin. and a human immunoglobulin binding column (e.g., Alb umin and IgG Depletion SpinTrap or LifeT Technologies Albumin / IgG Removal Kit) Collect the flow-through containing the cell-free DNA and transfer it to a new tube.
[0211] Prepared by centrifugation or albumin and immunoglobulin removal as described above The cell-free DNA was mixed with 5 μg of one or more anti-human histone antibodies and gently spun. Incubate at 4°C for 1 hour to overnight while stirring. The antibody binds to the nucleosomes and complexes them. A mixture of magnetic beads conjugated to anti-immunoglobulin antibodies is then injected into the reaction vessel. Anti-immunoglobulin (anti-immunoglobulin) is added to the reaction to sequester the human nucleosome-antibody complexes from the plasma sample. The globulin antibody binds to the isotope of the anti-human histone antibody added in the previous step. The mixture contains the antibody combination required for the following: Incubate at room temperature for 1 hour or at 4°C for 5 to 10 minutes. Incubate overnight. Magnetic beads with bound human nucleosome-antibody complexes Pellet the beads on a magnetic stand. Carefully collect the supernatant and discard any remnants of the magnetic bead fraction. Do not collect the supernatant. The pathogen was purified using the ng Nucleic Acid Kit (Qiagen). Concentrate cellular DNA.
[0212] [Example 8] Anti-human histone antibodies conjugated to protein A and / or G and magnetic beads Depletion of human nucleosome-associated DNA using azepam Cell-free DNA prepared using the albumin and immunoglobulin depletion method of Example 6 was The flow-through containing the antibody was mixed with 5 μg of one or more anti-human histone antibodies and slowly stirred. Incubate at 4°C for 1 hour to overnight with constant rotation. Add magnetic beads conjugated to protein A and / or G to the combined solution. Isolate the mixture from the solution by incubating at room temperature for 1 hour or at 4°C for 5 hours to overnight. The magnetic beads with the bound human nucleosome-antibody complexes are incubated with a magnetic stirrer. Pellet on a plate. Carefully collect the supernatant to avoid remnants of the magnetic bead fraction. The resulting supernatant was purified using a commercially available kit (e.g., QIAamp Circulating Nuclease). Pathogen cell-free DNA was purified using a cleic acid kit (Qiagen). Concentrate.
[0213] [Example 9] Depletion of human nucleosome-associated DNA using anti-human histone antibodies and spin columns Cell-free DNA containing material prepared using the anti-albumin immunoglobulin method of Example 6 The resulting flow-through was mixed with 5 μg of one or more anti-human histone antibodies and gently rotated. Incubate at 4°C for 1 hour to overnight while stirring. The solution was passed through a spin column functionalized with protein A / G or anti-immunoglobulin antibodies. The flow-through is purified using a commercially available kit (e.g., QIAamp Cir). Purified using a Calculating Nucleic Acid Kit (Qiagen). , to enrich for pathogen cell-free DNA.
[0214] [Example 10] Depletion of human nucleosome-associated DNA by DNA precipitation After plasma centrifugation using a DNA condensing agent (e.g., spermine), nucleosome-containing The cell-free DNA and nucleosome-bound cell-free DNA are precipitated. The pellet was then incubated with anti-human histone antibody against human nucleosomes. The resulting antigen-antibody complex is dissolved in a buffer optimized for antibody binding. bound to magnetic beads conjugated to either G or anti-immunoglobulin antibodies, The magnetic beads are isolated by pelleting them on a magnetic stand and collecting the supernatant. The supernatant was purified using a commercially available kit (e.g., QIAamp Circulating Nuclease). Purify and enrich pathogen cell-free DNA using the ELISA Kit (Qiagen). do.
[0215] [Example 11] Nucleotide domains in non-human collections of oligonucleotides Optimization of leotide length When selecting the length of the nucleotide domain in the non-human primer pool, some The parameters of the assay are optimized. The goal is to provide the greatest sensitivity for detecting non-human nucleic acids. The goal is to generate a pool with the minimum number of different sequences. As the length increases, domains of nucleotides with a length of k increase in size to 4 k has a permutation of Therefore, a larger number of different probes are needed to cover the sequence space. The number of probes in the chromatogram is increased so that individual sequences are present at higher concentrations and exhibit faster hybridization. This consideration is also important when choosing primers that are shorter than the primer length. At the same time, as the length of the primer increases, all possible sequences of that length are A smaller proportion of all sequences will bind to human sequences, which is a much larger proportion of the total genome space. are left in the pool of non-human primers, resulting in higher coverage of non-human sequences and This means that it provides high sensitivity, e.g., 3.5% of possible sequences of 13 nucleotides in length. Only (2.3 million) are not found in the human reference genome, which is a 13 nucleotide long non-human sequence. The pool of strings binds to only 3.5% of all possible 13-nucleotide stretches in non-human nucleic acids. In contrast, 15.2% (4070) of the possible sequences are 14 nucleotides long. 10,000) are not found in the human reference genome. 15.2% of the nucleotide extensions can be theoretically detected. 40.7 million probes may be too many to synthesize in one go. When present at 1 / 40.7 million of the probe concentration, the concentration of each individual probe is Unless other measures are taken to promote binding, they will bind favorably to their complementary sequences at equilibrium. This finding suggests that probe sets longer than 13 nucleotides may be too low for Alternatively, sequences shorter than 13 nucleotides in length can be excluded from the human reference genome. It is found so frequently in genomes that nearly all primers are from non-human pools. For example, at 12 nucleotides, 99.74% of the possible sequences are human. This is because a pool of non-human primers is used to identify 12 possible variants in the pathogen. It binds to only about 0.26% of individual nucleotide sequences, providing limited sensitivity. The 13 nucleotide domain represents the hybridization domain. The probes are sufficiently small (less than 10 million different probes) to have reasonable kinetics in the ionization reaction. The probes (including the nucleotide sequence) provide sufficient sensitivity (approximately one binding site per 30 nucleotides in the non-human genome). Provide.
[0216] [Example 12] Selective enrichment of short fragments in cell-free DNA The size enrichment protocol is optimized for the selection of cell-free DNA with relatively short fragment lengths. Cell-free DNA was purified from human erythrocytes using the Qiagen CNA kit with the following modifications: Cells were extracted from plasma: (a) using 3 volumes of ACB buffer, (b) ACW1 buffer Prepare ACW1 buffer according to the manufacturer's recommendations, adding an additional 100 µL of absolute ethanol per 600 µL of ACW1 buffer. (c) ACW2 buffer solution, supplemented with 1.75 mL of ethanol and 0.36 g of guanidinium chloride. Prepare according to the manufacturer's recommendations, using 1.05 mL of anhydrous ethanol per 750 µL of ACW2 buffer. The isolated cell-free DNA was then subjected to adapter ligation and and the 1.8x Ampure purification step after library amplification using NuGen's Ovation The library was then arrayed using the Ultralow V2 Library Kit. The sequences were determined and the reads were mapped to human and pathogen databases. The resulting reads were examined.
[0217] In particular, Figure 8A shows the sequence of host or human DNA (chr21) and non-host or pathogen DNA. For both, the amount of cfDNA (measured as a function of the number of sequence reads) versus the number of DNA sequences derived As shown, non-host DNA ranged from about 30 to about 100 bp. At fragment lengths of 100 bases, they enjoy a much more favorable ratio to the host DNA. Therefore, enrichment of fragments in this size range is expected to enrich for non-host versus host DNA. will be done.
[0218] Equimolar cell-free synthesis configured not to map to known human or pathogen sequences DNA size controls (e.g., 32, 52, 75, 100, 125, 150, 17 For DNA samples containing 5 and 350 bp fragments, a modified Qiagen CNA A purification kit was used to test the enrichment of fragments in the desired size range.
[0219] Figure 8B shows the results of the analysis of larger fragments (e.g., of the eight DNA size control fragment lengths). Typical CNA kits selected for length (100-200 bases with a peak enrichment of 175 bp) The fragment size profile of the sample when processed using the modified protocol is shown. The protocol used yielded peak concentrations at 75 bp among eight DNA size control fragment lengths. Fragments with a length of 30 to approximately 120 bases were selectively enriched.
[0220] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. The following claims define the scope of the invention, and it is intended that methods and structures within the scope of these claims and their equivalents be covered thereby. In one aspect, the present invention provides the following. [Item 1] 1. A method for priming or capturing non-host sequences in a nucleic acid sample from a host, comprising: (a) providing a nucleic acid sample from said host, said nucleic acid sample from said host comprising host nucleic acids and non-host nucleic acids; (b) mixing the nucleic acid sample from the host with a collection of oligonucleotides, thereby obtaining a mixture, wherein the collection of oligonucleotides comprises at least 1000 oligonucleotides having different nucleotide sequences, and the different nucleotide sequences are specifically selected to contain non-host nucleic acid sequences at least 10 nucleotides in length; (c) contacting the collection of oligonucleotides with the nucleic acid sample in the mixture, wherein said contacting causes non-host nucleic acids in the mixture to bind to said non-host nucleic acid sequences at least 10 nucleotides in length, thereby priming or capturing non-host nucleic acids, and wherein said contacting causes up to 10% of said host nucleic acids to bind to said non-host nucleic acid sequences at least 10 nucleotides in length; A method comprising: [Item 2] 2. The method of claim 1, further comprising preferentially amplifying the primed or captured non-host nucleic acid in a reaction. [Item 3] 2. The method of claim 1, further comprising sequencing the primed or captured non-host nucleic acid by performing a sequencing assay. [Item 4] 4. The method of claim 3, wherein the sequencing assay is a next-generation sequencing assay, a high-throughput sequencing assay, a massively parallel sequencing assay, a nanopore sequencing assay, or a Sanger sequencing assay. [Item 5] 2. The method of claim 1, further comprising preferentially isolating the primed or captured non-host nucleic acid. [Item 6] 6. The method of claim 5, wherein the preferentially isolating step comprises performing a pull-down assay. [Item 7] 2. The method of claim 1, further comprising the step of performing a primer extension reaction on the primed or captured non-host nucleic acid. [Item 8] 8. The method of claim 7, wherein the at least 1000 oligonucleotides having different nucleotide sequences contain a nucleic acid label. [Item 9] 2. The method of claim 1, wherein the primed or captured non-host nucleic acid is an RNA non-host nucleic acid. [Item 10] 10. The method of claim 9, further comprising the step of performing a polymerization reaction on the primed or captured RNA non-host nucleic acid. [Item 11] 11. The method of claim 10, wherein the polymerization reaction is carried out by a reverse transcriptase. [Item 12] 1. A method for sequencing a non-host sequence in a nucleic acid sample from a host, comprising: (a) providing a nucleic acid sample from the host, the nucleic acid sample from the host comprising host nucleic acid and non-host nucleic acid; (b) mixing the nucleic acid sample from the host with a collection of oligonucleotides, thereby obtaining a mixture, wherein the collection of oligonucleotides comprises at least 1000 oligonucleotides having different nucleotide sequences, and the different nucleotide sequences are specifically selected to contain non-host nucleic acid sequences at least 10 nucleotides in length; (c) contacting the collection of oligonucleotides with the nucleic acid ...
Claims
1. 1. A method for concentrating cell-free microbial nucleic acids, comprising: (a) identifying a nucleic acid length range that favors a ratio of cell-free microbial nucleic acid to cell-free host nucleic acid; (b) providing a host sample containing cell-free host nucleic acid and cell-free microbial nucleic acid selected from blood, plasma, serum, saliva, cerebrospinal fluid, synovial fluid, and lavage fluid; and (c) removing cell-free nucleic acids of nucleic acid length ranges outside of a preferred ratio of cell-free microbial nucleic acids to cell-free host nucleic acids from the host sample to enrich for nucleic acid lengths within said range; The method comprising:
2. 10. The method of claim 1, further comprising performing a sequencing assay on the cell-free host nucleic acid and the cell-free microbial nucleic acid from the host sample.
3. 10. The method of claim 1, wherein the host sample comprises at least five cell-free microbial nucleic acids, and further comprising detecting the five cell-free microbial nucleic acids.
4. The method of claim 1 , wherein the host is a human.
5. The method of claim 1, wherein the cell-free nucleic acid is DNA.