Methods and compositions for analyzing nucleic acids
A method for generating nucleic acid libraries using single-stranded nucleic acids and scaffold polynucleotides with specific hybridization regions addresses the inefficiencies of existing methods, enabling efficient and cost-effective library preparation for diverse sequencing platforms.
Patent Information
- Application Number
- JP2022579952
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-06-01
- Filing Date
- 2021-06-23
- Publication Date
- 2026-01-20
- Estimated Expiration
- 2041-06-23
AI Technical Summary
Existing methods for preparing nucleic acid libraries, particularly single-stranded nucleic acid libraries, are labor-intensive, expensive, and time-consuming, and require custom reagents, limiting their widespread use and efficiency.
A method involving a nucleic acid composition comprising single-stranded nucleic acids, first oligonucleotide species, and first scaffold polynucleotide species, where each polynucleotide has specific hybridization regions, allowing for the formation of hybridization products with adjacent oligonucleotide ends, thereby facilitating efficient library generation without custom reagents.
This approach enables the generation of high-quality nucleic acid libraries efficiently and cost-effectively, preserving valuable information at nucleic acid ends and supporting various sequencing platforms.
Smart Images

Figure 0007802707000001 
Figure 0007802707000002 
Figure 0007802707000003
Abstract
Description
[Technical Field]
[0001] Related patent applications This patent application claims the benefit of U.S. Provisional Patent Application No. 63 / 043,688, filed June 24, 2020, entitled "METHODS AND COMPOSITIONS FOR ANALYZING NUCLEIC ACID," naming Christopher J. Troll as inventor, and designated by attorney docket number CBS-2004-PV. This patent application claims the benefit of U.S. Provisional Patent Application No. 63 / 043,688, filed October 1, 2020, entitled "METHODS AND COMPOSITIONS FOR ANALYZING NUCLEIC ACID," naming Christopher J. Troll et al. as inventors, and designated by attorney docket number CBS-2004-PV2. day This patent application also claims the benefit of U.S. Provisional Patent Application No. 63 / 086,208, filed on March 10, 2021, entitled METHODS AND COMPOSITIONS FOR ANALYZING NUCLEIC ACID, naming Christopher J. Troll et al. as inventors, and designated by attorney docket number CBS-2004-PV3. day This patent application also claims the benefit of U.S. Provisional Patent Application No. 63 / 159,174, filed on June 1, 2021, entitled "METHODS AND COMPOSITIONS FOR ANALYZING NUCLEIC ACID," naming Christopher J. Troll et al. as inventors, and designated by attorney docket number CBS-2004-PV4. Note The entire contents of this patent application, including all text, tables, and drawings, are incorporated herein by reference for all purposes.
[0002] Field The present technology relates, in part, to methods and compositions for analyzing nucleic acids. In some aspects, the present technology relates to methods and compositions for preparing nucleic acid libraries from single-stranded nucleic acid fragments. [Background technology]
[0003] background raw thing The genetic information of organisms (e.g., animals, plants, and microorganisms) and other forms of replicating genetic information (e.g., viruses) is encoded in nucleic acids (i.e., deoxyribonucleic acid (DNA) or ribonucleic acid (RNA)). Genetic information corresponds to the chemical or hypothetical primary structure of nucleic acids. A series or modified nucleotides.
[0004] Various high-throughput sequencing platforms are used to analyze nucleic acids. For example, the ILLUMINA platform uses clonal amplification of adapter-ligated DNA fragments. involved in Another platform is based on nanopores, which rely on the translocation of nucleic acid molecules or individual nucleotides through small channels. Sequencing Library preparation for a particular sequencing platform often involves DNA fragmentation, modification of fragment ends, and adapter ligation, and may involve amplification of nucleic acid fragments (e.g., PCR amplification). Miteku do.
[0005] Choosing an appropriate sequencing platform for a particular type of nucleic acid analysis requires a detailed understanding of available technologies, including error sources, error rates, and sequencing speed and cost. Although sequencing costs have decreased, the throughput and cost of library preparation can be a limiting factor. One aspect of library preparation involves modifying the ends of nucleic acid fragments so that they are suitable for a particular sequencing platform. Nucleic acid ends can contain useful information. Therefore, methods for modifying nucleic acid ends (e.g., for library preparation) while preserving the information contained in the nucleic acid ends are useful for nucleic acid processing and analysis. There will be .
[0006] Another aspect of library preparation involves capturing single-stranded nucleic acid fragments. In certain cases, single-stranded library preparation methods can generate better and more complex libraries compared to traditional double-stranded DNA (dsDNA) preparation methods. Drawbacks of generating single-stranded DNA (ssDNA) libraries include labor-intensive, expensive, and time-consuming protocols, and exotic or special order (custom) These include the reagent requirements, which can result in labor-intensive, expensive, and time-consuming protocols and / or rare Alternatively, methods for capturing single-stranded nucleic acids (e.g., for library preparation) without the need for custom reagents may be used to capture nucleic acids (e.g., single-stranded nucleic acids, modified ... sex two It is useful for processing and analyzing double-stranded nucleic acids, or mixtures containing single-stranded nucleic acids. There will be . Summary of the Invention [Means for solving the problem]
[0007] overview In certain aspects, a method for generating a nucleic acid library is provided, comprising: (i) a nucleic acid composition comprising single-stranded nucleic acids (ssNAs); (ii) a plurality of first oligonucleotide species; and (iii) a plurality of first scaffold polynucleotide species. combine(a) each polynucleotide in the plurality of first scaffold polynucleotide species comprises a ssNA hybridization region and a first oligonucleotide hybridization region; and (b) each oligonucleotide in the plurality of first oligonucleotide species comprises a first flanking region and a second flanking region. is sandwiched between (c) the first oligonucleotide hybridization region comprises (i) a polynucleotide complementary to the first flanking region and (ii) a polynucleotide complementary to the second flanking region; (d) the nucleic acid composition, the plurality of first oligonucleotide species, and the plurality of first scaffold polynucleotide species are hybridized under conditions such that molecules of the first scaffold polynucleotide species hybridize to (i) the first ssNA terminal region and (ii) molecules of the first oligonucleotide species. Combination whereby a hybridization product is formed in which the end of the first oligonucleotide molecule is adjacent to the end of the first ssNA terminal region.
[0008] First adjacent region and second adjacent region sandwiched between Also provided is a composition comprising a plurality of first oligonucleotide species, each comprising a first unique molecular identifier (UMI) comprising a first flanking region and a plurality of first scaffold polynucleotide species, each comprising a ssNA hybridization region and a first oligonucleotide hybridization region, wherein the first oligonucleotide hybridization region comprises (i) a polynucleotide complementary to the first flanking region and (ii) a polynucleotide complementary to the second flanking region.
[0009] 1. A method for generating a nucleic acid library, comprising: (a) mixing single-stranded ribonucleic acid (ssRNA) and double-stranded deoxyribonucleic acid (dsDNA) in a first mixture containing ssRNA, a primer oligonucleotide, and a nucleic acid library containing a nucleic acid library having reverse transcriptase activity; agent(b) contacting the first oligonucleotide and the plurality of first scaffold polynucleotide species with the first oligonucleotide to generate a second mixture containing a complementary deoxyribonucleic acid (cDNA)-RNA duplex and dsDNA, wherein the primer oligonucleotide comprises an RNA-specific tag, the cDNA comprises an RNA-specific tag, and the dsDNA does not comprise an RNA-specific tag; (b) generating single-stranded cDNA (sscDNA) and single-stranded DNA (ssDNA) from the cDNA-RNA duplex and dsDNA, thereby generating a nucleic acid composition containing sscDNA and ssDNA; (c) contacting the nucleic acid composition with the first oligonucleotide and the plurality of first scaffold polynucleotide species with the first oligonucleotide and the plurality of first scaffold polynucleotide species. Combination (i) each polynucleotide in the plurality of first scaffold polynucleotide species comprises a sscDNA hybridization region or a ssDNA hybridization region and a first oligonucleotide hybridization region; and (ii) hybridizing the nucleic acid composition, the first oligonucleotide, and the plurality of first scaffold polynucleotide species under conditions such that molecules of the first scaffold polynucleotide species hybridize with (1) the first sscDNA terminal region or the first ssDNA and (2) molecules of the first oligonucleotide. Combination thereby forming a hybridization product in which the end of the first oligonucleotide molecule is adjacent to the first sscDNA terminal region or the end of the first ssDNA terminal region.
[0010] Also provided is a composition comprising: a nucleic acid composition comprising single-stranded complementary deoxyribonucleic acid (sscDNA) and single-stranded deoxyribonucleic acid (ssDNA), wherein the sscDNA comprises an RNA-specific tag; a first oligonucleotide; and a plurality of first scaffold polynucleotide species, each comprising an sscDNA hybridization region or an ssDNA hybridization region, and a first oligonucleotide hybridization region.
[0011] A method for generating a nucleic acid library, comprising: (i) a nucleic acid composition comprising single-stranded ribonucleic acid (ssRNA) and single-stranded deoxyribonucleic acid (ssDNA); (ii) a first oligonucleotide; (iii) a plurality of first scaffold polynucleotide species; (iv) a second oligonucleotide; and (v) a plurality of second scaffold polynucleotide species. Combination (a) the first oligonucleotide comprises an RNA-specific tag; (b) the second oligonucleotide comprises a DNA-specific tag; (c) each polynucleotide in the plurality of first scaffold polynucleotide species comprises an ssRNA hybridization region and a first oligonucleotide hybridization region; (d) each polynucleotide in the plurality of second scaffold polynucleotide species comprises an ssDNA hybridization region and a second oligonucleotide hybridization region; and (e) the nucleic acid composition, the first oligonucleotide, the plurality of first scaffold polynucleotide species, the second oligonucleotide, and the plurality of second scaffold polynucleotide species are hybridized to a first set of hybridization products in which molecules of the first scaffold polynucleotide species hybridize with (i) the first ssRNA terminal region and (ii) molecules of the first oligonucleotide, thereby forming a first set of hybridization products in which an end of a molecule of the first oligonucleotide is adjacent to an end of the first ssRNA terminal region. and and (ii) hybridizing molecules of a second scaffold polynucleotide species with (i) the first ssDNA terminal region and (ii) molecules of a second oligonucleotide, thereby forming a second set of hybridization products in which the ends of the molecules of the second oligonucleotide are adjacent to the ends of the first ssDNA terminal region. Combination A method is also provided.
[0012] Also provided are compositions comprising: a first oligonucleotide comprising an RNA-specific tag; a second oligonucleotide comprising a DNA-specific tag; a plurality of first scaffold polynucleotide species, each comprising an ssRNA hybridization region and a first oligonucleotide hybridization region; and a plurality of second scaffold polynucleotide species, each comprising an ssDNA hybridization region and a second oligonucleotide hybridization region.
[0013] A method for generating a nucleic acid library, comprising: (a) combining under extension conditions a first nucleic acid composition comprising a target nucleic acid and one or more specific nucleic acid sequences; Yes Nucleotide and elongation activity agent and (iii) contacting the target nucleic acids with one or more specific nucleic acids, thereby generating extended target nucleic acids, wherein (i) some or all of the target nucleic acids comprise double-stranded nucleic acids (dsNA) comprising overhangs, (ii) the extended target nucleic acids each comprise an extension region complementary to the overhangs, and (iii) the extension region comprises one or more specific nucleic acids. Yes (b) generating single-stranded nucleic acids (ssNAs) from the extended target nucleic acids, thereby generating a second nucleic acid composition comprising ssNAs; and (c) combining the second nucleic acid composition with the first oligonucleotide and a plurality of first scaffold polynucleotide species. Combination (i) each polynucleotide in the plurality of first scaffold polynucleotide species comprises a ssNA hybridization region and a first oligonucleotide hybridization region; and (ii) the second nucleic acid composition, the first oligonucleotide, and the plurality of first scaffold polynucleotide species are hybridized under conditions such that molecules of the first scaffold polynucleotide species hybridize with (1) the first ssNA terminal region and (2) molecules of the first oligonucleotide. Combination whereby a hybridization product is formed in which the end of the first oligonucleotide molecule is adjacent to the end of the first ssNA terminal region.
[0014] A method for generating a nucleic acid library, comprising: (a) reacting a nucleic acid composition comprising a target nucleic acid and one or more specific nucleic acids under extension conditions; Yes Nucleotide and elongation activity agent and (iii) contacting the target nucleic acids with one or more specific nucleic acids, thereby generating extended target nucleic acids, wherein (i) some or all of the target nucleic acids comprise double-stranded deoxyribonucleic acid (dsDNA) comprising overhangs, (ii) the extended target nucleic acids each comprise an extension region complementary to the overhangs, and (iii) the extension region comprises one or more specific nucleic acids. Yes nucleotide Including nothing (comprisesat one or more distinctive nucleotides) and (b) attaching an adaptor polynucleotide to the extended target nucleic acid, wherein the adaptor polynucleotide comprises one strand capable of forming a hairpin structure having a single-stranded loop and a double-stranded region, thereby generating a continuous strand-extended target nucleic acid comprising a single-stranded loop and a double-stranded region.
[0015] A method for generating a nucleic acid library, comprising: (a) reacting a nucleic acid composition comprising a target nucleic acid and one or more specific nucleic acids under extension conditions; Yes Nucleotide and elongation activity agent and (iii) contacting the target nucleic acids with one or more specific nucleic acids, thereby generating extended target nucleic acids, wherein (i) some or all of the target nucleic acids comprise double-stranded deoxyribonucleic acid (dsDNA) comprising overhangs, (ii) the extended target nucleic acids each comprise an extension region complementary to the overhangs, and (iii) the extension region comprises one or more specific nucleic acids. Yes nucleotide Including and (b) generating concatemers of the extended target nucleic acid, thereby generating concatemerized extended target nucleic acid.
[0016] A method for generating a nucleic acid library, comprising: (a) combining (i) a nucleic acid composition comprising a single-stranded nucleic acid (ssNA); (ii) a first oligonucleotide; and (iii) a plurality of first scaffold polynucleotide species. Combinationwherein each polynucleotide in the plurality of first scaffold polynucleotide species comprises a ssNA hybridization region and a first oligonucleotide hybridization region, and the nucleic acid composition, the first oligonucleotide, and the plurality of first scaffold polynucleotide species are hybridized with the nucleic acid composition, the first oligonucleotide, and the plurality of first scaffold polynucleotide species under conditions such that molecules of the first scaffold polynucleotide species hybridize with (1) the first ssNA terminal region and (2) molecules of the first oligonucleotide. Combination Also provided is a method comprising (a) (b) hybridizing a first oligonucleotide with a first ssNA terminal region to form a hybridization product in which the molecular end of the first oligonucleotide is adjacent to the end of the first ssNA terminal region; and (b) deaminating one or more unmethylated cytosine residues in the ssNA, thereby converting the one or more unmethylated cytosine residues to uracil.
[0017] 1. A method for generating a nucleic acid library, comprising: (a) providing a first mixture containing single-stranded ribonucleic acid (ssRNA) and double-stranded deoxyribonucleic acid (dsDNA), a priming polynucleotide, and a nucleic acid library containing a nucleic acid library having reverse transcriptase activity; agent (i) the priming polynucleotide comprises a primer, an RNA-specific tag, and a first oligonucleotide, (ii) the cDNA comprises an RNA-specific tag and a first oligonucleotide, and (iii) the dsDNA does not comprise an RNA-specific tag or a first oligonucleotide; (b) generating single-stranded cDNA (sscDNA) and single-stranded DNA (ssDNA) from the cDNA-RNA duplex and the dsDNA, thereby generating a nucleic acid composition comprising sscDNA and ssDNA. ;( c) combining a nucleic acid composition comprising sscDNA and ssDNA with a second oligonucleotide, a plurality of first scaffold polynucleotide species, a third oligonucleotide, and a plurality of second scaffold polynucleotide species; Combination(i) each polynucleotide in the plurality of first scaffold polynucleotide species comprises a sscDNA hybridization region or a ssDNA hybridization region and a second oligonucleotide hybridization region; (ii) each polynucleotide in the plurality of second scaffold polynucleotide species comprises a ssDNA hybridization region and a third oligonucleotide hybridization region; and (iii) a nucleic acid composition comprising sscDNA and ssDNA, a second oligonucleotide, a plurality of first scaffold polynucleotide species, a third oligonucleotide, and a plurality of second scaffold polynucleotide species is hybridized such that molecules of the first scaffold polynucleotide species hybridize with (1) the first sscDNA terminal region or the first ssDNA terminal region and (2) molecules of the second oligonucleotide. R, This results in the formation of a hybridization product in which the end of the second oligonucleotide molecule is adjacent to the end of the first sscDNA terminal region or the first ssDNA terminal region. and A molecule of a second scaffold polynucleotide species is hybridized with (1) the second ssDNA terminal region and (2) a molecule of a third oligonucleotide, thereby forming a hybridization product in which the end of the molecule of the third oligonucleotide is adjacent to the end of the second ssDNA terminal region. 、 Under conditions Combination Also provided is a method comprising the steps of:
[0018] Also provided is a method for differentially amplifying nucleic acids according to source, comprising: (I) generating a nucleic acid library according to the methods described herein; (II) amplifying nucleic acid molecules of the library, wherein the amplifying step comprises contacting the nucleic acid molecules of the library with a first amplification primer and a second amplification primer under amplification conditions, wherein nucleic acids from the first source and nucleic acids from the second source are differentially amplified, thereby producing differentially amplified products.
[0019] Also provided are compositions comprising: a nucleic acid composition comprising single-stranded complementary deoxyribonucleic acid (sscDNA) and single-stranded deoxyribonucleic acid (ssDNA), wherein the sscDNA comprises an RNA-specific tag and a first oligonucleotide; a second oligonucleotide; a plurality of first scaffold polynucleotide species, each comprising an sscDNA hybridization region or an ssDNA hybridization region and a second oligonucleotide hybridization region; a third oligonucleotide; and a plurality of second scaffold polynucleotide species, each comprising an ssDNA hybridization region and a third oligonucleotide hybridization region.
[0020] a priming polynucleotide comprising a primer, an RNA-specific tag, and a first oligonucleotide; a second oligonucleotide; a plurality of first scaffold polynucleotide species, each comprising a sscDNA hybridization region or a ssDNA hybridization region and a second oligonucleotide hybridization region; a third oligonucleotide; and a plurality of second scaffold polynucleotide species, each comprising a ssDNA hybridization region and a third oligonucleotide hybridization region; Instructions for A kit is also provided, comprising:
[0021] A method for generating a nucleic acid library, comprising the steps of: (a) covalently linking single-stranded ribonucleic acid (ssRNA) in a first mixture containing ssRNA and double-stranded deoxyribonucleic acid (dsDNA) to a first oligonucleotide, thereby generating a covalently linked ssRNA product; (b) combining the covalently linked ssRNA product with a primer oligonucleotide and a nucleic acid library containing a nucleic acid library containing a primer oligonucleotide and ... agent(c) contacting the cDNA-RNA duplex and the dsDNA to form a second mixture containing complementary deoxyribonucleic acid (cDNA)-RNA duplex and dsDNA, whereby a second mixture containing the primer oligonucleotide is produced, wherein the primer oligonucleotide comprises a first oligonucleotide hybridization region; (c) producing single-stranded cDNA (sscDNA) and single-stranded DNA (ssDNA) from the cDNA-RNA duplex and dsDNA, whereby a nucleic acid composition containing sscDNA and ssDNA is produced; and (d) contacting the nucleic acid composition containing sscDNA and ssDNA with a second oligonucleotide, a plurality of first scaffold polynucleotide species, a third oligonucleotide, and a plurality of second scaffold polynucleotide species. Combination (i) each polynucleotide in the plurality of first scaffold polynucleotide species comprises a sscDNA hybridization region or a ssDNA hybridization region and a second oligonucleotide hybridization region; (ii) each polynucleotide in the plurality of second scaffold polynucleotide species comprises a ssDNA hybridization region and a third oligonucleotide hybridization region; and (iii) a nucleic acid composition comprising sscDNA and ssDNA, a second oligonucleotide, a plurality of first scaffold polynucleotide species, a third oligonucleotide, and a plurality of second scaffold polynucleotide species is hybridized such that molecules of the first scaffold polynucleotide species hybridize with (1) the first sscDNA terminal region or the first ssDNA terminal region and (2) molecules of the second oligonucleotide. R, This results in the formation of a hybridization product in which the end of the second oligonucleotide molecule is adjacent to the end of the first sscDNA terminal region or the first ssDNA terminal region. ,and, A molecule of a second scaffold polynucleotide species is hybridized with (1) the second ssDNA terminal region and (2) a molecule of a third oligonucleotide, thereby forming a hybridization product in which the end of the molecule of the third oligonucleotide is adjacent to the end of the second ssDNA terminal region. 、 Under conditions CombinationAlso provided is a method comprising the steps of:
[0022] a first oligonucleotide; a primer oligonucleotide comprising a first oligonucleotide hybridization region; a second oligonucleotide; a plurality of first scaffold polynucleotide species, each comprising a sscDNA hybridization region or a ssDNA hybridization region and a second oligonucleotide hybridization region; a third oligonucleotide; and a plurality of second scaffold polynucleotide species, each comprising a ssDNA hybridization region and a third oligonucleotide hybridization region; use Instructions for A kit is also provided, comprising:
[0023] Certain implementations are further described in the following description, examples and claims, as well as in the drawings.
[0024] The drawings illustrate certain implementations of the present technology. Illustrate It is not intended to be limiting. Example For clarity and ease of illustration, the drawings are not made to scale. In some cases , various aspects depend on the understanding of a particular implementation. promotion In order to emphasis Or it may be shown enlarged. [Brief explanation of the drawings]
[0025] [Figure 1] Figure 1 shows an example of a scaffold adapter construct that includes an in-line random UMI flanked by non-random sequences (i.e., a non-random anchor sequence and a P7 adapter sequence). [Figure 2] Figure 2 shows examples of scaffold adapter configurations containing in-line random UMIs flanked by non-random sequences (i.e., non-random anchor sequences and P7 adapter sequences) that use different anchor sequences and / or varying UMI lengths to increase UMI complexity. [Figure 3] FIG. 3 shows the final library construct configuration using the in-line random UMI scaffold adapters described herein compared to existing adapters. [Figure 4A] Figures 4A and 4B show a comparison of molecular performance between a library generated using a standard scaffold adapter library (non-UMI; "SOP") and a library generated using an in-line UMI scaffold adapter. Size and fragment length distributions are shown by electrophoresis (Figure 4A) and trace (Figure 4B; Tapestation 4200). [Figure 4B] Figures 4A and 4B show a comparison of molecular performance between a library generated using a standard scaffold adapter library (non-UMI; "SOP") and a library generated using an in-line UMI scaffold adapter. Size and fragment length distributions are shown by electrophoresis (Figure 4A) and trace (Figure 4B; Tapestation 4200). [Figure 5] FIG. 5 shows an example of a data trimming scheme. [Figure 6] Figure 6 shows an example of a scaffold adapter configuration that includes an in-line non-random UMI flanked by non-random sequences (i.e., a GC anchor sequence and a P7 adapter sequence). [Figure 7] FIG. 7 shows an example of a workflow for processing a sample containing a mixture of DNA and RNA, with initial first strand synthesis. [Figure 8] FIG. 8 shows an example of a workflow for processing a sample containing a mixture of DNA and RNA, with an initial ligation step. [Figure 9A] 9A and 9B show an example of a method for processing RNA, involving initial first strand synthesis. [Figure 9B] 9A and 9B show an example of a method for processing RNA, involving initial first strand synthesis. [Figure 10A] 10A and 10B show an example of a method for processing RNA, involving an initial ligation step. [Figure 10B]10A and 10B show an example of a method for processing RNA, involving an initial ligation step. [Figure 11] FIG. 11 shows a schematic diagram of the adapter used in the experiments described in Example 2. [Figure 12] FIG. 12 provides a summary of the results of the experiment described in Example 2. [Figure 13] Figure 13 provides general metrics of the results of the experiment described in Example 2. Specifically, the table shows the number of read pairs sequenced for each sample, the amount of methylated CG dinucleotides, the amount of methylated other (non-human epigenetic) motifs, the percent of overlapping reads, the percent of aligned reads, the representative insert size, the amount of adaptor-containing reads (trimmed), and the GC content of the reads. [Figure 14] Figure 14 shows insert size versus read fraction for the four experimental conditions described in Example 2 that generated libraries. Each trace is labeled 1–4 from left to right: 1) ZYMO EZ DNA METHYLATION LIGHTNING Kit (bisulfite treatment) before scaffold adapter ligation (non-methyl-protected adapters); 2) methyl-protected scaffold adapters ligated to DNA before NEB Enzymatic Methylation Kit; 3) methyl-protected dsDNA adapters ligated to DNA before NEB Enzymatic Methylation Kit; and 4) scaffold adapter ligation (non-methyl-protected adapters) before NEB Enzymatic Methylation Kit. The transient spike at 150 bp is an artifact of the sequencing read length for this run (2 × 151). [Figure 15]Figure 15 shows the PreSeq complexity (total molecules vs. unique molecules) for the four experimental conditions described in Example 2 under which libraries were generated. Each trace is labeled 1-4 from left to right: 1) ZYMO EZ DNA METHYLATION LIGHTNING Kit (bisulfite treatment) before scaffold adapter ligation (non-methyl-protected adapters); 2) methyl-protected scaffold adapters ligated to DNA before NEB Enzymatic Methylation Kit; 3) NEB Enzymatic Methylation Kit before scaffold adapter ligation (non-methyl-protected adapters); and 4) methyl-protected dsDNA adapters ligated to DNA before NEB Enzymatic Methylation Kit. [Figure 16] Figure 16 shows the GC distribution (GC content vs. fraction of reads) for the four experimental conditions described in Example 2 under which libraries were generated. Each trace is labeled 1-4: 1) methyl-protected scaffold adapters ligated to DNA before NEB Enzymatic Methylation Kit; 2) methyl-protected dsDNA adapters ligated to DNA before NEB Enzymatic Methylation Kit; 3) ZYMO EZ DNA METHYLATION LIGHTNING Kit (bisulfite treatment) before scaffold adapter ligation (non-methyl-protected adapters); and 4) NEB Enzymatic Methylation Kit before scaffold adapter ligation (non-methyl-protected adapters). [Figure 17] Figure 17 shows an example workflow. Downstream of RNase H-based rRNA depletion, cDNA is generated using tagged differential P5 random hexamers. After heat denaturation, scaffold adapters are added to the mix to tag DNA-specific reads, and P7 adapters are attached to both cDNA and DNA molecules. Index PCR completes the library molecules, allowing differential amplification based on the P5 adapter sequences. [Figure 18] Figures 18A-18C show performance metrics for simultaneous DNA:RNA libraries: Figure 18A shows mapping metrics, Figure 18B shows insert size, and Figure 18C shows gene body coverage. [Figure 19] Figure 19 shows an example of a workflow for processing RNA, with an initial ligation step. DETAILED DESCRIPTION OF THE INVENTION
[0026] Detailed Description Provided herein are methods and compositions useful for the analysis of nucleic acids. Also provided herein are methods and compositions useful for generating nucleic acid libraries. Also provided herein are methods and compositions useful for the analysis of single-stranded nucleic acid fragments. In certain aspects, the methods involve the step of: exclusive Adapter and Combination In some embodiments, the method further comprises: exclusive The adaptor comprises a unique molecular identifier (UMI). exclusive The adaptor comprises a scaffold polynucleotide that can hybridize to the ends of single-stranded nucleic acids. The products of such hybridization can be used, for example, in generating a nucleic acid library and / or further become It may be useful for analysis or processing.
[0027] Scaffolding Adapter Certain methods herein involve combining a single-stranded nucleic acid (ssNA) with a scaffold adaptor or a component thereof. Combination The scaffold adaptor generally comprises a scaffold polynucleotide and an oligonucleotide. Thus, a "component" of a scaffold adaptor is a scaffold polynucleotide and / or an oligonucleotide, or secondary It can refer to a component or area. Can The oligonucleotides and / or scaffold polynucleotides can be composed of pyrimidine (C, T, U) and / or purine (A, G) nucleotides. addition Ingredients of or secondary The components may include one or more of an index polynucleotide, a unique molecular identifier (UMI), one or more regions adjacent to the unique molecular identifier (UMI), a primer binding site (e.g., a sequencing primer binding site, a P5 primer binding site, a P7 primer binding site), a flow cell binding region, etc., and their complements. A scaffold adapter containing a P5 primer binding site may be referred to as a P5 adapter or a P5 scaffold adapter. A scaffold adapter containing a P7 primer binding site may be referred to as a P7 adapter or a P7 scaffold adapter.
[0028] A scaffold polynucleotide is a single-stranded component of a scaffold adaptor. A polynucleotide, as used herein, generally refers to a single-stranded multimer of nucleotides, ranging from 5 to 500 nucleotides, e.g., 5 to 100 nucleotides. Polynucleotides can be synthetic or enzymatically generated, and in some embodiments, are about 5 to 50 nucleotides in length. Polynucleotides contain ribonucleotide monomers (i.e., can be polyribonucleotides or "RNA polynucleotides"), deoxyribonucleotide monomers (i.e., can be polydeoxyribonucleotides or "DNA polynucleotides"), or a combination thereof. Possible The polynucleotide may be, for example, 10 to 20, 20 to 30, 30 to 40, 40 to 50, 50 to 60, 60 to 70, 70 to 80, 80 to 100, 100 to 150, 150 to 200 nucleotides in length, or longest 500 Nucleotides Do The terms polynucleotide and oligonucleotide are used herein to refer to interchangeably It can be used.
[0029] The scaffold polynucleotide may comprise an ssNA hybridization region (also referred to as a scaffold, a scaffold region, a single-stranded scaffold, or a single-stranded scaffold region) and an oligonucleotide hybridization region. The ssNA hybridization region and the oligonucleotide hybridization region are the same as those of the scaffold polynucleotide. secondary Ingredients and do The ssNA hybridization region typically hybridizes to the ssNA terminal region, or And The oligonucleotide hybridization region comprises a polynucleotide that can hybridize to all or a portion of the oligonucleotide components of the scaffold adaptor, or And It typically comprises a polynucleotide that is capable of hybridizing.
[0030] The ssNA hybridization region of the scaffold polynucleotide may comprise a polynucleotide that is complementary or substantially complementary to an ssNA terminal region (e.g., an ssDNA terminal region, an sscDNA terminal region, an ssRNA terminal region). In some embodiments, the ssNA hybridization region is an ssDNA hybridization region, an sscDNA hybridization region, or an ssRNA hybridization region. In some embodiments, the sscDNA hybridization region of the scaffold polynucleotide comprises a polynucleotide or a polynucleotide that is complementary or substantially complementary to an RNA-specific tag (e.g., an RNA-specific tag described herein). secondary In some embodiments, the ssRNA hybridization region of the scaffold polynucleotide comprises a polynucleotide or a nucleotide sequence that is complementary or substantially complementary to an RNA-specific tag (e.g., an RNA-specific tag described herein). secondary In some embodiments, the ssDNA hybridization region of the scaffold polynucleotide comprises a polynucleotide or a secondaryIn some embodiments, the ssNA hybridization region comprises a random sequence. In some embodiments, the ssNA hybridization region comprises a sequence complementary to the desired ssNA terminal region sequence (e.g., target sequence). In certain embodiments, the ssNA hybridization region comprises one or more nucleotides, all of which are capable of non-specific base pairing with the bases in the ssNA. Nucleotides capable of non-specific base pairing are sometimes referred to as universal bases. A universal base is a base that can indiscriminately base pair with each of the four standard nucleotide bases: A, C, G, and T. Universal bases that can be incorporated into the ssNA hybridization region include, but are not limited to, inosine, deoxyinosine, 2'-deoxyinosine (dI, dInosine), nitroindole, 5-nitroindole, and 3-nitropyrrole. In certain embodiments, the ssNA hybridization region contains one or more degenerate / wobble bases that can replace two or three (but not all) of the four typical bases (e.g., the unnatural bases P and K).
[0031] The ssNA hybridization region of the scaffold polynucleotide can have any suitable length and sequence. In some embodiments, the length of the ssNA hybridization region is 10 nucleotides or less. In certain aspects, the ssNA hybridization region is 4 to 100 nucleotides in length, for example, about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 nucleotides in length. In certain aspects, the ssNA hybridization region is 4 to 20 nucleotides in length, for example, 5 to 15, 5 to 10, 5 to 9, 5 to 8, or 5 to 7 (e.g., 6 or 7) nucleotides in length. In some embodiments, the ssNA hybridization region is 7 nucleotides in length. In some embodiments, the ssNA hybridization regions comprise or consist of random nucleotide sequences, and thus, when multiple heterogeneous scaffold polynucleotides with different random ssNA hybridization regions are used, the group However, they can act as scaffold polynucleotides for a heterogeneous population of ssNAs regardless of the sequence of the terminal regions of the ssNAs. Each scaffold polynucleotide with a unique ssNA hybridization region sequence may be referred to as a scaffold polynucleotide species. a group of Multiple scaffold polynucleotides The seeds , may be referred to as multiple scaffold polynucleotide species (e.g., in the case of a scaffold polynucleotide designed to have seven random bases within the ssNA hybridization region, multiple scaffold polynucleotide species may be 4 7 (The scaffold adapter species may be a scaffold adapter species having a unique scaffold polynucleotide (i.e., a unique ssNA hybridization region sequence). Thus, each scaffold adapter having a unique scaffold polynucleotide (i.e., a unique ssNA hybridization region sequence) may be referred to as a scaffold adapter species. a group of Multiple Scaffolding Adapters The seeds, sometimes referred to as multiple scaffold adapter species. A scaffold polynucleotide species generally has characteristics that are unique relative to other scaffold polynucleotide species. For example, a scaffold polynucleotide species can have unique sequence characteristics. The unique sequence characteristics can include a unique sequence length, a unique nucleotide sequence (e.g., a unique random sequence, a unique target sequence), or a combination of a unique sequence length and nucleotide sequence.
[0032] The scaffold polynucleotide may be one or more polynucleotides, including an index polynucleotide, a unique molecular identifier (UMI), one or more regions adjacent to a unique molecular identifier (UMI), a primer binding site (e.g., a P5 primer binding site, a P7 primer binding site), a flow cell binding region, or the like, or complementary polynucleotides thereof. Additional Subordinates The scaffold polynucleotide may comprise a primer binding site (or a polynucleotide complementary to the primer binding site). A scaffold polynucleotide (or its complement) comprising a P5 primer binding site may be referred to as a P5 scaffold or P5 scaffold polynucleotide. A scaffold polynucleotide (or its complement) comprising a P7 primer binding site may be referred to as a P7 scaffold or P7 scaffold polynucleotide.
[0033] An oligonucleotide can be an additional single-stranded component of a scaffold adaptor. An oligonucleotide herein generally refers to a single-stranded multimer of nucleotides, between 5 and 500 nucleotides, e.g., between 5 and 100 nucleotides. An oligonucleotide can be synthetic or enzymatically produced, and in some embodiments, is between 5 and 50 nucleotides in length. An oligonucleotide can contain ribonucleotide monomers (i.e., can be oligoribonucleotides or "RNA oligonucleotides"), deoxyribonucleotide monomers (i.e., can be oligodeoxyribonucleotides or "DNA oligonucleotides"), or a combination thereof. PossibleThe oligonucleotides may have a length of, for example, 10 to 20, 20 to 30, 30 to 40, 40 to 50, 50 to 60, 60 to 70, 70 to 80, 80 to 100, 100 to 150, or 150–200 nucleotides, or longest 500 Nucleotides Do The terms oligonucleotide and polynucleotide are used herein to refer to interchangeably It can be used.
[0034] The oligonucleotide component of the scaffold adaptor generally comprises a nucleic acid sequence that is complementary or substantially complementary to an oligonucleotide hybridization region of the scaffold polynucleotide. The oligonucleotide component of the scaffold adaptor may comprise one or more nucleic acid sequences useful for one or more downstream applications, such as, for example, PCR amplification of the ssNA fragment or its derivatives, sequencing of the ssNA fragment or its derivatives, etc. secondary In some embodiments, the oligonucleotide secondaryThe component is a sequencing adapter. The sequencing adapter may be compatible with a sequencing platform of interest, such as a sequencing platform provided by Illumina® (e.g., HiSeq™, MiSeq™, and / or Genome Analyzer™ sequencing systems); a sequencing platform provided by Oxford Nanopore™ Technologies (e.g., MinION™ sequencing system); a sequencing platform provided by Ion Torrent™ (e.g., Ion PGM™ and / or Ion Proton™ sequencing systems); a sequencing platform provided by Pacific Biosciences (e.g., Sequel or PACBIO RS II sequencing systems); a sequencing platform provided by Life Technologies™ (e.g., SOLiD™ sequencing system); a sequencing platform provided by Roche (e.g., 454 GS FLX+ and / or GS "Genapsys Junior sequencing system"); a sequencing platform offered by Genapsys; a sequencing platform offered by BGI; or any sequencing platform of interest. "Genapsys" generally refers to one or more nucleic acid domains comprising at least a portion of a nucleotide sequence (or its complement) used by Genapsys Junior sequencing system; a sequencing platform offered by Genapsys; a sequencing platform offered by BGI; or any sequencing platform of interest.
[0035] In some embodiments, the oligonucleotide component of the scaffold adapter comprises a domain (e.g., a "capture site" or "capture sequence") that specifically binds to a surface-bound sequencing platform oligonucleotide (e.g., a P5 or P7 oligonucleotide bound to the surface of a flow cell in an Illumina® sequencing system); a sequencing primer binding domain (e.g., a domain to which the Read 1 or Read 2 primer of an Illumina® platform can bind); a unique identifier or index (e.g., to enable sample multiplexing by marking every molecule from a given sample with a specific barcode or "tag"); Sequencing a barcode or other domain that uniquely identifies the sample source of the ssNA to be sequenced; a barcode sequencing primer binding domain (which binds the barcode) Sequencing a domain to which the primers used to bind; a molecule of interest, e.g., Sequencing The nucleic acid domain may be or include a molecular identification domain or unique molecular identifier (UMI) (e.g., a molecular index tag, such as a randomized tag of 4, 6, or other nucleotides); a complement of any such domain; or any combination thereof, for uniquely marking the nucleic acid to determine expression levels based on the number of instances of the unique tag. In some embodiments, the oligonucleotide includes one or more regions adjacent to the unique molecular identifier (UMI). In some embodiments, the barcode domain (e.g., sample index tag) and the molecular identification domain (e.g., molecular index tag; UMI) can be included in the same nucleic acid. Sequencing platform oligonucleotides, sequencing primers, and their corresponding binding domains can be designed to be compatible with various available sequencing platforms and technologies, including but not limited to those discussed herein.
[0036] When the oligonucleotide component of the scaffold adaptor comprises one or a portion of a sequencing adaptor, various approaches can be used to synthesize one or more Additional Sequencing adapters and / or their sequencing adapters The rest can be added. For example, Additional parts of the sequencing adapter and / or The rest The portion can be added by any one of ligation, reverse transcription, PCR amplification, etc. In the case of PCR, a 3' hybridization region (e.g., for hybridizing with an adapter region of an oligonucleotide) and Additional parts of the sequencing adapter and / or The rest a first amplification primer comprising a 5' region comprising a portion, as well as a 3' hybridization region (e.g., for hybridizing with an adapter region of a second oligonucleotide added to the opposite end of the ssNA molecule), and optionally Additional parts of the sequencing adapter and / or The rest An amplification primer pair can be utilized that includes a 5' region containing the portion and a second amplification primer that includes the portion.
[0037] The oligonucleotide component of the scaffold adaptor may comprise one or more RNA-specific or DNA-specific tags. Additional Subordinates The RNA-specific tag can mark RNA fragments in a sample (e.g., a sample containing a mixture of RNA and DNA fragments). The DNA-specific tag can mark DNA fragments in a sample (e.g., a sample containing a mixture of RNA and DNA fragments). Typically, when RNA-specific tags and DNA-specific tags are used in the same library preparation, the RNA-specific tag can be distinguished from the DNA-specific tag. For example, the RNA-specific tag and the DNA-specific tag can contain different sequences, or the RNA-specific tag and the DNA-specific tag can be of different lengths. Consists of The RNA-specific tag and the DNA-specific tag may comprise different detectable markers, or may be any combination thereof. orA DNA-specific tag can comprise from about 5 to about 15 nucleotides. In some embodiments, an RNA-specific tag can comprise 9 nucleotides. In some embodiments, a DNA-specific tag can comprise 9 nucleotides. nothing In some embodiments, the RNA-specific tag or DNA-specific tag is located at the end of the oligonucleotide component of the scaffold adapter. In some embodiments, the RNA-specific tag or DNA-specific tag is located at the 5' end of the oligonucleotide component of the scaffold adapter. In some embodiments, the RNA-specific tag or DNA-specific tag is located at the 3' end of the oligonucleotide component of the scaffold adapter. In some embodiments, the RNA-specific tag or DNA-specific tag is located at the end of the oligonucleotide component of the scaffold adapter such that when the scaffold adapter is hybridized to ssRNA or ssDNA, the RNA-specific tag or DNA-specific tag is adjacent to the end of the ssRNA terminal region or the end of the ssDNA terminal region.
[0038] The oligonucleotide component of the scaffold adaptor may comprise one or more polynucleotides, including an index polynucleotide, a unique molecular identifier (UMI), one or more regions adjacent to a unique molecular identifier (UMI), a primer binding site (e.g., a P5 primer binding site, a P7 primer binding site), a flow cell binding region, or a sequencing adaptor, or complementary polynucleotides thereof. Additional Subordinates The oligonucleotide may contain a primer binding site (or a polynucleotide complementary to the primer binding site). An oligonucleotide containing a P5 primer binding site (or its complement) may be referred to as a P5 oligo or P5 oligonucleotide. An oligonucleotide containing a P7 primer binding site (or its complement) may be referred to as a P7 oligo or P7 oligonucleotide.
[0039] The oligonucleotide component of the scaffold adapter may comprise a guanine and cytosine (GC)-rich region. The GC-rich region may comprise at least about 50% guanine and cytosine nucleotides. For example, the GC-rich region may comprise about 60% guanine and cytosine nucleotides, about 70% guanine and cytosine nucleotides, about 80% guanine and cytosine nucleotides, about 90% guanine and cytosine nucleotides, or 100% guanine and cytosine nucleotides. In some embodiments, the GC-rich region comprises about 70% guanine and cytosine nucleotides. The oligonucleotide component of the scaffold adapter may comprise a guanine and cytosine (GC)-rich region at one end (e.g., at the 3' end or the 5' end). In some embodiments, the oligonucleotide component of the scaffold adapter comprises a guanine and cytosine (GC)-rich region at the end of the oligonucleotide that is bound to the ssNA fragment (i.e., at the oligonucleotide-ssNA junction or "ligation end"). The scaffold polynucleotide may include a corresponding region that is complementary to a GC-rich region in the oligonucleotide.
[0040] The scaffold polynucleotide can be hybridized with an oligonucleotide to form a duplex within the scaffold adapter. Thus, the scaffold adapter may be referred to as a scaffold duplex, duplex adapter, duplex oligonucleotide, or duplex polynucleotide. Each scaffold duplex having a unique scaffold polynucleotide (i.e., containing a unique ssNA hybridization region sequence) may be referred to as a scaffold duplex species. a group of Multiple scaffold duplexes The seeds , sometimes referred to as multiple scaffold duplex species. In some embodiments, the scaffold polynucleotide and oligonucleotide are on separate DNA strands. In some embodiments, the scaffold polynucleotide and oligonucleotide are on a single DNA strand (e.g., a single DNA strand that can form a hairpin structure).
[0041] The scaffold adapter may comprise DNA, RNA, or a combination thereof. The scaffold adapter may comprise a DNA scaffold polynucleotide and a DNA oligonucleotide, a DNA scaffold polynucleotide and an RNA oligonucleotide, an RNA scaffold polynucleotide and a DNA oligonucleotide, or an RNA scaffold polynucleotide and an RNA oligonucleotide. In one configuration, the scaffold adapter is coupled to an RNA sample nucleic acid. combination Examples of ligases for use with such adapter / sample configurations include T4 RNA ligase 2, T4 DNA ligase, truncated T4 RNA ligase 2, and thermostable 5'App DNA / RNA ligase. In another example of an adapter configuration, the scaffold adapter is used to bind to the RNA sample nucleic acid. combination In another example of an adapter configuration, the scaffold adapter comprises a DNA scaffold polynucleotide and an RNA oligonucleotide for ligation with an RNA sample nucleic acid. Examples of ligases for use with such an adapter / sample configuration include T4 RNA ligase 1, T4 RNA ligase 2, truncated T4 RNA ligase 2, and thermostable 5'App DNA / RNA ligase. combination Examples of ligases for use with such adapter / sample constructs include RNA scaffold polynucleotides and RNA oligonucleotides for the purpose of ligating the adaptor / sample construct include T4 RNA ligase 1, T4 RNA ligase 2, truncated T4 RNA ligase 2, and thermostable 5'App DNA / RNA ligase. In some cases The adaptor nucleotide composition is selected to provide uniformity between the sample nucleic acid and the scaffold adaptor nucleic acid (eg, so that at least the oligonucleotides are uniform with the sample nucleic acid). In some cases The adaptor nucleotide composition is selected to provide uniformity between the oligonucleotide and the sample nucleic acid, and heterogeneity between the scaffold polynucleotide and the sample nucleic acid.
[0042] Unique Molecular Identifier (UMI) In some embodiments, the scaffold adaptor comprises a unique molecular identifier (UMI). In some embodiments, an oligonucleotide (e.g., an oligonucleotide component of a scaffold adaptor) comprises a unique molecular identifier (UMI). A unique molecular identifier (UMI), which may also be referred to as a molecular barcode, barcode, molecular identification domain, molecular index tag, sequence tag, and / or tag, generally identifies an input nucleic acid molecule (Multiple options possible) of identification UMIs are short sequences (e.g., about 3 to about 10 nucleotides in length) that can be added to nucleic acid fragments during nucleic acid library preparation to identify or mark them. In certain applications, UMIs can be used to identify or mark nucleic acid fragments, e.g., , solid Ari tags is sequenced It can be useful to uniquely mark molecules of interest to determine expression levels based on the number of instances. UMIs are usually added before an amplification step (e.g., PCR amplification), and can be useful, for example, to reduce errors and quantity bias introduced by amplification. Street A scaffold adaptor and / or an oligonucleotide component of a scaffold adaptor that includes a UMI may be said to contain an "in-line" UMI. An in-line UMI is a UMI that is present in the ssNA fragment ligated to the oligonucleotide component of the scaffold adaptor. Sequencing When a scaffold adapter contains an in-line UMI, library generation generally refers to a UMI sequence that is a component of a scaffold adapter and / or oligonucleotide described herein that becomes part of the sequence reads generated by: Certain additional No processing steps (e.g., addition of a UMI to the adaptor by an extension step using a strand-displacing polymerase) may be required.
[0043] In some embodiments, a UMI comprises a random sequence. In some embodiments, a UMI comprises a non-random sequence. In some embodiments, a UMI comprises one or more universal bases. In some embodiments, a UMI consists of a random sequence. In some embodiments, a UMI consists of a non-random sequence. In some embodiments, a UMI consists of universal bases. A UMI can be of any suitable length. In some embodiments, a UMI comprises between 3 and 10 nucleotides. For example, a UMI can comprise 3 nucleotides, 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, or 10 nucleotides. In some embodiments, a UMI comprises 5 nucleotides. In some embodiments, a UMI comprises 5 random nucleotides. In some embodiments, a UMI comprises 5 non-random nucleotides. In some embodiments, a UMI comprises 5 universal bases.
[0044] In some embodiments, an oligonucleotide (e.g., an oligonucleotide component of a scaffold adaptor) comprises one or two flanking regions. sandwiched between The adjacent region contains a unique molecular identifier (UMI). sandwiched between A UMI is usually adjacent to two adjacent regions. sandwiched between The UMIs are usually adjacent to each flanking region, and the UMIs are located between two flanking regions. The flanking regions, also called anchor sequences, are the regions where the complex is formed. RThe flanking region may be located at the oligonucleotide end adjacent to the ssNA end (i.e., adjacent to the oligonucleotide-ssNA junction or "ligation end") in some cases. The flanking region generally comprises a non-random sequence. In some embodiments, the flanking region comprises a non-random sequence species from a pool of non-random sequence species. In some embodiments, the pool of non-random sequence species comprises two or more non-random sequence species. In some embodiments, the pool of non-random sequence species comprises three or more non-random sequence species. In some embodiments, the pool of non-random sequence species comprises four or more non-random sequence species. In some embodiments, the pool of non-random sequence species comprises five or more non-random sequence species. In some embodiments, the pool of non-random sequence species comprises six or more non-random sequence species. In some embodiments, the pool of non-random sequence species comprises four non-random sequence species. The flanking region can be of any suitable length. In some embodiments, the flanking region comprises between 8 and 15 nucleotides. For example, the flanking region can comprise 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides. In some embodiments, the flanking region comprises 10 nucleotides. The combination of a UMI sequence (e.g., 5 random bases) with a particular flanking sequence species (e.g., 10 non-random bases from a pool of four possible flanking sequence species) can serve as a molecular identifier and can be considered a "UMI."
[0045] The flanking region can be designed to have a suitable melting temperature (Tm). As described herein, melting temperature generally refers to the temperature at which half of the flanking region / [polynucleotide complementary to the flanking region] remains hybridized and half of the flanking region / [polynucleotide complementary to the flanking region] dissociates into single-stranded strands. A suitable melting temperature can be higher than the temperature at which a ligation reaction (e.g., a ligation reaction described herein) is performed. For example, if the ligation reaction is performed at 37°C, the suitable melting temperature for the flanking region is higher than 37°C. If the ligation reaction is performed at 16°C, the suitable melting temperature is higher than 16°C. In some embodiments, the suitable melting temperature is equal to or higher than about 37°C. For example, a suitable melting temperature can be equal to or greater than about 38° C., 39° C., 40° C., 41° C., 42° C., 43° C., 44° C., 45° C., 46° C., 47° C., 48° C., 49° C., or 50° C. In some embodiments, a suitable melting temperature is equal to or greater than about 38° C. In some embodiments, a suitable melting temperature is equal to or greater than about 45° C.
[0046] In certain configurations, the flanking region can be designed to be of sufficient length to have sufficient guanine and cytosine content and / or to contain one or more modified nucleotides (e.g., locked nucleic acid (LNA) bases) to have a suitable melting temperature (Tm). Generally, increasing the flanking region length can compensate for a lower GC content, and increasing the GC content can compensate for a shorter flanking region (i.e., to obtain a flanking region with a suitable Tm). For example, the flanking region can contain 10 nucleotides, in which case 70% of these nucleotides are guanine or cytosine for a Tm greater than 45°C. In another example, the flanking region can contain 18 nucleotides, in which case 50% of these nucleotides are guanine or cytosine for a Tm greater than 45°C. In the above example, the flanking region can be shorter and / or have a lower GC content if one or more modified nucleotides (e.g., LNA bases) that increase the Tm are included in the flanking region.
[0047] The flanking region may be guanine and cytosine (GC) rich. The GC-rich flanking region may contain at least about 50% guanine and cytosine nucleotides. For example, the GC-rich flanking region may contain about 60% guanine and cytosine nucleotides, about 70% guanine and cytosine nucleotides, about 80% guanine and cytosine nucleotides, about 90% guanine and cytosine nucleotides, or 100% guanine and cytosine nucleotides. In some embodiments, the GC-rich flanking region contains about 70% guanine and cytosine nucleotides. In some embodiments, the flanking region contains about 90% guanine and cytosine nucleotides. In some embodiments, the flanking region contains about 90% guanine and cytosine nucleotides and has a Tm of about 38°C. In some embodiments, the flanking region contains the following polynucleotide sequence: GGCCCGACGG.
[0048] The oligonucleotide may contain an additional flanking region. The additional flanking region may be located distal to the oligonucleotide end adjacent to the ssNA end when the complex is formed (i.e., distal to the oligonucleotide-ssNA junction or "ligation end"). The additional flanking region generally comprises a non-random sequence. The additional flanking region may include any of the features of the flanking region or anchor sequence described herein. In some configurations, the additional flanking region may be located distal to one or more of the oligonucleotide components of the scaffold adaptor. Additional Subordinates For example, the additional flanking region may include one or more of a primer binding domain, a sequencing adaptor, or a portion thereof, and an index (e.g., a sample identification index).
[0049] In some embodiments, the oligonucleotide comprises, starting from the oligonucleotide-ssNA junction end, a flanking region, followed by a UMI, followed by a further flanking region.In some embodiments, the oligonucleotide comprises, starting from the oligonucleotide-ssNA junction end, a non-random flanking region, followed by a random UMI, followed by a further non-random flanking region.In some embodiments, the oligonucleotide comprises, starting from the oligonucleotide-ssNA junction end, a non-random flanking region, followed by a non-random UMI, followed by a further non-random flanking region.
[0050] In some embodiments, the scaffold polynucleotide comprises an oligonucleotide hybridization region comprising a polynucleotide complementary to an adjacent region in the oligonucleotide. In some embodiments, the scaffold polynucleotide comprises an oligonucleotide hybridization region comprising a polynucleotide complementary to an adjacent region in the oligonucleotide and a polynucleotide complementary to a further adjacent region in the oligonucleotide. In some embodiments, the scaffold polynucleotide comprises an oligonucleotide hybridization region comprising a region corresponding to a UMI in the oligonucleotide. The region corresponding to a UMI in the oligonucleotide may comprise a sequence complementary to the UMI, or may comprise a sequence that is not complementary to the UMI. When the oligonucleotide comprises a random UMI sequence, the region corresponding to the UMI may also comprise a random sequence, and therefore the UMI and the region corresponding to the UMI are generally not complementary. The random UMI sequence and the region corresponding to the UMI may contain the same number of nucleotides or may contain different numbers of nucleotides. When the oligonucleotide comprises a non-random UMI sequence, the region corresponding to the UMI may also comprise a non-random sequence, and the UMI and the region corresponding to the UMI are designed to be complementary. When an oligonucleotide includes a UMI that includes a universal base, the region corresponding to the UMI may also include a universal base. In some embodiments, the scaffold polynucleotide includes a polynucleotide that is complementary to a flanking region in the oligonucleotide and a polynucleotide that is complementary to a further flanking region in the oligonucleotide. sandwiched between The region corresponds to the UMI in the oligonucleotide.
[0051] have a unique UMI configuration (i.e., a unique UMI sequence and / or specific flanking sequence species) Combination Each oligonucleotide (containing a unique UMI sequence) may be referred to as an oligonucleotide species, a group of Multiple oligonucleotides The seeds, may be referred to as multiple oligonucleotide species (e.g., for an oligonucleotide designed with five random base UMIs, multiple oligonucleotide species may be referred to as 4 5 Thus, it has a unique oligonucleotide (i.e., a unique UMI sequence and / or a specific flanking sequence species). Combination Each scaffold adapter having a unique UMI sequence (i.e., a unique ssNA hybridization region sequence) and / or a unique scaffold polynucleotide (i.e., a unique ssNA hybridization region sequence) may be referred to as a scaffold adapter species; a group of Multiple Scaffolding Adapters The seeds , sometimes referred to as multiple scaffold adapter species. An oligonucleotide species generally has characteristics that are unique to other oligonucleotide species. For example, an oligonucleotide species may have a unique sequence characteristic. The unique sequence characteristic may include a unique sequence length, a unique nucleotide sequence (e.g., a unique random sequence), or a combination of a unique sequence length and nucleotide sequence.
[0052] Interaction of scaffold adaptors or their components with ssNAs combination The methods herein include combining one or more scaffold adaptors or components thereof with a composition comprising single-stranded nucleic acids (ssNAs). Combination The scaffold polynucleotide may be used in complex formation. When The oligonucleotide component is designed for simultaneous hybridization with the ssNA fragment and the oligonucleotide component, such that the end of the oligonucleotide component is adjacent to the end of the terminal region of the ssNA fragment. When The 5' end of the oligonucleotide component is adjacent to the 3' end of the terminal region of the ssNA, or the 5' end of the oligonucleotide component is adjacent to the 3' end of the terminal region of the ssNA. WhenIn cases where scaffold adaptors are attached to both ends of the ssNA fragment, the 5' end of one oligonucleotide component is adjacent to the 3' end of one terminal region of the ssNA, and the 5' end of the second oligonucleotide component is adjacent to the 3' end of the second terminal region of the ssNA.
[0053] In some embodiments, the method comprises combining an ssNA composition, an oligonucleotide, Not yet and a plurality of heterogeneous scaffold polynucleotides having various random ssNA hybridization regions that can act as scaffolds for a heterogeneous population of ssNAs having terminal regions of defined sequences. Combination In some embodiments, the method comprises combining an ssNA composition with a plurality of heterogeneous oligonucleotides having different UMI configurations to form a complex. Not yet and a plurality of heterogeneous scaffold polynucleotides having various random ssNA hybridization regions that can act as scaffolds for a heterogeneous population of ssNAs having terminal regions of defined sequences. Combination In some embodiments, the method comprises combining an ssNA composition with an oligonucleotide or a plurality of heterogeneous oligonucleotides having different UMI configurations, and a plurality of heterogeneous scaffold polynucleotides, the scaffold polynucleotides being provided in an amount that exceeds the amount of the oligonucleotides, to form a complex. Combination In some embodiments, the scaffold polynucleotide and oligonucleotide are provided in a ratio of at least 1.1 to 1 (scaffold polynucleotide to oligonucleotide). For example, the scaffold polynucleotide and oligonucleotide may be provided in a ratio of at least 1.2 to 1, 1.3 to 1, 1.4 to 1, 1.5 to 1, 1.6 to 1, 1.7 to 1, 1.8 to 1, 1.9 to 1, or 2 to 1. In some embodiments, the scaffold polynucleotide and oligonucleotide are provided in a ratio of 1.4 to 1 (scaffold polynucleotide to oligonucleotide). For example, the method may involve combining the ssNA composition with 14 μM of scaffold polynucleotide and 10 μM of oligonucleotide. Combination The method may include a step of:
[0054] In some embodiments, the ssNA hybridization region comprises a known sequence designed to hybridize with the ssNA terminal region of a known sequence. In some embodiments, two or more heterogeneous scaffold polynucleotides having different ssNA hybridization regions of known sequences are designed to hybridize with each ssNA terminal region of a known sequence. Embodiments in which the ssNA hybridization region has a known sequence can be useful, for example, for generating a nucleic acid library from a subset of ssNAs having terminal regions of known sequences. Thus, in certain embodiments, the methods herein involve combining an ssNA composition, an oligonucleotide, and one or more heterogeneous scaffold polynucleotides having one or more different ssNA hybridization regions of known sequences that can act as a scaffold for one or more ssNAs having one or more terminal regions of known sequences. Combination The method includes a step of forming a complex by allowing the mixture to react with the catalyst.
[0055] The ssNA fragments, oligonucleotides and scaffold polynucleotides can be synthesized by various methods. Combination In some configurations, Combination The hybridization step comprises hybridizing 1) a complex comprising a scaffold polynucleotide hybridized to an oligonucleotide component by an oligonucleotide hybridization region, and 2) a ssNA fragment. Combination In another configuration, Combination The hybridization step comprises: 1) a complex comprising a scaffold polynucleotide hybridized to an ssNA fragment by an ssNA hybridization region; and 2) an oligonucleotide component. Combination In another configuration, Combination The process involves combining 1) ssNA fragments, 2) oligonucleotides, and 3) scaffold polynucleotides. Combination In this case, all three components are Combination It is not pre-complexed or hybridized with another component prior to being mixed with the antibody.
[0056] The binding can be carried out under hybridization conditions that allow the formation of a complex comprising the scaffold polynucleotide hybridized to the terminal region of the ssNA fragment via the ssNA hybridization region and the scaffold polynucleotide hybridized to the oligonucleotide component via the oligonucleotide hybridization region.Whether specific hybridization occurs can be determined by factors such as the degree of complementarity between the hybridization regions of the scaffold polynucleotide, the terminal region of the ssNA fragment, and the oligonucleotide component, as well as their lengths, salt concentration, GC content, and the temperature at which hybridization occurs, and the melting temperature (Tm) of the relevant region can provide information on the temperature at which hybridization occurs.
[0057] The complex can be formed such that the ends of the oligonucleotide components are adjacent to the ends of the terminal regions of the ssNA fragment. .next "Adjacent" refers to the terminal nucleotide of the end of the oligonucleotide and the terminal nucleotide end of the terminal region of the ssNA fragment being close enough to each other that these terminal nucleotides can be covalently linked, for example, by chemical ligation, enzymatic ligation, etc. In some embodiments, the ends are adjacent to each other because the terminal nucleotide of the end of the oligonucleotide and the terminal nucleotide end of the terminal region of the ssNA are hybridized to adjacent nucleotides of the scaffold polynucleotide. The scaffold polynucleotide can be designed to ensure that the end of the oligonucleotide is adjacent to the end of the terminal region of the ssNA fragment.
[0058] In some embodiments, the complex can be formed such that the end of the RNA-specific tag in the oligonucleotide component is adjacent to the end of the terminal region of the ssRNA fragment. .nextAdjacent refers to the terminal nucleotide of the end of the RNA-specific tag and the terminal nucleotide end of the terminal region of the ssRNA fragment are close enough to each other so that these terminal nucleotides can be covalently linked, for example, by chemical ligation, enzymatic ligation, etc. In some embodiments, the ends are adjacent to each other because the terminal nucleotide of the end of the RNA-specific tag and the terminal nucleotide end of the terminal region of the ssRNA are hybridized with the adjacent nucleotide of the scaffold polynucleotide.The scaffold polynucleotide can be designed so that the end of the RNA-specific tag is adjacent to the end of the terminal region of the ssRNA fragment.
[0059] In some embodiments, the complex can be formed such that the end of the DNA-specific tag in the oligonucleotide component is adjacent to the end of the terminal region of the ssDNA fragment. .next "Adjacent" refers to the terminal nucleotide of the end of the DNA-specific tag and the terminal nucleotide end of the terminal region of the ssDNA fragment being close enough to each other that these terminal nucleotides can be covalently linked, for example, by chemical ligation, enzymatic ligation, etc. In some embodiments, the ends are adjacent to each other because the terminal nucleotide of the end of the DNA-specific tag and the terminal nucleotide end of the terminal region of the ssDNA are hybridized with adjacent nucleotides of the scaffold polynucleotide. The scaffold polynucleotide can be designed to ensure that the end of the DNA-specific tag is adjacent to the end of the terminal region of the ssDNA fragment.
[0060] The scaffold polynucleotide can be designed with one or more uracil bases instead of thymine. In some embodiments, one of the strands in the scaffold adaptor duplex can be degraded by generating multiple cleavage sites at the uracil bases, for example, by using uracil-DNA glycosylase and an endonuclease.
[0061] Scaffold adapters containing the inline UMI designs described herein can be configured to be attached to one or both ends of ssNA fragments. In some configurations, scaffold adapters are designed so that the adapter species attached to the 5' end of the ssNA contains the inline UMI designs described herein. In some configurations, scaffold adapters are designed so that the adapter species attached to the 3' end of the ssNA contains the inline UMI designs described herein. In some configurations, scaffold adapters are designed so that the adapter species attached to the 5' end of the ssNA contains the inline UMI designs described herein, and the adapter species attached to the 3' end of the ssNA does not contain an inline UMI. In some configurations, scaffold adapters are designed so that the adapter species attached to the 3' end of the ssNA contains the inline UMI designs described herein, and the adapter species attached to the 5' end of the ssNA does not contain an inline UMI. In some configurations, the scaffold adapter is designed such that the adapter species attached to the 5' end of the ssNA includes an in-line UMI design described herein, and such that the adapter species attached to the 3' end of the ssNA also includes an in-line UMI design described herein.
[0062] The scaffold adaptor, oligonucleotide component, and scaffold polynucleotide may be referred to herein as a first scaffold adaptor (or first scaffold duplex), a first oligonucleotide component (or first oligonucleotide), a first unique molecular identifier (UMI), and a first scaffold polynucleotide; or a second scaffold adaptor (or second scaffold duplex), a second oligonucleotide component (or second oligonucleotide), a second unique molecular identifier (UMI), and a second scaffold polynucleotide. The terms first and second generally refer to the scaffold adaptor, or a component thereof, hybridized to and / or covalently linked to the first and second ends (i.e., the 5' and 3' ends) of the ssNA fragment ends. The terms first and second ends do not necessarily refer to a particular orientation of the ssNA fragment. Thus, the first end of the ssNA terminus can be the 5' end or the 3' end, and the second end of the ssNA terminus can be the 5' end or the 3' end. The first scaffold adaptor, or a component thereof, can refer to a P5 adaptor, or a component thereof, or a P7 adaptor, or a component thereof. The second scaffold adaptor, or a component thereof, can refer to a P5 adaptor, or a component thereof, or a P7 adaptor, or a component thereof.
[0063] In some cases, a scaffold adapter, an oligonucleotide component, and a scaffold polynucleotide may be referred to herein as (i) a first scaffold adapter (or first scaffold duplex), a first oligonucleotide component (or first oligonucleotide), and a first scaffold polynucleotide; (ii) a second scaffold adapter (or second scaffold duplex), a second oligonucleotide component (or second oligonucleotide), and a second scaffold polynucleotide; (iii) a third scaffold adapter (or third scaffold duplex), a third oligonucleotide component (or third oligonucleotide), and a third scaffold polynucleotide; or (iv) a fourth scaffold adapter (or fourth scaffold duplex), a fourth oligonucleotide component (or fourth oligonucleotide), and a fourth scaffold polynucleotide. case (e.g., the scaffold adaptor or its components may be used with a mixture of ssRNA and ssDNA. Combination When used herein, the terms first and second generally refer to a scaffold adaptor, or a component thereof, that hybridizes to and / or is covalently linked to the first ends (i.e., 5' and 3' ends) of the ssRNA fragments and the first ends (i.e., 5' and 3' ends) of the ssDNA fragments, respectively. The terms third and fourth generally refer to a scaffold adaptor, or a component thereof, that hybridizes to and / or is covalently linked to the second ends (i.e., 5' and 3' ends) of the ssRNA fragments and the second ends (i.e., 5' and 3' ends) of the ssDNA fragments, respectively.
[0064] The regions flanking the first unique molecular identifier (UMI) may be referred to as the first flanking region and the second flanking region. The first flanking region is a region where the complex is formed. R The term generally refers to a region in a first oligonucleotide that is proximal to the end of the oligonucleotide that is adjacent to the ssNA end (i.e., adjacent to the oligonucleotide-ssNA junction or "ligation end") when the complex is formed. R The term generally refers to a region in the first oligonucleotide that is distal to the end of the oligonucleotide adjacent to the ssNA end in the case of a second unique molecular identifier (UMI). The regions adjacent to the second unique molecular identifier (UMI) are sometimes referred to as the third and fourth flanking regions. The third flanking region is a region adjacent to the second unique molecular identifier (UMI) when the complex is formed. R The term generally refers to a region in the second oligonucleotide that is proximal to the end of the oligonucleotide that is adjacent to the ssNA end (i.e., adjacent to the oligonucleotide-ssNA junction or "ligation end") when the complex is formed. R In this specification, the term "flanking region" generally refers to the region in the second oligonucleotide that is distal to the oligonucleotide end adjacent to the ssNA end.The terms "first flanking region," "second flanking region," "third flanking region," and "fourth flanking region" do not necessarily refer to a specific orientation of the components in the oligonucleotide.The first flanking region and the third flanking region may be referred to herein as flanking region or anchor sequence.The second flanking region and the fourth flanking region may be referred to herein as further flanking region.
[0065] Depending on the situation The method comprises combining a scaffold adaptor or a component thereof with a nucleic acid sample containing ssNA. Combination Prior to cleaving, the nucleic acid sample can be treated with a nuclease to remove unwanted nucleic acids. For example, as disclosed herein, a double-strand specific nuclease (e.g., T7 nuclease) can be used to digest some or all double-stranded DNA, followed by the addition of scaffold adapters. (scaffolding adapter) A sequencing library of the remaining nucleic acids can be prepared using double-stranded nucleases. In one example, double-stranded-specific nucleases are used to digest double-stranded nucleic acids in a sample, digesting double-stranded DNA from host organisms and / or bacteria but leaving intact single-stranded nucleic acids, such as those from single-stranded DNA viruses, single-stranded RNA viruses, and single-stranded DNA (e.g., damaged DNA).
[0066] Interaction of scaffold adaptors or their components with ssRNA and / or sscDNA combination The methods herein include combining one or more scaffold adaptors or components thereof with a composition comprising single-stranded ribonucleic acid (ssRNA) and / or single-stranded complementary deoxyribonucleic acid (sscDNA) to form one or more complexes. Combination The scaffold polynucleotide may comprise a step of: Described in As shown, complex formation When The oligonucleotide moiety is designed for simultaneous hybridization with the ssRNA or sscDNA fragment and the oligonucleotide moiety such that the termini of the oligonucleotide moiety are adjacent to the termini of the terminal regions of the ssRNA or sscDNA fragment.
[0067] In some embodiments, the nucleic acid composition comprises sscDNA. Combination The method includes generating sscDNA from single-stranded ribonucleic acid (ssRNA) prior to the step of generating sscDNA. Typically, when the nucleic acid composition includes sscDNA, the methods herein use first-strand cDNA and do not require the generation of second-strand cDNA. Thus, in some embodiments, the nucleic acid composition includes first-strand sscDNA. In some embodiments, the nucleic acid composition consists essentially of first-strand sscDNA. A nucleic acid composition "consisting essentially of" first-strand sscDNA generally includes first-strand sscDNA, Additional It does not contain any protein or nucleic acid components. A nucleic acid composition consisting essentially of first-strand sscDNA generally does not contain second-strand sscDNA. In addition, for example, a nucleic acid composition "consisting essentially of" first-strand sscDNA may also contain double-stranded cDNA (dscDNA). Can be excluded A nucleic acid composition "consisting essentially of" first strand sscDNA may contain a low percentage of dscDNA (e.g., less than 10% dscDNA, less than 5% dscDNA, less than 1% dscDNA). Can be excluded For example, a nucleic acid composition "consisting essentially of" first-strand sscDNA does not contain single-strand binding protein (SSB) or other proteins useful for stabilizing first-strand sscDNA. Can be excludedA nucleic acid composition that "consists essentially of" first strand sscDNA does not contain chemical components that are normally present in nucleic acid compositions, e.g., buffer , salts, alcohol, crowding agents (e.g., PEG), etc., and may include residual components (e.g., nucleic acids (e.g., residual RNA), proteins, cell membrane components) from the nucleic acid source (e.g., sample), from nucleic acid extraction, or from cDNA synthesis. A nucleic acid composition "consisting essentially of" first-strand sscDNA may include first-strand sscDNA fragments having one or more phosphates (e.g., terminal phosphate, 5'-terminal phosphate). A nucleic acid composition "consisting essentially of" first-strand sscDNA may include first-strand sscDNA fragments containing one or more modified nucleotides.
[0068] In some embodiments, the step of generating sscDNA comprises using ssRNA, a primer, and a nucleic acid having reverse transcriptase activity. agent In some embodiments, the step of generating sscDNA comprises contacting the DNA-RNA duplex with a nucleic acid having RNAse activity. agent To bring into contact with moreover Contains Miteku In some embodiments, the RNA may be digested to generate a sscDNA product. agent is a reverse transcriptase or RNA-dependent DNA polymerase (i.e., an enzyme used to generate complementary DNA (cDNA) from an RNA template by reverse transcription). Examples of reverse transcriptases include HIV-1 reverse transcriptase, M-MLV reverse transcriptase, and AMV reverse transcriptase. In some embodiments, a nucleotide sequence having reverse transcriptase activity is agent also has RNAse activity. Thus, in some embodiments, reverse transcription and RNAse digestion are combined in one step. In some embodiments, agent is M-MuLV reverse transcriptase (also called M-MLV reverse transcriptase).
[0069] The primer(s), sometimes referred to as primer oligonucleotides, may be any primer(s) suitable for use with reverse transcriptase. can be mentionedThe primer(s) can be selected from one or more of a random primer (e.g., a random n-mer, a random hexamer primer, a random octamer primer), and a poly(T) primer. The sscDNA product can be purified by a suitable purification or washing method, such as those described herein. In some embodiments, the primer oligonucleotide comprises a priming region and an RNA-specific tag. In some embodiments, the primer may be referred to as a priming polynucleotide. The priming polynucleotide may comprise a primer, an RNA-specific tag, and an oligonucleotide (e.g., a sequencing adapter or portion thereof; an amplification priming site). The RNA-specific tag may comprise about 5 to about 15 nucleotides. In some embodiments, the RNA-specific tag comprises 9 nucleotides. In some embodiments, the RNA-specific tag is located at the end of the primer oligonucleotide. In some embodiments, the RNA-specific tag is located at the 5' end of the primer oligonucleotide. The priming region in the primer oligonucleotide may comprise a sequence that hybridizes to an RNA fragment. The priming region in the primer oligonucleotide may comprise a sequence that hybridizes to an end region of the RNA fragment. The priming region in the primer oligonucleotide may comprise a sequence that hybridizes with the 3'-end region of the RNA fragment. The priming region may comprise a random primer (e.g., a random n-mer, a random hexamer primer, or a random octamer primer). In some embodiments, the priming region hybridizes with the RNA fragment, and the RNA-specific tag does not hybridize with the RNA fragment. Thus, in some embodiments, the method herein comprises generating a single-stranded cDNA (sscDNA) comprising a sequence complementary to the RNA fragment and an additional sequence comprising the RNA-specific tag. In some embodiments, the RNA-specific tag is located at the end of the sscDNA. In some embodiments, the RNA-specific tag is located at the 5'-end of the sscDNA.In the case of a nucleic acid composition comprising a mixture of ssRNA and dsDNA, the RNA-specific tag may be attached to the cDNA derived from the ssRNA and to any of the dsDNA. chain In the case of a nucleic acid composition comprising a mixture of cDNA and dsDNA, the cDNA may comprise an RNA-specific tag and the dsDNA may not comprise an RNA-specific tag. In the case of a nucleic acid composition comprising a mixture of sscDNA and ssDNA, the sscDNA may comprise an RNA-specific tag and the ssDNA may not comprise an RNA-specific tag.
[0070] In some embodiments, the nucleic acid composition comprises a mixture of single-stranded complementary deoxyribonucleic acid (sscDNA) and single-stranded deoxyribonucleic acid (ssDNA). In some embodiments, the sscDNA is derived from a cDNA-RNA duplex (e.g., Described in In some embodiments, ssDNA includes, but is not limited to, sscDNA (produced by reverse transcription as described herein). For example, sscDNA can be obtained from a cDNA-RNA duplex that is denatured (e.g., heat-denatured and / or chemically denatured) or subjected to RNAse treatment to produce sscDNA. In some embodiments, ssDNA includes, but is not limited to, ssDNA derived from double-stranded DNA (dsDNA). For example, ssDNA can be obtained from double-stranded DNA that is denatured (e.g., heat-denatured and / or chemically denatured) to produce ssDNA. In some embodiments, the methods herein involve combining sscDNA and ssDNA with a scaffold adapter, or a component thereof, as described herein. Combination The method may further comprise generating sscDNA from the cDNA-RNA duplex and generating ssDNA from the dsDNA prior to the denaturing step. In some embodiments, the sscDNA and ssDNA may be generated by denaturing the cDNA-RNA duplex and the dsDNA.
[0071] In some embodiments, the nucleic acid composition comprises ssRNA. In such embodiments, the scaffold adaptor can hybridize directly to the ssRNA fragment, and the oligonucleotide component (Multiple options possible) are covalently linked to one or more ends of the ssRNA termini, thereby forming a hybridization product containing one or more scaffold adapters and the ssRNA fragments. In some embodiments, the method further comprises generating a single-stranded ligation product from the hybridization product (e.g., by denaturing the hybridization product). In such embodiments, the single-stranded ligation product comprises the ssRNA fragments covalently linked to one or more oligonucleotide components. In some embodiments, the method further comprises denaturing the single-stranded ligation product with a primer and a nucleotide sequence having reverse transcriptase activity. agent In some embodiments, the method further comprises contacting the DNA-RNA duplex with a nucleic acid having RNAse activity. agent In some embodiments, the method further comprises contacting a cDNA having reverse transcriptase activity with a cDNA having reverse transcriptase activity, thereby digesting the RNA and generating a single-stranded cDNA (sscDNA) product. agent also has RNAse activity. Thus, in some embodiments, reverse transcription and RNAse digestion are combined in one step. In some embodiments, agent is M-MuLV reverse transcriptase (also called M-MLV reverse transcriptase). The primer can be any primer suitable for use with a reverse transcriptase. In some embodiments, the primer comprises a nucleotide sequence complementary to a sequence in the oligonucleotide component (i.e., the oligonucleotide component covalently linked to the ssRNA fragment). The sscDNA product can be purified by a suitable purification or washing method, for example, a purification or washing method described herein.
[0072] In some embodiments, the oligonucleotide can be covalently linked to the ssRNA (without prior hybridization to a scaffold adapter). The covalently linked ssRNA product can be synthesized by hybridizing the oligonucleotide with a primer oligonucleotide and a nucleotide sequence having reverse transcriptase activity, as described herein. agent The oligonucleotide can be contacted with ssRNA and the oligonucleotide to generate cDNA. The primer oligonucleotide typically comprises an oligonucleotide hybridization region. The oligonucleotide may comprise RNA or may consist of RNA. In some embodiments, the oligonucleotide comprises an RNA-specific tag. In some embodiments, the oligonucleotide comprises a sequencing adapter, or a portion thereof, or a primer binding site. In some embodiments, the oligonucleotide can be contacted with ssRNA and the oligonucleotide and one or more ligases having ligase activity. agent and one or more oligonucleotides having ligase activity, which are covalently ligated to the ssRNA by contacting the oligonucleotide with the ssRNA terminal region under conditions in which the end of the ssRNA terminal region is covalently ligated to the end of the oligonucleotide. agent may be selected from, for example, T4 RNA ligase 1, T4 RNA ligase 2, truncated T4 RNA ligase 2, and thermostable 5'App DNA / RNA ligase.
[0073] In some embodiments, the sscDNA product is amplified. The sscDNA product can be amplified by a suitable amplification method, such as the amplification methods described herein. In some embodiments, amplifying the sscDNA product is performed in conjunction with generating a DNA-RNA duplex and / or generating the sscDNA product. combine (e.g., in a single step, reaction, vessel, and / or volume) combine ) can be used. Therefore, a reagent for generating a DNA-RNA duplex (e.g., one or more nucleic acids having reverse transcriptase activity) can be used. agent ), a reagent for generating a sscDNA product (e.g., one or more agent) and reagents for amplifying the sscDNA product (e.g., primers, agent ) can be combined for use in a single step, reaction, vessel, and / or volume. In some embodiments, the reagents for amplifying the sscDNA product include amplification primers that hybridize to a component (e.g., a first oligonucleotide) of the scaffold adapter described herein. The amplification primers can be any primers suitable for use with a polymerase. In some embodiments, each primer includes a nucleotide sequence complementary to a sequence in the sscDNA product corresponding to the oligonucleotide component (i.e., the oligonucleotide component covalently linked to the ssRNA fragment). The amplified sscDNA product can be purified by a suitable purification or wash method, such as those described herein.
[0074] In some embodiments, the methods herein involve combining ssRNA and a scaffold adaptor or component thereof. Combination The method may further comprise fragmenting the ssRNA prior to the step of forming the ssRNA or prior to the step of generating the sscDNA, thereby generating ssRNA fragments. Any suitable fragmentation method can be used, such as the fragmentation methods described herein. In some embodiments, the method herein comprises fragmenting the ssRNA with a scaffold adaptor or a component thereof. Combination Before the step of inducing or before the step of generating sscDNA, a step of depleting ribosomal RNA (rRNA) and / or a step of depleting messenger RNA (mRNA) may be performed. Enrichment For example, the rRNA depletion methods and / or mRNA depletion methods described herein may be used. Enrichment Any suitable rRNA depletion method and / or mRNA depletion method, such as a method Enrichment The method can be used.
[0075] The methods herein include combining one or more scaffold adaptors or components thereof with a composition comprising a mixture of single-stranded ribonucleic acid (ssRNA) and single-stranded complementary deoxyribonucleic acid (ssDNA), or a mixture of single-stranded complementary DNA (sscDNA) and ssDNA, to form one or more complexes. Combination The scaffold polynucleotide may comprise a step of: Described in As shown, complex formation When The oligonucleotide moiety may be designed for simultaneous hybridization with an ssRNA, sscDNA, or ssDNA fragment and an oligonucleotide moiety such that the termini of the oligonucleotide moiety are adjacent to the termini of the terminal region of the ssRNA, sscDNA, or ssDNA fragment.
[0076] Figure 7 shows an example workflow for processing a sample containing a mixture of DNA and RNA. First, the RNA is subjected to first-strand cDNA synthesis (where the cDNA is tagged as described herein). receive The double-stranded DNA can be either left unchanged or tagged, for example with a barcode. The scaffold adapter can then be attached to the nucleic acid. Combination The adaptor-ligated nucleic acid can then be amplified, for example, by index PCR. The nucleic acid can then be optionally amplified by Enrichment (e.g., for a target of interest) and / or Sequencing It is possible to deconvolute DNA and RNA sequences.
[0077] 9A and 9B show an example of a method for processing RNA with initial first-strand synthesis (e.g., as in the workflow shown in FIG. 7). A fragmented sample containing both DNA and RNA is first subjected to reverse transcription and RNA tagging (sometimes called RNA bar-coating), for example, using primers containing tags and random n-mer (e.g., random hexamer) sequences. The DNA is then denatured. Sex DNA can be stabilized in single-stranded form, for example, by using a single-stranded enhancer such as single-stranded binding protein (SSB). Next, the scaffold adapter is contacted with nucleic acid (including original sample DNA and cDNA) and ligated. Amplification such as index PCR is performed. Tags that do not hybridize with the target (e.g., human) genome or transcriptome can be selected. Tags that do not use promoters (e.g., T7 promoters) that will be used in other parts of the workflow can be selected. Sequencing The RNA-specific tags can then be used to identify or deconvolute the RNA sequence.
[0078] Figure 8 shows another example of a workflow for processing a sample containing a mixture of DNA and RNA. First, both the RNA and DNA in the sample are bound to a scaffold adapter containing tags (RNA-specific and DNA-specific tags). Combination Then, one-step PCR is performed. The nucleic acid is then optionally Enrichment (e.g., for a target of interest) and / or Sequencing It is possible to deconvolute DNA and RNA sequences.
[0079] Figures 10A and 10B show an example of a method for processing RNA (e.g., as in the workflow shown in Figure 8) with an initial ligation step. A fragmented sample containing both DNA and RNA is first subjected to DNA denaturation. Sex DNA can be stabilized in single-stranded form by using a single-stranded enhancer such as a single-stranded binding protein (SSB). Next, a scaffold adapter (some containing RNA-specific tags and some containing DNA-specific tags) is contacted with and ligated to nucleic acids (DNA and RNA). The specificity of the RNA and DNA adapters that bind to ssRNA and ssDNA, respectively, can be achieved by selecting the adapter or its components and / or the enzyme (e.g., ligation enzyme) used. In some embodiments, the oligonucleotide component of the scaffold adapter that ligates to the RNA fragment is made of RNA. In some embodiments, the oligonucleotide component of the scaffold adapter that ligates to the RNA fragment is made of RNA, and the scaffold polynucleotide is made of RNA or DNA. In some embodiments, the oligonucleotide component of the scaffold adapter that ligates to the DNA fragment is made of DNA. In some embodiments, the oligonucleotide component of the scaffold adapter that ligates to the DNA fragment is made of DNA, and the scaffold polynucleotide is made of DNA. Enzymes with RNA or DNA specificity (e.g., T4 RNA ligase 2 for RNA, T4 DNA ligase for DNA) can be used, at least for the 5' end of the target nucleic acid. These enzymes will not ligate the "wrong" type of nucleic acid onto the 5' end. The 3' end of the target nucleic acid. About Ligation can be more flexible in some embodiments, as the adapter that hybridizes to the 3' end of the target nucleic acid can be the same for the RNA or DNA fragment. After ligation, a one-step PCR is performed, resulting in the generation of cDNA from the RNA and the resulting modified DNA. Sex D NA About No. 2 The strand is synthesized. Sequencing The DNA and RNA sequences can then be deconvoluted using the RNA-specific and DNA-specific tags.
[0080] Applications of these methods include analysis of cfDNA, single-cell analysis, and analysis of human samples.
[0081] Hybridization and ligation Nucleic acid fragments (e.g., ssNA fragments) are linked to scaffold adaptors or components thereof. Combination By doing Combined Generate a product can The ssNA fragment and the scaffold adaptor or its components are Combination Bringing can include hybridization and / or ligation (eg, ligation of hybridization products). Combined The product may include an ssNA fragment connected to (e.g., hybridized to and / or ligated to) a scaffold adaptor or component thereof at one or both ends of the ssNA fragment. Combined The product comprises an ssNA fragment hybridized to a scaffold adapter or component thereof at one or both ends of the ssNA fragment. Mite, Ha These are sometimes called hybridization products. Combined The product comprises an ssNA fragment ligated to a scaffold adapter or component thereof at one or both ends of the ssNA fragment. Mite, La In some embodiments, the products from the cleavage step (i.e., cleavage products) are ligated to the scaffold adaptors or components thereof. Combination By doing Combined Certain methods herein can produce a product by: Combined A set of products (e.g., Combined a first set of products and Combined In some embodiments, the method further comprises generating a second set of products. Combined The first set of products is First set of In some embodiments, the scaffold adaptor comprises an ssNA connected to (e.g., hybridized to and / or ligated to) a scaffold adaptor or a component thereof. Combined The second set of products is The second set Scaffold adapter, or its components Minutesconnected to (e.g., hybridized to and / or ligated to) the scaffold adapter or a component thereof Combined The first set of products includes:
[0082] A hybridization product can be generated by binding the ssNA to the scaffold adapter or its components under hybridization conditions. In some embodiments, the scaffold adapter is provided as a pre-hybridized product, and the hybridization step comprises hybridizing the scaffold adapter to the ssNA. In some embodiments, the scaffold adapter components (i.e., the oligonucleotide and the scaffold polynucleotide) are provided as individual components, and the hybridization step comprises hybridizing the scaffold adapter components 1) to each other and 2) to the ssNA. In some embodiments, the scaffold adapter components (i.e., the oligonucleotide and the scaffold polynucleotide) are provided sequentially as individual components, and the hybridization step comprises 1) hybridizing the scaffold polynucleotide to the ssNA, and then 2) hybridizing the oligonucleotide to the oligonucleotide hybridization region of the scaffold polynucleotide. combination The conditions during the step of specifically hybridizing the scaffold adapter, or a component thereof (e.g., a single-stranded scaffold region), are such that the scaffold adapter specifically hybridizes to an ssNA having a terminal region(s) that is complementary in sequence to the single-stranded scaffold region. combination The conditions during the step of hybridizing can also include conditions under which the components of the scaffold adaptor (e.g., oligonucleotides and oligonucleotide hybridization regions in the scaffold polynucleotide) specifically hybridize or remain hybridized to one another.
[0083] Specific hybridization occurs between the single-stranded scaffold region and the ssNA terminal region. (Multiple options possible) with Between or between the oligonucleotide and the oligonucleotide hybridization region Betweenthe degree of complementarity, their length, and The melting temperature (Tm) of the single-stranded scaffold region can be used to inform The temperature at which hybridization occurs and other factors affect the law of nature influence can be affected or affected obtain The melting temperature generally refers to the temperature at which half of the single-stranded scaffold region / ssNA-end region remains hybridized and half of the single-stranded scaffold region / ssNA-end region dissociates into single strands. The Tm of a duplex can be determined experimentally or calculated using the formula Tm = 81.5 + 16.6(log 10 It can be predicted using the formula [Na+]) + 0.41(fraction G+C) - (60 / N) where N is the chain length and [Na+] is less than 1M. It depends on various parameters Additional Models can be used to predict the Tm of related regions that depend on various hybridization conditions. Approaches to achieving specific nucleic acid hybridization are described, for example, in Tijssen, Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization with Nucleic Acid Probes, part I, chapter 2, "Overview of principles of hybridization and the strategy of nucleic acid probe assays," Elsevier (1993).
[0084] In some embodiments, the methods herein include exposing the hybridization product to conditions under which the end of the ssNA is bound to the end of the scaffold adaptor to which it is hybridized. In particular, the methods herein may include exposing the hybridization product to conditions under which the end of the ssNA is bound to the end of the oligonucleotide component of the scaffold adaptor to which it is hybridized. Binding can be achieved by any suitable approach that allows the ssNA to be covalently bound to the scaffold adaptor and / or the oligonucleotide component of the scaffold adaptor to which it is hybridized. When one end of the ssNA is bound to the end of the scaffold adaptor and / or the oligonucleotide component of the scaffold adaptor to which it is hybridized, one of the following two binding events usually occurs: 1) the 3' end of the ssNA is bound to the 5' end of the oligonucleotide component of the scaffold adaptor, or 2) the 5' end of the ssNA is bound to the 3' end of the oligonucleotide component of the scaffold adaptor. When both ends of the ssNA are each bound to the ends of the scaffold adapter and / or oligonucleotide component of the scaffold adapter to which it is hybridized, two binding events typically occur: 1) a binding event of the 3' end of the ssNA to the 5' end of the oligonucleotide component of the first scaffold adapter, and 2) a binding event of the 5' end of the ssNA to the 3' end of the oligonucleotide component of the second scaffold adapter.
[0085] In some embodiments, the methods herein involve combining the hybridization product with a ligase having ligase activity. agent and a step of contacting the target nucleic acid (ssNA) with the target nucleic acid (ssNA) under conditions in which the end of the ssNA is covalently ligated to the end of the scaffold adaptor and / or the end of the oligonucleotide component of the scaffold adaptor to which the target nucleic acid (ssNA) is hybridized. The ligase activity can be, for example, blunt-end ligase activity, nick-sealing ligase activity, adhesiveThe ligase activity may include terminal ligase activity, circularization ligase activity, sticky end ligase activity, DNA ligase activity, RNA ligase activity, single-stranded ligase activity, and double-stranded ligase activity. The ligase activity may include ligating the 5' phosphorylated end of one polynucleotide to the 3' OH end of another polynucleotide (5'P to 3'OH). The ligase activity may include ligating the 3' phosphorylated end of one polynucleotide to the 5' OH end of another polynucleotide (3'P to 5'OH). The ligase activity may include ligating the 5' end of an ssNA to the 3' end of a scaffold adaptor hybridized thereto, and / or is a foot The ligase activity can include ligating the 3' end of the ssNA to the 5' end of the scaffold adapter to which it is hybridized, and / or to the 3' end of the oligonucleotide component of the scaffold adapter in a ligation reaction. is a footThe method may involve ligating the 5'-end of the oligonucleotide component of the field adapter to the 5'-end of the oligonucleotide component of the field adapter in a ligation reaction. Suitable reagents (e.g., ligases) and kits for performing ligation reactions are known and available. For example, the Instant Sticky-end Ligase Master Mix available from New England Biolabs (Ipswich, MA) can be used. Ligases that can be used include, but are not limited to, T3 ligase, T4 DNA ligase (e.g., at low or high concentrations), T7 DNA ligase, E. coli DNA ligase, ElectroLigase®, RNA ligase, T4 RNA ligase 1, T4 RNA ligase 2, truncated T4 RNA ligase 2, thermostable 5'App DNA / RNA ligase, Splint® ligase, RtcB ligase, Taq ligase, etc., and combinations thereof. If required, a phosphate group can be added to the 5' end of an oligonucleotide component or ssNA fragment using a suitable kinase, such as, for example, T4 polynucleotide kinase (PNK). Such kinases and guidance for using them to phosphorylate 5' ends are available, for example, from New England BioLabs, Inc. (Ipswich, MA).
[0086] In some embodiments, the method includes covalently linking adjacent ends of the oligonucleotide component and the ssNA terminal region, thereby generating a covalently linked hybridization product. In some embodiments, the covalently linking step is performed by covalently linking the hybridization product (e.g., at least one scaffold adapter herein). to Hybridize death ssNA fragment) and a ligase activity agentand contacting the oligonucleotide component under conditions in which the terminus of the ssNA-terminal region is covalently linked to the terminus of the oligonucleotide component. In some embodiments, the method comprises covalently linking adjacent termini of a first oligonucleotide component and the first ssNA-terminal region, and covalently linking adjacent termini of a second oligonucleotide component and the second ssNA-terminal region, thereby producing a covalently linked hybridization product. In some embodiments, the covalently linking step is performed by covalently linking the hybridization product (e.g., two scaffold adapters herein). to Each hybridizes did ssNA fragment) and has ligase activity agent and under conditions such that an end of the first ssNA terminal region is covalently linked to an end of the first oligonucleotide component and an end of the second ssNA terminal region is covalently linked to an end of the second oligonucleotide component. In some embodiments, the method comprises contacting a nucleic acid having ligase activity with a nucleic acid having ligase activity. agent is T4 DNA ligase. In some embodiments, T4 DNA ligase is used in an amount between about 1 unit / μl and about 50 units / μl. In some embodiments, T4 DNA ligase is used in an amount between about 5 units / μl and about 30 units / μl. In some embodiments, T4 DNA ligase is used in an amount between about 5 units / μl and about 15 units / μl. In some embodiments, T4 DNA ligase is used in an amount of about 10 units / μl. In some embodiments, T4 DNA ligase is used in an amount less than 25 units / μl. In some embodiments, T4 DNA ligase is used in an amount less than 20 units / μl. In some embodiments, T4 DNA ligase is used in an amount less than 15 units / μl. In some embodiments, T4 DNA ligase is used in an amount less than 10 units / μl.
[0087] In some embodiments, the hybridization product comprises a first ligase having a first ligase activity. agent and a second ligase having a second ligase activity different from the first ligase activity. agentFor example, the first ligase activity and the second ligase activity may be a blunt-end ligase activity, a nick-sealing ligase activity, adhesive Terminal ligase activity, circularization ligase activity, and sticky end ligase activity, double-stranded ligase activity, single-stranded ligase activity, 3'OH of 5'P to Ligase activity, and 5'OH of 3'P to The ligase activity can be selected independently.
[0088] In some embodiments, the methods herein involve attaching ssNAs to a scaffold adaptor and / or an oligonucleotide component of a scaffold adaptor by biocompatible attachment. CombinationThe method may include, for example, click chemistry or tagging, which includes biocompatible reactions useful for attaching biomolecules. In some embodiments, the terminus of each of the oligonucleotide components includes a first chemically reactive moiety, and the terminus of each of the ssNAs includes a second chemically reactive moiety. In such embodiments, the first chemically reactive moiety is typically capable of reacting with a second chemically reactive moiety to form a covalent bond between the oligonucleotide component of the scaffold adaptor and the ssNA to which the scaffold adaptor is hybridized. In some embodiments, the method herein includes contacting the ssNA with one or more chemical agents under conditions in which a second chemically reactive moiety is incorporated into the terminus of each of the ssNA fragments. In some embodiments, the method herein includes exposing the hybridization product to conditions in which the first chemically reactive moiety reacts with the second chemically reactive moiety to form a covalent bond between the oligonucleotide component and the ssNA to which the scaffold adaptor is hybridized. In some embodiments, the first chemically reactive moiety can react with the second chemically reactive moiety to form a 1,2,3-triazole between the oligonucleotide component and the ssNA to which the scaffold adaptor is hybridized. In some embodiments, the first chemically reactive moiety can react with the second chemically reactive moiety under copper-containing conditions. The first and second chemically reactive moieties can comprise any suitable pairing. For example, the first chemically reactive moiety can be selected from an azide-containing moiety and 5-octadiynyldeoxyuracil, and the second chemically reactive moiety can independently be selected from an azide-containing moiety, hexynyl, and 5-octadiynyldeoxyuracil. In some embodiments, the azide-containing moiety is an N-hydroxysuccinimide (NHS) ester-azide.
[0089] Covalently linking the adjacent ends of the oligonucleotide and the ssNA fragment produces a covalently linked product, which may be referred to as a ligation product. A covalently linked product comprising an ssNA fragment covalently linked to an oligonucleotide component that remains hybridized to the scaffold polynucleotide may be referred to as a covalently linked hybridization product. The covalently linked hybridization product may be denatured (e.g., heat denatured) to separate the ssNA fragment covalently linked to the oligonucleotide component from the scaffold polynucleotide. A covalently linked product comprising an ssNA fragment covalently linked to an oligonucleotide component that is no longer hybridized to the scaffold polynucleotide (after denaturation) may be referred to as a single-stranded ligation product. Depending on the situation is a fragment of a scaffold polynucleotide, for example, a fragment of a uracil base or bases in the scaffold polynucleotide. Hey Uracil-DNA glycosylase and by cleaving using an endonuclease and / or deterioration It can be done.
[0090] The covalently linked hybridization products and / or single-stranded ligation products are then subjected to downstream applications of interest (e.g., amplification; Sequencing ) can be purified before use as input. For example, the covalently linked hybridization products and / or single-stranded ligation products can be purified by Combination from certain components present during the covalent bonding, hybridization and / or ligation steps (e.g., solid-phase reversible immobilization). Fixed( SPRI), column purification etc. It can be purified by
[0091] In some embodiments, the methods herein comprise combining an ssNA composition with a scaffold adaptor or component thereof described herein. Combinationand covalently linking the adjacent ends of the oligonucleotide component and the ssNA fragment, Combination The total duration of the covalently linking and covalently linking steps can be 4 hours or less, 3 hours or less, 2 hours or less, or 1 hour or less. Combination The total duration of the bonding and covalent linking steps is less than 1 hour.
[0092] In some embodiments, the methods herein comprise: In In some embodiments, the ssNA composition and the scaffold adaptor or components thereof described herein are combined in a single container, a single chamber, and / or a single volume (i.e., a continuous volume), including, but not limited to, Combination The steps of covalently linking adjacent ends of the oligonucleotide component and the ssNA fragment are carried out by a microfluidic device. In In some embodiments, the methods herein are performed in a single container, a single chamber, and / or a single volume (i.e., a continuous volume), including, but not limited to, a microfluidic device. In In some embodiments, the method is performed in a series of wells, droplets, emulsions, compartments, or other reaction volumes, including, but not limited to, a series of wells, droplets, emulsions, compartments, or other reaction volumes. Combination The steps of covalently linking adjacent ends of the oligonucleotide component and the ssNA fragment are carried out by a microfluidic device. In The reaction may be performed in a series of wells, droplets, emulsions, compartments or other reaction volumes, including but not limited to: Depending on the situation In this case, a series of reaction volumes are prepared such that most or all of the reaction volumes contain at most one ssNA. Depending on the situationA series of reaction volumes are prepared such that most or all of the reaction volumes contain at most 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 100000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000 or more ssNAs. The distribution of one or a limited number of ssNAs into the reaction volumes can result in favorable reaction kinetics, such as increased library conversion of rare species of sample nucleic acids.
[0093] Adapters for epigenetic analysis The adaptors described herein can be used in epigenetics (or epigenome) analysis. For example, the adaptors described herein can be used in methylation analysis (e.g., methylome analysis). DNA methylation is a type of epigenetic modification that can affect certain developmental processes. Methylation abnormalities, such as hypomethylation or hypermethylation of cytosine-guanine (CpG) dinucleotides, can cause problems such as genomic instability and / or transcriptional silencing, which can lead to the development of various mental disorders or diseases, such as cancer, diabetes, cardiovascular disease, and inflammatory disease.
[0094] Methylation analysis can include methylation sequencing (Methyl-Seq). Methylation sequencing typically involves a treatment to deaminate cytosines in sample nucleic acids. Deamination refers to the removal of an amino group from a molecule. Such treatment produces two distinct outcomes based on the methylation status of the cytosines: 1) unmethylated cytosine residues are converted to uracil, and 2) methylated cytosine (5'-methylcytosine, 5-mC, 5-hmC) residues remain unmodified by the treatment. In some assays, the deamination treatment is followed by nucleic acid amplification (e.g., PCR) and / or nucleic acid amplification to reveal the methylation status of cytosine residues in gene-specific or whole-genome analyses. Sequencing (e.g., massively parallel sequencing) can be performed. Unmethylated cytosine residues converted to uracil are usually amplified as thymine residues in subsequent amplification reactions, while methylated cytosine residues are amplified as cytosine residues. Between Comparison of sequence information can provide information about cytosine methylation patterns.
[0095] Deamination treatments can include chemical-based and / or enzyme-based treatments. Chemical-based treatments can include sodium bisulfite treatment, also called bisulfite conversion (e.g., ZYMO's EZ METHYLATION-LIGHTNING Kit). Enzyme-based treatments can include deaminase enzymes (e.g., sucralose ... Chiji Deaminase enzymes may include the use of NEBNext® Enzymatic Methyl-seq (EM-seq™) (NEB #E7120). Chiji The enzyme-based treatment may include APOBEC (apolipoprotein B mRNA editing enzyme, catalytic polypeptide-like), a member of the deaminase family. Bisulfite treatment is generally considered harsh and often results in denaturation, shearing, and / or loss of sample nucleic acid, whereas enzyme-based treatments are considered gentler than bisulfite treatment and can minimize damage to sample nucleic acid. Without being limited by theory, bisulfite treatment may be suitable for sample nucleic acid containing short nucleic acid fragments (e.g., fragments less than about 250 bp), and in certain cases, the treatment results in little shearing and / or loss.
[0096] A method for generating a nucleic acid library, comprising: (a) combining a nucleic acid composition comprising single-stranded nucleic acids (ssNAs) with a plurality of scaffold adaptors, or components thereof, as described herein; Combinationand (b) deaminating one or more unmethylated cytosine residues in the ssNA, thereby converting the one or more unmethylated cytosine residues to uracil. In some embodiments, the scaffold adapter comprises an in-line UMI as described herein. In some embodiments, the scaffold adapter does not comprise an in-line UMI. In some embodiments, (b) in The deamination step comprises: (a) Combinations in In some embodiments, (b) in The deamination step comprises: (a) Combinations in In some embodiments, the scaffold adaptor, or one or more components thereof, comprises one or more methylated cytosine residues. case , scaffold adapter, or one or more components thereof may be referred to herein as a methylation adapter or methylation component. In some embodiments, the oligonucleotide component of the scaffold adapter comprises one or more methylated cytosine residues (methylated oligonucleotide). In some embodiments, the scaffold polynucleotide component of the scaffold adapter comprises one or more methylated cytosine residues (methylated scaffold polynucleotide). In some embodiments, the deaminating step comprises the use of sodium bisulfite. In some embodiments, the deaminating step comprises the use of a deaminase.
[0097] The library can be prepared according to the method for methylation sequencing herein.In some embodiments, the library is prepared for methylation sequencing of nucleic acid compositions comprising genomic nucleic acid (e.g., gDNA).In some embodiments, the library is prepared for methylation sequencing of nucleic acid compositions comprising cell-free nucleic acid (e.g., cfDNA).In some embodiments, the library is prepared for methylation sequencing of ancient ofIn some embodiments, the library is prepared for methylation sequencing of nucleic acid compositions comprising nucleic acids (aDNA). In some embodiments, the library is prepared for methylation sequencing of nucleic acid compositions comprising nucleic acids from forensic samples. In some embodiments, the library is prepared for methylation sequencing of nucleic acid compositions comprising synthetic nucleic acids (e.g., synthetic oligonucleotides).
[0098] In some embodiments, a library is prepared for methylation sequencing of a nucleic acid composition comprising nucleic acid fragments having a representative, average, median, or mode length less than a certain threshold or cutoff length. In some embodiments, a library is prepared for methylation sequencing of a nucleic acid composition comprising nucleic acid fragments having a representative, average, median, or mode length less than a certain threshold or cutoff length, where the nucleic acids are treated with sodium bisulfite. In some embodiments, a library is prepared for methylation sequencing of a nucleic acid composition comprising nucleic acid fragments having a representative, average, median, or mode length less than a certain threshold or cutoff length, where the nucleic acids are treated with sodium bisulfite after binding to a scaffold adapter or a component thereof (e.g., a methylation adapter or a methylation component thereof) as described herein. In some embodiments, the nucleic acid composition comprises nucleic acid fragments having a representative, average, median, or mode length less than about 250 bp. For example, a nucleic acid composition may contain nucleic acid fragments having a representative, average, median, or mode length of about 250 bp, 200 bp, 150 bp, 100 bp, or less than 50 bp. In some embodiments, a nucleic acid composition contains nucleic acid fragments having a representative, average, median, or mode length of between about 30 bp and about 250 bp. For example, a nucleic acid composition may contain nucleic acid fragments having a representative, average, median, or mode length of about 50 bp, about 60 bp, about 70 bp, about 80 bp, about 90 bp, about 100 bp, about 110 bp, about 120 bp, about 130 bp, about 140 bp, about 150 bp, about 160 bp, about 170 bp, about 180 bp, about 190 bp, or about 200 bp. In some embodiments, a nucleic acid composition contains nucleic acid fragments having a representative, average, median, or mode length of about 75 bp. In some embodiments, a nucleic acid composition contains nucleic acid fragments having a mode length of about 75 bp. In some embodiments, the nucleic acid composition comprises nucleic acid fragments having a representative, mean, median, or mode length of about 167 bp. In some embodiments, the nucleic acid composition comprises nucleic acid fragments having a mode length of about 167 bp.
[0099] Adapter Dimer In some embodiments, the methods herein include one or more modifications and / or modifications to prevent, reduce or eliminate adapter dimers. Additional The method includes the steps of: (a) forming a dimer of an adapter with a dimer of a base pair; (b) forming a dimer of an adapter with a dimer of a base pair; (c) forming a dimer of an adapter with a dimer of a base pair; (d) forming a dimer of an adapter with a dimer of a base pair; (e) forming a dimer of an adapter with a dimer of a base pair; (f) forming a dimer of an adapter with a dimer of a base pair; (g) forming a dimer of an adapter with a dimer of a base pair; (h) forming a dimer of an adapter with a dimer of a base pair; (i) forming a dimer of an adapter with a dimer of a base pair; (ii) forming a dimer of an adapter with a dimer of a base pair; (iii) forming a dimer of an adapter with a dimer of a base pair; (iv) forming a dimer of an adapter with a dimer of a base pair; (v) forming a dimer of an adapter with a dimer of a base pair; (
[0100] In certain embodiments, the scaffold adaptor, or a component thereof, is modified to prevent adaptor dimer formation. fart Examples of modifications include modified nucleotides that are capable of interrupting the covalent linkage of a scaffold adaptor, oligonucleotide component, or scaffold polynucleotide to another oligonucleotide, polynucleotide, or nucleic acid molecule (e.g., another scaffold adaptor, oligonucleotide component, and / or scaffold polynucleotide). Examples of modified nucleotides are listed below. Described in Scaffolding adapter fart Other / Additional Modifications include structures such as Y structures or hairpin structures, which are described in more detail below. description In some embodiments, the scaffold adaptor, oligonucleotide component, and / or scaffold polynucleotide comprises a phosphorothioate backbone modification (e.g., a phosphorothioate linkage between the last two nucleotides on the strand). Miteku do.
[0101] In some embodiments, the method comprises a dephosphorylation step to prevent or reduce adapter dimer formation. In some embodiments, the method comprises dephosphorylating a scaffold adapter or a component thereof with a ssNA. Combination Before the step of forming a scaffold adaptor, an oligonucleotide component and / or a scaffold polynucleotide, a phosphatase-active agentand a scaffold adaptor under conditions in which the scaffold adaptor, oligonucleotide component and / or scaffold polynucleotide are dephosphorylated, thereby producing a dephosphorylated scaffold adaptor, dephosphorylated oligonucleotide component, and / or dephosphorylated scaffold polynucleotide.
[0102] In some embodiments, the method comprises one or more stepwise ligation approaches to prevent or reduce adapter dimer formation. In some embodiments, the method comprises a method comprising: agent of Attachment Delaying the addition of the second scaffold adaptor or a component thereof (e.g., until the hybridization product has finished forming) and / or Attachment For example, the method may include a stepwise ligation method in which the addition of the oligonucleotide components is delayed after the step of forming the hybridization product. (Multiple options possible) ssNA-terminal region (Multiple options possible) prior to the step of covalently linking the oligonucleotide component to (Multiple options possible) and has phosphoryl transfer activity agent and a 5′ phosphate are added to the 5′ end of the oligonucleotide component. In another example, the method can include contacting a first set of scaffold adaptors with ssNA. Combination The first set of scaffold adaptors may include an oligonucleotide component having a 3'OH. The first set of scaffold adaptors is hybridized to the ssNA, and the 3'OH of the oligonucleotide component is covalently linked to the 5' end (e.g., the 5' phosphorylated end) of the ssNA terminal region. To do The products of this first round of covalent linkage are sometimes referred to as covalently linked hybridization product intermediates. The covalently linked hybridization product intermediates are then coupled to a second set of scaffold adaptors. CombinationThe second set of scaffold adaptors may comprise oligonucleotide components having 5'-ends that can be phosphorylated as described herein. The second set of scaffold adaptors is hybridized to the covalently linked hybridization product intermediate, and the 5'-phosphorylated end of the oligonucleotide component is covalently linked to the 3'-end of the ssNA-terminal region.
[0103] In some embodiments, the method comprises stepwise ligation, including the use of scaffold adapters or components thereof having adenylation modifications. For example, a first set of scaffold adapters may comprise a 5′ end of an oligonucleotide component. in The first set of scaffold adaptors may contain an adenylation modification (5'App). The first set of scaffold adaptors is hybridized to the ssNA, and the 5'App of the oligonucleotide component is covalently linked to the 3' end of the ssNA terminal region. Covalent linkage may be performed in the absence of ATP. Hybridization To do The products of this first round of covalent linkage are sometimes referred to as covalently linked hybridization product intermediates. The covalently linked hybridization product intermediates are then coupled to a second set of scaffold adaptors. Combination The second set of scaffold adaptors may comprise oligonucleotide components having 3' OH ends. The second set of scaffold adaptors is hybridized to the covalently linked hybridization product intermediate, and the 3' OH ends of the oligonucleotide components are covalently linked (by the addition of ATP) to the 5' ends (e.g., 5' phosphorylated ends) of the ssNA-terminal region. In one variation, the first set of scaffold adaptors and the second set of scaffold adaptors are simultaneously hybridized with the ssNA in the absence of ATP. Combination Ligation of the first set of scaffold adaptors can proceed in the absence of ATP, and ligation of the second set of scaffold adaptors can proceed in the presence of ATP. Attachment Until it is added only progress obtain .
[0104] In some embodiments, the method includes stepwise ligation, including the use of oligonucleotides (i.e., single-stranded oligonucleotides) having a 3' phosphorylated end, as described herein for the oligonucleotide component of the scaffold adaptor. secondary The oligonucleotide having a 3' phosphorylated end is generally single-stranded and not hybridized to the scaffold polynucleotide. In one example, the method comprises hybridizing the scaffold adaptor or a component thereof with the ssNA. Combination Before the step of adding ssNA and an oligonucleotide containing a phosphate at the 3' end, Combination and covalently linking the 3' phosphorylated end of the oligonucleotide to the 5' end (e.g., the 5' non-phosphorylated end) of the ssNA terminal region. In some embodiments, prior to covalently linking the oligonucleotide to the ssNA, the ssNA has phosphatase activity. agent and under conditions that dephosphorylate the ssNA, thereby producing a dephosphorylated ssNA. In some embodiments, the step of covalently linking the oligonucleotide to the ssNA comprises contacting the ssNA and the oligonucleotide with a single-stranded ligase having single-stranded ligase activity. agent and a ligase having ligase activity under conditions in which the 5' end of the ssNA is covalently linked to the 3' end of the oligonucleotide. agent The RtcB ligase is the RtcB ligase. The product of such a covalent ligation step is sometimes referred to as a covalently ligated product intermediate. The covalently ligated product intermediate is then ligated with a set of scaffold adaptors. Combination The set of scaffold adaptors can include an oligonucleotide component having a 5' phosphorylated end. The set of scaffold adaptors is hybridized to the covalently linked product intermediate, and the 5' phosphorylated end of the oligonucleotide component is covalently linked to the 3' end of the ssNA-terminal region.
[0105] In some embodiments, the methods include using an oligonucleotide capable of hybridizing to an oligonucleotide dimer product to reduce or eliminate adapter dimers. The oligonucleotide dimer product may be a component of a scaffold adapter dimer and may contain an oligonucleotide component from a first scaffold adapter covalently linked to an oligonucleotide component from a second scaffold adapter. The methods herein may include a denaturation step that can release the oligonucleotide dimer product from the scaffold adapter dimer. The oligonucleotide dimer product can hybridize to an oligonucleotide having a sequence complementary to the oligonucleotide dimer product or a portion thereof, thereby forming an oligonucleotide dimer hybridization product. In some embodiments, the oligonucleotide dimer hybridization product includes a cleavage site. In some embodiments, the cleavage site is a restriction enzyme recognition site. In some embodiments, the methods herein further include contacting the oligonucleotide dimer hybridization product with a cleavage agent (e.g., a restriction enzyme, a rare-cutter restriction enzyme).
[0106] In some embodiments, the methods include purifying or washing the nucleic acid products at various stages of library preparation to reduce or eliminate adapter dimers. Depending on the situation Purifying or washing nucleic acid products can reduce or eliminate adapter dimers. For example, covalently linked hybridization products (i.e., hybridized to a scaffold adapter and covalently linked to an oligonucleotide component, ssNA), single-stranded ligation products (i.e., denatured NAs), and the like can be purified or washed to reduce or eliminate adapter dimers. death Alternatively, the covalently linked hybridization products (ssNAs) or amplification products thereof, which are covalently linked to the oligonucleotide components and can no longer hybridize to the scaffold polynucleotide, can be purified or washed by any suitable purification or washing method. In some embodiments, the purifying or washing step can be performed using a solid-phase reversible immobilization method. Fixed( This includes the use of SPRI beads. For example, SPRI beads can be suspended in a DNA binding buffer containing about 2.5 M to about 5 M NaCl, about 0.1 mM to about 1 M EDTA, about 10 mM Tris, about 0.01% to about 0.05% TWEEN®-20, and about 8% to about 38% PEG-8000. For example, 1 ml of SPRI bead suspension can be prepared by mixing 2.5 M NaCl, 10 mM Tris, 1 mM EDTA, 0.05% Tween-20, and 20% PEG-8000. Combination In some embodiments, SPRI includes sequential SPRI (washes performed consecutively) and / or sequential SPRI (washes involving sequential addition and incubation of SPRI beads). Sequential SPRI includes multiple sequential (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, or more sequential washes). One after the other) washes. Sequential SPRI can involve multiple sequential additions of SPRI beads (intervening incubations), which can include 2, 3, 4, 5, 6, 7, 8, 9, 10, or more sequential additions of SPRI beads. In some embodiments, the amount of SPRI beads used in SPRI purification can include between 0.1× and 3× SPRI beads (× is the ratio of beads to nucleic acid (e.g., bead volume to reaction volume)). For example, the amount of SPRI beads used in SPRI purification may include about 0.1x, 0.2x, 0.3x, 0.4x, 0.5x, 0.6x, 0.7x, 0.8x, 0.9x, 1.0x, 1.1x, 1.2x, 1.3x, 1.4x, 1.5x, 1.6x, 1.7x, 1.8x, 1.9x, 2.0x, 2.1x, 2.2x, 2.3x, 2.4x, 2.5x, 2.6x, 2.7x, 2.8x, 2.9x, or 3.0x SPRI beads. In some embodiments, the amount of SPRI beads used in SPRI purification is 1.2x. In some embodiments, the amount of SPRI beads used in SPRI purification is 1.5x. In some embodiments, the purifying or washing step comprises column purification (e.g., column chromatography). In some embodiments, the purifying or washing step does not include column purification (e.g., column chromatography). In some embodiments, the covalently linked hybridization products, single-stranded ligation products, and / or their amplification products are not purified or washed.
[0107] SPRI purification is usually performed in the presence of a buffer. Any suitable buffer, such as Tris buffer or water of a similar pH, can be used. SPRI purification beads can be added directly to a sample solution (e.g., a sample solution containing covalently linked hybridization products (ligation products) or their amplified products). In certain cases, buffer can be added to increase the reaction volume, allowing additional beads to be added. In some embodiments, the SPRI bead solution is composed of carboxylated magnetic beads added to PEG 8000 dissolved in water, NaCl, Tris, and EDTA. The amount of PEG typically determines the PEG percentage of the SPRI bead solution. For example, adding 9 g of PEG 8000 to a 50 ml SPRI bead solution can be referred to as "18% SPRI." In another example, adding 19 g of PEG 8000 to a 50 ml SPRI solution can be referred to as "38% SPRI." Generally, the higher the PEG ratio, the smaller the size of the DNA fragments retained.
[0108] In some embodiments, the purification process involves the covalently linked hybridization products (ligation products) and the solid phase reversible immobilization. Fixed(The method includes contacting SPRI beads and a buffer. In some embodiments, some or all of the SPRI buffer is replaced with isopropanol. In some embodiments, the SPRI buffer includes isopropanol. In some embodiments, the SPRI buffer is completely replaced with isopropanol. In some embodiments, the SPRI buffer includes about 5% volume / volume (v / v) isopropanol to about 50% v / v isopropanol. In some embodiments, the SPRI buffer includes about 10% v / v isopropanol to about 40% v / v isopropanol. For example, the SPRI buffer can include about 10% v / v isopropanol, 15% v / v isopropanol, 20% v / v isopropanol, 25% v / v isopropanol, 30% v / v isopropanol, 35% v / v isopropanol, or 40% v / v isopropanol. In some embodiments, the SPRI buffer includes about 20% v / v isopropanol.
[0109] In some embodiments, the purification or washing step may be performed for a particular length or A series of long Sawo a nucleic acid fragment having the nucleic acid fragment or an amplification product thereof Enrichment In some embodiments, SPRI purification can be performed on a specific length or A series of long Sawo a nucleic acid fragment having the nucleic acid fragment or an amplification product thereof Enrichment In some embodiments, the amount of PEG 8000 in the SPRI bead solution used in SPRI purification can be Enrichment The length of the fragment to be A series of long Sani For example, SPRI purification at 1.5x v / v ratio can recover more fragments in the range of <100 bases than SPRI purification at 1.2x because the final concentration of PEG 8000 is higher at 1.5x than at 1.2x. In some embodiments, the methods herein can be used to recover fragments of a desired length or A series of long Enriching In some embodiments, the methods herein include adjusting the SPRI ratio to achieve a desired fragment length or A series of long Enriching In some embodiments, the methods herein comprise adjusting the amount of isopropanol in the SPRI purification to achieve a desired fragment length or fragment length while minimizing the amount of undesired artifacts (e.g., adapter dimers). A series of long Enriching For example, the methods herein can be used to obtain a desired fragment length or A series of long Enriching In another example, the method herein comprises adjusting the amount of isopropanol in the SPRI purification so that the amount of adapter dimer recovered is less than about 10% of the total nucleic acid recovered. A series of long Enriching The method includes adjusting the amount of isopropanol in the SPRI purification so that the amount of adapter dimer recovered is less than about 5% of the total nucleic acid recovered.
[0110] In some embodiments, the methods described herein (e.g., combining an ssNA with a scaffold adaptor or component thereof) are CombinationThe steps of covalently linking, hybridizing, and covalently linking can be carried out in a suitable reaction volume and / or using a suitable amount of ssNA and / or a suitable ratio of ssNA to scaffold adaptor (or components thereof). A suitable reaction volume and / or a suitable amount of ssNA and / or a suitable ratio of ssNA to scaffold adaptor (or components thereof) can include a reaction volume, amount of ssNA, and / or ratio of ssNA to scaffold adaptor that reduces or prevents adapter-dimer formation. In some embodiments, a suitable amount of ssNA can range from about 250 pg to about 5 ng of ssNA. For example, a suitable amount of ssNA can be about 250 pg, 500 pg, 750 pg, 1 ng, 1.5 ng, 2 ng, 2.5 ng, 3 ng, 3.5 ng, 4 ng, 4.5 ng, or 5 ng. In some embodiments, a suitable amount of ssNA can be about 1 ng of ssNA. In some embodiments, for a final reaction volume of 25 μl, 1 ng of ssNA is hybridized with between about 1.0-2.0 pmoles of each scaffold adapter (i.e., about 1.0-2.0 pmoles of scaffold adapter hybridized to the 5' end of the ssNA terminal region (for a pool of scaffold adapters containing multiple scaffold adapter species) and about 1.0-2.0 pmoles of scaffold adapter hybridized to the 3' end of the ssNA terminal region (for a pool of scaffold adapters containing multiple scaffold adapter species)). Combination For example, for a final reaction volume of 25 μl, 1 ng of ssNA can be combined with approximately 1.0, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, or 2.0 pmoles of each scaffold adaptor. Combination In some embodiments, for a final reaction volume of 25 μl, 1 ng of ssNA can be hybridized with about 1.6 pmoles of each scaffold adaptor (i.e., about 1.6 pmoles of scaffold adaptor hybridizing to the 5' end of the ssNA-terminal region and about 1.6 pmoles of scaffold adaptor hybridizing to the 3' end of the ssNA-terminal region). Combination For larger reaction volumes, the relative amounts of ssNA and scaffold adapter are preserved. as long asFor smaller reaction volumes, the relative amounts of ssNA and scaffold adapter are preserved. as long as In some embodiments, the scaffold adaptors herein can be synthesized with ssNAs at a molar ratio of between about 5:1 (scaffold adaptor to ssNAs) and about 50:1 (scaffold adaptor to ssNAs). Combination For example, the scaffold adaptor may be mixed with the ssNA at a molar ratio of about 5:1 (scaffold adaptor to ssNA), about 10:1 (scaffold adaptor to ssNA), about 15:1 (scaffold adaptor to ssNA), about 20:1 (scaffold adaptor to ssNA), about 25:1 (scaffold adaptor to ssNA), about 30:1 (scaffold adaptor to ssNA), about 35:1 (scaffold adaptor to ssNA), about 40:1 (scaffold adaptor to ssNA), about 45:1 (scaffold adaptor to ssNA), or about 50:1 (scaffold adaptor to ssNA). Combination It can be (may combined with) In some embodiments, the scaffold adaptor is combined with the ssNA at a molar ratio of about 15:1 (scaffold adaptor to ssNA). Combination In some embodiments, the scaffold adaptor is mixed with the ssNA at a molar ratio of about 30:1 (scaffold adaptor to ssNA). Combination will be done.
[0111] In some embodiments, the methods herein include the use of a crowding agent. A suitable amount of crowding agent can be used to reduce or prevent adapter dimer formation. Examples of crowding agents include Ficoll 70, Dextran 70, polyethylene glycol (PEG) 2000, and polyethylene glycol (PEG) 8000. In some embodiments, the methods herein include the use of polyethylene glycol (PEG) 8000. PEG can be used, for example, in an amount between about 15% and about 20%, where this percentage refers to the final concentration of PEG in the ligation reaction. For example, PEG can be used at about 15%, 15.5%, 16%, 16.5%, 17%, 17.5%, 18%, 18.5%, 19%, 19.5%, or 20%. In some embodiments, 18.5% PEG is used. In some embodiments, 18% PEG is used.
[0112] During purification, SPRI bead solution can be added to the sample solution, and there are often instructions for the v / v ratio. For example, 1.2x 18% SPRI is added to a 50 μl sample. Yeah This means adding 60 μl (50 × 1.2) of 18% SPRI beads when using a ligation kit. This v / v ratio, assuming no PEG in the sample solution, results in a final PEG concentration of 9.8%. However, in many cases, there is already a certain amount of PEG present in the sample solution (i.e., the ligation product) after ligation. Therefore, users can adjust the amount of SPRI beads added to reach the desired final PEG concentration. The desired final PEG concentration can range from about 5% final PEG to about 15% final PEG. For example, the desired final PEG concentration can be about 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, or 15%. In some embodiments, the desired final PEG concentration is about 10% (e.g., for hair samples and cfDNA samples). In some embodiments, the desired final PEG concentration is about 12% (e.g., for formalin-fixed, paraffin-embedded (FFPE) samples and samples with large template fragments).
[0113] Y adapter In some embodiments, a scaffold adaptor described herein comprises two strands, with a single-stranded scaffold region at a first end and two non-complementary strands at a second end. Such scaffold adaptors may be referred to as Y scaffold adaptors, Y adaptors, Y-shaped scaffold adaptors, Y-shaped adaptors, Y duplexes, Y-shaped duplexes, Y scaffold duplexes, Y scaffold duplexes, etc. Scaffold adaptors with a Y-shaped structure generally comprise a double-stranded duplex region, two single-stranded "arms" at one end, and a single-stranded scaffold region at the other end.
[0114] The Y scaffold adaptor contains multiple nucleic acid components and secondary In some embodiments, the Y scaffold adaptor comprises a first nucleic acid strand and a second nucleic acid strand. In some embodiments, the first nucleic acid strand is complementary to the second nucleic acid strand. In some embodiments, a portion of the first nucleic acid strand is complementary to a portion of the second nucleic acid strand. In some embodiments, the first nucleic acid strand comprises a first region that is complementary to a first region in the second nucleic acid strand, and the first polynucleotide comprises a second region that is not complementary to a second region in the second polynucleotide. The complementary regions often form duplex regions of the Y scaffold adaptor, and the non-complementary regions often form arms or portions thereof of the Y scaffold adaptor. The first and second nucleic acid strands are secondary components (e.g., scaffold polynucleotides) secondary Components, oligonucleotides secondary components, as well as the sequencing adapters described herein. secondary In some embodiments, the first and second nucleic acid strands may comprise components such as a UMI, a UMI flanking region, an amplification priming site, and / or individual sequencing adapters (e.g., P5, P7 adapters), etc. In some embodiments, the first and second nucleic acid strands may comprise certain of the sequencing adapters described herein. secondary It does not include components such as amplification priming sites and / or individual sequencing adapters (e.g., P5, P7 adapters).
[0115] In some embodiments, the Y scaffold adapter comprises a single-stranded scaffold region (ssNA hybridization region). The single-stranded scaffold region of the Y scaffold adapter is typically located adjacent to the double-stranded duplex portion, at the opposite end of the non-complementary strand (or "arm") portion. The single-stranded scaffold region of the Y scaffold adapter is typically complementary to a terminal region of the target nucleic acid (e.g., a terminal region of a single-stranded nucleic acid).
[0116] hairpin In some embodiments, a scaffold adaptor comprises one strand capable of forming a hairpin structure with a single-stranded loop. In some embodiments, a scaffold adaptor consists of one strand capable of forming a hairpin structure with a single-stranded loop. A scaffold adaptor having a hairpin structure generally comprises a double-stranded "stem" region and a single-stranded "loop" region. In some embodiments, a scaffold adaptor comprises one strand (i.e., one continuous strand) capable of adopting a hairpin structure. In some embodiments, a scaffold adaptor consists essentially of one strand (i.e., one continuous strand) capable of adopting a hairpin structure. Consisting essentially of one strand means that the scaffold adaptor is not part of a continuous strand of nucleic acid (e.g., hybridized to the scaffold adaptor). Additional "Hairpin" refers to a scaffold adapter that does not contain any strands. Thus, "consisting essentially of" herein refers to the number of strands in the scaffold adapter, and the scaffold adapter may include other features that are not essential to the number of strands (e.g., it may include a detectable label; it may include other regions). A scaffold adapter that includes or consists essentially of one strand that can form a hairpin structure may be referred to herein as a hairpin, hairpin scaffold adapter, or hairpin adapter.
[0117] Hairpin scaffold adaptors contain multiple nucleic acid components and secondaryIn some embodiments, the hairpin scaffold adaptor comprises an oligonucleotide and a scaffold polynucleotide. In some embodiments, the oligonucleotide is complementary to an oligonucleotide hybridization region in the scaffold polynucleotide. In some embodiments, a portion of the oligonucleotide is complementary to a portion of the oligonucleotide hybridization region in the scaffold polynucleotide. In some embodiments, the hairpin scaffold adaptor comprises a complementary region and a non-complementary region. The complementary region often forms the stem of the hairpin adaptor, and the non-complementary region often forms the loop or a portion of the loop of the hairpin scaffold adaptor. The oligonucleotide and scaffold polynucleotide secondary components (e.g., scaffold polynucleotides) secondary Components, oligonucleotides secondary components, as well as the sequencing adapters described herein. secondary In some embodiments, the oligonucleotide and scaffold polynucleotide may comprise certain of the sequencing adapters described herein. secondary It does not include components such as amplification priming sites and individual sequencing adapters (e.g., P5, P7 adapters).
[0118] The hairpin scaffold adaptor contains one or more cleavage sites that can be cleaved under cleavage conditions. MitekuIn some embodiments, the cleavage site is located between the oligonucleotide and the scaffold polynucleotide. Cleavage at the cleavage site often generates two separate strands from the hairpin scaffold adaptor. In some embodiments, cleavage at the cleavage site generates a partially double-stranded scaffold adaptor with two unpaired strands forming a "Y" structure. The cleavage site may include any suitable cleavage site, such as, for example, the cleavage sites described herein. In some embodiments, the cleavage site includes RNA nucleotides and may be cleaved using, for example, an RNAse. In some embodiments, the cleavage site includes uracil and / or deoxyuridine and may be cleaved using, for example, a DNA glycosylase, an endonuclease, an RNAse, or the like, and combinations thereof. In some embodiments, the cleavage site does not include uracil and / or deoxyuridine. In some embodiments, the methods herein involve separating a hairpin scaffold adaptor from a single-stranded nucleic acid. Combination After the step of allowing, the method includes exposing the one or more cleavage sites to cleavage conditions, whereby the scaffold adaptor is cleaved.
[0119] In some embodiments, the hairpin scaffold adapter comprises a single-stranded scaffold region (ssNA hybridization region). The single-stranded scaffold region of the hairpin scaffold adapter is typically located adjacent to the double-stranded stem portion and at the opposite end of the loop portion. The single-stranded scaffold region of the hairpin scaffold adapter is typically complementary to the terminal region of the target nucleic acid (e.g., the terminal region of the single-stranded nucleic acid).
[0120] In some embodiments, the hairpin scaffold adapter comprises, in a 5' to 3' direction, an oligonucleotide; one or more cleavage sites; and a scaffold polynucleotide comprising an oligonucleotide hybridization region and a scaffold region (ssNA hybridization region). In some embodiments, the hairpin oligonucleotide comprises, in a 5' to 3' direction, a scaffold polynucleotide comprising a scaffold region (ssNA hybridization region) and an oligonucleotide hybridization region; one or more cleavage sites; and an oligonucleotide. In some embodiments, multiple hairpin scaffold adapters seed , or hairpin scaffold adaptors seed The pool comprises a mixture of: 1) hairpin scaffold adaptors comprising, in a 5' to 3' direction, an oligonucleotide; one or more cleavage sites; and a scaffold polynucleotide comprising an oligonucleotide hybridization region and a scaffold region (ssNA hybridization region); and 2) hairpin scaffold adaptors comprising, in a 5' to 3' direction, a scaffold polynucleotide comprising a scaffold region (ssNA hybridization region) and an oligonucleotide hybridization region; one or more cleavage sites; and an oligonucleotide.
[0121] Modified Nucleotides In some embodiments, the scaffold adaptor, or a component thereof, comprises one or more modified nucleotides. In some embodiments, the UMI and / or the adjacent region adjacent to the UMI comprises one or more modified nucleotides. Modified nucleotides may be referred to as modified bases or non-standard bases, and may include, for example, nucleotides conjugated to members of binding pairs, blocked nucleotides, non-natural nucleotides, nucleotide analogs, peptide nucleic acid (PNA) nucleotides, morpholino nucleotides, locked nucleic acid (LNA) nucleotides, bridged nucleic acid (BNA) nucleotides, glycol nucleic acid (GNA) nucleotides, threose nucleic acid (TNA) nucleotides, etc., and combinations thereof. In certain configurations, the scaffold adaptor, or a component thereof (e.g., the UMI and / or the adjacent region adjacent to the UMI), may comprise an amino modifier, biotinylation, thiol, alkyne, 2'-O-methoxy-ethyl base (2'-MOE), RNA, fluoro base, iso (iso-dG, iso-DC), reverse (inverted) The nucleotides include one or more nucleotides having modifications selected from one or more of the following: methyl, nitro, phos, and the like.
[0122] In some embodiments, the scaffold adaptor, or a component thereof (e.g., the UMI and / or an adjacent region adjacent to the UMI), comprises one or more modified nucleotides within the duplex region, within the scaffold region, at one end of the scaffold adaptor, or a component thereof, or at both ends. In some embodiments, the scaffold adaptor, or a component thereof, comprises one or more unpaired modified nucleotides. In some embodiments, the scaffold adaptor, or a component thereof, comprises one or more unpaired modified nucleotides at one end of the adaptor. In some embodiments, the scaffold adaptor, or a component thereof, comprises one or more unpaired modified nucleotides at the end of the adaptor opposite the end that hybridizes to the target nucleic acid (e.g., the end comprising the single-stranded scaffold region). The modified nucleotides may be present at the end of the strand having a 3' end or at the end of the strand having a 5' end.
[0123] In some embodiments, the oligonucleotide component comprises one or more modified nucleotides. In some embodiments, the one or more modified nucleotides can block the covalent linkage of the oligonucleotide component to another oligonucleotide, polynucleotide, or nucleic acid molecule. In some embodiments, the oligonucleotide component comprises one or more modified nucleotides at the end not adjacent to the ssNA. In some embodiments, the scaffold polynucleotide comprises one or more modified nucleotides. In some embodiments, the one or more modified nucleotides can block the covalent linkage of the scaffold polynucleotide to another oligonucleotide, polynucleotide, or nucleic acid molecule. The scaffold polynucleotide may comprise one or more modified nucleotides at one or both ends of the polynucleotide. In some embodiments, the one or more modified nucleotides comprise a modification that blocks ligation.
[0124] In some embodiments, a scaffold adaptor or component thereof comprises one or more blocked nucleotides. In one example, a scaffold adaptor or component thereof can comprise one or more modified nucleotides capable of blocking hybridization with nucleotides in another scaffold adaptor or component thereof. Depending on the situation In some cases, the one or more modified nucleotides can block ligation to nucleotides in another scaffold adapter or component thereof. In another example, the scaffold adapter or component thereof can include one or more modified nucleotides that can block hybridization with nucleotides in a target nucleic acid (e.g., ssNA). Depending on the situationIn some embodiments, one or more modified nucleotides can block ligation to a nucleotide in a target nucleic acid. In some embodiments, one or both ends of a scaffold polynucleotide contain a blocking modification, and / or the end of an oligonucleotide component that is not adjacent to an ssNA fragment can contain a blocking modification. A blocking modification refers to a modified end that cannot be ligated to the end of another nucleic acid component using the approach used to covalently link the adjacent ends of an oligonucleotide component and an ssNA fragment. In certain embodiments, the blocking modification is a ligation-blocking modification. Examples of blocking modifications that may be included on one or both ends of a scaffold polynucleotide and / or on the end of an oligonucleotide component that is not adjacent to an ssNA fragment include an absent 3' OH and an inaccessible 3' OH. Non-limiting examples of blocking modifications that have an inaccessible 3' OH include amino modifiers, amino linkers, spacers, isodeoxy bases, dideoxy bases, inverted dideoxy bases, 3' phosphates, etc. In some embodiments, a scaffold adaptor or component thereof contains one or more modified nucleotides that cannot bind to natural nucleotides.
[0125] In some embodiments, one or more modified nucleotides comprise an isodeoxy-base. In some embodiments, one or more modified nucleotides comprise isodeoxy-guanine (iso-dG). In some embodiments, one or more modified nucleotides comprise isodeoxy-cytosine (iso-dC). Iso-dC and iso-dG are chemical variants of cytosine and guanine, respectively. Iso-dC can hydrogen bond with iso-dG, but cannot with unmodified guanine (natural guanine). Iso-dG can base pair with iso-dC, but cannot with unmodified cytosine (natural cytosine). An iso-dC-containing backbone adaptor or a component thereof can be designed to hybridize with a complementary oligo containing iso-dG, but cannot hybridize with any naturally occurring nucleic acid sequence.
[0126] In some embodiments, the one or more modified nucleotides comprise an epigenetic-related modification, including, but not limited to, methylation, hydroxymethylation, and carboxylation. Examples of epigenetic-related modifications include carboxycytosine, 5-methylcytosine (5mC) and its oxidized derivatives (e.g., 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), and 5-carboxymethylcytosine (5fC). Kishi These include N(6)-methylcytosine (arboxylcytosine (5caC)), N(6)-methyladenine (6mA), N4-methylcytosine (4mC), N(6)-methyladenosine (m(6)A), pseudouridine (Ψ), 5-methylcytidine (m(5)C), hydroxymethyluracil, 2'-O-methylation of the 3' end, tRNA modifications, miRNA modifications, and snRNA modifications.
[0127] In some embodiments, one or more modified nucleotides comprise a dideoxy base. In some embodiments, one or more modified nucleotides comprise a dideoxy cytosine. In some embodiments, one or more modified nucleotides comprise an inverted dideoxy base. In some embodiments, one or more modified nucleotides comprise an inverted dideoxy thymine. For example, an inverted dideoxy thymine located at the 5' end of a sequence can prevent undesired 5' ligation.
[0128] In some embodiments, one or more modified nucleotides comprise a spacer. In some embodiments, one or more modified nucleotides comprise a C3 spacer. A C3 spacer phosphoramidite can be incorporated internally or at the 5' end of the oligonucleotide. Multiple C3 spacers can be added to either end of the backbone adaptor or its components to introduce a long hydrophilic spacer arm (e.g., for attachment of a fluorophore or other pendant group). Other spacers include, for example, a photocleavable (PC) spacer, hexanediol, spacer 9, spacer 18, 1',2'-dideoxyribose (dSpacer), and the like.
[0129] In some embodiments, the modified nucleotide comprises an amino linker or an amino blocker. In some embodiments, the modified nucleotide comprises an amino linker C6 (e.g., a 5' amino linker C6 or a 3' amino linker C6). In one example, an amino linker C6 can be used to incorporate an active primary amino group at the 5' end of the oligonucleotide, which can then be conjugated to a ligand. The amino group is then attached to the 5'-terminal ligand. On the other hand, inside The amino group is separated from the 5'-terminal nucleotide base by a six-carbon spacer arm to reduce steric interactions between the amino group and the oligo. separation In some embodiments, the modified nucleotide comprises an amino linker C12 (e.g., a 5' amino linker C12 or a 3' amino linker C12). In one example, an amino linker C12 can be used to incorporate an activated primary amino group at the 5' end of the oligonucleotide. The amino group can be used to bond the amino group to the oligonucleotide. Between A 12-carbon spacer arm separates the 5'-terminal nucleotide base from the base to minimize steric interactions. separation will be done.
[0130] In some embodiments, the modified nucleotide comprises a member of a binding pair, such as, for example, antibody / antigen, antibody / antibody, antibody / antibody fragment, antibody / antibody receptor, antibody / protein A or protein G, hapten / anti-hapten, biotin / avidin, biotin / streptavidin, folate / folate binding protein, vitamin B12 / intrinsic factor, chemically reactive group / complementary chemically reactive group, digoxigenin moiety / anti-digoxigenin antibody, fluorescein moiety / anti-fluorescein antibody, steroid / steroid binding protein, operator / repressor, nuclease / nucleotide, lectin / polysaccharide, active compound / active compound receptor, hormone / hormone receptor, enzyme / substrate, oligonucleotide or polynucleotide / its corresponding phase In some embodiments, the modified nucleotide comprises biotin.
[0131] In some embodiments, the modified nucleotide comprises a first member of a binding pair (e.g., biotin), and the second member of the binding pair (e.g., streptavidin) is conjugated to a solid support or substrate. The solid support or substrate can be any physically separable solid capable of directly or indirectly binding a member of the binding pair, including, but not limited to, surfaces provided by microarrays and wells, and particles, such as beads (e.g., paramagnetic beads, magnetic beads, microbeads, nanobeads), microparticles, and nanoparticles. The solid support can be, for example, a chip, a column, an optical fiber, a wipe, or the like. Material, filters (e.g., flat filters), one or more capillaries, glass and modified or functionalized glass (e.g., controlled pore glass (CPG)), quartz, mica, diazotized membranes (paper or nylon), polyformaldehyde, cellulose, cellulose acetate, paper, ceramics, metals, metalloids, semiconductor materials, quantum dots, coated beads or particles, other chromatographic materials, magnetic particles; plastics (including acrylics, polystyrene, copolymers of styrene or other materials, polybutylene, polyurethane, TEFLON®, polyethylene, polypropylene, polyamide, polyester, polyvinylidene fluoride (PVDF), etc.), polysaccharides, The solid support or substrate may also include nylon or nitrocellulose, resins, silica, or silica-based materials, including silicon, silica gels, and modified silicon, Sephadex®, Sepharose®, carbon, metals (e.g., steel, gold, silver, aluminum, silicon, and copper), inorganic glass, conductive polymers (including polymers such as polypyrrole and polyindole); micro- or nanostructured surfaces, such as nucleic acid tiling arrays, nanotubes, nanowires, or nanoparticle-decorated surfaces; or porous surfaces, or gels, such as methacrylate, acrylamide, sugar polymers, cellulose, silicates, or other fibrous or chain-like polymers. In some embodiments, the solid support or substrate may be coated using a passivation or chemically derivatized coating with any number of materials, including polymers such as dextran, acrylamide, gelatin, or agarose. The beads and / or particles may be Freedom from each other or each other Related In some embodiments, the solid support can be a collection of particles. In some embodiments, the particles can comprise silica, which can comprise silicon dioxide. In some embodiments, the silica can be porous, and in certain embodiments, the silica can be non-porous. In some embodiments, the particles can be formed using a material that confers paramagnetic properties to the particles. agent In certain embodiments, agentcomprises a metal, and in certain embodiments, agent is oxidation metal (e.g., iron or iron oxide, where iron oxide contains a mixture of Fe2+ and Fe3+.) The binding pair members can be linked to the solid support by covalent or non-covalent interactions, and can be linked to the solid support directly or indirectly (e.g., by a spacer molecule or an intermediary such as biotin).
[0132] In some embodiments, the scaffold polynucleotide, the oligonucleotide component (e.g., the UMI and / or the flanking region adjacent to the UMI), or both, comprise one or more non-natural nucleotides, also referred to as nucleotide analogs. Non-limiting examples of non-natural nucleotides that may be included in the scaffold polynucleotide, the oligonucleotide component, or both, include LNA (locked nucleic acid), PNA (peptide nucleic acid), FANA (2'-deoxy-2'-fluoroarabinonucleotide), GNA (glycol nucleic acid), TNA (threose nucleic acid), 2'-O-Me RNA, 2'-fluoro RNA, morpholino nucleotides, and any combination thereof.
[0133] End Treatment In some embodiments, the methods herein include a nucleic acid composition comprising a single-stranded nucleic acid (ssNA) and a nucleic acid having end-processing activity. agent and single-stranded nucleic acid (ssNA) molecules end The method includes contacting the ssNA composition under conditions for treatment, thereby producing a terminally treated ssNA composition. Terminal treatment can include, but is not limited to, phosphorylation, dephosphorylation, methylation, demethylation, oxidation, deoxidation, base modification, extension, polymerization, and combinations thereof. Terminal treatment can be performed using enzymes, including, but not limited to, ligase, polynucleotide kinase (PNK), terminal transferase, methyltransferase, methylase (e.g., 3' methylase, 5' methylase), polymerase (e.g., polyA polymerase), oxidase, and combinations thereof.
[0134] In some embodiments, the methods herein include a nucleic acid composition comprising a single-stranded nucleic acid (ssNA) and a nucleic acid having phosphatase activity. agent and a single-stranded nucleic acid (ssNA) molecule under conditions in which the single-stranded nucleic acid (ssNA) molecule is dephosphorylated, thereby producing a dephosphorylated ssNA composition. In some embodiments, the method herein comprises contacting a scaffold adaptor or a component thereof with a single-stranded nucleic acid (ssNA) molecule having phosphatase activity. agent and a scaffold adaptor or a component thereof under conditions in which the scaffold adaptor or a component thereof is dephosphorylated, thereby producing a dephosphorylated scaffold adaptor or a component thereof (e.g., a dephosphorylated oligonucleotide; a dephosphorylated scaffold polynucleotide). Generally, the ssNA composition, and / or the scaffold adaptor or a component thereof, Combination The ssNA is dephosphorylated prior to the step of cleaving (i.e., prior to hybridization). Combination Before the hybridization step (i.e., before hybridization), the protein is dephosphorylated and then phosphorylated. profit Scaffold adaptors, or components thereof, may be: Combination Before the hybridization step (i.e., before hybridization), the protein is dephosphorylated and then phosphorylated. profit Scaffold adaptors, or components thereof, may be: Combination Before the step of hybridization (i.e., before hybridization), the protein is dephosphorylated and then phosphorylated. profit Na stomach. Scaffold adaptors, or components thereof, Combination Before the step of allowing the nucleotides to react (i.e., before hybridization), the nucleotides are dephosphorylated, unphosphorylated, and then Combination phosphorylated after the hybridization step (i.e., after hybridization) and before or during the ligation step profitReagents and kits for dephosphorylating nucleic acids are known and available. For example, a target nucleic acid (e.g., ssNA) and / or a scaffold adaptor or a component thereof can be treated with a phosphatase (i.e., an enzyme that uses water to cleave phosphate monoesters to phosphate ions and alcohols).
[0135] In some embodiments, the methods herein include a nucleic acid composition comprising a single-stranded nucleic acid (ssNA) and a nucleic acid having phosphoryl transfer activity. agent In some embodiments, the method herein comprises contacting a dephosphorylated ssNA composition with a 5' phosphate-containing ssNA having phosphoryl transfer activity under conditions in which a 5' phosphate is added to the 5' end of the ssNA. agent In some embodiments, the method herein comprises contacting a scaffold adaptor or a component thereof with a 5' phosphate-transferase-active nucleotide under conditions in which a 5' phosphate is added to the 5' end of the ssNA. agent and a dephosphorylated scaffold adaptor or component thereof under conditions whereby a 5' phosphate is added to the 5' end of the scaffold adaptor or component thereof. In some embodiments, the methods herein comprise contacting a dephosphorylated scaffold adaptor or component thereof with a phosphorylated Lunar Transfer Has transfer activity agent and a 5' phosphate under conditions whereby a 5' phosphate is added to the 5' end of the scaffold adaptor or component thereof. In certain instances, the ssNA composition, and / or the scaffold adaptor or component thereof, comprises: Combination The 5' phosphorylation of the nucleic acid can be achieved by various techniques. For example, the ssNA composition and / or the scaffold adapter or components thereof can be treated with polynucleotide kinase (PNK) (e.g., T4 PNK), which catalyzes the transfer and exchange of Pi from the γ-position of ATP to the 5'-hydroxyl terminus and nucleoside 3'-monophosphate of polynucleotides (double-stranded and single-stranded DNA and RNA). Suitable reaction conditions include, for example, treating the nucleic acid with PNK in 1x PNK reaction buffer. liquidIncubation at 37°C for 30 minutes in a buffer (e.g., 70 mM Tris-HCl, 10 mM MgCl2, 5 mM DTT, pH 7.6 at 25°C); and T4 DNA ligase buffer for nucleic acids and PNK. liquid (e.g., 50 mM Tris-HCl, 10 mM MgCl, 1 mM ATP, 10 mM DTT, pH 7.5 at 25°C) for 30 minutes at 37°C. If necessary, after the phosphorylation reaction, PNK may be heat inactivated, for example, at 65°C for 20 minutes.
[0136] In some embodiments, the methods herein involve the use of a method comprising administering to a subject a subject having phosphoryl transfer activity. agent In some embodiments, the method does not include the use of a 5'-phosphorylated ssNA by phosphorylating the 5'-end of the ssNA from the nucleic acid sample. In certain cases, the nucleic acid sample contains ssNA with a naturally phosphorylated 5'-end. In some embodiments, the method does not include the use of a 5'-phosphorylated scaffold adapter or a component thereof by phosphorylating the 5'-end of the scaffold adapter or a component thereof.
[0137] Disconnect In some embodiments, the ssNA, scaffold adapter, and / or hybridization product (e.g., a scaffold adapter hybridized with an ssNA) is cleaved or sheared before, during, or after the methods described herein. In some embodiments, the ssNA, scaffold adapter, and / or hybridization product is cleaved or sheared at a cleavage site. In some embodiments, the scaffold adapter and / or hybridization product is cleaved or sheared at a cleavage site within a hairpin loop. In some embodiments, the scaffold adapter and / or hybridization product is cleaved or sheared at a cleavage site at an internal position of the scaffold adapter (e.g., within the duplex region of the scaffold adapter). In some embodiments, the scaffold adapter is cleaved at a cleavage site (e.g., uracil) that is present only on the scaffold polynucleotide and not on the complementary oligonucleotide component. Thus, in some embodiments, the scaffold polynucleotide comprises one or more uracil bases, and the oligonucleotide component does not comprise a uracil base. In some embodiments, the circular hybridization product is cleaved or sheared before, during, or after the methods described herein. In some embodiments, nucleic acids, e.g., circular nucleic acids and / or large fragments (e.g., greater than 500 base pairs in length), are cleaved or sheared before, during, or after the methods described herein. Large fragments may be referred to as high molecular weight (HMW) nucleic acids, HMW DNA, or HMW RNA. HMW nucleic acid fragments may include fragments greater than about 500 bp, about 600 bp, about 700 bp, about 800 bp, about 900 bp, about 1000 bp, about 2000 bp, about 3000 bp, about 4000 bp, about 5000 bp, about 10,000 bp, or longer. The terms "shearing" or "cleaving" refer to the process of separating a nucleic acid molecule into two (or more) smaller nucleic acid molecules. Detach Such shearing or cleavage can be sequence-specific, base-specific, or non-specific, and can involve a variety of methods, reagents, or conditions, including, for example, chemical, enzymatic, and physical (e.g., physical fragmentation). EitherThe sheared or cleaved nucleic acids can have a nominal, representative, or average length of about 5 to about 10,000 base pairs, about 100 to about 1,000 base pairs, about 100 to about 500 base pairs, or about 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, or 9000 base pairs.
[0138] Sheared or cleaved nucleic acids can be produced by any suitable method, including, but not limited to, physical methods (e.g., shearing, e.g., sonication, ultrasonication, French press, heat, UV irradiation, etc.), enzymatic processes (e.g., enzymatic cleavage, etc.), and the like. Agent (e.g., suitable nucleases, suitable restriction enzymes), chemical methods (e.g., alkylation, DMS, piperidine, acid hydrolysis, base hydrolysis, heat, etc., or a combination thereof), ultraviolet (UV) light (e.g., a photocleavable moiety (e.g., comprising a photocleavable spacer)), or a combination thereof. The representative, average, or nominal length of the resulting nucleic acid fragments can be controlled by selecting an appropriate fragment generation method.
[0139] The term "cleavage agent" (cleavage agent) " generally refers to a nucleic acid that can cleave a nucleic acid at one or more specific or non-specific sites. agent , sometimes refers to a chemical substance or enzyme. Often, a specific cleaving agent specifically cleaves at a specific site, sometimes called a cleavage site, according to a particular nucleotide sequence. Cleavage agents include enzymatic cleaving agents, chemical cleaving agents, and light (e.g., ultraviolet (UV)) light ).
[0140] enzymatic cleavage AgentExamples include, but are not limited to, endonucleases; deoxyribonucleases (DNases; e.g., DNase I, II); ribonucleases (RNases; e.g., RNAse A, RNAse E, RNAse F, RNAse H, RNAse III, RNAse L, RNAse P, RNAse PhyM, RNAse T1, RNAse T2, RNAse U2, and RNAse V); endonuclease VIII; CLEAVASE enzymes; TAQ DNA polymerase; E. coli DNA polymerase I; eukaryotic structure-specific endonucleases; mouse FEN-1 endonuclease; nicking enzymes; type I, II, or III restriction endonucleases (i.e., restriction enzymes), such as Acc I, Aci I, Afl III, Alu I, Alw44 I, Apa I, Asn I, Ava I, Ava II, BamH I, Ban II, Bcl I, Bgl I.Bgl II, Bln I, Bsm I, BssH II, BstE II, BstUI, Cfo I, CIa I, Dde I, Dpn I, Dra I, EcIX I, EcoR I, EcoR I, EcoR II, EcoR V, Hae II, Hae II, HhaI, Hind II, Hind III, Hpa I, Hpa II, Kpn I, Ksp I, MaeII, McrBC, Mlu I, MIuN I, Msp I, Nci I, Nco I, Nde I, Nde II, Nhe I, Not I, Nru I, Nsi I, Pst I, Pvu I, Pvu II, Rsa I, Sac I, Sal I, Sau3A I, Sca I, ScrF I, Sfi I, Sma I, Spe I, Sph I, Ssp I, Stu I, Sty I, Swa I, Taq I, Xba I, Xho I; glycosylases (e.g., uracil-DNA glycosylase (UDG), 3-methyladenine DNA glycosylase, 3-methyladenine DNA glycosylase II, pyrimidine hydrate-DNA glycosylase, FaPy-DNA glycosylase, thymine mismatch-DNA glycosylases (e.g., hypoxanthine-DNA glycosylase, uracil-DNA glycosylase (UDG), 5-hydroxymethyluracil-DNA glycosylase (HmUDG), 5-hydroxymethylcytosine-DNA glycosylase, or 1,N6-etheno- exonucleases (e.g., exonuclease I, exonuclease II, exonuclease III, exonuclease IV, exonuclease V, exonuclease VI, exonuclease VII, exonuclease VIII); 5' to 3' exonucleases (e.g., exonuclease II); 3' to 5' exonucleases (e.g., exonuclease I); poly(A)-specific 3' to 5' exonucleases; ribozymes; DNAzymes; and the like, and combinations thereof.
[0141] In some embodiments, the cleavage site comprises a restriction enzyme recognition site. Agentcomprises a restriction enzyme. In some embodiments, the cleavage site comprises a rare-cutter restriction enzyme recognition site (e.g., a NotI recognition sequence). In some embodiments, the cleavage site Agent These include rare-cutter enzymes (e.g., rare-cutter restriction enzymes). Rare-cutter enzymes generally refer to restriction enzymes with recognition sequences that are extremely rare in genomes (e.g., the human genome). One example is NotI, which cuts after the first GC in the 5'-GCGGCCGC-3' sequence. Restriction enzymes with 7 and 8 base pair recognition sequences are often considered rare-cutter enzymes.
[0142] Procedures for selecting cleavage methods and restriction enzymes for cleaving DNA at specific sites are well known to those skilled in the art. Many restriction enzyme suppliers, including, for example, New England BioLabs, Pro-Mega Biochems, Boehringer-Mannheim, etc., provide information on the conditions and types of DNA sequences cleaved by specific restriction enzymes. Enzymes are those that allow cleavage of DNA with about 95% to 100% efficiency, preferably about 98% to 100% efficiency. Article It is often used under the following circumstances.
[0143] In some embodiments, the cleavage site comprises one or more ribonucleic acid (RNA) nucleotides. In some embodiments, the cleavage site comprises a single-stranded portion comprising one or more RNA nucleotides. In some embodiments, the single-stranded portion comprises a double-stranded portion. sandwiched betweenIn some embodiments, the single-stranded portion is a hairpin loop. In some embodiments, the cleavage site comprises one RNA nucleotide. In some embodiments, the cleavage site comprises two RNA nucleotides. In some embodiments, the cleavage site comprises three RNA nucleotides. In some embodiments, the cleavage site comprises four RNA nucleotides. In some embodiments, the cleavage site comprises five RNA nucleotides. In some embodiments, the cleavage site comprises more than five RNA nucleotides. In some embodiments, the cleavage site comprises one or more RNA nucleotides selected from adenine (A), cytosine (C), guanine (G), and uracil (U). In some embodiments, the cleavage site comprises one or more RNA nucleotides selected from adenine (A), cytosine (C), and guanine (G). In some embodiments, the cleavage site does not comprise uracil (U). In some embodiments, the cleavage site comprises one or more RNA nucleotides that comprise guanine (G). In some embodiments, the cleavage site comprises one or more RNA nucleotides that consist of guanine (G). In some embodiments, the cleavage site comprises one or more RNA nucleotides that comprise cytosine (C). In some embodiments, the cleavage site comprises one or more RNA nucleotides consisting of cytosine (C). In some embodiments, the cleavage site comprises one or more RNA nucleotides consisting of adenine (A). In some embodiments, the cleavage site comprises one or more RNA nucleotides consisting of adenine (A). In some embodiments, the cleavage site comprises one or more RNA nucleotides consisting of adenine (A), cytosine (C), and guanine (G). In some embodiments, the cleavage site comprises one or more RNA nucleotides consisting of adenine (A) and cytosine (C). In some embodiments, the cleavage site comprises one or more RNA nucleotides consisting of adenine (A) and guanine (G). In some embodiments, the cleavage site comprises one or more RNA nucleotides consisting of cytosine (C) and guanine (G). In some embodiments, the cleavage site comprises one or more RNA nucleotides consisting of cytosine (C) and guanine (G). Agentcomprises a ribonuclease (RNAse). In some embodiments, the RNAse is an endoribonuclease. The RNAse can be selected from one or more of RNAse A, RNAse E, RNAse F, RNAse H, RNAse III, RNAse L, RNAse P, RNAse PhyM, RNAse T1, RNAse T2, RNAse U2, and RNAse V.
[0144] In some embodiments, the cleavage site comprises a photocleavable spacer or a photocleavable modification, for example, a photocleavable spacer or a photocleavable modification that can be cleaved by ultraviolet (UV) light of a specific wavelength (e.g., 300-350 nm). light An example of a photocleavable spacer (available from Integrated DNA Technologies; product no. 1707) is cleavable by UV irradiation within the appropriate spectral range. light The photocleavable spacer is a 10-atom linker arm that can be cleaved only when exposed to ultraviolet (UV) light. Oligonucleotides containing a photocleavable spacer can have a 5' phosphate group available for subsequent ligase reactions. The photocleavable spacer can be placed between the DNA bases or between the oligo and the terminal modification (e.g., a fluorophore). In such embodiments, ultraviolet (UV) light can be used to cleave the oligo. light Disconnect Agent It can be thought of as follows.
[0145] In some embodiments, the cleavage site comprises a diol. For example, the cleavage site may comprise vicinal diols incorporated at a 5'-to-5' linkage. A cleavage site comprising a diol can be chemically cleaved, for example, using periodate. In some embodiments, the cleavage site comprises a blunt-end restriction enzyme recognition site. A cleavage site comprising a blunt-end restriction enzyme recognition site can be cleaved by a blunt-end restriction enzyme.
[0146] Nick seal and fill-in In some embodiments, the methods herein include performing a nick-sealing reaction (e.g., using a DNA ligase or other suitable enzyme, and in certain cases, a kinase configured to 5' phosphorylate the nucleic acid (e.g., using polynucleotide kinase (PNK)). In some embodiments, the methods herein include performing a fill-in reaction. For example, if the scaffold adapter exists as a duplex, some or all of the duplex may include an overhang at the end of the duplex opposite the end that hybridizes to the ssNA. If such a duplex overhang is present, the methods herein include: Combination After the step of hybridizing, the method may further comprise a step of filling in the overhang formed by the duplex. In some embodiments, the fill-in reaction is performed to generate a blunt-ended hybridization product. Any suitable reagents can be used to perform the fill-in reaction. Suitable polymerases for performing the fill-in reaction include, for example, DNA polymerase I, the large (Klenow) fragment of DNA polymerase I, T4 DNA polymerase, Bacillus stearothermophilus (Bst) DNA polymerase, and thermostable DNA polymerases (e.g., superoxide dismutase). good from a thermophilic marine archaea), 9°N™ DNA polymerase (GENBANK accession number AAA88769.1), THERMINATOR polymerase (strange In some embodiments, a strand-displacing polymerase is used (e.g., Bst DNA polymerase).
[0147] Exonuclease treatment In some embodiments, nucleic acids (e.g., RNA-DNA duplexes, hybridization products, circularized hybridization products) are treated with an exonuclease. In some embodiments, RNA in an RNA-DNA duplex (e.g., an RNA-DNA duplex generated by first-strand cDNA synthesis) is treated with an exonuclease. An exonuclease is an enzyme that acts by cleaving nucleotides one by one from the end of a polynucleotide chain by a hydrolysis reaction that cleaves phosphodiester bonds at either the 3' or 5' end. Exonucleases include, for example, DNAses, RNAses (e.g., RNAse H), 5' to 3' exonucleases (e.g., exonuclease II), 3' to 5' exonucleases (e.g., exonuclease I), and poly(A)-specific 3' to 5' exonucleases. In some embodiments, the exonuclease activity is provided by a reverse transcriptase (e.g., RNAse activity provided by M-MLV reverse transcriptase with a fully functional RNAse H domain). In some embodiments, the hybridization product is treated with an exonuclease to remove contaminating nucleic acids, such as, for example, single-stranded oligonucleotides, nucleic acid fragments, or RNA from an RNA-DNA duplex. In some embodiments, the circularized hybridization product is any The nucleic acid is treated with an exonuclease to remove non-circularized hybridization products, unhybridized oligonucleotides, unhybridized target nucleic acids, oligonucleotide dimers, etc., and combinations thereof.
[0148] sample Methods and compositions for processing and / or analyzing nucleic acids are provided herein. The nucleic acids or nucleic acid mixtures utilized in the methods and compositions described herein can be isolated from a sample obtained from a subject (e.g., a test subject). A subject can be any living organism, including, but not limited to, a human, a non-human animal, a plant, a bacterium, a fungus, a protist, or a pathogen. thing or non-living thingAny human or non-human animal can be selected, including, for example, mammals, reptiles, birds, amphibians, fish, ungulates, ruminants, bovine animals (e.g., cows), equid animals (e.g., horses), goats, etc. (caprine) and sheep (ovine) (e.g. sheep (sheep) ,goat (goat) ), Swine (e.g., pigs (pig) ), camel kind (e.g., camels, llamas, alpacas), monkeys, apes (e.g., gorillas, chimpanzees), bears Department( The subject may be a mammal, such as a bear, poultry, dog, cat, mouse, rat, fish, dolphin, whale, or shark. The subject may be male or female (e.g., a woman or a pregnant woman). The subject may be of any age (e.g., embryo, fetus, infant, child, adult). The subject may be a cancer patient, a patient suspected of having cancer, a patient in remission, a patient with a family history of cancer, and / or a subject undergoing cancer screening. The subject may be a cancer patient, a patient suspected of having cancer, a patient in remission, a patient with a family history of cancer, and / or a subject undergoing cancer screening. Dyeing or infectious diseases have The subject may be a patient infected with a pathogen (e.g., bacteria, virus, fungus, protozoan, etc.), a patient suffering from an infection or infectious disease or suspected of being infected with a pathogen, a patient who has recovered from an infection, infectious disease, or pathogen infection, a patient with a history of an infection, infectious disease, or pathogen infection, and / or a subject undergoing infectious disease or pathogen screening. The subject may be a transplant recipient. The subject may be a patient undergoing microbiome analysis. In some embodiments, the subject is female. In some embodiments, the subject is a human woman In some embodiments, the subject is male. In some embodiments, the subject is human. male is.
[0149] A nucleic acid sample can be isolated or obtained from any type of suitable biological specimen or sample (e.g., a test sample). A nucleic acid sample can be isolated or obtained from a single cell, multiple cells (e.g., cultured cells), cell culture medium, conditioned medium, tissue, organ, or organism (e.g., bacteria, yeast, etc.). In some embodiments, a nucleic acid sample is isolated or obtained from cells of an animal (e.g., an animal subject). (Multiple options possible) , tissues, organs etc. In some embodiments, the nucleic acid sample is isolated or obtained from a source such as bacteria, yeast, insects (e.g., Drosophila), mammals, amphibians (e.g., frogs (e.g., Xenopus)), viruses, plants, or any other mammalian or non-mammalian nucleic acid sample source.
[0150] The nucleic acid sample can be isolated or obtained from a living organism or animal. Depending on the situation The nucleic acid sample can be isolated or obtained from an extinct (or "archaic") organism or animal (e.g., an extinct mammal; an extinct mammal from the genus Homo). Depending on the situation The nucleic acid sample may be obtained as part of a diagnostic assay.
[0151] Depending on the situation Nucleic acid samples can be obtained as part of forensic analysis. In some embodiments, the single-stranded nucleic acid library preparation (ssPrep) method described herein is applied to a forensic sample or specimen. A forensic sample or specimen can include any biological material containing nucleic acids. For example, a forensic sample or specimen can include blood, semen, hair, skin, sweat, saliva, decomposed tissue, bone, fingernail scrapings, licked stamps / envelopes, exfoliation material, touch DNA, razor residue, etc.
[0152] The sample or test sample may be from a subject or part thereof (e.g., a human subject, a pregnant female, a cancer patient, a patient with an infection or infectious disease, haveThe sample can be any specimen isolated or obtained from a patient, transplant recipient, fetus, tumor, infected organ or tissue, transplanted organ or tissue, microbiome, or any other organ or tissue during pregnancy (e.g., for human subjects, No. 1 , second or third trimester The sample may be from a pregnant female subject carrying a fetus with all chromosomes euploid, or from a subject after birth. The sample may be from a pregnant subject carrying a fetus with all chromosomes euploid, or from a pregnant subject carrying a fetus with a chromosomal aneuploidy (e.g., 1, 3 (i.e., trisomy (e.g., T21, T18, T13)) or 4 copies of a chromosome) or other genetic variation. Non-limiting examples of specimens include blood or blood product (e.g., serum, plasma, etc.), umbilical cord blood, chorionic villi, amniotic fluid, cerebrospinal fluid, spinal fluid, lavage fluid (lavage fluid) (e.g., bronchoalveolar, gastric, peritoneal, ductal (ductal) , ear, arthroscopic), biopsy samples (e.g., from preimplantation embryos; cancer biopsies), body cavity aspiration samples (celocentesis sample) , cells (blood cells, placental cells, embryonic or fetal cells, fetal nucleated cells or fetal cells) Remnant , normal cells, abnormal cells (e.g., cancer cells)) or their parts (e.g., mitochondrial, nuclear, extracts, etc.), washings of the female reproductive tract, urine, feces, sputum, saliva, nasal mucosa, prostatic fluid, lavage lavage In some embodiments, the biological sample may be a bodily fluid or tissue from a subject, including, but not limited to, semen, lymph, bile, tears, sweat, breast milk, milk, etc., or a combination thereof. In some embodiments, the biological sample is a cervical swab from a subject. The bodily fluid or tissue sample from which nucleic acids are extracted may be acellular (e.g., acellular). In some embodiments, the bodily fluid or tissue sample may contain cellular elements or cells. Remnant In some embodiments, the sample may contain fetal cells or cancer cells.
[0153] The sample is a liquid sample. Great deal The liquid sample contains extracellular nucleic acids (e.g., circulating cell-free DNA). MitekuExamples of liquid samples include blood or blood product (e.g., serum, plasma, etc.), urine, cerebrospinal fluid, saliva, sputum, a biopsy sample (e.g., a liquid biopsy for the detection of cancer), a liquid sample of the above, etc., or a combination thereof. In certain embodiments, the sample is a liquid biopsy, and the liquid biopsy is used to detect a disease. (e.g., cancer) The existence, non-existence, exacerbation Liquid biopsy generally refers to the evaluation of a liquid sample from a subject for progression or remission. Liquid biopsy may be used in conjunction with or in place of a solid biopsy (e.g., tumor biopsy). Replacement In certain instances, extracellular nucleic acids are analyzed in liquid biopsies.
[0154] In some embodiments, the biological sample may be blood, plasma, or serum. The term "blood" encompasses the conventional definition of whole blood, blood products, or any fraction of blood, such as serum, plasma, buffy coat, etc. Blood or fractions thereof often contain nucleosomes. Nucleosomes contain nucleic acids and may be acellular or intracellular. Blood also includes buffy coats. Buffy coats may be isolated by utilizing a Ficoll gradient. Buffy coats contain white blood cells. (white blood cell) (e.g., white blood cells (leukocyte) Plasma refers to the fraction of whole blood obtained as a result of centrifugation of blood that has been treated with an anticoagulant. Serum refers to the watery part of the fluid that remains after a blood sample has clotted. Body fluid or tissue samples are typically collected by hospitals or clinics. Engaged According to standard protocols collection In the case of blood, an appropriate amount of peripheral blood (e.g., between 3 and 40 milliliters, between 5 and 50 milliliters) is collection The formulation may be prepared in a variety of ways, and may be stored according to standard procedures before or after preparation.
[0155] found in the subject's blood Served Analysis of the nucleic acids found in maternal blood can be performed, for example, using whole blood, serum, or plasma. ServedAnalysis of fetal DNA can be performed, for example, using whole blood, serum, or plasma. Served Analysis of the tumor or cancer DNA can be performed, for example, using whole blood, serum, or plasma. Served Analysis of pathogen DNA can be performed, for example, using whole blood, serum, or plasma. For example, Served Analysis of the transferred DNA can be performed using, for example, whole blood, serum, or plasma. Methods for preparing serum or plasma from blood obtained from a subject (e.g., a maternal subject; a patient; a cancer patient) are known. For example, the subject's blood (e.g., a pregnant woman's blood; a patient's blood; a cancer patient's blood) can be placed in a tube containing EDTA or a dedicated commercial product, such as Cell-Free DNA BCT (Streck, Omaha, NE) or Vacutainer SST (Becton Dickinson, Franklin Lakes, NJ), to prevent blood clotting, and plasma can then be obtained from the whole blood by centrifugation. Serum can be obtained with or without centrifugation after blood clotting. When centrifugation is used, the centrifugation is usually, but not exclusively, performed at an appropriate speed, e.g., 1,500 to 3,000 x g. × The plasma or serum may be subjected to an additional centrifugation step before being transferred to a fresh tube for nucleic acid extraction. In addition to the acellular portion of whole blood, the cellular fraction may be separated into the buffy coat portion, which can be obtained after centrifugation of a whole blood sample from a subject and removal of the plasma. Enriched Nucleic acids can also be recovered.
[0156] The sample may be a tumor nucleic acid sample (i.e., a nucleic acid sample isolated from a tumor). The term "tumor" generally refers to neoplastic cell growth and proliferation, whether malignant or benign, and may include pre-cancerous and cancerous cells and tissues. The terms "cancer" and "cancerous" generally refer to the physiological condition in mammals that is typically characterized by unregulated cell growth / proliferation. Examples of cancer include: cancer,These include, but are not limited to, lymphoma, blastoma, sarcoma, leukemia, squamous cell carcinoma, small cell lung cancer, non-small cell lung cancer, lung adenocarcinoma, lung squamous cell carcinoma, peritoneal cancer, hepatocellular carcinoma, gastrointestinal cancer, pancreatic cancer, glioblastoma, cervical cancer, ovarian cancer, liver cancer, bladder cancer, hepatoma, breast cancer, colon cancer, colorectal cancer, endometrial or uterine cancer, salivary gland cancer, kidney cancer, liver cancer, prostate cancer, vulvar cancer, thyroid cancer, liver cancer, and various types of head and neck cancer.
[0157] A sample can be heterogeneous, for example, a sample can contain more than one cell type and / or one or more nucleic acid species. Depending on the situation In some embodiments, the sample may contain (i) fetal and maternal cells, (ii) cancer and non-cancerous cells, and / or (iii) pathogenic and host cells. Depending on the situation The sample may be divided into (i) cancer and non-cancer nucleic acids, (ii) pathogen and host nucleic acids, (iii) fetal and maternal nucleic acids, and / or more generally, (iv )strange It may contain variant nucleic acids and wild-type nucleic acids. Depending on the situation The samples are described in more detail below. description A little like number nucleus acid species (minority nucleic acid species) and many number nucleus acid species (majority nucleic acid species) It may include: Depending on the situation A sample may contain cells and / or nucleic acid from a single subject, or may contain cells and / or nucleic acid from multiple subjects.
[0158] nucleic acid Methods and compositions for processing and / or analyzing nucleic acids are provided herein. (Multiple options possible) , nucleic acid molecule (Multiple options possible) , nucleic acid fragment (Multiple options possible) , target nucleic acid (Multiple options possible) , nucleic acid template (Multiple options possible) , template nucleic acid (Multiple options possible) , nucleic acid target (Multiple options possible) , target nucleic acid (Multiple options possible) , polynucleotides (Multiple options possible) , polynucleotide fragment (Multiple options possible) , target polynucleotide (Multiple options possible), polynucleotide target (Multiple options possible) Terms such as interchangeably These terms may be used to refer to DNA (e.g., complementary DNA (cDNA; synthesized from any RNA or DNA of interest), genomic DNA (gDNA), genomic DNA fragments, mitochondrial DNA (mtDNA), recombinant DNA (e.g., plasmid DNA), etc.), RNA (e.g., messenger RNA (mRNA), Small molecule inhibition RNA (siRNA), ribosomal RNA (rRNA), transfer RNA (tRNA), microRNA, trans-acting small interfering RNA (ta-siRNA), naturally occurring small interfering RNA (nat-siRNA), small nucleolar RNA (snoRNA), small nuclear RNA (snRNA), long non-coding RNA (lncRNA), non-coding RNA (ncRNA), transfer messenger RNA (tmRNA), precursor messenger RNA (pre-mRNA), small Cajal body-specific RNA (scaRNA), piwi interaction RNA (piRNA), endoribonuclease-prepared siRNA (esiRNA), Temporary low molecular RNA (stRNA), signal recognition RNA, telomeric RNA, RNA highly expressed by the fetus or placenta, etc.), and / or DNA or RNA analogs (e.g., containing base analogs, sugar analogs, and / or non-native backbones, etc.), RNA / DNA hybrids, and polyamide nucleic acids (PNAs), theAll of these may be in single-stranded or double-stranded form, and unless otherwise limited, refer to nucleic acids of any composition, including known analogs of natural nucleotides that can function similarly to naturally occurring nucleotides. Nucleic acids may be or be derived from plasmids, phages, viruses, bacteria, autonomously replicating sequences (ARS), mitochondria, centromeres, artificial chromosomes, chromosomes, or in certain embodiments, other nucleic acids that can replicate or be replicated in vitro or in the nucleus or cytoplasm of a host cell, cell, cell. In some embodiments, the template nucleic acid may be derived from a single chromosome (e.g., a nucleic acid sample may be derived from one chromosome of a sample obtained from a diploid organism). Unless otherwise limited, this term encompasses nucleic acids that have similar binding properties to the reference nucleic acid and are metabolized similarly to naturally occurring nucleotides, including known analogs of natural nucleotides. Unless otherwise indicated, a particular nucleic acid sequence implicitly encompasses not only the sequence explicitly indicated, but also its conservatively modified variants (e.g., degenerate codon substitutions), alleles, orthologs, single nucleotide polymorphisms (SNPs), and complementary sequences. Specifically, degenerate codon substitutions can be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and / or deoxyinosine residues. The term nucleic acid refers to a locus, a gene, a cDNA, and an mRNA encoded by a gene. interchangeablyThe term may also be used interchangeably to refer to RNA or DNA derivatives, variants, and analogs synthesized from nucleotide analogs, single-stranded ("sense" or "antisense," "plus" or "minus" strands, "forward" or "reverse" reading frames), and double-stranded polynucleotides. The term "gene" refers to a segment of DNA involved in producing a polypeptide chain; it generally includes regions preceding and following the coding region (leader and trailer) involved in the transcription / translation of the gene product and in the regulation of transcription / translation, as well as intervening sequences (introns) between individual coding regions (exons). Nucleotides or bases generally refer to the purine and pyrimidine molecular units of nucleic acids (e.g., adenine (A), thymine (T), guanine (G), and cytosine (C)). For RNA, the base thymine is replaced by uracil. Nucleic acid length or size can be expressed in terms of the number of bases.
[0159] The target nucleic acid can be any nucleic acid of interest. The nucleic acid can be a polymer of any length composed of deoxyribonucleotides (i.e., DNA bases), ribonucleotides (i.e., RNA bases), or a combination thereof, for example, 10 bases or more, 20 bases or more, 50 bases or more, 100 bases or more, 200 bases or more, 300 bases or more, 400 bases or more, 500 bases or more, 1000 bases or more, 2000 bases or more, 3000 bases or more, 4000 bases or more, 5000 bases or more. In certain embodiments, a nucleic acid is a polymer made up of deoxyribonucleotides (i.e., DNA bases), ribonucleotides (i.e., RNA bases), or combinations thereof, e.g., 10 or fewer bases, 20 or fewer bases, 50 or fewer bases, 100 or fewer bases, 200 or fewer bases, 300 or fewer bases, 400 or fewer bases, 500 or fewer bases, 1000 or fewer bases, 2000 or fewer bases, 3000 or fewer bases, 4000 or fewer bases, or 5000 or fewer bases.
[0160] Nucleic acid is a single With chains There may be one or two With chains For example, single-stranded DNA (ssDNA) can be generated by denaturing double-stranded DNA, for example, by heating or by treating with alkali. Thus, in some embodiments, ssDNA is derived from double-stranded DNA (dsDNA). In some embodiments, the methods herein involve combining a nucleic acid composition comprising dsDNA with a scaffold adapter or a component thereof as described herein. Combination The method includes a step of denaturing the dsDNA prior to the step of denaturing the dsDNA, thereby producing ssDNA.
[0161] In certain embodiments, the nucleic acid is in a D-loop structure formed by strand invasion of a double-stranded DNA molecule by an oligonucleotide or DNA-like molecule, such as a peptide nucleic acid (PNA). For example, D-loop formation can be inhibited by the addition of E. coli RecA protein and / or by altering salt concentration using methods known in the art. promotion It is possible.
[0162] Nucleic acids (e.g., nucleic acid targets, single-stranded nucleic acids (ssNAs), oligonucleotides, overhangs, scaffold polynucleotides, and their hybridization regions (e.g., ssNA hybridization regions, oligonucleotide hybridization regions)) may be described herein as being complementary to, having a region of complementarity with, being capable of hybridizing to, or having a hybridization region for, another nucleic acid. The terms "complementary" or "complementarity" or "hybridization" generally refer to nucleotide sequences that noncovalently base pair to a region of a nucleic acid (e.g., the nucleotide sequence of an ssNA hybridization region that hybridizes with a terminal region of an ssNA fragment and the nucleotide sequence of an oligonucleotide hybridization region that hybridizes with an oligonucleotide component of a scaffold adapter). In standard Watson-Crick base pairing, in DNA, adenine (A) base pairs with thymine (T), and guanine (G) base pairs with cytosine (C). In the case of RNA, thymine (T) is replaced by uracil (U). Thus, A is complementary to T and G is complementary to C. In the case of RNA, A is complementary to U, and vice versa. In the case of a DNA-RNA duplex, A (in the DNA strand) is complementary to U (in the RNA strand). In some embodiments, one or more thymine (T) bases are replaced by uracil (U) in the scaffold adapter or a component thereof, which is complementary to adenine (A). Typically, "complementary" or "complementarity" or "capable of hybridizing" refers to nucleotide sequences that are at least partially complementary. These terms can also encompass duplexes that are fully complementary, such that every nucleotide in one strand is complementary to or hybridizes with every nucleotide at the corresponding position in the other strand.
[0163] In certain cases, the nucleotide sequence may be partially complementary to the target, in which case all nucleotides are At all corresponding positions in target nucleic acid NoaFor example, the ssNA hybridization region may not be completely complementary to the target ssNA terminal region. all Alternatively, the ssNA hybridization region may be completely complementary (i.e., 100%) to the all In another example, the oligonucleotide hybridization region may share some degree of complementarity, but not all (e.g., 70%, 75%, 85%, 90%, 95%, 99%). all Alternatively, the oligonucleotide hybridization region may be completely complementary (i.e., 100%) to the target site. all They may share some degree of complementarity (e.g., 70%, 75%, 85%, 90%, 95%, 99%), but not all.
[0164] The percent identity of two nucleotide sequences can be determined by aligning the sequences for optimal comparison (e.g., gaps can be introduced into the sequence of the first sequence for optimal alignment). In that case, nucleotides at corresponding positions are compared, and the percent identity between the two sequences is a function of the number of identical positions shared by the sequences (i.e., % identity = number of identical positions / total number of positions × 100). When a position in one sequence is occupied by the same nucleotide as the corresponding position in the other sequence, then the molecules are identical at that position.
[0165] In some embodiments, nucleic acids in a mixture of nucleic acids are analyzed. The mixture of nucleic acids can include two or more nucleic acid species having the same or different nucleotide sequences, different lengths, different origins (e.g., genomic origin, fetal vs. maternal origin, cell or tissue origin, cancer vs. non-cancer origin, tumor vs. non-tumor origin, host vs. pathogen, host vs. graft, host vs. microbiome, sample origin, subject origin, etc.), different overhang lengths, different overhang types (e.g., 5' overhang, 3' overhang, no overhang), or combinations thereof. In some embodiments, the mixture of nucleic acids includes single-stranded nucleic acids and double-stranded nucleic acids. In some embodiments, the mixture of nucleic acids includes DNA and RNA. In some embodiments, the mixture of nucleic acids includes ribosomal RNA (rRNA) and messenger RNA (mRNA). The processes described herein can be used to analyze nucleic acids. proposal The provided nucleic acids may be from one sample or from two or more samples (e.g., one or It's more , 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more, or 20 or more samples).
[0166] In some embodiments, the target nucleic acid (e.g., ssNA) comprises degraded DNA. Degraded DNA may be referred to as low-quality DNA or highly degraded DNA. Degraded DNA may be highly fragmented and may contain damage such as base analogs and abasic sites that are prone to miscoding and / or intermolecular crosslinking. For example, sequencing errors resulting from deamination of cytosine residues may be present in certain sequences obtained from degraded DNA (e.g., miscoding C to T and G to A). In some embodiments, the target nucleic acid (e.g., ssNA) is derived from a nicked double-stranded nucleic acid fragment. The nicked double-stranded nucleic acid fragment can be denatured (e.g., heat denatured) to generate ssNA fragments.
[0167] Nucleic acids can be obtained from one or more sources (e.g., biological samples, blood, cells, serum, plasma, buffy coat, urine, lymph, skin, hair, soil, etc.) by methods known in the art. Any suitable method can be used to obtain nucleic acids from biological samples (e.g., blood or blood samples). productNon-limiting examples of such methods include DNA preparation methods (e.g., as described in Sambrook and Russell, Molecular Cloning: A Laboratory Manual, 3rd ed., 2001), various commercially available reagents or kits, such as DNeasy®, RNeasy®, QIAprep®, QIAquick®, and QIAamp® (e.g., QIAamp® Circulating Nucleic Acid Kit, QiaAmp® DNA Mini Kit, or QiaAmp® DNA Blood Mini Kit) nucleic acid isolation / purification kits by Qiagen, Inc. (Germantown, Md); GenomicPrep™ Blood DNA Isolation Kit (Promega, Madison, Wis.); GFX™ Genomic Blood DNA Purification Kit (Amersham, Piscataway, NJ); Life DNAzol®, ChargeSwitch®, Purelink®, GeneCatcher® nucleic acid isolation / purification kits from Biosciences Technologies, Inc. (Carlsbad, CA); NucleoMag®, NucleoSpin®, and NucleoBond® nucleic acid isolation / purification kits from Clontech Laboratories, Inc. (Mountain View, CA); etc., or combinations thereof. In certain embodiments, nucleic acids are isolated from fixed biological samples, such as formalin-fixed, paraffin-embedded (FFPE) tissue.Genomic DNA from FFPE tissue can be isolated using commercially available kits, such as the AllPrep® DNA / RNA FFPE kit from Qiagen, Inc. (Germantown, Md.), the RecoverAll® Total Nucleic Acid Isolation kit for FFPE from Life Technologies, Inc. (Carlsbad, Calif.), and the NucleoSpin® FFPE kit from Clontech Laboratories, Inc. (Mountain View, Calif.).
[0168] In some embodiments, nucleic acid is extracted from cells using cell lysis procedures.Cell lysis procedures and reagents are known in the art, and can generally be carried out by chemical methods (such as detergents, hypotonic solutions, enzymatic procedures, etc., or a combination thereof), physical methods (such as French press, sonication, etc.), or electrolytic lysis.Any suitable lysis procedure can be used.For example, chemical methods generally use a lysis agent to disrupt cells, extract nucleic acid from cells, and then treat with chaotropic salts.Physical methods, such as freeze / thaw followed by crushing, use of cell press, etc., are also useful. Depending on the situation High salt and / or alkaline lysis procedures may be utilized. Depending on the situation The lysis procedure involves a lysis step with EDTA / proteinase K, a binding buffer step using a large amount of salt (e.g., guanidine hydrochloride (GuHCl), sodium acetate) and isopropanol, and then loading the DNA in this solution onto a silica-based column. Combination The method may include a step of: Depending on the situation The lysis protocol includes certain steps described in Dabney et al., Proceedings of the National Academy of Sciences 110, no. 39 (2013): 15758-15763.
[0169] Nucleic acids can include extracellular nucleic acids in certain embodiments. The term "extracellular nucleic acid," as used herein, can refer to nucleic acids isolated from a source that is substantially cell-free, and can also be referred to as "cell-free" nucleic acids (cell-free DNA, cell-free RNA, or both), "circulating cell-free nucleic acids" (e.g., CCF fragments, ccfDNA), and / or "cell-free circulating vesicle nucleic acids." Extracellular nucleic acids can be present in blood, and extracellular nucleic acids can be obtained from blood (e.g., from the blood of a human subject). Extracellular nucleic acids often do not contain detectable cells and are free of cellular elements or cellular components. Remnant Non-limiting examples of cell-free sources of extracellular nucleic acid are blood, plasma, serum, and urine. In certain embodiments, the cell-free nucleic acid may be obtained from whole blood, plasma, serum, amniotic fluid, saliva, urine, breast water , bronchial lavage liquid , bronchial aspirate, breast milk, colostrum, tears, semen, ascites, breasts water As used herein, the term "obtaining cell-free circulating sample nucleic acid" refers to directly obtaining a sample (e.g., by extracting a sample, e.g., a test sample) from a bodily fluid sample selected from stool, feces, and the like. collection (or the sample) collection The extracellular nucleic acid may be a product of cellular secretion and / or nucleic acid release (e.g., DNA release). The extracellular nucleic acid may be, for example, in any form. The fine cell death It may be the product of Depending on the situation Extracellular nucleic acids are mitotic, swelling Extracellular nucleic acids are products of any form of type I or type II cell death, including apoptotic, toxic, ischemic, etc., and combinations thereof. Without being limited by theory, extracellular nucleic acids may be products of cellular apoptosis and cell destruction, which provides the basis for extracellular nucleic acids that often have a range of lengths across a spectrum (e.g., a "ladder"). Depending on the situation Extracellular nucleic acids are involved in cell necrosis, necroptosis, Oncosis, entosis, pyroptosis, etc., and combinations thereof. In some embodiments, the sample nucleic acid from the subject is circulating cell-free nucleic acid. In some embodiments, the circulating cell-free nucleic acid is from plasma or serum from the subject. In some aspects, the cell-free nucleic acid is degraded. In some embodiments, the cell-free nucleic acid comprises cell-free fetal nucleic acid (e.g., cell-free fetal DNA). In certain aspects, the cell-free nucleic acid comprises circulating cancer nucleic acid (e.g., cancer DNA). In certain aspects, the cell-free nucleic acid comprises circulating tumor nucleic acid (e.g., tumor DNA). In some embodiments, the cell-free nucleic acid comprises infectious agent nucleic acid (e.g., pathogen DNA). In some embodiments, the cell-free nucleic acid comprises nucleic acid (e.g., DNA) from a transplant. In some embodiments, the cell-free nucleic acid comprises nucleic acid (e.g., DNA) from the microbiome (e.g., gut microbiome, blood microbiome, oral microbiome, spinal fluid microbiome, fecal microbiome).
[0170] Cell-free DNA (cfDNA) can be derived from degraded sources and often provides only small amounts of DNA when extracted. The methods described herein for generating single-stranded DNA (ssDNA) libraries can capture large amounts of short DNA fragments from cfDNA. For example, cfDNA from cancer samples tends to have a higher population of short fragments. In certain cases, the short fragments in cfDNA can be associated with transcription factors rather than nucleosomes. Origin About the fragments Enrichment It is possible.
[0171] Extracellular nucleic acids can contain different nucleic acid species and are therefore referred to herein as "heterogeneous" in certain embodiments. For example, serum or plasma from a person suffering from a tumor or cancer can contain nucleic acids from tumor or cancer cells (e.g., neoplasms) and nucleic acids from non-tumor or non-cancer cells. In another example, serum or plasma from a pregnant female can contain maternal and fetal nucleic acids. In another example, serum or plasma from a pregnant female can contain maternal and fetal nucleic acids. dyed or infectious diseases have Serum or plasma from a patient receiving a transplant may contain host nucleic acid and infectious agent or pathogen nucleic acid. In another example, a sample from a subject who has undergone a transplant may contain host nucleic acid and nucleic acid from the donor organ or tissue. Depending on the situation In some cases, cancer, tumor, fetal, pathogen, or graft nucleic acids may represent about 5% to about 50% of the total nucleic acid (e.g., about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, or 49% of the total nucleic acid is cancer, tumor, fetal, pathogen, graft, or microbiome nucleic acid). In another example, heterogeneous nucleic acid may include nucleic acids from two or more subjects (e.g., a sample from a crime scene).
[0172] At least two different nucleic acid species may be present in different amounts in the extracellular nucleic acid, and these nucleic acid species may be present in small amounts. Several types and many Several types In certain cases, a small amount of nucleic acid Several types are from diseased cell types (e.g., cancer cells, wasted cells, cells attacked by the immune system). In certain embodiments, genetic mutations or genetic alterations (e.g., copy number alterations, copy number variants) (copy number variation) , single base changes, single base mutations, chromosomal changes, and / or translocations) are rare. number nucleus In certain embodiments, genetic variations or genetic alterations are determined for multiple nucleic acid species. number" Or "Many number" is not intended to be rigidly defined in any respect. number" The nucleic acids considered to be at least about 0.1% of the total nucleic acids in the sample can have an abundance of at least about 0.1% of the total nucleic acids in the sample to less than 50% of the total nucleic acids in the sample. number nucleus The acid can have an abundance of at least about 1% of the total nucleic acids in the sample up to about 40% of the total nucleic acids in the sample. number nucleus The acid can have an abundance of at least about 2% of the total nucleic acids in the sample up to about 30% of the total nucleic acids in the sample. number nucleus The acid can have an abundance of at least about 3% of the total nucleic acids in the sample up to about 25% of the total nucleic acids in the sample. number nucleus The acid can have an abundance of about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29% or 30% of the total nucleic acids in the sample. Depending on the situation The small amount of extracellular nucleic acid Several types may represent about 1% to about 40% of the total nucleic acids (e.g., about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, or 40% of the nucleic acids may be small). Several types In some embodiments, at least number nucleus In some embodiments, the acid is extracellular DNA. number nucleus The acid is extracellular DNA from apoptotic tissue. number nucleus The acid is Inside the cell Some cells undergo apoptosis received In some embodiments, the extracellular DNA is from a tissue containing at least number nucleus In some embodiments, the nucleic acid is extracellular DNA from necrotic tissue. Inside the cell Some cells in the received Necrosis is extracellular DNA from damaged tissue. Necrosis may, in certain cases, refer to the post-mortem process following cell death. In some embodiments, number nucleus The acid is extracellular DNA from tissue affected by a cell proliferative disorder (e.g., cancer). number nucleus The acid is extracellular DNA from tumor cells. number nucleusThe acid is extracellular fetal DNA. number nucleus The acid is extracellular DNA from a pathogen. number nucleus The acid is extracellular DNA from the graft. number nucleus The acid is extracellular DNA from the microbiome.
[0173] In another embodiment, e.g., number" The nucleic acids considered to be mutated can have an abundance of greater than 50% of the total nucleic acids in the sample up to about 99.9% of the total nucleic acids in the sample. number nucleus The acid can have an abundance of at least about 60% of the total nucleic acids in the sample up to about 99% of the total nucleic acids in the sample. number nucleus The acid can have an abundance of at least about 70% of the total nucleic acids in the sample up to about 98% of the total nucleic acids in the sample. number nucleus The nucleic acid can have an abundance of at least about 75% of the total nucleic acids in the sample up to about 97% of the total nucleic acids in the sample. number nucleus The nucleic acid can have an abundance of at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the total nucleic acids in the sample. number nucleus The acid is extracellular DNA. number nucleus The acid is extracellular maternal DNA. number nucleus The acid is DNA from healthy tissue. number nucleus The DNA is from a non-tumor cell. number nucleus The acid is DNA from the host cell.
[0174] In some embodiments, the amount of extracellular nucleic acid Several types The DNA fragment is about 500 base pairs or less in length (e.g., Several types(About 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the nucleic acids are about 500 base pairs or less in length.) In some embodiments, at least some of the extracellular nucleic acids Several types The DNA fragment is about 300 base pairs or less in length (e.g., Several types (About 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the nucleic acids are about 300 base pairs or less in length.) In some embodiments, at least some of the extracellular nucleic acids Several types The DNA fragment is about 250 base pairs or less in length (e.g., Several types (About 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the nucleic acids are about 250 base pairs or less in length.) In some embodiments, at least some of the extracellular nucleic acids Several types The DNA fragment is about 200 base pairs or less in length (e.g., Several types (About 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the nucleic acids are about 200 base pairs or less in length.) In some embodiments, at least some of the extracellular nucleic acids Several types The DNA fragment is about 150 base pairs or less in length (e.g., Several types (About 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the nucleic acids are about 150 base pairs or less in length.) In some embodiments, at least some of the extracellular nucleic acids Several types The DNA fragments are of about 100 base pairs or less in length (e.g., Several types (About 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the nucleic acids are about 100 base pairs or less in length.) In some embodiments, at least some of the extracellular nucleic acids Several types The DNA fragment is about 50 base pairs or less in length (e.g., Several types About 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or 100% of the nucleic acids are about 50 base pairs or less in length).
[0175] Nucleic acid is a sample containing nucleic acid. (Multiple options possible) The performance of the methods described herein, with or without treatment with For to proposal In some embodiments, the nucleic acid can be provided in a nucleic acid-containing sample. (Multiple options possible) After the treatment of the method described herein For to proposal For example, nucleic acids are used as a sample. (Multiple options possible) The nucleic acid may be extracted, isolated, purified, partially purified, or amplified from a nucleic acid. The term "isolated," as used herein, refers to a nucleic acid that has been removed from its original environment (e.g., the natural environment if it is naturally occurring, or a host cell if exogenously expressed), and thus has been altered from its original environment by human intervention (e.g., "by the hand of man"). The term "isolated nucleic acid," as used herein, can refer to a nucleic acid that has been removed from a subject (e.g., a human subject). Isolated nucleic acids can be provided that have less non-nucleic acid components (e.g., proteins, lipids) than the amount of components present in the source sample. A composition comprising isolated nucleic acids has between about 50% and 99% Super A composition containing isolated nucleic acid may be about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99% free of non-nucleic acid components. Super The term "purified" as used herein refers to a nucleic acid that contains less non-nucleic acid components (e.g., proteins, lipids, carbohydrates) than the amount of non-nucleic acid components present before the nucleic acid is subjected to a purification procedure. Ta A composition containing purified nucleic acid can be about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99% Super The term "purified," as used herein, refers to a nucleic acid sample that contains fewer nucleic acid species than in the sample source from which the nucleic acid is derived. Ta A composition containing purified nucleic acid can be about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99% purified nucleic acid. Super The nucleic acid may be free of other nucleic acid species. For example, fetal nucleic acid can be purified from a mixture containing maternal and fetal nucleic acids. In certain examples, small fragments of nucleic acid (e.g., 30-500 bp fragments) can be purified or partially purified from a mixture containing nucleic acid fragments of different lengths. In certain examples, nucleosomes containing smaller fragments of nucleic acid can be purified from a mixture of larger nucleosome complexes containing larger fragments of nucleic acid. In certain examples, larger nucleosome complexes containing larger fragments of nucleic acid can be purified from nucleosomes containing smaller fragments of nucleic acid. In certain examples, small fragments of fetal nucleic acid (e.g., 30-500 bp fragments) can be purified or partially purified from a mixture containing both fetal and maternal nucleic acid fragments. In certain examples, nucleosomes containing smaller fragments of fetal nucleic acid can be purified from a mixture of larger nucleosome complexes containing larger fragments of maternal nucleic acid. In certain examples, cancer cell nucleic acid can be purified from a mixture containing cancer cell nucleic acid and non-cancer cell nucleic acid. In certain instances, nucleosomes containing smaller fragments of cancer cell nucleic acid can be purified from a mixture of larger nucleosome complexes containing larger fragments of non-cancer nucleic acid. (Multiple options possible) The implementation of the methods described herein without prior treatment of For to proposal For example, nucleic acids can be analyzed directly from a sample without prior extraction, purification, partial purification, and / or amplification.
[0176] Nucleic acids can be amplified under amplification conditions. The term "amplified" or "amplification" or "amplification conditions," as used herein, refers to subjecting a target nucleic acid (e.g., ssNA) in a sample or a nucleic acid product produced by a method herein to a process that linearly or exponentially produces amplicon nucleic acids having the same or substantially the same nucleotide sequence as the target nucleic acid (e.g., ssNA) or a portion thereof. In certain embodiments, the term "amplified" or "amplification" or "amplification conditions" refers to a method that includes polymerase chain reaction (PCR). In certain cases, the amplification product can contain one or more more nucleotides than the amplified nucleotide region of the nucleic acid template sequence (e.g., a primer can contain "extra" nucleotides, such as a transcription initiation sequence, in addition to nucleotides complementary to the nucleic acid template gene molecule, "extra" nucleotides or does not correspond to the amplified nucleotide region of the nucleic acid template gene molecule. Dog Contains nucleotides yields amplification products ).
[0177] The nucleic acid may be subjected to the methods described herein. proposal Before being subjected to the treatment, the nucleic acid may be exposed to a process that modifies certain nucleotides in the nucleic acid. For example, a process that selectively modifies nucleic acids based on the methylation state of internal nucleotides may be applied to the nucleic acid. In addition, conditions such as high temperature, ultraviolet irradiation, and X-ray irradiation may alter the sequence of nucleic acid molecules. to The nucleic acid can be provided in any suitable form that is useful for performing sequence analysis.
[0178] In some embodiments, the target nucleic acid (e.g., ssNA) is linked to a scaffold adaptor or component thereof as described herein. Combination In some embodiments, the target nucleic acid (e.g., ssNA) is not modified prior to the step of reacting with a scaffold adaptor or a component thereof as described herein. Combination "Unmodified" in this context refers to a target nucleic acid that is isolated from a sample and then coupled to a scaffold adaptor or components thereof without modifying either the length or the composition of the target nucleic acid. CombinationFor example, target nucleic acids (e.g., ssNAs) may not be shortened (e.g., they are not contacted with a restriction enzyme, a nuclease, or physical conditions that reduce their length (e.g., shearing conditions, cleavage conditions)), or may not be increased in length by one or more nucleotides (e.g., the ends are not shortened by an overhang). In (No fill-in; no nucleotides added to the ends). The addition of a phosphate group or a chemically reactive group to one or both ends of a target nucleic acid (e.g., ssNA) is generally not considered a modification of the nucleic acid or a modification of the length of the nucleic acid. The denaturation of a double-stranded nucleic acid (dsNA) fragment to produce an ssNA fragment is generally not considered a modification of the nucleic acid or a modification of the length of the nucleic acid.
[0179] In some embodiments, one or both native ends of a target nucleic acid (e.g., ssNA) are cleaved to form a nucleotide sequence such that the ssNA is cleaved to a scaffold adaptor or component thereof as described herein. Combination Native ends generally refer to unmodified fragments of a nucleic acid fragment. In some embodiments, the native ends of a target nucleic acid (e.g., ssNA) are present when the target nucleic acid is modified with a scaffold adaptor or component thereof as used herein. Combination "Unmodified" in this context refers to a target nucleic acid that is isolated from a sample and then coupled to a scaffold adaptor or component thereof without modifying the length of the native ends of the target nucleic acid. Combination For example, a target nucleic acid (e.g., ssNA) may be shortened in length by one or more nucleotides. profit (e.g., they have not been contacted with restriction enzymes, nucleases, or length-reducing physical conditions (e.g., shearing conditions, cleavage conditions) to generate non-native ends) and profit (e.g., the native ends are not filled in with an overhang; nucleotides are not added to the native ends.) The addition of phosphate groups or chemically reactive groups to one or both of the native ends of a target nucleic acid is generally not considered a modification of the length of the nucleic acid.
[0180] In some embodiments, the target nucleic acid (e.g., ssNA) is linked to a scaffold adaptor or component thereof as described herein. Combination In some embodiments, the target nucleic acid is not contacted with a cleaving agent (e.g., an endonuclease, an exonuclease, a restriction enzyme) and / or a polymerase prior to the step of cleaving. Combination In some embodiments, the target nucleic acid is not subjected to mechanical shearing (e.g., sonication (e.g., the Adaptive Focused Acoustics™ (AFA) process by Covaris)) prior to the step of shearing. ... Combination In some embodiments, the target nucleic acid is not contacted with an exonuclease (e.g., DNAse) prior to the step of cleaving. Combination In some embodiments, the target nucleic acid is not amplified prior to the step of reacting with a scaffold adaptor or a component thereof as described herein. Combination In some embodiments, the target nucleic acid is not bound to a solid support prior to the step of reacting with a scaffold adaptor or a component thereof as described herein. Combination In some embodiments, the target nucleic acid is not conjugated to another molecule prior to the step of conjugating the target nucleic acid to a scaffold adaptor or component thereof as described herein. Combination In some embodiments, the target nucleic acid is not cloned into a vector prior to the step of reacting with a scaffold adaptor or a component thereof as described herein. Combination In some embodiments, the target nucleic acid may be dephosphorylated prior to the step of reacting with a scaffold adaptor or a component thereof as described herein. Combination The protein may be phosphorylated prior to the step of cleaving.
[0181] In some embodiments, a target nucleic acid (e.g., ssNA) and a scaffold adaptor herein or a component thereof are Combination The step of allowing the target nucleic acid to react with the scaffold adaptor or a component thereof can include isolating the target nucleic acid and reacting the isolated target nucleic acid with the scaffold adaptor or a component thereof. CombinationIn some embodiments, the target nucleic acid is linked to a scaffold adaptor or a component thereof as described herein. Combination The step of allowing the target nucleic acid to react with the target nucleic acid includes isolating the target nucleic acid, phosphorylating the isolated target nucleic acid, and reacting the phosphorylated target nucleic acid with the scaffold adaptor or a component thereof described herein. Combination In some embodiments, the target nucleic acid and the scaffold adaptor herein or a component thereof are intercalated. Combination The step of causing the target nucleic acid to undergo dephosphorylation includes isolating the target nucleic acid, dephosphorylating the scaffold adaptor herein or a component thereof, and combining the isolated target nucleic acid with the dephosphorylated scaffold adaptor herein or a dephosphorylated component thereof. Combination In some embodiments, the target nucleic acid is linked to a scaffold adaptor or a component thereof as described herein. Combination The step of reacting includes isolating the target nucleic acid, dephosphorylating the isolated target nucleic acid, phosphorylating the dephosphorylated target nucleic acid, and reacting the phosphorylated target nucleic acid with the scaffold adaptor herein or a component thereof. Combination In some embodiments, the target nucleic acid is linked to a scaffold adaptor or a component thereof as described herein. Combination The step of causing the target nucleic acid to undergo dephosphorylation includes isolating the target nucleic acid, dephosphorylating the isolated target nucleic acid, phosphorylating the dephosphorylated target nucleic acid, dephosphorylating the scaffold adaptor or a component thereof, and combining the phosphorylated target nucleic acid with the dephosphorylated scaffold adaptor herein or a dephosphorylated component thereof. Combination This includes making
[0182] In some embodiments, a target nucleic acid (e.g., ssNA) and a scaffold adaptor herein or a component thereof are Combination The step of allowing the target nucleic acid to react with the scaffold adaptor or a component thereof can include isolating the target nucleic acid and reacting the isolated target nucleic acid with the scaffold adaptor or a component thereof. Combination In some embodiments, the target nucleic acid is linked to a scaffold adaptor or a component thereof as described herein. CombinationThe step of allowing the target nucleic acid to react with the target nucleic acid includes isolating the target nucleic acid, phosphorylating the isolated target nucleic acid, and reacting the phosphorylated target nucleic acid with the scaffold adaptor or a component thereof described herein. Combination In some embodiments, the target nucleic acid is linked to a scaffold adaptor or a component thereof as described herein. Combination The step of allowing to occur includes isolating the target nucleic acid, dephosphorylating the scaffold adaptor or a component thereof, and combining the isolated target nucleic acid with the dephosphorylated scaffold adaptor herein or a dephosphorylated component thereof. Combination In some embodiments, the target nucleic acid is linked to a scaffold adaptor or a component thereof as described herein. Combination The step of reacting includes isolating the target nucleic acid, dephosphorylating the isolated target nucleic acid, phosphorylating the dephosphorylated target nucleic acid, and reacting the phosphorylated target nucleic acid with the scaffold adaptor herein or a component thereof. Combination In some embodiments, the target nucleic acid is linked to a scaffold adaptor or a component thereof as described herein. Combination The step of causing the target nucleic acid to undergo dephosphorylation includes isolating the target nucleic acid, dephosphorylating the isolated target nucleic acid, phosphorylating the dephosphorylated target nucleic acid, dephosphorylating the scaffold adaptor or a component thereof, and combining the phosphorylated target nucleic acid with the dephosphorylated scaffold adaptor herein or a dephosphorylated component thereof. Combination It consists of making it possible to
[0183] overhang A target nucleic acid may include one overhang (e.g., at one end of a nucleic acid fragment) or two overhangs (e.g., at both ends of a nucleic acid fragment). The nucleic acid overhangs may include different overhang lengths and / or different overhang types (e.g., a 5' overhang, a 3' overhang, or no overhang). A target nucleic acid may include two overhangs, one overhang and one blunt end, two blunt ends, or a combination thereof. A target nucleic acid may include two 3' overhangs, two 5' overhangs, one 3' overhang and one 5' overhang, one 3' overhang and one blunt end, one 5' overhang and one blunt end, two blunt ends, or a combination thereof. In some cases, overhangs in double-stranded nucleic acids may be extended (i.e., filled in) before further processing (e.g., before a denaturing step).
[0184] In some embodiments, the overhang in the target nucleic acid is a native overhang. In some embodiments, the overhang in the target nucleic acid before extension is a native overhang. In some embodiments, the target nucleic acid end is a native blunt end. The native overhang and the native blunt end may be a blunt end before extension, before denaturation, and / or before use with a scaffold adaptor or component thereof described herein. Combination The fragments are unmodified (e.g., unextended, unfilled, uncut or digested (e.g., by an endonuclease or exonuclease), uncut or undigested (e.g., by an endonuclease or exonuclease), unadded or uncleaved (e.g., by an exonuclease), uncleaved or undigested (e.g., by an endonuclease or ... Added The term "native overhang" generally refers to an overhang or blunt end (not necessarily a blunt end), overhang, or blunt end. Often, native overhangs and native blunt ends refer to ends prior to extension, prior to denaturation, and / or prior to interaction with a scaffold adaptor or component thereof as described herein. CombinationThe nucleic acid sequence may be modified ex vivo (e.g., not extended ex vivo, not filled in ex vivo, not cleaved or digested ex vivo (e.g., by an endonuclease or exonuclease)) prior to the step of adding the nucleic acid sequence. Added In certain instances, native overhangs and native blunt ends refer to ends prior to extension, prior to denaturation, and / or prior to use with a scaffold adaptor or component thereof as described herein. Combination The nucleic acid sequence may be modified after collection from the subject or source (e.g., not extended after collection from the subject or source, not filled in after collection from the subject or source, not cleaved or digested (e.g., by an endonuclease or exonuclease) after collection from the subject or source, or added after collection from the subject or source, even if the nucleic acid sequence is modified after collection from the subject or source, prior to the step of adding the nucleic acid sequence to the nucleic acid sequence. Added Native overhangs and native blunt ends generally refer to overhangs / ends (e.g., ends that are not fragmented or fragmented), overhangs, and blunt ends. Native overhangs and native blunt ends generally do not include overhangs / ends created by contacting an isolated sample with a cleavage agent (e.g., an endonuclease, an exonuclease, a restriction enzyme) and / or a polymerase. Native overhangs and native blunt ends generally do not include overhangs / ends created by mechanical shearing (e.g., ultrasonication (e.g., the Adaptive Focused Acoustics™ (AFA) process by Covaris)). Native overhangs and native blunt ends generally do not include overhangs / ends created by contacting an isolated sample with an exonuclease (e.g., a DNAse). Native overhangs and native blunt ends generally do not include overhangs / ends created by amplification (e.g., polymerase chain reaction). Native overhangs and native blunt ends generally refer to overhangs / ends created by the synthesis of a fragment of another molecule attached to a solid support. toIt generally does not include overhangs / ends that are conjugated or cloned into a vector. In some embodiments, native overhangs and native blunt ends may be dephosphorylated and may be referred to as dephosphorylated native overhangs and dephosphorylated native blunt ends. In some embodiments, native overhangs and native blunt ends may be phosphorylated and may be referred to as phosphorylated native overhangs and phosphorylated native blunt ends.
[0185] In some embodiments, the methods herein involve combining a nucleic acid composition comprising a target nucleic acid with one or more specific nucleic acids under elongation conditions. Yes Nucleotide and elongation activity agent The extension conditions include contacting the nucleic acid with an enzyme suitable for extending the nucleic acid, buffer , reagents and temperature. agent polymerases (e.g., DNA polymerase I, large (Klenow) fragment of DNA polymerase I, T4 DNA polymerase, Bacillus stearothermophilus (Bst) DNA polymerase, thermostable DNA polymerases (e.g., good from a thermophilic marine archaea), 9°N™ DNA polymerase (GENBANK accession number AAA88769.1), THERMINATOR polymerase (strange In some embodiments, the polymerase may have extension activity, such as 9°N™ DNA polymerase, which has mutations D141A, E143A, and A485L. agent is a THERMINATOR polymerase. In some embodiments, agent teeth, Exo In some embodiments, the polymerase has an extension activity. agentis a polymerase that does not have 3' to 5' exonuclease activity. Thus, in some embodiments, a polymerase that does not have exonuclease activity is selected to fill in target nucleic acid overhangs without digesting any single-stranded portions in the target nucleic acid.
[0186] Some or all of the target nucleic acids may comprise double-stranded nucleic acids (dsNA) containing overhangs. Some or all of the target nucleic acids may comprise double-stranded DNA (dsDNA) containing overhangs. Target nucleic acids containing overhangs may comprise a duplex region and a single-stranded overhang. A target nucleic acid having at least one overhang may be extended such that the overhang is filled in and a blunt end is generated. The extended target nucleic acid may comprise an extension region complementary to the overhang (i.e., the overhang present in the target nucleic acid prior to extension). In some embodiments, the extension region comprises one or more specific Yes Contains nucleotides.
[0187] Overhang, especially Yes Nucleotides can be used to fill in. Yes Nucleotides (e.g. Yes The base (also referred to as a base) generally refers to any suitable nucleotide that can be distinguished from the nucleotides in the target nucleic acid. Yes Non-limiting examples of nucleotides include universal bases (e.g., inosine, deoxyinosine, 2'-deoxyinosine (dI, dInosine), nitroindole, 5-nitroindole, and 3-nitroindole), modified bases (e.g., modified nucleotides described herein), methylated bases (e.g., methylcytosine), nucleic acid analogs or artificial nucleic acids (e.g., xenonucleic acid (XNA), peptide nucleic acid (PNA), morpholino, locked nucleic acid (LNA), glycol nucleic acid (GNA), threose nucleic acid (TNA)), or otherwise detectably labeled bases. Yes The use of nucleotides is Which region Was it filled in?This can allow for the subsequent identification of overhang regions (e.g., native overhangs) and thereby allow for the detection of overhang regions (e.g., native overhangs). Yes Nucleotides Sequencing The specific sequences can be detected during sequencing (e.g., by nanopore sequencing). Yes Nucleotides can be incorporated.
[0188] In some embodiments, the extension region comprises one or more specific Yes In some embodiments, the extension region comprises a Yes In such embodiments, the overhang consists of all specific nucleotides. Yes In some embodiments, the extension region is filled in with all specific nucleotides. Yes It is not a nucleotide, but it contains one or more specific Yes In such embodiments, one or more, but not all, species of bases may be specifically Yes Filled in with nucleotides (e.g., cytosines only, e.g., methylcytosines). Yes The use of bases, in certain embodiments, allows for accurate single-base resolution of the overhang region. In Noh All particular Yes It is not a nucleotide, but it contains one or more specific Yes The use of nucleotides can, in certain embodiments, be closest Special Yes Space to base resolution may allow for the identification of overhang regions.
[0189] The nucleic acid having the filled-in overhangs can be prepared by, for example, the methods discussed herein. Sequencing In some cases, the nucleic acid with the filled-in overhangs can be prepared for nanopore sequencing. Nanopore sequencing preparation can include: SequencingConcatamerization involves the concatenation of multiple nucleic acids into longer nucleic acids for purposes of identification. Concatamerization can involve the use of adapters or spacers that represent or highlight different sample nucleic acids. Alternatively, concatemerization can directly connect sample nucleic acids, allowing different sample nucleic acids in the same concatemer to be distinguished by detection of overhangs (e.g., specific Yes The sequences can be deconvoluted by DNA sequencing (by base detection) or other informatics means. Nanopore sequencing preparation can include attaching nanopore sequencing adapters, such as hairpin adapters. The use of hairpin adapters can connect both strands, thereby allowing for easy association of two single-stranded sequences - for example, by using a hairpin adapter that specifically identifies a universal base (e.g., inosine). Yes When used as a base, connecting the two strands may allow the overhang sequence to be determined from the corresponding complementary sequence. Sequencing They can then be linked informatically, for example, based on sequence and / or length matches.
[0190] single stranded nucleic acid Provided herein are methods and compositions for capturing single-stranded nucleic acids (ssNAs) using dedicated adapters (e.g., for generating sequencing libraries). Single-stranded nucleic acids or ssNAs are nucleic acids that are single-stranded over 70% or more of their length. With chains In some embodiments, ssNAs generally refer to a group of polynucleotides that are single stranded (i.e., have no intermolecular or intramolecular hybridization). In some embodiments, ssNAs are single stranded over 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, or 99% or more of the length of the polynucleotide. With chains In certain embodiments, the ssNA is a single stranded polynucleotide. With chains The single-stranded nucleic acid may be referred to herein as the target nucleic acid.
[0191] ssNA can include single-stranded deoxyribonucleic acid (ssDNA). In some embodiments, ssDNA includes, but is not limited to, ssDNA derived from double-stranded DNA (dsDNA). For example, ssDNA can be denatured (e.g., heat-denatured) to generate ssDNA. sex double-stranded DNA (which is denatured and / or chemically denatured) can be derived from In some embodiments, the methods herein involve combining ssDNA with a scaffold adaptor described herein or a component thereof. Combination The method further comprises the step of generating ssDNA by denaturing the dsDNA prior to the step of denaturing the dsDNA.
[0192] In some embodiments, the ssNA comprises a single-stranded ribonucleic acid (ssRNA). RNA can be, for example, messenger RNA (mRNA), microRNA (miRNA), small interfering RNA (siRNA), trans-acting small interfering RNA (ta-siRNA), naturally occurring small interfering RNA (nat-siRNA), ribosomal RNA (rRNA), transfer RNA (tRNA), small nucleolar RNA (snoRNA), small nuclear RNA (snRNA), long non-coding RNA (lncRNA), non-coding RNA (ncRNA), transfer messenger RNA (tmRNA), precursor messenger RNA (pre-mRNA), small Cajal body-specific RNA (scaRNA), piwi, etc. interaction RNA (piRNA), endoribonuclease-prepared siRNA (esiRNA), Temporary low It may comprise molecular RNA (stRNA), signal recognition RNA, telomere RNA, ribozyme, or a combination thereof. In some embodiments, when ssRNA is ssRNA, ssRNA is mRNA. In some embodiments, ssNA comprises single-stranded complementary DNA (cDNA).
[0193] In some embodiments, the methods herein comprise contacting ssNA with a single-stranded nucleic acid binding agent. In some embodiments, the methods herein comprise contacting ssNA with a single-stranded nucleic acid binding protein (SSB) to generate SSB-bound ssNA. In some embodiments, the methods herein comprise contacting sscDNA with a single-stranded nucleic acid binding protein (SSB) to generate SSB-bound sscDNA. In some embodiments, the methods herein comprise contacting ssDNA with a single-stranded nucleic acid binding protein (SSB) to generate SSB-bound ssDNA. In some embodiments, the methods herein comprise contacting ssRNA with a single-stranded nucleic acid binding protein (SSB) to generate SSB-bound ssRNA. SSBs are generally cooperation SSBs bind homogeneously to single-stranded nucleic acids (ssDNA) and generally do not bind well to double-stranded nucleic acids (dsDNA). Upon binding to ssDNA, SSBs destabilize the helical duplex. SSBs can be prokaryotic SSBs (e.g., bacterial or archaeal SSBs) or eukaryotic SSBs. Examples of SSBs include E. coli SSB, E. coli RecA, highly thermostable single-stranded DNA-binding protein (ET SSB), Thermus thermophilus (Tth) RecA, T4 gene 32 protein, and replication protein A (RPA - eukaryotic SSB). ET SSB, Tth RecA, E. coli RecA, T4 gene 32 protein, On top of that For preparing SSB-bound ssNA using such an SSB, buffer and detailed protocols are commercially available (New England Biolabs, Inc. (Ipswich, Mass.)).
[0194] In some embodiments, the methods herein do not include a step of contacting an ssNA with a single-stranded nucleic acid binding protein (SSB) to generate an SSB-bound ssNA. Thus, the methods herein can omit the step of generating an SSB-bound ssNA. For example, the methods herein can involve contacting an ssNA with a scaffold adaptor or a component thereof described herein without contacting the ssNA with an SSB. Combination Such a step may include case The methods herein may be referred to as "SSB-free" methods for generating nucleic acid libraries. Certain SSB-free methods described herein can generate libraries with parameters similar to those of libraries prepared using SSBs, as shown in the figures and discussed in the examples. In some embodiments, the methods herein include contacting ssNAs with a single-stranded nucleic acid binder other than SSB. Such single-stranded nucleic acid binders can stably bind to single-stranded nucleic acids, prevent or reduce the formation of nucleic acid duplexes, still allow ligation or otherwise end-modification of the bound nucleic acids, and be thermostable. Examples of single-stranded nucleic acid binders include, but are not limited to, topoisomerases, helicases, domains thereof, and fusion proteins containing these domains.
[0195] In some embodiments, the methods herein involve combining a nucleic acid composition comprising a single-stranded nucleic acid (ssNA) with a scaffold adaptor or component thereof described herein. Combination In some embodiments, the methods herein include combining a nucleic acid composition consisting of a single-stranded nucleic acid (ssNA) with a scaffold adaptor or a component thereof as described herein. Combination In some embodiments, the methods herein include combining a nucleic acid composition consisting essentially of single-stranded nucleic acid (ssNA) with a scaffold adaptor or component thereof described herein. CombinationA nucleic acid composition "consisting essentially of" single-stranded nucleic acid (ssNA) generally comprises ssNA, Additional It does not contain protein or nucleic acid components. For example, a nucleic acid composition that "consists essentially of" single-stranded nucleic acid (ssNA) does not contain double-stranded nucleic acid (dsNA). Can be excluded A nucleic acid composition "consisting essentially of" single-stranded nucleic acids (ssNA) may contain a low percentage of dsNA (e.g., less than 10% dsNA, less than 5% dsNA, less than 1% dsNA). Can be excluded For example, a nucleic acid composition "consisting essentially of" a single-stranded nucleic acid (ssNA) does not contain a single-stranded binding protein (SSB) or other protein useful for stabilizing the ssNA. Can be excluded A nucleic acid composition "consisting essentially of" single-stranded nucleic acid (ssNA) may include chemical components normally present in nucleic acid compositions, such as buffers, salts, alcohols, crowding agents (e.g., PEG), and may also include residual components (e.g., nucleic acids, proteins, cell membrane components) from the nucleic acid source (e.g., sample) or nucleic acid extraction. A nucleic acid composition "consisting essentially of" single-stranded nucleic acid (ssNA) may include ssNA fragments having one or more phosphates (e.g., terminal phosphate, 5'-terminal phosphate). A nucleic acid composition "consisting essentially of" single-stranded nucleic acid (ssNA) may also include ssNA fragments containing one or more modified nucleotides.
[0196] Nucleic acid Enrichment In some embodiments, the nucleic acids (e.g., extracellular nucleic acids) are identified as subpopulations or species of nucleic acids. Enrichment or relatively Enrichment Nucleic acid subpopulations can include, for example, fetal nucleic acids, maternal nucleic acids, cancer nucleic acids, tumor nucleic acids, patient nucleic acids, host nucleic acids, pathogen nucleic acids, graft nucleic acids, microbiome nucleic acids, nucleic acids of a particular length or A series of long Sano The nucleic acids may include nucleic acids comprising fragments or nucleic acids from a particular genomic region (e.g., a single chromosome, a set of chromosomes, and / or a particular chromosomal region). EnrichmentSuch samples can be used with the methods provided herein. Thus, in certain embodiments, the methods of the present technology provide for the determination of a subpopulation of nucleic acids in a sample. Enrichment do Additional In certain embodiments, nucleic acids from normal tissue (e.g., non-cancer cells, host cells) are selectively (partially, substantially, nearly completely, or completely) removed from the sample. In certain embodiments, maternal nucleic acids are selectively (partially, substantially, nearly completely, or completely) removed from the sample. In certain embodiments, specific low copy number of seed of Nucleic acids (e.g., cancer, tumor, fetal, pathogen, transplant, microbiome nucleic acids) Enrichment The sensitivity of quantification can be improved by analyzing the sample for specific species of nucleic acid. Enrichment Methods for this purpose are described, for example, in U.S. Pat. No. 6,927,028, International Patent Application Publication No. WO2007 / 140417, International Patent Application Publication No. WO2007 / 147063, International Patent Application Publication No. WO2009 / 032779, International Patent Application Publication No. WO2009 / 032781, International Patent Application Publication No. WO2010 / 033639, International Patent Application Publication No. WO2011 / 034631, International Patent Application Publication No. WO2006 / 056480, and International Patent Application Publication No. WO2011 / 143659, the entire contents of each of which are incorporated herein by reference, including all text, tables, formulas, and drawings.
[0197] In some embodiments, the nucleic acid is Enrichment In certain embodiments, the nucleic acid is Described in and detecting specific nucleic acid fragment lengths or fragment sizes using one or more length-based separation methods. A series of piece Long Follow Enrichment In certain embodiments, nucleic acids are isolated for fragments from selected genomic regions (e.g., chromosomes) using one or more sequence-based isolation methods described herein and / or known in the art. Enrichment will be done.
[0198] Nucleic acid subpopulations in the sample Enrichment do for Non-limiting examples of methods include methods that utilize epigenetic differences between nucleic acid species (e.g., methylation-based fetal nucleic acid sequencing, as described in U.S. Patent Application Publication No. 2010 / 0105049, incorporated herein by reference). Enrichment methods); approaches that enhance polymorphic sequences with restriction endonucleases (such as those described in U.S. Patent Application Publication No. 2009 / 0317818, incorporated herein by reference); selective enzymatic digestion approaches; massively parallel signature sequencing (MPSS) approaches; amplification (e.g., PCR)-based approaches (e.g., locus-specific amplification methods, multiplex SNP allele PCR approaches; universal amplification methods); pull-down approaches (e.g., biotinylated ultramer pull-down); extension and ligation-based methods (e.g., molecular inversion probe (MIP) extension and ligation); and combinations thereof.
[0199] In some embodiments, the modified nucleic acid Enrichment Nucleic acid modifications include, but are not limited to, carboxycytosine, 5-methylcytosine (5mC) and its oxidized derivatives (e.g., 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), and 5-arboxylcytosine (5caC)), N(6)-methyladenine (6mA), N4-methylcytosine (4mC), N(6)-methyladenosine (m(6)A), pseudouridine (Ψ), 5-methylcytidine (m(5)C), hydroxymethyluracil, 2'-O-methylation of the 3' end, tRNA modifications, miRNA modifications, and snRNA modifications. Nucleic acids containing one or more modifications can be isolated by various methods, including, but not limited to, antibody-based pull-down. Enrichment The modified nucleic acid Enrichment This can be done before or after denaturing the dsDNA. Enrichment of the complementary strand, which may lack modifications. Enrichment may also result in Enrichment represents the complementary strand lacking the modification. Enrichment do not.
[0200] In some embodiments, the nucleic acid is isolated for fragments from selected genomic regions (e.g., chromosomes) using one or more sequence-based isolation methods described herein. Enrichment Sequence-based separation is performed by determining the presence of nucleotides in the fragments of interest (e.g., target and / or reference fragments). death , generally based on nucleotide sequences that are substantially absent or present in only small amounts (e.g., 5% or less) of other fragments of the sample. In some embodiments, sequence-based separation can result in separated target fragments and / or separated reference fragments. The separated target fragments and / or separated reference fragments are often isolated from fragments remaining in the nucleic acid sample. In certain embodiments, the separated target fragments and the separated reference fragments are also isolated from each other (e.g., isolated in separate assay compartments). In certain embodiments, the separated target fragments and the separated reference fragments are isolated together (e.g., isolated in the same assay compartment). In some embodiments, the unbound fragments are separated into a plurality of fragments. Distinguish Able to be removed or broken down or digested.
[0201] In some embodiments, the scaffold adaptor binds the target nucleic acid. Enrichment For example, scaffold adapters can be designed such that some or all of the bases in the ssNA hybridization region are defined or known bases. These scaffold adapters can preferentially hybridize to target nucleic acids having sequences complementary to the defined or known bases in the scaffold adapter ssNA hybridization region, thereby ensuring that target nucleic acids in the resulting library are Enrichment For example, a GC dinucleotide in a ssNA hybridization region Include to identify target nucleic acids with terminal CG (also called CpG) dinucleotides. EnrichmentPart or all of the length of the scaffold adaptor ssNA hybridization region can be used to similarly target any other defined sequence, including, but not limited to, nuclease cleavage sites, gene promoter regions, pathogen sequences, tumor-associated sequences, and other motifs. In one example, the library is Not enriched Scaffold adaptors and CG dinucleotides Enrichment Prepared using scaffold adaptors. Enrichment For libraries prepared without , 1.7% of reads started with a CG and 1.1% of reads ended with a CG. Enrichment For libraries prepared using , 5.2% of reads started with CG and 19.6% of reads ended with CG. In another example, a sample containing RNA (e.g., host and pathogen RNA) is reverse transcribed using primers specific for the pathogen RNA of interest to generate cDNA, which is then purified and amplified using standard scaffold adapters, or reverse transcription primers. Enrichment The pathogen DNA is prepared using either a scaffold adapter with a ssNA hybridization region targeted to the target region, or a single-stranded library preparation method as discussed herein. Enrichment It is possible.
[0202] Depending on the situation In some cases, the target nucleic acid sequence at the 5' or 3' nucleic acid end is defined or known. In other cases, a scaffold adapter can be used to identify a novel target of interest at the 5' or 3' nucleic acid end. The nucleic acid sequence or pattern of interest can be Enrichment can be characterized from the scaffold adapter library output with or without Depending on the situation The scaffold adapters can be used to identify known or novel target sequences at the nucleic acid termini between samples and controls, for example, cell-free DNA from cancer patients and healthy controls. (Multiple options possible) Determine the presence or relative abundance of Weigh These data can be used to learn the relationship between sequence information at the DNA ends and a given state. By training with well-characterized datasets of patient and healthy samples, in one example, analytical methods or algorithms can be used to predict states or transitions through states. For example, the inventors have observed an increase in AT dinucleotides and a decrease in CpG dinucleotides at the 5' and 3' DNA ends in cfDNA from patients with acute myeloid leukemia (AML) compared to non-AML patient samples. In this example, analytical tools can be used on cfDNA end sequence information to predict the risk of developing AML. people It is possible to predict the risk of
[0203] In some embodiments, selective nucleic acid capture process is used to separate target and / or reference fragments from nucleic acid samples.Commercially available nucleic acid capture systems include, for example, Nimblegen sequence capture system (Roche NimbleGen, Madison, WI); ILLUMINA BEADARRAY platform (Illumina, San Diego, CA); Affymetrix GENECHIP platform (Affymetrix, Santa Clara, CA); Agilent SureSelect Target Enrichment System (Agilent Technologies, Santa Clara, CA); and related platforms.Such methods usually involve hybridization of capture oligonucleotides with part or all of the nucleotide sequence of target or reference fragments, and may include the use of solid-phase (e.g., solid-phase array) and / or solution-based platforms.Capture oligonucleotides (sometimes referred to as "bait") can be selected or designed so that they preferentially hybridize with nucleic acid fragments from selected genomic regions or loci, or with specific sequences in nucleic acid targets. In certain embodiments, hybridization-based methods (e.g., using oligonucleotide arrays) are used to identify fragments containing certain nucleic acid sequences. Enrichment Thus, in some embodiments, a nucleic acid sample can be optionally fragmented, for example, by capturing a subset of fragments using capture oligonucleotides complementary to selected sequences in the sample nucleic acid. Enrichment In certain cases, the captured fragments are amplified. For example, the captured fragments containing an adapter can be amplified using primers complementary to the adapter sequence to form a collection of amplified fragments indexed according to the adapter sequence. In some embodiments, the nucleic acid comprises a region of interest. (several) or part(s) thereof A fragment containing in Distribution In a rowFor fragments from selected genomic regions (e.g., chromosomes, genes) by amplification of one or more regions of interest using complementary oligonucleotides (e.g., PCR primers) Enrichment will be done.
[0204] In some embodiments, nucleic acids are separated using one or more length-based separation methods to identify specific nucleic acid fragment lengths below or above a particular threshold or cutoff, A series of long difference, Or about length(s) Enrichment Nucleic acid fragment length typically refers to the number of nucleotides in the fragment. Nucleic acid fragment length may also be referred to as nucleic acid fragment size. In some embodiments, length-based separation methods are performed without measuring the length of the individual fragments. In some embodiments, length-based separation methods are performed without determining the length of the individual fragments. for In some embodiments, length-based separation refers to a size fractionation procedure, which can isolate (e.g., retain) and / or analyze all or a portion of the fractionated pool. Size fractionation procedures are known in the art (e.g., array separation, molecular sieve separation, gel electrophoresis separation, column chromatography (e.g., size exclusion column) separation, and microfluidics-based approaches). In certain cases, length-based separation approaches can include, for example, selective sequence tagging approaches, fragment circularization, chemical treatments (e.g., formaldehyde, polyethylene glycol (PEG) precipitation), mass spectrometry, and / or size-specific nucleic acid amplification.
[0205] In some embodiments, the nucleic acid comprises fragments associated with one or more nucleic acid binding proteins. Enrichment will be done. Enrichment Examples of methods include chromatin immunoprecipitation (ChIP), cross-linking ChIP (XCHIP), native ChIP (NChIP), bead-free ChIP, carrier ChIP (CChIP), fast ChIP (qChIP), and rapid and quantitative ChIP (QChIP). 2These include, but are not limited to, microchip-ChIP, microchip (µChIP), matrix-ChIP, pathogen-ChIP (PAT-ChIP), ChIP-exo, ChIP-on-chip, RIP-ChIP, HiChIP, ChIA-PET, and HiChIRP.
[0206] In some embodiments, the methods herein involve isolating RNA species in a mixture of RNA species. Enrichment For example, the methods herein include steps of: Enrichment Any suitable mRNA may be used. Enrichment Methods can be used, such as rRNA depletion and / or mRNA depletion. Enrichment Methods such as magnetic bead rRNA depletion (e.g., depleting rRNA from a sample and thus mRNA) Enrichment To achieve this, rRNA depletion probes are used in combination with magnetic beads (Ribo-zero™, Ribominus™, and MICROBExpress™), oligo(dT)-based poly(A) Enrichment (e.g., BioMag® Oligo(dT)20), nuclease-based rRNA depletion (e.g., Terminator™ 5'-phosphate-dependent exonuclease rRNA by (digestion of ribosomal RNA), as well as combinations thereof.
[0207] EnrichmentThe strategy can increase the relative abundance of the target nucleic acid (e.g., as assessed by percent of sequencing reads) by at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000%, 3000%, 4000%, 5000%, 6000%, 7000%, 8000%, 9000%, 10000%, or more.
[0208] Length-based separation In some embodiments, the methods herein include separating target nucleic acids (e.g., ssNAs) according to fragment length. For example, the target nucleic acids (e.g., ssNAs) can be separated into specific nucleic acid fragment lengths (singular) below or above a certain threshold or cutoff using one or more length-based separation methods. A series of long difference, Or about length(s) Enrichment Nucleic acid fragment length typically refers to the number of nucleotides in a fragment. Nucleic acid fragment length may also be referred to as nucleic acid fragment size. In some embodiments, length-based separation methods are performed without measuring the length of individual fragments. In some embodiments, length-based separation methods are performed in conjunction with methods for determining the length of individual fragments. In some embodiments, length-based separation refers to a size fractionation procedure, in which all or a portion of the fractionated pool can be isolated (e.g., retained) and / or analyzed. Size fractionation procedures are known in the art (e.g., array separation, molecular sieve separation, gel electrophoresis separation, column chromatography (e.g., size exclusion column), and microfluidics-based approaches). In some embodiments, length-based separation approaches may include, for example, fragment circularization, chemical treatment (e.g., formaldehyde, polyethylene glycol (PEG)), mass spectrometry, and / or size-specific nucleic acid amplification. In some embodiments, length-based separation is performed using solid-phase reversible immobilization. Fixed( This is done using SPRI beads.
[0209] In some embodiments, a particular length below or above a particular threshold or cutoff, A series of long difference, Nucleic acid fragments of a particular length or length are isolated from the sample. In some embodiments, fragments having a length below a particular threshold or cutoff (e.g., 500 bp, 400 bp, 300 bp, 200 bp, 150 bp, 100 bp) are referred to as "short" fragments, and fragments having a length above a particular threshold or cutoff (e.g., 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1000 bp) are referred to as "long," large, and / or high molecular weight (HMW) fragments. In some embodiments, fragments of a particular length or length below or above a particular threshold or cutoff are referred to as "long," large, and / or high molecular weight (HMW) fragments. A series of long difference, or length(s) are retained for analysis, while fragments of different lengths, or lengths above or below a threshold or cutoff, are retained for analysis. A series of long difference,Or fragments of length(s) are not retained for analysis. In some embodiments, fragments less than about 500 bp are retained. In some embodiments, fragments less than about 400 bp are retained. In some embodiments, fragments less than about 300 bp are retained. In some embodiments, fragments less than about 200 bp are retained. In some embodiments, fragments less than about 150 bp are retained. For example, fragments less than about 190 bp, 180 bp, 170 bp, 160 bp, 150 bp, 140 bp, 130 bp, 120 bp, 110 bp, or 100 bp are retained. In some embodiments, fragments between about 100 bp and about 200 bp are retained. For example, fragments less than about 190 bp, 180 bp, 170 bp, 160 bp, 150 bp, 140 bp, 130 bp, 120 bp, or 110 bp are retained. In some embodiments, fragments within the range of about 100 bp to about 200 bp are retained, for example, fragments within the range of about 110 bp to about 190 bp, 130 bp to about 180 bp, 140 bp to about 170 bp, 140 bp to about 150 bp, 150 bp to about 160 bp, or 145 bp to about 155 bp are retained.
[0210] In some embodiments, target nucleic acids (e.g., ssNAs) having a fragment length of less than about 1000 bp are cleaved with a plurality of scaffold adaptor species, or a pool of scaffold adaptor species, or components of a scaffold adaptor species, as described herein. Combination In some embodiments, target nucleic acids (e.g., ssNAs) having a fragment length of less than about 500 bp are ligated with a plurality of scaffold adaptor species, or a pool of scaffold adaptor species, or components of a scaffold adaptor species, as described herein. Combination In some embodiments, target nucleic acids (e.g., ssNAs) having a fragment length of less than about 400 bp are purified with a plurality of scaffold adaptor species, or a pool of scaffold adaptor species, or components of a scaffold adaptor species, as described herein. Combination In some embodiments, target nucleic acids (e.g., ssNAs) having a fragment length of less than about 300 bp are ligated with a plurality of scaffold adaptor species, or a pool of scaffold adaptor species, or components of a scaffold adaptor species, as described herein. CombinationIn some embodiments, target nucleic acids (e.g., ssNAs) having a fragment length of less than about 200 bp are ligated with a plurality of scaffold adaptor species, or a pool of scaffold adaptor species, or components of a scaffold adaptor species, as described herein. Combination In some embodiments, target nucleic acids (e.g., ssNAs) having a fragment length of less than about 100 bp are ligated with a plurality of scaffold adaptor species, or a pool of scaffold adaptor species, or components of a scaffold adaptor species, as described herein. Combination will be done.
[0211] In some embodiments, target nucleic acids (e.g., ssNAs) having a fragment length of about 100 bp or greater are ligated with a plurality of scaffold adaptor species, or a pool of scaffold adaptor species, or components of a scaffold adaptor species, as described herein. Combination In some embodiments, target nucleic acids (e.g., ssNAs) having a fragment length of about 200 bp or greater are subjected to cleavage with a plurality of scaffold adaptor species, or a pool of scaffold adaptor species, or components of a scaffold adaptor species, as described herein. Combination In some embodiments, target nucleic acids (e.g., ssNAs) having a fragment length of about 300 bp or greater are subjected to cleavage with a plurality of scaffold adaptor species, or a pool of scaffold adaptor species, or components of a scaffold adaptor species, as described herein. Combination In some embodiments, target nucleic acids (e.g., ssNAs) having a fragment length of about 400 bp or greater are subjected to cleavage with a plurality of scaffold adaptor species, or a pool of scaffold adaptor species, or components of a scaffold adaptor species, as described herein. Combination In some embodiments, target nucleic acids (e.g., ssNAs) having a fragment length of about 500 bp or greater are subjected to cleavage with a plurality of scaffold adaptor species, or a pool of scaffold adaptor species, or components of a scaffold adaptor species, as described herein. Combination In some embodiments, target nucleic acids (e.g., ssNAs) having a fragment length of about 1000 bp or greater are subjected to cleavage with a plurality of scaffold adaptor species, or a pool of scaffold adaptor species, or components of a scaffold adaptor species, as described herein. Combination will be done.
[0212] In some embodiments, any fragment length or fragment length Any The target nucleic acid (e.g., ssNA) having a combination of scaffold adaptor species, or a pool of scaffold adaptor species, or components of a scaffold adaptor species, as described herein, can be used. Combination For example, target nucleic acids (e.g., ssNAs) having fragment lengths of less than 500 bp and fragment lengths of 500 bp or longer can be purified by cleaving multiple scaffold adapter species, or a pool of scaffold adapter species, or components of scaffold adapter species, as described herein, Combination It can be done.
[0213] Certain length-based separation methods that can be used with the methods described herein, for example, utilize selective sequence tagging approaches. Such methods selectively tag fragment size species (e.g., short fragments) of nucleic acids in a sample containing long and short nucleic acids. do. Such methods typically involve performing a nucleic acid amplification reaction using a set of nested primers, including an inner primer and an outer primer. Accompany In some embodiments, one or both of the internal primers can be tagged, thereby introducing a tag onto the target amplification product. The external primers generally do not anneal to short fragments that have the (internal) target sequence. The internal primers can anneal to short fragments and generate amplification products that have the tag and the target sequence. Typically, tagging of long fragments is inhibited by a combination of mechanisms, including, for example, blocking extension of the internal primers by pre-annealing and extension of the external primers. Enrichment can be accomplished by any of a variety of methods, including, for example, exonuclease digestion of single-stranded nucleic acid and amplification of tagged fragments using amplification primers specific for at least one tag.
[0214] Another length-based separation method that can be used with the methods described herein involves subjecting the nucleic acid sample to polyethylene glycol (PEG) precipitation. Accompany Exemplary methods include those described in International Patent Application Publication Nos. WO2007 / 140417 and WO2010 / 115016. The methods generally involve contacting a nucleic acid sample with PEG in the presence of one or more monovalent salts under conditions sufficient to substantially precipitate large nucleic acids without substantially precipitating small (e.g., less than 300 nucleotides) nucleic acids.
[0215] Another length-based method that can be used with the methods described herein Enrichment The method involves circularization by ligation, for example using circligase. Accompany Typically, short nucleic acid fragments can be circularized more efficiently than long fragments. Non-circularized sequences can be separated from circularized sequences, Enrichment The resulting short fragments can be used for further analysis.
[0216] Nucleic Acid Library The methods herein may include preparing a nucleic acid library and / or modifying nucleic acids for the nucleic acid library. In some embodiments, the ends of nucleic acid fragments are modified to allow the fragments or their amplification products to be incorporated into a nucleic acid library. Generally, nucleic acid libraries are prepared by immobilizing the fragments on a solid phase (e.g., a solid support, a flow cell, beads), non-limiting examples of which include the immobilization of the fragments on a solid phase (e.g., a solid support, a flow cell, beads). fixed, enriched for specific processes, including amplification, cloning, detection, and / or nucleic acid Sequencing It refers to a plurality of polynucleotide molecules (e.g., nucleic acid sample) that are prepared, assembled, and / or modified for sequencing. In certain embodiments, nucleic acid library is prepared before or during sequencing process. Nucleic acid library (e.g., sequencing library) can be prepared by suitable method as known in the art. Nucleic acid library can be prepared by targeted or non-targeted preparation process.
[0217] In some embodiments, the library of nucleic acids comprises immobilizing nucleic acids on a solid support. Fixed In some embodiments, the nucleic acid is modified to include a chemical moiety (e.g., a functional group) configured for of The library may be prepared by immobilizing the library on a solid support. Fixed modified to include a biomolecule (e.g., a functional group) and / or a member of a binding pair configured for So Non-limiting examples include thyroxine-binding globulin, steroid-binding proteins, antibodies, antigens, haptens, enzymes, lectins, nucleic acids, repressors, protein A, protein G, avidin, streptavidin, biotin, complement component C1q, nucleic acid-binding proteins, receptors, carbohydrates, oligonucleotides, polynucleotides, complementary nucleic acid sequences, and the like, and combinations thereof. Some examples of specific binding pairs include, but are not limited to, an avidin moiety and a biotin moiety; an antigen epitope and an antibody or immunoreactive fragment thereof; an antibody and a hapten; a digoxigenin moiety and an anti-digoxigenin antibody; a fluorescein moiety and an anti-fluorescein antibody; an operator and a repressor; a nuclease and a nucleotide; a lectin and a polysaccharide; a steroid and a steroid-binding protein; an active compound and an active compound receptor; a hormone and a hormone receptor; an enzyme and a substrate; an immunoglobulin and protein A; an oligonucleotide or polynucleotide and its corresponding phase complement; etc., or combinations thereof.
[0218] In some embodiments, a library of nucleic acids is modified to include one or more polynucleotides of known composition, including, but not limited to, an identifier (e.g., a tag, an indexing tag), a capture sequence, a label, an adapter, a restriction enzyme site, a promoter, an enhancer, an origin of replication, a stem-loop, a complementary sequence (e.g., a primer binding site, an annealing site), a suitable integration site (e.g., a transposon, a viral integration site), a modified nucleotide, a unique molecular identifier (UMI) as described herein, a palindromic sequence as described herein, or the like, or a combination thereof. The polynucleotide of known sequence is added at a suitable position, for example, at the 5' end, the 3' end, or within the nucleic acid sequence. vinegar Polynucleotides of known sequence can be identified by the same sequence. But often , or a different sequence It is okay In some embodiments, the polynucleotide of known sequence is immobilized on a surface (e.g., the surface of a flow cell). Determined The nucleic acid molecules are configured to hybridize with one or more oligonucleotides that have been hybridized with the 5'-known sequence. For example, a nucleic acid molecule comprising a 5'-known sequence can hybridize with a first plurality of oligonucleotides, while a 3'-known sequence can hybridize with a second plurality of oligonucleotides. In some embodiments, the library of nucleic acids can include chromosome-specific tags, capture sequences, labels, and / or adapters (e.g., oligonucleotide adapters described herein). In some embodiments, the library of nucleic acids includes one or more detectable labels. In some embodiments, one or more detectable labels can be incorporated at the 5'-end of the nucleic acid library, at the 3'-end, and / or at any nucleotide position within the nucleic acids in the library. In some embodiments, the library of nucleic acids includes hybridized oligonucleotides. In certain embodiments, the hybridized oligonucleotides are labeled probes. In some embodiments, the library of nucleic acids includes hybridized oligonucleotide probes that have been immobilized on a solid phase. Fixed Including before.
[0219] In some embodiments, the polynucleotide of known sequence comprises a universal sequence. A universal sequence is a specific nucleotide sequence incorporated into two or more nucleic acid molecules or into two or more subsets of nucleic acid molecules, and the universal sequence is the same for all molecules or subsets of molecules into which it is incorporated. Universal sequences are often designed to hybridize with and / or amplify multiple different sequences using a single universal primer that is complementary to the universal sequence. In some embodiments, two (e.g., a pair) or more universal sequences and / or universal primers are used. The universal primer often comprises a universal sequence. In some embodiments, an adapter (e.g., a universal adapter) comprises a universal sequence. In some embodiments, one or more universal sequences are used to capture, identify, and / or detect multiple species or subsets of nucleic acids.
[0220] In certain embodiments of nucleic acid library preparation (e.g., in certain single-base synthesis procedures), nucleic acids are size-selected and / or fragmented (e.g., in preparation for library generation) to lengths of a few hundred base pairs or less. In some embodiments, library preparation is performed without fragmentation (e.g., when using cell-free DNA).
[0221] In certain embodiments, ligation-based library preparation methods are used (e.g., ILLUMINA TRUSEQ, Illumina, San Diego, CA). Ligation-based library preparation methods can incorporate index sequences (e.g., sample index sequences for identifying the sample origin for a nucleic acid sequence) in the first ligation step, and often use adapter (e.g., methylated adapter) designs that can be used to prepare samples for single-read sequencing, paired-end sequencing, and multiplexed sequencing. For example, nucleic acids (e.g., fragmented nucleic acids or cell-free DNA) can be end-repaired by fill-in reaction, exonuclease reaction, or a combination thereof. In some embodiments, the resulting blunt-end repaired nucleic acids are then repaired by adding the 3' end of the adapter / primer to the 3' end of the adapter / primer. in single One Any nucleotide can be extended by a single nucleotide that is complementary to the nucleotide overhang. For In some embodiments, end repair is omitted and scaffold adaptors (e.g., scaffold adaptors described herein) are ligated directly to the native ends of nucleic acids (e.g., single-stranded nucleic acids, fragmented nucleic acids, and / or cell-free DNA).
[0222] In some embodiments, nucleic acid library preparation includes ligating a scaffold adapter, such as a scaffold adapter described herein, or a component thereof (e.g., to a sample nucleic acid, to a sample nucleic acid fragment, to a template nucleic acid, to a target nucleic acid, to an ssNA). The scaffold adapter, or a component thereof, can include a sequence complementary to a flow cell anchor, allowing the nucleic acid library to be anchored to a solid support, such as the interior surface of a flow cell. DetermineIn some embodiments, a scaffold adapter, or a component thereof, comprises an identifier, one or more sequencing primer hybridization sites (e.g., sequences complementary to a universal sequencing primer, a single-end sequencing primer, a paired-end sequencing primer, a multiplexed sequencing primer, etc.), or a combination thereof (e.g., adapter / sequencing, adapter / identifier, adapter / identifier / sequencing). In some embodiments, a scaffold adapter, or a component thereof, comprises a primer annealing polynucleotide (e.g., for annealing to a flow cell-bound oligonucleotide and / or to a free amplification primer), an index polynucleotide (e.g., a sample index sequence for tracking nucleic acids from different samples; also referred to as a sample ID), a barcode polynucleotide (e.g., a sequence complementary to a universal sequencing primer, a single-end sequencing primer, a paired-end sequencing primer, a multiplexed sequencing primer, etc.), or a combination thereof (e.g., adapter / sequencing, adapter / identifier, adapter / identifier / sequencing), also referred to herein as a priming sequence or primer-binding domain. Sequencing In some embodiments, the scaffold adapter or a component thereof includes one or more of a single molecule barcode (SMB; also called a molecular barcode or unique molecular identifier (UMI)) for tracking individual molecules of the sample nucleic acid prior to amplification. In some embodiments, the primer annealing component (or priming sequence or primer binding domain) of the scaffold adapter or a component thereof includes one or more universal sequences (e.g., sequences complementary to one or more amplification primers). In some embodiments, the index polynucleotide (e.g., sample index; sample ID) is a component of the scaffold adapter or a component thereof. In some embodiments, the index polynucleotide (e.g., sample index; sample ID) is a component of a universal amplification primer sequence.
[0223] In some embodiments, the scaffold adapter or its components, when used in combination with an amplification primer (e.g., a universal amplification primer), are designed to generate a library construct that includes one or more of a universal sequence, a molecular barcode (UMI), a UMI flanking sequence, a sample ID sequence, a spacer sequence, and a sample nucleic acid sequence (e.g., an ssNA sequence). In some embodiments, the scaffold adapter or its components, when used in combination with a universal amplification primer, are designed to generate a library construct that includes an ordered combination of one or more of a universal sequence, a molecular barcode (UMI), a sample ID sequence, a spacer sequence, and a sample nucleic acid sequence (e.g., an ssNA sequence). For example, a library construct may include a first universal sequence, followed by a second universal sequence, followed by a first molecular barcode (UMI), followed by a spacer sequence, followed by a template sequence (e.g., a sample nucleic acid sequence; an ssNA sequence), followed by a spacer sequence, followed by a second molecular barcode (UMI), followed by a third universal sequence, followed by a sample ID, followed by a fourth universal sequence. In some embodiments, the scaffold adapter, or a component thereof, is designed to generate a library construct for each strand of a template molecule (e.g., a sample nucleic acid molecule; ssNA molecule) when used in combination with an amplification primer (e.g., a universal amplification primer). In some embodiments, the scaffold adapter is a double-stranded adapter.
[0224] The identifier may be a suitable detectable label incorporated into or attached to a nucleic acid (e.g., a polynucleotide) that allows for detection and / or identification of the nucleic acid comprising the identifier. In some embodiments, the identifier is Sequencing In some embodiments, the identifier is incorporated into or attached to the nucleic acid during the method (e.g., by a polymerase). SequencingThe identifier is incorporated into or attached to the nucleic acid prior to the method (e.g., by extension reaction, amplification reaction, ligation reaction). Non-limiting examples of identifiers include nucleic acid tags, nucleic acid indexes or barcodes, radiolabels (e.g., isotopes), metal labels, fluorescent labels, chemiluminescent labels, phosphorescent labels, fluorophore quenchers, dyes, proteins (e.g., enzymes, antibodies or portions thereof, linkers, binding pair members), etc., or combinations thereof. In some embodiments, the identifier (e.g., nucleic acid index or barcode) is a unique, known, and / or identifiable sequence of nucleotides or nucleotide analogs. In some embodiments, the identifier is six or more consecutive nucleotides. Numerous fluorophores with a variety of different excitation and emission spectra are available. Any suitable type and / or number of fluorophores can be used as identifiers. In some embodiments, one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, twenty or more, thirty or more, or fifty or more different identifiers are used in a method described herein (e.g., nucleic acid detection and / or Sequencing In some embodiments, one or two types of identifiers (e.g., fluorescent labels) are linked to each nucleic acid in the library. The amount, can be performed by any suitable method, device or machine, non-limiting examples of which include flow cytometry, quantitative polymerase chain reaction (qPCR), gel electrophoresis, luminometer, fluorometer, spectrophotometer, suitable gene chip or microarray analysis, Western blot, mass spectrometry, chromatography, cytofluorometry, fluorescence microscopy, suitable fluorescence or digital imaging methods, confocal laser scanning microscopy, laser scanning cytometry, affinity chromatography, manual batch mode separation, electric field suspension, suitable nucleic acid Sequencing Methods and / or Nucleic Acids Sequencing devices, etc., as well as combinations thereof.
[0225] In some embodiments, the identifier, sequencing-specific index / barcode, and sequencer-specific flow cell-bound primer site are incorporated into the nucleic acid library by single primer extension (e.g., by a strand-displacing polymerase).
[0226] In some embodiments, the nucleic acid library, or a portion thereof, is amplified under amplification conditions (e.g., amplified by a PCR-based method). Sequencing The method includes amplifying a nucleic acid library by immobilizing the nucleic acid library on a solid support (e.g., a solid support in a flow cell). Fixed Nucleic acid amplification can be performed by amplifying and / or cleaving nucleic acid templates present (e.g., in a nucleic acid library). phase The number of complements was determined by comparing the number of template and / or phase The term "amplification" includes the process of amplifying or increasing by generating one or more copies of the complement. Amplification can be carried out by any suitable method. The nucleic acid library can be amplified by thermocycling or by isothermal amplification. In some embodiments, rolling circle amplification is used. In some embodiments, amplification involves amplifying the nucleic acid library or a portion thereof by adding a nucleic acid library to a solid support. Determined The procedure is carried out on a solid support (e.g., in a flow cell) that is fitted with a particular SequencingIn the method, a nucleic acid library is added to a flow cell and immobilized by hybridization to the anchors under suitable conditions. Determined This type of nucleic acid amplification is often referred to as solid-phase amplification. In some embodiments of solid-phase amplification, all or a portion of the amplification product is deposited on a solid phase. Determined Solid-phase amplification reactions involve the synthesis of oligonucleotides by extension initiated from the...
Claims
1. 1. A method for generating a nucleic acid library, comprising: (i) a nucleic acid composition comprising a single-stranded nucleic acid (ssNA); (ii) a plurality of first oligonucleotide species; and (iii) a plurality of first scaffold polynucleotide species; (a) each polynucleotide in said plurality of first scaffold polynucleotide species comprises a ssNA hybridization region and a first oligonucleotide hybridization region; (b) each oligonucleotide in the plurality of first oligonucleotide species comprises a first unique molecular identifier (UMI) flanked by a first flanking region and a second flanking region, (i) the first unique molecular identifier (UMI) comprises a random sequence, (ii) the first flanking region for each of the first oligonucleotide species comprises a non-random sequence species from a pool of non-random sequence species, and (iii) the second flanking region comprises a non-random sequence; (c) the first oligonucleotide hybridization region comprises (i) a polynucleotide complementary to the first flanking region, and (ii) a polynucleotide complementary to the second flanking region; (d) combining the nucleic acid composition, the plurality of first oligonucleotide species, and the plurality of first scaffold polynucleotide species under conditions in which molecules of the first scaffold polynucleotide species hybridize with (i) a first ssNA terminal region and (ii) molecules of the first oligonucleotide species, thereby forming hybridization products in which the ends of the molecules of the first oligonucleotides are adjacent to the ends of the first ssNA terminal region.
2. 2. The method of claim 1, wherein the first oligonucleotide hybridization region comprises (iii) a region corresponding to the first UMI.
3. the second flanking region for each of the first oligonucleotide species comprises one or more features selected from: (1) a first primer binding domain; (2) a first sequencing adaptor, or a portion thereof; and (3) an index. The method according to claim 1 or claim 2.
4. further comprising combining the nucleic acid composition with (iv) a second oligonucleotide, and (v) a plurality of second scaffold polynucleotide species; (e) each polynucleotide in said plurality of second scaffold polynucleotide species comprises a ssNA hybridization region and a second oligonucleotide hybridization region; (f) combining the nucleic acid composition, the second oligonucleotide, and the plurality of second scaffold polynucleotide species under conditions in which molecules of the second scaffold polynucleotide species hybridize with (i) a second ssNA terminal region and (ii) molecules of the second oligonucleotide, thereby forming a hybridization product in which an end of the molecule of the second oligonucleotide is adjacent to an end of the second ssNA terminal region; 4. The method according to any one of claims 1 to 3.
5. combining the nucleic acid composition with (iv) a plurality of second oligonucleotide species, and (v) a plurality of second scaffold polynucleotide species; (e) each polynucleotide in said plurality of second scaffold polynucleotide species comprises a ssNA hybridization region and a second oligonucleotide hybridization region; (f) each oligonucleotide in said plurality of second oligonucleotide species comprises a second unique molecular identifier (UMI) flanked by a third flanking region and a fourth flanking region; (g) the second oligonucleotide hybridization region comprises (i) a polynucleotide complementary to the third flanking region, and (ii) a polynucleotide complementary to the fourth flanking region; (h) combining the nucleic acid composition, the plurality of second oligonucleotide species, and the plurality of second scaffold polynucleotide species under conditions in which molecules of the second scaffold polynucleotide species hybridize with (i) a second ssNA terminal region and (ii) molecules of the second oligonucleotide species, thereby forming hybridization products in which the termini of the second oligonucleotide molecules are adjacent to the termini of the second ssNA terminal region; 4. The method according to any one of claims 1 to 3.
6. The method of claim 5, wherein the second oligonucleotide hybridization region comprises (iii) a region corresponding to the second UMI.
7. 7. The method of claim 5 or claim 6, wherein the third flanking region for each of the second oligonucleotide species comprises a non-random sequence species from a pool of non-random sequence species.
8. 8. The method of any one of claims 5 to 7, wherein the fourth flanking region for each of the second oligonucleotide species comprises one or more features selected from: (1) a non-random sequence; (2) a second primer binding domain; (3) a second sequencing adaptor, or portion thereof; and (4) an index.
9. 9. The method of claim 1, wherein the ssNA is not modified prior to the combining step, and / or one or both native ends of the ssNA are present when the ssNA is combined with the plurality of first oligonucleotide species and the plurality of first scaffold polynucleotide species.
10. 10. The method of claim 1, wherein the ssNA is from cell-free nucleic acid.
11. a plurality of first oligonucleotide species, each comprising a first unique molecular identifier (UMI) flanked by a first flanking region and a second flanking region, wherein (i) the first unique molecular identifier (UMI) comprises a random sequence, (ii) the first flanking region for each of the first oligonucleotide species comprises a non-random sequence species from a pool of non-random sequence species, and (iii) the second flanking region comprises a non-random sequence; a plurality of first scaffold polynucleotide species, each of which comprises a ssNA hybridization region and a first oligonucleotide hybridization region; wherein the first oligonucleotide hybridization region comprises (i) a polynucleotide complementary to the first flanking region, and (ii) a polynucleotide complementary to the second flanking region.
12. The composition of claim 11 , wherein the first oligonucleotide hybridization region comprises (iii) a region corresponding to the first UMI.
13. 13. The composition of claim 11 or claim 12, wherein the second flanking region for each of the first oligonucleotide species comprises one or more features selected from: (1) a first primer binding domain; (2) a first sequencing adaptor, or portion thereof; and (3) an index.
14. a second oligonucleotide; and a plurality of second scaffold polynucleotide species, each of which comprises a ssNA hybridization region and a second oligonucleotide hybridization region; 14. The composition of any one of claims 11 to 13, further comprising:
15. a plurality of second oligonucleotide species, each comprising a second unique molecular identifier (UMI) flanked by a third flanking region and a fourth flanking region; a plurality of second scaffold polynucleotide species, each of which comprises a ssNA hybridization region and a second oligonucleotide hybridization region; wherein the second oligonucleotide hybridization region comprises (i) a polynucleotide complementary to the third flanking region, and (ii) a polynucleotide complementary to the fourth flanking region.
14. The composition of any one of claims 11 to 13.
16. 16. The composition of claim 15, wherein the second oligonucleotide hybridization region comprises (iii) a region corresponding to the second UMI.
17. 17. The composition of claim 15 or claim 16, wherein the third flanking region for each of the second oligonucleotide species comprises a non-random sequence species from a pool of non-random sequence species.
18. 18. The composition of any one of claims 15 to 17, wherein the fourth flanking region for each of the second oligonucleotide species comprises one or more features selected from: (1) a non-random sequence; (2) a second primer binding domain; (3) a second sequencing adaptor, or portion thereof; and (4) an index.
19. 19. A kit comprising the composition of any one of claims 11 to 18 and instructions for use.
Citation Information
Patent Citations
Immuno-pete
US20180087108A1