Methods for increasing the yield of sequencing libraries

The adapter exchange method addresses inefficiencies in sequencing library generation by ensuring both ends of nucleic acids have distinct adapters, improving yield and data quality across diverse sequencing applications.

JP7802295B2Active Publication Date: 2026-01-20ILLUMINA INC +2
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022567511
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-06-09
Filing Date
2021-06-09
Publication Date
2026-01-20
Estimated Expiration
2041-06-09

AI Technical Summary

Technical Problem

Existing methods for generating sequencing libraries are inefficient due to the difficulty in selectively targeting different adapters to each end of a DNA fragment, resulting in reduced yield and theoretical efficiency of 50% when using tagmentation reactions.

Method used

A method involving adapter exchange to generate sequencing libraries with both forward and reverse adapters on the top and bottom strands of nucleic acids, allowing for single-cell combinatorial indexing and improved data quality.

Benefits of technology

Enhances data quality and yield by achieving near theoretical maximum efficiency in sequencing library generation, applicable to various sequencing methods including whole genome sequencing and single-cell assays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007802295000002
    Figure 0007802295000002
  • Figure 0007802295000003
    Figure 0007802295000003
  • Figure 0007802295000004
    Figure 0007802295000004
Patent Text Reader

Abstract

The present disclosure relates to compositions and methods for preparing sequencing libraries. In one embodiment, the method involves generating a library of target nucleic acids with the same adapter at each end, and then switching the identity of one adapter to result in target nucleic acids flanked by different adapters.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of U.S. Provisional Patent Application No. 63 / 036,710, filed June 9, 2020, which is incorporated herein by reference in its entirety.

[0002] (Government investment) This invention was made with government support under R35GM124704 awarded by the National Institutes of Health. The government has certain rights in this invention.

[0003] (Sequence Listing) This application contains a Sequence Listing that has been submitted electronically to the U.S. Patent and Trademark Office via EFS-Web as an ASCII text file entitled "2021-06-08-SequenceListing_ST25.txt," 2 kilobytes in size, created on June 8, 2021. The information contained in the Sequence Listing is incorporated herein by reference.

[0004] FIELD OF THE INVENTION Embodiments of the present disclosure relate to preparing nucleic acids for sequencing. In particular, embodiments of the methods, compositions, systems, and kits provided herein relate to converting a nucleic acid library from fragments containing symmetric universal sequences to fragments containing asymmetric universal sequences, and obtaining sequence data therefrom. [Background technology]

[0005] Next-generation sequencing (NGS) technology is revolutionizing genomic research. One approach to NGS that has proven effective is the generation of sequencing libraries, in which fragments are processed to have different adapters at each end. Paired-end sequencing is then used to obtain sequence information from both strands. The advantage of the paired-end approach is that sequencing two "n" bases from a single template yields significantly more information than sequencing "n" bases from each of two independent templates in a random fashion. However, methods for adding different adapters to each end are often inefficient due to the difficulty of selectively targeting one adapter to one end of a DNA fragment and a second adapter to the other end of the same DNA fragment. For example, while sequencing libraries can be generated using highly efficient tagmentation, viable sequencing library molecules are only generated if different adapters are incorporated at each end of the molecule in the form of forward or reverse primary sequences. During some tagmentation reactions, each of the two sequences has an equal probability of being incorporated, thus resulting in half of the molecules having a forward-forward or reverse-reverse adapter combination, thereby reducing the theoretical yield to 50%. Summary of the Invention [Means for solving the problem]

[0006] Presented herein are methods and compositions for efficiently converting nucleic acids into sequencing libraries. The methods presented herein include an alternative strategy that uses adapter exchange to generate a library of target nucleic acids tagged with both forward and reverse adapters for the top strand of nucleic acid, the bottom strand of nucleic acid, or both the top and bottom strands of nucleic acid. The methods are useful for a wide range of sequencing library preparation methods, including but not limited to whole genome sequencing, genome conformation capture, circular DNA sequencing, targeted sequencing, co-assays of two or more analytes, such as RNA and ATAC or DNA and RNA, and single-cell genomes. Furthermore, this format allows for the use of one or more index sequences embedded within the adapters, enabling the application of single-cell combinatorial indexing (sci) (e.g., Cusanovich, et al., Science 348, 910-914 (2015); Vitak et al., Nat. Methods 14, 302-308 (2017); Mulqueen et al., Nat. Biotechnol. 36, 428-431 (2018)). The methods provided herein provide improved data quality, e.g., when compared to sci-HiC, significant improvements over known methods in terms of pass reads obtained per cell without sacrificing signal enrichment in the case of s3-ATAC, coverage uniformity in the case of s3-WGS, and improved chromatin contacts obtained per cell in the case of s3-GCC. s3-ATAC, s3-WGS, and s3-GCC are described herein.

[0007] definition Terms used herein will be understood to take their ordinary meaning in the relevant art unless otherwise specified. Some terms used herein and their meanings are set forth below.

[0008] As used herein, the terms "organism" and "subject" are used interchangeably and refer to microorganisms (e.g., prokaryotes or eukaryotes), animals, and plants. An example of an animal is a mammal, such as a human.

[0009] As used herein, the term "target nucleic acid," when used with respect to a nucleic acid, is intended as a semantic identifier of the nucleic acid in the context of a method or composition described herein and does not necessarily limit the structure or function of the nucleic acid other than as otherwise expressly indicated. A target nucleic acid can be essentially any nucleic acid of known or unknown sequence. For example, it can be a genomic DNA fragment (e.g., chromosomal DNA), extrachromosomal DNA such as a plasmid, circulating DNA or circulating RNA, nucleic acid from one or more cells, cell-free DNA, RNA (e.g., mRNA), or cDNA. Sequencing can result in determining the sequence of all or part of a target molecule. Targets can be derived from primary nucleic acid samples, such as nuclei. In one embodiment, targets can be processed into templates suitable for amplification by placing universal sequences at the ends of each target fragment. Targets can also be obtained from primary RNA samples by reverse transcription into cDNA. In one embodiment, targets are used with respect to a subset of DNA or RNA within a cell. Targeted sequencing typically uses selection and isolation of genes of interest, either by PCR amplification (e.g., region-specific primers) or hybridization-based capture methods or antibodies. Target enrichment can be performed at various stages of the method. For example, target RNA representation can be achieved using target-specific primers in a reverse transcription step or by hybridization-based enrichment of subsets from more complex libraries. Examples include exome sequencing or the L1000 assay (Subramanian et al., 2017, Cell, 171; 1437-1452). Target sequencing can include any enrichment process known to those skilled in the art. Target nucleic acids with universal sequences at one or both ends may be referred to as modified target nucleic acids. References to nucleic acids, such as target nucleic acids, include both single-stranded and double-stranded nucleic acids unless otherwise specified. For example, symmetric and asymmetric target nucleic acids can be double-stranded, single-stranded, or partially double-stranded and single-stranded in some respects in the methods of the present disclosure.

[0010] As used herein, the term "adapter" and its derivatives, such as universal adapter, generally refer to any linear oligonucleotide that can be added to a target nucleic acid. Adapters can be single-stranded or double-stranded DNA, or can contain both double-stranded and single-stranded regions. Adapters can include a sequence that is substantially identical to or substantially complementary to at least a portion of a primer, such as a universal primer; an index (also referred to herein as a barcode or tag) to aid in downstream error correction, identification, or sequencing; and / or a UMI. In some embodiments, adapters are substantially non-complementary to the 3' or 5' end of any target sequence present in a sample. In some embodiments, suitable adapter lengths range from about 6 to 100 nucleotides, about 12 to 60 nucleotides, or about 15 to 50 nucleotides in length. For example, the terms "adaptor" and "adapter" are used interchangeably.

[0011] As used herein, the term "universal," when used to describe a nucleotide sequence, refers to a region of sequence common to two or more nucleic acid molecules, with the molecules also having regions of sequence that differ from one another. Universal sequences present in different members of a nucleic acid collection can be used as "landing pads" in subsequent steps to anneal nucleotide sequences that can be used as primers for adding additional nucleotide sequences, such as indexes, to target nucleic acids. Universal sequences present in different members of a nucleic acid collection can capture multiple different nucleic acids using a population of universal capture nucleic acids, e.g., capture oligonucleotides complementary to a portion of the universal sequence, e.g., universal capture sequences. Non-limiting examples of universal capture sequences include sequences identical to or complementary to P5 and P7 primers. Similarly, universal sequences present in different members of a collection of molecules can replicate (e.g., sequence) or amplify multiple different nucleic acids using a population of universal primers complementary to a portion of the universal sequence, e.g., universal anchor sequences. The terms "A14" and "B15" can be used to refer to universal anchor sequences. The terms "A14'" (A14 prime) and "B15'" (B15 prime) refer to the complements of A14 and B15, respectively. It will be understood that any suitable universal anchor sequence may be used in the methods presented herein, and that the use of A14 and B15 is only an exemplary embodiment. In one embodiment, a universal anchor sequence is used as the site to which a universal primer (e.g., a sequencing primer for Read 1 or Read 2) anneals for sequencing. Thus, the capture oligonucleotide or universal primer comprises a sequence that can specifically hybridize to a universal sequence.

[0012] The terms "P5" and "P7" may be used to refer to universal capture sequences or capture oligonucleotides. The terms "P5'" (P5 prime) and "P7'" (P7 prime) refer to the complements of P5 and P7, respectively. It will be understood that any suitable universal capture sequence or capture nucleotide can be used in the methods presented herein, and the use of P5 and P7 is only an exemplary embodiment. The use of capture nucleotides such as P5 and P7 or their complements on flow cells is known in the art, as exemplified by the disclosures of WO 2007 / 010251, WO 2006 / 064199, WO 2005 / 065814, WO 2015 / 106941, WO 1998 / 044151, and WO 2000 / 018957. For example, any suitable forward amplification primer, whether immobilized or in solution, can be useful in the methods provided herein for amplifying complementary sequences and sequences. Similarly, any suitable reverse amplification primer, whether immobilized or in solution, can be useful in the methods provided herein for amplifying complementary sequences and sequences. Those skilled in the art will understand how to design and use suitable primer sequences for capturing and / or amplifying nucleic acids as provided herein.

[0013] As used herein, the term "primer" and its derivatives generally refer to any nucleic acid capable of hybridizing to a target sequence of interest. Typically, a primer serves as a substrate onto which nucleotides can be polymerized by a polymerase or to which a polynucleotide can be ligated; however, in some embodiments, a primer can be incorporated into a synthesized nucleic acid strand to provide a site to which another primer can hybridize and prime synthesis of a new strand complementary to the synthesized nucleic acid molecule. A primer can comprise any combination of nucleotides or their analogs. In some embodiments, a primer is a single-stranded oligonucleotide or polynucleotide. The terms "polynucleotide" and "oligonucleotide" are used interchangeably herein to refer to polymeric forms of nucleotides of any length and can include ribonucleotides, deoxyribonucleotides, their analogs, or mixtures thereof. It should be understood that these terms include, as equivalents, analogs of any of DNA, RNA, cDNA, or antibody-oligoconjugates made from nucleotide analogs, and are applicable to single-stranded (such as sense or antisense) and double-stranded polynucleotides. As used herein, the term also encompasses cDNA, which is complementary or copy DNA produced from an RNA template, for example, by the action of reverse transcriptase. The term refers only to the primary structure of the molecule. Thus, the term includes triple-, double-, and single-stranded deoxyribonucleic acid ("DNA"), as well as triple-, double-, and single-stranded ribonucleic acid ("RNA").

[0014] As used herein, an "index" (also referred to as an "index region," "index adapter," "tag," or "barcode") refers to a unique nucleic acid tag that can be used to identify a sample or source of nucleic acid material, or a compartment in which a target nucleic acid is present. The index can be in solution or on a solid support, or can be attached or bound to a solid support and released into a solution or compartment. When nucleic acid samples are derived from multiple sources, the nucleic acids in each nucleic acid sample can be tagged with a different nucleic acid tag so that the source of the sample can be identified. Any suitable index or set of indexes can be used, as known in the art and as exemplified by the disclosures of U.S. Pat. No. 8,053,192, WO 05 / 068656, and U.S. Patent Application Publication No. 2013 / 0274117. In some embodiments, the index may include a 6-base index 1 (i7) sequence, an 8-base index 1 (i7) sequence, an 8-base index 2 (i5e) sequence, a 10-base index 1 (i7) sequence, or a 10-base index 2 (i5) sequence from Illumina (San Diego, CA).

[0015] As used herein, the term "unique molecular identifier" or "UMI" refers to a molecular tag that can be attached to a nucleic acid, either randomly, non-randomly, or semi-randomly. When incorporated into a nucleic acid, the unique molecular identifier (UMI) can be used to correct for subsequent amplification bias by directly counting the UMI after amplification and sequencing. The UMI can bind to similar nucleic acids, such as adapters, making each nucleic acid unique.

[0016] As used herein, the term "amplicon," when used with reference to a nucleic acid, refers to the product of nucleic acid copying, which product has a nucleotide sequence identical to or complementary to at least a portion of the nucleotide sequence of the nucleic acid. Amplicons can be generated by any of a variety of amplification methods using a nucleic acid or its amplicon as a template, including, for example, polymerase extension, polymerase chain reaction (PCR), rolling circle amplification (RCA), ligation extension, or ligation chain reaction. An amplicon can be a nucleic acid molecule having a single copy of a particular nucleotide sequence (e.g., a PCR product) or multiple copies of a nucleotide sequence (e.g., a concatemeric product of RCA). A first amplicon of a target nucleic acid is typically a complementary copy. Subsequent amplicons are copies made from the target nucleic acid or the first amplicon after the generation of the first amplicon. Subsequent amplicons can have a sequence that is substantially complementary to or substantially identical to the target nucleic acid.

[0017] As used herein, "amplifying," "amplification," or "amplification reaction," and derivatives thereof, generally refer to any act or process in which at least a portion of a nucleic acid molecule is duplicated or copied onto at least one additional nucleic acid molecule. The additional nucleic acid molecule optionally comprises a sequence that is substantially identical to or substantially complementary to at least a portion of a template nucleic acid molecule. The template nucleic acid molecule may be single-stranded or double-stranded, and the additional nucleic acid molecules may independently be single-stranded or double-stranded. Amplification optionally involves linear or exponential replication of the nucleic acid molecule. In some embodiments, such amplification can be performed using isothermal conditions; in other embodiments, such amplification can involve thermal cycling. In some embodiments, amplification is a multiplex amplification that involves simultaneous amplification of multiple target sequences in a single amplification reaction. In some embodiments, "amplification" includes amplifying at least a portion of DNA- and RNA-based nucleic acids, alone or in combination. The amplification reaction can involve any amplification process known to those of skill in the art. In some embodiments, the amplification reaction involves polymerase chain reaction (PCR).

[0018] As used herein, the term "polymerase chain reaction" ("PCR") refers to the Mullis method (U.S. Pat. Nos. 4,683,195 and 4,683,202), which describes a method for increasing the concentration of a polynucleotide segment of interest in a mixture of genomic DNA without cloning or purification. This process for amplifying a polynucleotide of interest involves introducing a large excess of two oligonucleotide primers into a DNA mixture containing the desired polynucleotide of interest, followed by a series of thermal cycling reactions in the presence of a DNA polymerase. The two primers are complementary to each strand of the double-stranded polynucleotide of interest. The mixture is first denatured at a higher temperature, and then the primers are annealed to complementary sequences within the polynucleotide of the molecule of interest. After annealing, the primers are extended with a polymerase to form a new pair of complementary strands. Denaturation, primer annealing, and polymerase extension can be repeated multiple times (called thermal cycling) to obtain a highly concentrated amplified segment of the desired polynucleotide of interest. The length of the amplified segment (amplicon) of the desired target polynucleotide is determined by the relative positions of the primers relative to each other, and therefore this length is a controllable parameter. By repeating this process, the method is called PCR. Because the desired amplified segment of the target polynucleotide becomes the predominant nucleic acid sequence (in terms of concentration) in the mixture, it is said to be "PCR amplified." In a modification of the above method, the target nucleic acid molecule can be PCR amplified using multiple different primer pairs, and in some cases, one or more primer pairs can be used per target nucleic acid molecule of interest, thereby forming a multiplex PCR reaction.

[0019] As used herein, "amplification conditions" and its derivatives generally refer to conditions suitable for amplifying one or more nucleic acid sequences. Such amplification can be linear or exponential. In some embodiments, amplification conditions can include isothermal conditions, or can include thermal cycling conditions, or a combination of isothermal and thermal cycling conditions. In some embodiments, conditions suitable for amplifying one or more nucleic acid sequences include polymerase chain reaction (PCR) conditions. Typically, amplification conditions refer to a reaction mixture sufficient to amplify a nucleic acid, such as one or more target sequences flanked by a universal sequence or target-specific primer, or to amplify an amplification target sequence flanked by one or more adapters. Generally, amplification conditions include a catalyst for amplification or nucleic acid synthesis, e.g., a polymerase, primers having a degree of complementarity to the nucleic acid to be amplified, and nucleotides such as deoxyribonucleotide triphosphates (dNTPs) to facilitate primer extension when hybridized to the nucleic acid. Amplification conditions can require hybridization or annealing of primers to the nucleic acid, extension of the primers, and denaturation, in which the extended primers are separated from the nucleic acid sequence undergoing amplification. Typically, although not necessarily, amplification conditions may include thermal cycling, but in some embodiments, amplification conditions include multiple cycles in which the steps of annealing, extension, and separation are repeated. Typically, amplification conditions include Mg 2+ or Mn 2+ and may also include various modifiers of ionic strength.

[0020] As defined herein, "multiplex amplification" refers to the selective, non-random amplification of two or more target sequences within a sample using at least one target-specific primer. In some embodiments, multiplex amplification is performed such that some or all of the target sequences are amplified in a single reaction vessel. The "plex" of a given multiplex amplification generally refers to the number of different target-specific sequences amplified during that single multiplex amplification. In some embodiments, the plex can be about 12-plex, 24-plex, 48-plex, 96-plex, 192-plex, 384-plex, 768-plex, 1536-plex, 3072-plex, 6144-plex, or more. The amplified target sequences can be analyzed by several different methodologies (e.g., gel electrophoresis followed by densitometry, quantification by bioanalyzer or quantitative PCR, hybridization with labeled probes, incorporation of biotinylated primers followed by avidin-enzyme conjugate detection, or detection of the amplified target sequences). 32 Detection is also possible by incorporation of P-labeled deoxynucleotide triphosphates.

[0021] As used herein, the term "amplification site" refers to a site within or on an array where one or more amplicons can be generated. An amplification site can be further configured to contain, retain, or attach at least one amplicon generated at that site.

[0022] As used herein, the term "array" refers to a collection of sites that can be distinguished from one another according to their relative positions. Different molecules at different sites of an array can be distinguished from one another according to the site's position within the array. Each site of an array can contain one or more molecules of a particular type. For example, a site can contain a single target nucleic acid molecule having a particular sequence, or a site can contain several nucleic acid molecules having the same sequence (and / or its complementary sequence). The sites of an array can be different features located on the same substrate. Exemplary features include, but are not limited to, droplets of liquid, wells in a substrate, beads (or other particles) in or on a substrate, protrusions from a substrate, bumps on a substrate, or channels within a substrate. The sites of an array can be separate substrates, each with a different molecule. The different molecules attached to the separate substrates can be identified according to the position of the substrate on a surface to which the substrates are associated, or according to the position of the substrate within a liquid or gel. An exemplary array in which separate substrates are located on a surface includes, but is not limited to, beads in wells.

[0023] As used herein, the term "compartment" is intended to mean an area or volume that separates or isolates something from another. Exemplary compartments include, but are not limited to, vials, tubes, wells, droplets, boluses, beads, containers, surface features, flow cells, or areas or volumes separated by physical forces such as fluid flow, magnetism, or electric current. In one embodiment, a compartment is a well of a multiwell plate, such as a 96- or 384-well plate. As used herein, a droplet may include hydrogel beads, which are beads for encapsulating one or more nuclei or cells and comprise a hydrogel composition. In some embodiments, the droplets are homogenous droplets of hydrogel material or hollow droplets with a polymer hydrogel shell. Whether homogenous or hollow, the droplets may be capable of encapsulating one or more nuclei or cells. In some embodiments, the droplets are surfactant-stabilized droplets. In some embodiments, a single cell or nucleus is present per compartment. In some embodiments, two or more cells or nuclei are present per compartment. In some embodiments, each compartment comprises a compartment-specific index. In some embodiments, the index is in solution or attached or bound to a solid phase within each compartment.

[0024] As used herein, the term "flow cell" refers to a chamber that includes a solid surface through which one or more fluidic reagents can flow. Examples of flow cells and related fluidic systems and detection platforms that can be easily used in the methods of the present disclosure are described, for example, in Bentley et al., Nature 456:53-59 (2008), WO 04 / 018497, U.S. Patent No. 7,057,026, WO 91 / 06678, WO 07 / 123744, U.S. Patent Nos. 7,329,492, 7,211,414, 7,315,019, 7,405,281, and U.S. Patent Application Publication No. 2008 / 0108082.

[0025] As used herein, the term "clonal population" refers to a population of nucleic acids that are homogeneous with respect to a particular nucleotide sequence. Homogeneous sequences are typically at least 10 nucleotides in length, but may be longer, e.g., at least 50, 100, 250, 500, or 1000 nucleotides in length. A clonal population can be derived from a single target or template nucleic acid. Typically, all nucleic acids in a clonal population have the same nucleotide sequence. It will be understood that minor variations (e.g., due to amplification artifacts) can occur without deviating from clonality.

[0026] As used herein, the term "each," when used in reference to a collection of items, is intended to identify each individual item in the set, but does not necessarily refer to every item in the set, unless the context clearly dictates otherwise.

[0027] As used in this specification and the appended claims, the term "or" is generally used in its inclusive sense unless the content clearly dictates otherwise. The term "and / or" means one or all of the listed elements or a combination of any two or more of the listed elements. The use of "and / or" in some instances does not imply that the use of "or" cannot mean "and / or" in other instances.

[0028] The words "preferred" and "preferably" refer to embodiments of the present disclosure that may offer certain benefits, under particular circumstances. However, other embodiments may also be preferred, under the same or other circumstances. Furthermore, the recitation of one or more preferred embodiments does not imply that other embodiments are not useful, and is not intended to exclude other embodiments from the scope of the present disclosure.

[0029] As used herein, the words "have," "has," "having," "include," "includes," "including," "comprise," "comprises," "comprising," and the like are used in their inclusive, open-ended sense and generally mean "include, but not limited to," "includes, but not limited to," or "including, but not limited to."

[0030]

[0013] Wherever something is described herein using words such as "have," "has," "having," "include," "includes," "including," "comprise," "comprises," "comprising," and the like, it is understood that similar embodiments are also provided that are otherwise described using the terms "consisting of" and / or "consisting essentially of." The term "consisting of" is intended to include what follows the phrase "consisting of." That is, "consisting" indicates that the listed elements are required or essential, and that no other elements may be present. The term "consisting essentially of" indicates that any elements listed after the phrase are included, and that other elements other than those listed may be included, so long as those elements do not interfere with or contribute to the activity or function specified in the disclosure of the listed elements.

[0031] Unless otherwise noted, "a," "an," "the," and "at least one" are used interchangeably and mean one or more than one.

[0032] Conditions that are "favorable" or "favorable" for an event to occur are conditions that do not prevent such an event from occurring. Thus, these conditions enable, enhance, facilitate, and / or are conducive to the event.

[0033] As used herein, "providing," for example, in the context of a composition or nucleic acid, means making the composition or nucleic acid, purchasing the composition or nucleic acid, or otherwise obtaining the compound or nucleic acid.

[0034] References to "one embodiment," "an embodiment," "particular embodiments," or "some embodiments" mean that the particular feature, configuration, composition, or characteristic described in connection with this embodiment is included in at least one embodiment of the present disclosure. Thus, the appearances of such phrases in various places throughout this specification do not necessarily refer to the same embodiment of the present disclosure. Furthermore, the particular features, configurations, compositions, or characteristics may be combined in any suitable manner in one or more embodiments.

[0035] Various aspects of the present disclosure may be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the present disclosure. Accordingly, the description of a range should be considered to have specifically disclosed all possible subranges and individual numerical values ​​within that range. For example, the description of a range such as 1 to 6 should be considered to have specifically disclosed subranges such as 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, etc., as well as individual numbers within that range, such as 1, 2, 2, 7, 3, 4, 5, 5, 3, and 6. This applies regardless of the breadth of the range.

[0036] In any method disclosed herein that includes separate steps, the steps may be performed in any practicable order, and, suitably, any combination of two or more steps may be performed simultaneously. The following detailed description of the exemplary embodiments of the present disclosure can be best understood when read in conjunction with the following drawings. [Brief explanation of the drawings]

[0037] [Figure 1] FIG. 1 shows a general block diagram of one embodiment of a general exemplary method for generating a library for sequencing according to the present disclosure. [Figure 2] 1A-D show schematic diagrams of embodiments for converting a target nucleic acid from symmetric to asymmetric according to various aspects of the present disclosure presented herein. For simplicity, only one target nucleic acid is shown. [Figure 3] 1A-D show schematic diagrams of embodiments of converting a target nucleic acid from symmetric to asymmetric and adding another adaptor according to various aspects of the present disclosure presented herein. For simplicity, only one target nucleic acid is shown. [Figure 4A] 1 shows a schematic diagram of an embodiment of converting a target nucleic acid from symmetric to asymmetric and adding another adapter according to various aspects of the disclosure presented herein. For simplicity, only one target nucleic acid is shown. [Figure 4B] 1 shows a schematic diagram of an embodiment of converting a target nucleic acid from symmetric to asymmetric and adding another adapter according to various aspects of the disclosure presented herein. For simplicity, only one target nucleic acid is shown. [Figure 4C] 1 shows a schematic diagram of an embodiment of converting a target nucleic acid from symmetric to asymmetric and adding another adapter according to various aspects of the disclosure presented herein. For simplicity, only one target nucleic acid is shown. [Figure 4D] 1 shows a schematic diagram of an embodiment of converting a target nucleic acid from symmetric to asymmetric and adding another adapter according to various aspects of the disclosure presented herein. For simplicity, only one target nucleic acid is shown. [Figure 4E] 1 shows a schematic diagram of an embodiment of converting a target nucleic acid from symmetric to asymmetric and adding another adapter according to various aspects of the disclosure presented herein. For simplicity, only one target nucleic acid is shown. [Figure 4F]1 shows a schematic diagram of an embodiment of converting a target nucleic acid from symmetric to asymmetric and adding another adapter according to various aspects of the disclosure presented herein. For simplicity, only one target nucleic acid is shown. [Figure 5] FIG. 1 shows a general block diagram of a general exemplary method for single-cell combinatorial indexing according to the present disclosure. [Figure 6] 1 shows a schematic diagram of an embodiment of the present disclosure, in which whole cell genomic DNA is converted into symmetric target nucleic acid, and then converted into asymmetric target nucleic acid (s3-WGS).For simplicity, only one target nucleic acid is shown. [Figure 7] 1 shows a schematic diagram of an embodiment of converting accessible genomic DNA into a symmetric target nucleic acid and then into an asymmetric target nucleic acid (s3-ATAC) according to various aspects of the present disclosure presented herein. For simplicity, only one target nucleic acid is shown. [Figure 8A] 1 shows a schematic diagram of an embodiment of the processing of mRNA nucleic acids to DNA and subsequent processing to result in three populations of asymmetric nucleic acids. For simplicity, only one mRNA nucleic acid is shown. [Figure 8B] 1 shows a schematic diagram of an embodiment of the processing of mRNA nucleic acids to DNA and subsequent processing to result in three populations of asymmetric nucleic acids. For simplicity, only one mRNA nucleic acid is shown. [Figure 8C] 1 shows a schematic diagram of an embodiment of the processing of mRNA nucleic acids to DNA and subsequent processing to result in three populations of asymmetric nucleic acids. For simplicity, only one mRNA nucleic acid is shown. [Figure 8D] 1 shows a schematic diagram of an embodiment of the processing of mRNA nucleic acids to DNA and subsequent processing to result in three populations of asymmetric nucleic acids. For simplicity, only one mRNA nucleic acid is shown. [Figure 9]1 shows a schematic diagram of an embodiment of a simultaneous assay for converting total cellular genomic DNA into a symmetric target nucleic acid and then into an asymmetric target nucleic acid (s3-GCC) according to various aspects of the present disclosure presented herein. For simplicity, only one target nucleic acid is shown. [Figure 10] FIG. 1 shows a schematic of an embodiment of a protocol for plate-based combinatorial indexing. [Figure 11] The effect of DNA lesion size on library generation is shown. [Figure 12] 1 shows the effect of modified nucleotides on extension to add adapters. [Figure 13] Modified nucleotides that enhance the second extension are shown. [Figure 14] The effect of annealing temperature is shown. [Figure 15] 1 shows the experimental layout of a barnyard experiment demonstrating indexing at both the tagmentation and PCR stages of multiple 96-well plates using nuclei from flash-frozen human cortex and mouse whole brain samples. [Figure 16] Boxplot of projected library complexity by unique reads per cell. s3ATAC outperforms all other published single-cell ATAC sequence libraries on flash-frozen mouse cortex based on predicted unique library molecules. [Figure 17] Comparison of human and mouse reads per cell in a "true barnyard" (left, mixed-species tagmentation wells) and a PCR barnyard (right; species mixed at the PCR stage) is shown, demonstrating little to no cell-to-cell exchange of library molecules. The 5.12% index collision rate in the true barnyard suggests an optimal 15 nuclei per well for an acceptable collision rate. [Figure 18] UMAP projection of the human nucleus is shown. [Figure 19-1] We show that canonical markers of macroscopic cell types within the cortex reveal distinct cell populations. [Figure 19-2]We show that canonical markers of macroscopic cell types within the cortex reveal distinct cell populations. [Figure 19-3] We show that canonical markers of macroscopic cell types within the cortex reveal distinct cell populations. [Figure 20] UMAP projections of mouse nuclei are shown. [Figure 21-1] We show that canonical markers of macroscopic cell types in the mouse brain reveal distinct cell populations. [Figure 21-2] We show that canonical markers of macroscopic cell types in the mouse brain reveal distinct cell populations. [Figure 21-3] We show that canonical markers of macroscopic cell types in the mouse brain reveal distinct cell populations. [Figure 22] 1 shows the experimental layout of PDAC low passage patient-derived lines to generate s3-WGS libraries showing indexing at both the tagmentation and PCR stages of multiple 96-well plates. [Figure 23] Boxplots of library complexity measured by unique reads per cell as well as projection onto library saturation are shown. [Figure 24] Box plots of mean absolute deviation (MAD) scores for unbiased genome coverage across bins. Key continued from Figure 23. [Figure 25] 1 shows the experimental layout of PDAC low-passage patient-derived lines for the generation of s3-GCC libraries, demonstrating indexing at both the tagmentation and PCR stages of multiple 96-well plates. [Figure 26] Boxplots of library complexity measured by unique reads per cell and projection to 50% and 95% library saturation are shown. Top: total reads, middle: distal (>1 kbp mapping) intrachromosomal reads, bottom: reads mapped to chromosomes. [Figure 27] 10 shows a density plot of the mapped read length distribution showing distal region capture. [Figure 28] Clustering of single-cell GCC libraries on shared topology domains: cell lines (left) and K-means defined clusters (right). [Figure 29] Exemplary nucleotide sequences of the first strand of the transposon, the second index sequence, and oligonucleotides containing P5, i5, P7, i7, ME, A14, and B15 are shown (SEQ ID NOs: 1-9, respectively). DETAILED DESCRIPTION OF THE INVENTION

[0038] The schematic diagrams are not necessarily to scale. Like numbers used in the figures refer to like components, steps, etc. However, it will be understood that the use of a number to refer to a component in a given figure is not intended to limit the component in another figure labeled with the same number. Furthermore, the use of different numbers to refer to a component is not intended to indicate that the differently numbered component may not be the same as or similar to the other numbered component.

[0039] Presented herein are methods, compositions, systems, and kits related to performing nucleic acid sequencing and / or assays. The present disclosure provides a method for significantly increasing the number of target nucleic acids present in a sequencing library. Figure 1 shows a general overview of one exemplary embodiment of the method. In this exemplary embodiment, the method includes providing a target nucleic acid modified to contain identical adapters at each end (Figure 1, Block 10), herein referred to as a target nucleic acid with symmetric adapters. The source of the target nucleic acid is not intended to be limiting, and the target nucleic acid can be derived from DNA or RNA converted to DNA. Similarly, the method used to add adapters to the ends of the target nucleic acid is not intended to be limiting, and can include, for example, transposition, fragmentation followed by ligation, ligation, or extension and ligation. The method further includes modifying one of the symmetric adapters to convert the symmetrically modified target nucleic acid into an asymmetrically modified target nucleic acid (Figure 1, Block 12), a target nucleic acid containing different adapters at each end. The adapters can include an index sequence, a UMI, a universal sequence, and / or a sequence derived from a primer. Optionally, the asymmetric target nucleic acid can be amplified ( FIG. 1 , block 14). Amplification of the asymmetric target nucleic acid can include adding one or more index sequences, UMI sequences, universal sequences, or other useful sequences, including but not limited to, sequences derived from primers, to one or both ends.

[0040] The inventors have surprisingly and unexpectedly observed that during conversion of a symmetric target nucleic acid to an asymmetric target nucleic acid, the modified target nucleic acid can be exposed to conditions that significantly increase the yield of the asymmetrically modified target nucleic acid to near the theoretical maximum yield. This can be used with any source of target nucleic acid and is particularly useful for methods where high-efficiency library generation is advantageous, including methods using limited input primary nucleic acid. Any sequencing library method can benefit from high-efficiency generation, including, but not limited to, whole genome sequencing, targeted sequencing, methylation sequencing, genomic conformation capture (GCC), e.g., HiC, chromatin conformation, single-cell assays, single-cell combinatorial indexing, RNA-seq and ATAC-seq methods, simultaneous assays, e.g., DNA and RNA, embodiments in which the source is cell-free DNA or RNA, and liquid biopsies. High-efficiency conversion assays are also useful in detecting the presence of analytes, for example, by increasing sensitivity. Examples of detection or screening assays include, but are not limited to, PCR, qPCR, digital PCR, DNA or RNA or antibody or protein detection assays, or general analyte detection assays. Examples of analytes include, but are not limited to, DNA, RNA, and proteins.

[0041] target nucleic acid The target nucleic acids used in the methods, compositions, systems, and kits provided herein are typically derived from primary nucleic acids present in a sample. The primary nucleic acids may be derived from the sample in double-stranded DNA (dsDNA) form (e.g., genomic DNA fragments, amplification products, etc.), or may be derived from the sample in single-stranded form as DNA or RNA and converted to dsDNA form. For example, during the methods described herein, mRNA molecules can be copied into double-stranded cDNA using standard techniques known in the art. The exact sequence of polynucleotide molecules from a primary nucleic acid sample is generally not critical to the present disclosure and may be known or unknown.

[0042] In one embodiment, the primary nucleic acid comprises a DNA molecule. The primary nucleic acid molecule may represent the entire gene complement of an organism, for example, a genomic DNA molecule including both intron and exon sequences, as well as non-coding regulatory sequences such as promoter and enhancer sequences. In one embodiment, a specific subset of genomic DNA can be used, such as one or more specific sequences, such as a specific chromosome, DNA associated with open chromatin, DNA associated with closed chromatin, or a region of a specific gene (e.g., targeted sequencing).

[0043] In one embodiment, the primary nucleic acid comprises an RNA molecule. The primary nucleic acid molecule may represent the entire transcriptome or one or more cells of the sample, for example, mRNA molecules. The primary nucleic acid molecule may represent one or more cells of the sample, for example, non-coding RNA, such as microRNA or small interfering RNA. In one embodiment, a specific subset of RNA molecules can be used, for example, one or more specific sequences, such as regions encoded by specific genes.

[0044] Samples may include nucleic acid molecules obtained from biopsies, tumors, scrapings, swabs, blood, mucus, urine, plasma, semen, hair, laser capture microdissections, surgical resections, and other clinically or laboratory-derived samples. In some embodiments, samples may be epidemiological, agricultural, forensic, or pathogenic samples. In some embodiments, samples may include cultured cells. In some embodiments, samples may include nucleic acid molecules obtained from animals, such as humans or mammalian sources. In other embodiments, samples may include nucleic acid molecules obtained from non-mammalian sources, such as plants, bacteria, viruses, or fungi. In some embodiments, the source of the nucleic acid molecules may be preserved or extinct samples or species.

[0045] Additionally, the methods, compositions, systems, and kits disclosed herein may be useful for amplifying nucleic acid samples containing low-quality nucleic acid molecules, such as degraded and / or fragmented genomic DNA from forensic samples. In one embodiment, a forensic sample may include nucleic acids obtained from a crime scene, from a missing persons DNA database, from a laboratory associated with a forensic investigation, or from a forensic sample obtained by a law enforcement agency, one or more military services, or personnel thereof. A nucleic acid sample may be a purified sample or crude lysate containing nucleic acids from, for example, a buccal swab, paper, cloth, or other substrate impregnated with saliva, blood, or other bodily fluid. Thus, in some embodiments, a nucleic acid sample may contain small amounts of DNA or fragmented portions of DNA, such as genomic DNA. In some embodiments, the target nucleic acid may be present in one or more bodily fluids, including, but not limited to, blood, sputum, plasma, semen, urine, and serum. In some embodiments, the target sequence may be obtained from a victim's hair, skin, tissue sample, autopsy, or corpse. In some embodiments, nucleic acids comprising one or more target sequences can be obtained from a deceased animal or human. In some embodiments, the target sequences can comprise nucleic acids obtained from non-human DNA, such as microbial, plant, or entomological DNA. In some embodiments, the target sequences are for purposes of human identification, such as forensic samples.

[0046] Further non-limiting examples of sources of biological samples include whole organisms, as well as samples obtained from patients. Biological samples can be obtained from any biological fluid or tissue and can be in a variety of forms, including liquid fluids and tissues, solid tissues, and preserved forms such as dried, frozen, and fixed forms. Samples can be of any biological tissue, cell, or bodily fluid. Such samples include, but are not limited to, sputum, blood, serum, plasma, blood cells (e.g., white blood cells), ascites, urine, saliva, tears, sputum, vaginal fluid, saliva, tears, vaginal fluid, saliva, tears, vaginal fluid (secretion), washings obtained during medical procedures (e.g., pelvic or other washings obtained during biopsy, endoscopy, or surgery), tissue, nipple aspirate, core or fine needle biopsy sample, cell-containing bodily fluids, peritoneal fluid, and pleural fluid, or cells therefrom, as well as free-floating nucleic acids such as cell-free circulating DNA. Biological samples can also include sections of tissue, such as frozen or fixed sections taken for histological purposes or microdissected cells or their extracellular portions. In some embodiments, the sample can be a blood sample, such as a whole blood sample. In another example, the sample is an unprocessed dried blood spot (DBS) sample. In yet another example, the sample is a formalin-fixed paraffin-embedded (FFPE) sample. In yet another example, the sample is a saliva sample. In yet another example, the sample is a dried saliva spot (DSS) sample.

[0047] Exemplary biological samples from which target nucleic acids can be derived include, for example, those from eukaryotes, e.g., mammals such as rodents, mice, rats, rabbits, guinea pigs, ungulates, horses, sheep, pigs, goats, cows, cats, dogs, primates, humans, or non-human primates; plants, e.g., Arabidopsis thaliana, corn, sorghum, oats, wheat, rice, canola, or soybeans; algae, e.g., Chlamydomonas reinhardtii; nematodes, e.g., Caenorhabditis elegans; insects, e.g., Drosophila melanogaster, mosquitoes, fruit flies, honeybees, or spiders; fish, e.g., zebrafish; reptiles, amphibians, e.g., frogs or Xenopus laevis; Dictyostelium discoideum; Pneumocystis carinii; Takifugu rubripes; yeast, Saccharomoyces cerevisiae; or Schizosaccharomyces Examples of target nucleic acids include fungi such as Escherichia coli, Staphylococci, or Mycoplasma pneumoniae, prokaryotes such as Escherichia coli, Staphylococci, or Mycoplasma pneumoniae, archaea, viruses such as Hepatitis C virus or human immunodeficiency virus, or viroids. Target nucleic acids can be derived from homogenous cultures or populations of organisms as described herein, or alternatively, can be derived from a collection of several different organisms, for example, in a community or ecosystem.

[0048] In some embodiments, the sample comprises tissue that is processed to obtain the desired primary nucleic acid. In some embodiments, cells are used to obtain the desired primary nucleic acid. In some embodiments, nuclei are used to obtain the desired primary nucleic acid. The method may further include dissociating the cells and / or isolating the nuclei. Methods for isolating cells and nuclei from tissue are available (WO 2019 / 236599).

[0049] In some embodiments, nucleic acids present in tissues, cells, or isolated nuclei can be processed according to the desired readout. For example, nucleic acids can be fixed during processing, and useful fixation methods are available (WO 2019 / 236599). Fixation can be useful for preserving samples or maintaining the continuity of analytes from samples, cells, or nuclei. Fixation methods preserve and stabilize tissue, cell, and nuclear morphology and architecture, inactivate proteolytic enzymes, and strengthen samples, cells, and nuclei so they can withstand further processing and staining and protect them from contamination. Examples of methods in which fixation can be useful include, but are not limited to, whole genome sequencing of isolated nuclei and chromosome conformation capture methods such as Hi-C. Common fixation methods include perfusion, immersion, freezing, and desiccation (Srinivasanet al., Am J Pathol. 2002 Dec;161(6):1961-1971. doi:10.1016 / S0002-9440(10)64472-0).

[0050] In some embodiments, such as whole genome sequencing, methods are available for treating isolated nuclei to dissociate nucleosomes from DNA while leaving the nuclei intact, generating nucleosome-free nuclei (WO 2018 / 018008). In one embodiment, a detergent-based nucleosome method is used (Example 2). In some embodiments, such as chromosome conformation capture methods, nucleic acids present in tissues, cells, or isolated nuclei can be fragmented, for example, by restriction endonuclease digestion. Fragmentation is described in more detail herein. In some embodiments, such as chromosome conformation capture methods, nucleic acids present in tissues, cells, or isolated nuclei can be exposed to conditions for proximity-based ligation, such as blunt-end ligation.

[0051] In some embodiments, for example, primary nucleic acids in bulk from multiple cells can be used to generate the sequencing libraries described herein. In other embodiments, individual cells or nuclei can be used as a source of primary nucleic acids, allowing sequence information to be obtained from single cells and nuclei. Many different single-cell library preparation methods are known in the art, including, but not limited to, (Hwang et al. Experimental & Molecular Medicine, vol. 50, Article number: 96 (2018)), Drop-seq, Seq-well, and single-cell combinatorial indexing ("sci-") methods. Companies providing single-cell products and related technologies include 10X Genomics, Takara biosciences, BD biosciences, Biorad, 1cellbio, IsoPlexis, CellSee, nanoselect, and Dolomite. Biotechnology-based methods include, but are not limited to, SCI-seq. SCI-seq is a methodological framework that uses split-pool barcoding to uniquely label the nucleic acid content of a large number of single cells or nuclei. Generally, the number of nuclei or cells can be at least two. The upper limit depends on the practical limitations of the equipment used in other steps of the methods described herein (e.g., multi-well plates, number of indexes). The number of nuclei or cells that can be used is not intended to be limiting and can reach billions. For example, in one embodiment, the number of nuclei or cells can be 100,000,000 or less, 10,000,000 or less, 1,000,000 or less, 10,000,000 or less, 1,000,000 or less, 100,000 or less, 10,000,000 or less, 1,000,000 or less, 100,000 or less, 10,000 or less, 1,000 or less, 500 or less, or 50 or less.

[0052] adapter The disclosed method can include adding adapters to both ends of a target nucleic acid. Many adapters are known for use in preparing sequencing libraries, and essentially any adapter can be used. For example, an adapter can be single-stranded, double-stranded, or contain a double-stranded region and a single-stranded region. In one embodiment, the single-stranded region of an adapter having both a single-stranded and a double-stranded region can be used as a "sticky end" to aid in binding the adapter to a target nucleic acid having a complementary single-stranded region at each end. In one embodiment, an adapter having both a single-stranded and a double-stranded region is also called a forked or mismatch adapter, the general characteristics of which are known (Gormley et al., U.S. Pat. No. 7,741,463; Bignell et al., U.S. Pat. No. 8,053,192). In one embodiment, the adapter is present as part of a transposome complex. Transposome complexes are described in detail herein.

[0053] One or both ends of an adapter used to attach to both ends of a target nucleic acid can be modified to alter the interaction of the adapter with other nucleic acids. In one embodiment, one 3' end of the adapter can be blocked to reduce the ligation efficiency of that particular end. In one embodiment, the addition of an adapter, such as a double-stranded adapter, to each end of a target nucleic acid results in a gap in one strand of the resulting modified target nucleic acid. In one embodiment, the gap is at least one nucleotide. In one embodiment, the gap is located between the 3' end of the target nucleic acid and the 5' end of the adapter attached to the target nucleic acid.

[0054] An adapter can include one or more index sequences, one or more UMIs, one or more universal sequences, one or more DNA lesions, or a combination thereof. As described in more detail herein, the presence of an index sequence in an adapter can aid in sci-based applications, sample indexing, or single-cell identification.

[0055] When used as a template by a DNA polymerase during DNA synthesis, the DNA lesion nucleotide has a structure that causes certain DNA polymerases to reduce their activity and stall or terminate DNA synthesis at the DNA lesion. This type of DNA polymerase is referred to herein as a "damage-intolerant polymerase." Examples of nucleotides that can be used as DNA lesions are known to those skilled in the art and include, but are not limited to, abasic sites, modified bases, mismatches, single-strand breaks, or bridged nucleotides. Examples of modified bases include, but are not limited to, methylated bases (e.g., N3-methyladenine, N7-O6-methylguanine, N3-methylcytosine, O4-methylthymine), O6-alkylguanine, O4-alkylthymine, hypoxanthine, xanthine, and uracil. Modified bases can also include oxidized bases, including, but not limited to, FapyTA, 8-oxo-G, and thymine glycol. Examples of bridged nucleotides include, but are not limited to, thymine dimers.

[0056] Damage-intolerant polymerases are known to those skilled in the art (Heyn et al., Nucleic Acids Res. 2010 Sep;38(16):e161; Sikorsky et al., Biochem Biophys Res Commun. 2007 Apr 6;355(2):431-437; and Gruz et al., Nucleic Acids Res. 2003-Jul-15;31(14):4024-4030). Examples of useful damage-intolerant polymerases are listed in Table 1.

[0057] [Table 1]

[0058] The method of the present disclosure can include a step of using a damage-intolerant polymerase, and can also include another step of using a DNA polymerase with unreduced activity using DNA damage as a template. When DNA damage is used as a template, the polymerase with unreduced activity is referred to herein as a "damage-tolerant polymerase." Damage-tolerant polymerases are known to those skilled in the art and include, but are not limited to, those listed in Table 1. The use of a damage-tolerant polymerase can occur during the conversion of a symmetrically modified target nucleic acid to an asymmetrically modified target nucleic acid, typically resulting in the loss of DNA damage in the resulting amplicon. The use of a damage-tolerant polymerase during conversion is described herein.

[0059] DNA damage can comprise one or more nucleotides that have the activity of reducing DNA polymerase activity.For example, the number of nucleotides that constitute DNA damage can be at least 1, at least 2, at least 3, at least 4, or at least 5.In one embodiment, the number of nucleotides that constitute DNA damage can be 5 or less, 4 or less, 3 or less, or 2 or less.In one embodiment, DNA damage is 2, 3, or 4 uracil nucleotides.When DNA damage comprises two or more nucleotides, the nucleotides of DNA damage are typically consecutive.

[0060] The DNA damage is typically present on one strand of the adaptor present at each end of the target nucleic acid. In one embodiment, when the adaptor contains a DNA damage and a gap is present on one strand where the adaptor is bound to the target nucleic acid, the DNA damage and the gap are located on different strands.

[0061] The adapter may also comprise a capture agent. As used herein, the term "capture agent" refers to a material, chemical, molecule, or portion thereof that can attach, retain, or bind to a nucleic acid (e.g., an adapter strand). Exemplary capture agents include, but are not limited to, members of receptor-ligand binding pairs (e.g., avidin, streptavidin, biotin, lectins, carbohydrates, nucleic acid-binding proteins, epitopes, antibodies, etc.) that can bind to a member of the receptor-ligand binding pair, or chemical reagents that can form a covalent bond with a linking moiety. In one embodiment, the capture agent is biotin. The capture agent can bind to the adapter strand and bind to the end of the adapter so as not to interfere with binding of the adapter to the target nucleic acid. For example, the 5' end of the adapter can comprise a capture agent, or the 3' end of the adapter can comprise a capture agent. In one embodiment, the capture agent is attached to the 5' end of the transposon strand or the 3' end of the other transposon strand. Capture agents are useful for attaching adapters to solid surfaces such as beads or wells.

[0062] The adaptor may also include a cleavable linker between the capture agent and the adaptor. Examples of cleavable linkers include, but are not limited to, disulfide bonds, which can be cleaved with, for example, dithiothreitol to release the capture agent. Capture agents with cleavable linkers, including biotin-labeled nucleotides with cleavable linkers, are commercially available.

[0063] Generation of target nucleic acids with symmetric adapters The methods, compositions, systems, and kits provided herein may optionally include processing of primary nucleic acids to obtain modified target nucleic acids with the same adapter at each end, thereby achieving a length suitable for sequencing and symmetry. The primary nucleic acid sample may include high-molecular-weight material, such as genomic DNA, or low-molecular-weight material, such as nucleic acid molecules obtained from liquid biopsies or by converting RNA to DNA. Various methods are known for processing nucleic acids present in bulk, isolated nuclei, or isolated cells into nucleic acid fragments. In one embodiment, a transposome complex is used to add adapters. In another embodiment, DNA is fragmented, for example, by enzymatic or mechanical methods, and then adapters are added to the ends of the fragments. In another embodiment, RNA molecules, such as mRNA, are converted to cDNA, and adapters are added to the ends.

[0064] The transposome complex typically contains a transposase bound to a transposon sequence containing a transposase recognition site, which can then be inserted into a target nucleic acid within a DNA molecule in a process sometimes referred to as "tagmentation." Tagmentation combines a single-step fragmentation and ligation process to add universal adapters (Gunderson et al., International Publication No. 2016 / 130704). Those skilled in the art will recognize that tagmentation is typically used to generate nucleic acid fragments containing different adapters at each end, as the generation of asymmetric target nucleic acids is easily and efficiently achieved by transposition and is ready for sequencing. While useful, tagmentation methods for generating asymmetric target nucleic acids are inefficient, typically reducing theoretical yields to 50%. In contrast, when used in the methods disclosed herein, tagmentation generates nucleic acid fragments containing the same nucleotide sequence at each end, increasing theoretical yields to nearly 100%.

[0065] In some embodiments, one strand of the transposon may be transferred, e.g., covalently linked, to the 5' end of the target nucleic acid during an insertion event. Such a strand is referred to as the "transfer strand." The transposon sequence may include an adapter, which may include one or more index sequences, one or more UMIs, one or more universal sequences, one or more DNA lesions, or a combination thereof. In one embodiment, the universal sequence is a transposase recognition site. Examples of transposase recognition sites include, but are not limited to, mosaic elements (MEs). In one embodiment, an adapter, e.g., one or more index sequences, one or more UMIs, one or more universal sequences, one or more DNA lesions, or a combination thereof, is present on the transferred strand. In some embodiments, one strand of the transposon is not transferred, e.g., covalently linked, to the 3' end of the target nucleic acid during an insertion event. Such a strand is referred to as the "transfer strand." The presence of the non-transferred strand may result in the generation of overlapping nucleotides of the target nucleic acid during the transposition reaction, causing a gap between the 5' end of the adapter sequence and the 3' end of the target nucleic acid. The size of the gap can vary and typically depends on the transposon system used, for example, gaps introduced by Tn5-based systems are typically 9 bases.

[0066] Some embodiments may include the use of hyperactive Tn5 transposase and Tn5-type transposase recognition sites (Goryshin and Reznikoff, J. Biol. Chem., 273:7367 (1998)), or MuA transposase and Mu transposase recognition sites containing R1 and R2 end sequences (Mizuuchi, K., Cell, 35:785, 1983; Savilahti, H. et al., EMBO J., 14:4893, 1995). Tn5 Mosaic End (ME) sequences, transposase recognition sites, can also be used, as optimized by those skilled in the art.

[0067] Further examples of transposition systems that can be used with certain embodiments of the methods, compositions, systems, and kits provided herein include Staphylococcus aureus Tn552 (Colegio et al., J. Bacteriol., 183:2384-8, 2001; Kirby C et al., Mol. Microbiol., 43:173-86, 2002), Ty1 (Devine & Boeke, Nucleic Acids Res., 22:3765-72, 1994, and WO 95 / 23875), transposon Tn7 (Craig, NL, Science. 271:1512, 1996; reviewed in Craig, NL, Curr. Top Microbiol. Immunol., 204:27-48, 1996), Tn / O, and IS10 (Kleckner N, et al., Curr. Top Microbiol. Immunol., 204:49-82, 1996), Mariner transposase (Lampe DJ, et al., EMBO J., 15:5470-9, 1996), Tc1 (Plasterk RH, Curr. Topics Microbiol. Immunol., 204:125-43, 1996), P element (Gloor, GB, Methods Mol. Biol., 260:97-114, 2004), Tn3 (Ichikawa & Ohtsubo, J Biol. Chem. 265:18829-32, 1990), bacterial insertion sequences (Ohtsubo & Sekine, Curr. Top. Microbiol. Immunol. 204:1-26, 1996), retrovirus (Brown, et al., Proc. Natl. Acad. Sci. USA, 86:2525-9, 1989), and yeast retrotransposons (Boeke & Corces, Annu Rev Microbiol. 43:403-34, 1989).Other examples include IS5, Tn10, Tn903, IS911, and engineered forms of transposase family enzymes (Zhang et al., (2009) PLoS Genet. 5:e1000689. Epub 2009 Oct 16; Wilson C. et al. (2007) J. Microbiol. Methods 71:332-5).

[0068] Other examples of integrases that can be used with the methods and compositions provided herein include retroviral integrases and integrase recognition sequences for such retroviral integrases, e.g., integrases from HIV-1, HIV-2, SIV, PFV-1, and RSV.

[0069] Transposon sequences useful in the methods and compositions described herein are provided in U.S. Patent Application Publication No. 2012 / 0208705, U.S. Patent Application Publication No. 2012 / 0208724, and WO 2012 / 061832.

[0070] Various transposome complex configurations are known in the art. In one embodiment, a transposome complex comprises a dimeric transposase having two subunits and two discontinuous transposon sequences. Examples of such transposomes are known in the art (see, e.g., U.S. Patent Application Publication No. 2010 / 0120098). In some embodiments, a transposome complex comprises a transposon sequence nucleic acid that links two transposase subunits to form a "looped complex" or "looped transposome." In one example, a transposome comprises a dimeric transposase and a transposon sequence. The looped complex can ensure that the transposon is inserted into the target DNA without fragmenting the target DNA and while maintaining the sequence information of the original target DNA. As will be appreciated, the looped structure can insert a desired adapter sequence into the target nucleic acid while maintaining the physical connectivity of the target nucleic acid. In some embodiments, the transposon sequence of a looped transposome complex can include a fragmentation site so that the transposon sequence can be fragmented to create a transposome complex containing two transposon sequences. Such transposome complexes are useful for ensuring that adjacent target DNA fragments into which the transposon is inserted receive barcode combinations that can be unambiguously assembled at later stages of the assay.

[0071] Fragmentation sites can be introduced into target nucleic acids by using transposome complexes. In one embodiment, after fragmentation of the nucleic acid, the transposase remains bound to the nucleic acid fragments, so that nucleic acid fragments derived from the same genomic DNA molecule remain physically linked (Adey et al., 2014, Genome Res., 24:2041-2049). Cleavage can be performed by biochemical, chemical, or other means. In some embodiments, the fragmentation site can include nucleotides or nucleotide sequences that can be fragmented by various means. Examples of fragmentation sites include, but are not limited to, restriction endonuclease sites, at least one ribonucleotide cleavable by an RNAse, a nucleotide analogue cleavable in the presence of a particular chemical agent, a diol bond cleavable by treatment with periodate, a disulfide group cleavable by a chemical reducing agent, a cleavable moiety that can be subjected to photochemical cleavage, and a peptide cleavable by a peptidase enzyme or other suitable means (see, e.g., U.S. Patent Application Publication No. 2012 / 0208705, U.S. Patent Application Publication No. 2012 / 0208724, and WO 2012 / 061832).

[0072] In embodiments where the primary nucleic acid is DNA, the result of transposition is a library of modified target nucleic acids, each fragment containing a symmetric adapter at each end. In contrast, in embodiments where the primary nucleic acid is RNA, the result of transposition is up to three different types of modified target nucleic acids. The first population contains a library of modified target nucleic acids, each fragment containing a symmetric adapter at each end. The second and third populations each contain an adapter introduced by the transposon at one end, and the other end, i.e., the end corresponding to either the 3' or 5' end of the RNA, is added by an alternative method such as a template switch primer, a random primer, or polyT.

[0073] Instead of transposition, target nucleic acids can be obtained by fragmentation. Fragmentation of primary nucleic acids from a sample can be achieved in any order by enzymatic, chemical, or mechanical methods, and then adapters are added to the ends of the fragments. Examples of enzymatic fragmentation include CRISPR and Talen-like enzymes, as well as enzymes that unwind DNA (e.g., helicases) to create single-stranded regions to which DNA fragments can hybridize and initiate extension or amplification. For example, helicase-based amplification can be used (Vincent et al., 2004, EMBO Rep., 5(8):795-800). In one embodiment, extension or amplification is initiated using random primers. Examples of mechanical fragmentation include nebulization or sonication.

[0074] Fragmentation of primary nucleic acids by mechanical means results in fragments with a heterogeneous mixture of blunt ends, 3' overhanging ends, and 5' overhanging ends. Therefore, it is desirable to repair the fragment ends using methods known in the art, for example, to generate ends that are optimal for adding adapters to the blunt sites. In certain embodiments, the fragment ends of the nucleic acid population are blunt. More specifically, the fragment ends are blunt and phosphorylated. Phosphate moieties can be introduced by enzymatic treatment, for example, using polynucleotide kinase.

[0075] In one embodiment, fragmented nucleic acids are prepared with overhanging nucleotides. For example, a single overhanging nucleotide can be added by the activity of certain types of DNA polymerases, such as Taq polymerase or Klenow exo-minus polymerase, which have non-templated terminal transferase activity that adds a single deoxynucleotide, e.g., adding an "A" nucleotide to the 3' end of a DNA molecule. Using such enzymes, a single "A" nucleotide can be added to the 3' end of the blunt end of each strand of a double-stranded nucleic acid fragment. Thus, an "A" can be added to the 3' end of each end-repaired strand of a double-stranded target fragment by reaction with Taq or Klenow exo-minus polymerase, while the adapter can be a T construct with a compatible "T" overhang present at the 3' end of each region of the double-stranded nucleic acid of the universal adapter. In one example, terminal deoxynucleotidyl transferase (TdT) can be used to add multiple "T" nucleotides (Swift Biosciences, Ann Arbor, MI). This type of end modification also prevents self-ligation of both the vector and the target, so that there is a bias towards forming target nucleic acids with the same adapter at each end.

[0076] Adapters can be added to the ends of fragmented DNA or asymmetric DNA target nucleic acids by various methods, including, for example, ligation of double-stranded adapters to the ends of fragments or extension of annealed primers. Ligation of double-stranded adapters to the ends of fragments can be blunt-ended or assisted by using overhangs present at the ends of the fragments. Adapters can also be added using single-stranded or double-stranded adapters (e.g., TdT labeling) involving ligation or polymerization. In one embodiment, the adapter is configured to introduce a gap in one strand of the resulting modified target nucleic acid. In one embodiment, the gap is at least one nucleotide. In one embodiment, the gap is located between the 3' end of the target nucleic acid and the 5' end of the adapter attached to the target nucleic acid.

[0077] In embodiments in which the primary nucleic acid is RNA, generating a target nucleic acid with symmetric adapters typically involves converting the RNA to DNA, with the optional introduction of adapters at one or both ends. Various methods can be used to add adapters to the 3' end of mRNA. For example, adapters can be added using routine methods used to generate cDNA. A primer with a poly-T sequence at its 3' end and an adapter upstream of the poly-T sequence can be annealed to the mRNA molecule and extended using reverse transcriptase. This results in a one-step conversion of the mRNA to DNA, and optionally, a one-step conversion of a universal sequence at the 3' end. In one embodiment, the primer can also include one or more index sequences, one or more UMIs, one or more universal sequences, or a combination thereof. In one embodiment, random primers are used.

[0078] Non-coding RNA can also be converted to DNA and, optionally, modified to contain a universal sequence using various methods. For example, an adapter can be added using a first primer containing a random sequence and a template-switch primer, and either primer can contain an adapter. A reverse transcriptase with terminal transferase activity can be used to add non-templated nucleotides to the 3' end of the synthesized strand, and the template-switch primer contains nucleotides that anneal with the non-templated nucleotides added by the reverse transcriptase. An example of a useful reverse transcriptase is Moloney murine leukemia virus reverse transcriptase. In certain embodiments, the SMARTer™ reagent (Cat. No. 634926) available from Takara Bio USA, Inc. is used to add an index to the non-coding RNA and, if necessary, to the mRNA. Optionally, a template-switch primer can be used with mRNA in conjunction with a primer containing a poly-T sequence to add universal sequences to both ends of the DNA target nucleic acid generated from the RNA. In one embodiment, the same adapter is added to both ends.

[0079] A population of target nucleic acids can have an average chain length desirable or appropriate for a particular application of a method or composition described herein. For example, the average chain length of members used in one or more steps of a method described herein or present in a particular composition, system, or kit can be less than about 100,000 nucleotides, 50,000 nucleotides, 10,000 nucleotides, 5,000 nucleotides, 1,000 nucleotides, 500 nucleotides, 100 nucleotides, or 50 nucleotides. Alternatively, or in addition, the average chain length can be greater than about 10 nucleotides, 50 nucleotides, 100 nucleotides, 500 nucleotides, 1,000 nucleotides, 5,000 nucleotides, 10,000 nucleotides, 50,000 nucleotides, or 100,000 nucleotides. The average chain length of a population of target nucleic acids can be within a range between the maximum and minimum values ​​listed above. It will be understood that amplicons generated at an amplification site (or otherwise produced or used herein) can have an average chain length ranging between an upper and lower limit selected from those exemplified above.

[0080] In some embodiments, the target nucleic acid is sized relative to the area of ​​the amplification site, for example, to facilitate exclusion amplification. For example, the area of ​​each of the array sites can be larger than the diameter of the exclusion volume of the target nucleic acid to achieve exclusion amplification. For example, in embodiments using surface features of an array, the area of ​​each feature can be larger than the diameter of the exclusion volume of the target nucleic acid transported to the amplification site. The exclusion volume and its diameter of the target nucleic acid can be determined, for example, from the length of the target nucleic acid. Methods for determining the exclusion volume and diameter of the exclusion volume of a nucleic acid are described, for example, in U.S. Pat. No. 7,785,790, Rybenkov et al. Proc. Natl. Acad. Sci. USA 90:5307-5311 (1993), Zimmerman et al., J. Mol. Biol. 222:599-620 (1991), or Sobel et al., Biopolymers 31:1559-1564 (1991).

[0081] A cleanup process to increase molecular purity can be followed by tagmentation or fragmentation and processing of the target nucleic acid to generate primary nucleic acid fragments. Any suitable cleanup process, such as electrophoresis or size exclusion chromatography, can be used. In some embodiments, solid-phase reversibly immobilized paramagnetic beads can be used to separate the desired DNA molecules from unincorporated primers and select nucleic acids based on size, for example. Solid-phase reversibly immobilized paramagnetic beads are commercially available from Beckman Coulter (Agencourt AMPure XP), Thermo Fisher (MagJet), Omega Biotech (Mag-Bind), Promega Beads, and Kapa Biosystems (Kapa Pure Beads).

[0082] Conversion of target nucleic acids from symmetric to asymmetric The methods, compositions, systems, and kits provided herein involve converting a symmetric target nucleic acid to a target nucleic acid with an asymmetric adapter. As discussed herein, in some embodiments, the addition of an adapter to each end of the target nucleic acid results in a gap in each strand of the resulting modified target nucleic acid. In one embodiment, the gap is located between the 3' end of the target nucleic acid and the 5' end of the adapter attached to each end of the target nucleic acid. In one embodiment, the gap can be filled with nucleotides and ligated using the 3' end of the target nucleic acid as a primer. For example, in some embodiments using transposome complexes, a 9-bp target sequence overlap created by Tn5-based transposon insertion is extended. In one embodiment, extension results in displacement of the upstream sequence using a strand-displacing polymerase. In one embodiment, the target sequence overlap created by transposition is not extended. In one embodiment, ligation is used. When extension is used to fill the gap, a damage-intolerant or damage-tolerant polymerase can be used.

[0083] In one embodiment, when an adapter contains a DNA lesion and a gap exists in one strand where the adapter is bound to the target nucleic acid, the DNA lesion and the gap are located on different strands. The polymerase used to fill the gap by extension uses the DNA lesion in the template strand, and if the polymerase is damage-intolerant, extension will terminate. Thus, in this configuration, the use of a damage-intolerant polymerase results in the retention of only a portion of the adapter sequence of the adapter downstream of the gap. This results in the modification of one adapter of the target nucleic acid and the generation of an asymmetric target nucleic acid. Those skilled in the art will recognize that asymmetric target nucleic acids can be used in sequencing reactions, including paired-end sequencing reactions. However, the method of the present disclosure provides further advantages, as described herein.

[0084] An example of a possible structure for one embodiment in which a symmetric target nucleic acid is generated and then one adapter is modified to result in an asymmetric target nucleic acid is shown in Figure 2. An exemplary target nucleic acid 20 is shown with a symmetric adapter 22, as shown in Figure 2A. In this exemplary embodiment, the symmetric adapter contains a DNA lesion (denoted by U). The 3' end of one strand is blocked (denoted by *), and the 3' end of the other strand contains an overhang. The adapter may contain one or more universal sequences, one or more index sequences, one or more UMIs, or a combination thereof. After attachment of an adapter to each end of target nucleic acid 20, modified target nucleic acid 23 contains a gap 24 at the 3' end of the original target nucleic acid 20. Extension of modified target nucleic acid 23 using a damage-intolerant polymerase begins at the 3' end of gap 24 and terminates at the DNA lesion U, resulting in modified target nucleic acid 25, as shown in Figure 2C. Denaturation of the modified target nucleic acid 25 results in an asymmetric target nucleic acid 26 in which the nucleic acid comprises a strand of a symmetric adapter 22 with a DNA lesion at one end and a portion of the symmetric adapter sequence 27 at the other end located between the gap and the DNA lesion.

[0085] After modifying the symmetric target nucleic acid to a target nucleic acid having an asymmetric adapter, the asymmetric target nucleic acid can be further modified. For example, a sequence can be added by specifically targeting one of the ends, for example, by adding nucleotides to the first adapter added to the target nucleic acid, or by adding nucleotides to the adapter modified to produce an asymmetric target nucleic acid. In one embodiment, the modification can include using a primer in an extension reaction to add a second adapter to the adapter modified to produce an asymmetric target nucleic acid (e.g., modifying adapter 22 to 27, as illustrated in Figure 2D).

[0086] The primer used for modification can contain at least two domains. The first domain is located at the 3' end of the primer and contains a sequence that anneals to a portion of the modified adaptor to produce an asymmetric target nucleic acid. The first domain is also referred to herein as the annealing domain. Those skilled in the art will recognize that a primer is useful in this method if the first domain has a sufficient length for specific annealing. Those skilled in the art will also recognize that a primer is useful in this method if the nucleotide to which the primer anneals includes the 3' nucleotide of the asymmetric adaptor, thereby making the 3' nucleotide a suitable initiation site for extension using the second domain of the primer as a template. The 3' end of the asymmetric target nucleic acid can also be modified using ligation.

[0087] In one embodiment, one or more nucleotides of the annealing domain are modified nucleotides. Modified nucleotides are nucleotides that denature at a higher melting temperature than the corresponding natural DNA nucleotides, e.g., nucleotide hydrogen bonds with the complementary natural nucleotides A, T, G, or C with greater strength than the corresponding natural DNA nucleotides. Examples of modified nucleotides include, but are not limited to, locked nucleic acids (LNA), bridged nucleic acids (BNA), pseudo-complementary bases, peptide nucleic acids (PNA), 2,6-diaminopurine, 5' methyl dC, SuperT, RNA nucleotides, or any nucleotide or base known in the art that essentially increases the melting temperature. The number of modified nucleotides in the first domain of the primer can be at least one, at least two, at least three, at least four, or at least five. In some embodiments, a combination of natural and modified nucleotides is used. In one embodiment, the modified nucleotides are at least 5, at least 10, or at least 15 nucleotides away from the polymerase start site. In one embodiment, the concentration of primer useful for extension can be determined by routine titration.

[0088] In one embodiment, the 3' end of a primer or adapter is blocked to prevent incorporation of a nucleotide on the 3' end of the primer by a DNA polymerase. Examples of methods for blocking the 3' end of a primer include, but are not limited to, removal of the 3'-OH group or the presence of a nucleotide, such as a dideoxynucleotide (ddNTP), a reverse base, an additional base not containing its complement, or a mismatch base, at the 3' end of the primer.

[0089] The second domain of the primer has a nucleotide sequence containing an adapter. The adapter may include one or more index sequences, one or more UMIs, one or more universal sequences, or a combination thereof. Typically, any index sequence, UMI, and universal sequence present in the adapter is unique compared to any index sequence, UMI, and universal sequence already present in the asymmetric target nucleic acid. In some embodiments, if present, the universal sequence may be located at the 5' end of the primer, and any optional sequence, such as an index or UMI, may be present between the first domain and the universal sequence.

[0090] A primer is used to extend or ligate the 3' end of a single-stranded asymmetric target nucleic acid that has a symmetric adapter at one end and an asymmetric adapter at the other end.

[0091] In some embodiments, the effectiveness of extension depends on the annealing temperature, and one skilled in the art can easily identify useful annealing temperatures using temperature titration and amplification, such as qPCR. In one embodiment, a damage-intolerant DNA polymerase is used for extension. The result of extension is an asymmetric target nucleic acid that retains a symmetric adapter at one end, and the asymmetric adapter at the other end has been modified to contain a different adapter.

[0092] Natural nucleotides A, T, G, and C can be used for extension. In some embodiments, non-natural nucleotides are used. For example, methylated cytosine can be used. Methylated cytosine is advantageous in methylation sequencing applications (WO 2017 / 106481) because adapter primers typically are not converted during cytosine to uracil conversion.

[0093] In one embodiment, the extension reaction is repeated. The inventors have surprisingly and unexpectedly found that the use of multiple extension cycles with a two-domain primer having at least one modified nucleotide increases the yield of asymmetrically modified target nucleic acid to near the theoretical maximum yield. In one embodiment, the number of extensions can be at least 1, at least 3, at least 5, at least 7, at least 9, or at least 10. In one embodiment, the number of extensions can be 15 or less, 13 or less, or 11 or less. In one embodiment, the number of extensions is 10.

[0094] Another example of a possible structure in one embodiment in which a symmetric target nucleic acid is generated by tagmentation and then one adapter is modified to yield an asymmetric target nucleic acid is shown in Figure 3. An exemplary modified target nucleic acid 33 is shown in Figure 3A along with a target nucleic acid 30 and a symmetric adapter 32. The adapter can include one or more universal sequences, one or more index sequences, one or more UMIs, or a combination thereof. In this exemplary embodiment, the symmetric adapter 32 includes a DNA lesion (denoted U), a gap 34, and a universal sequence, such as a transposase recognition domain 35. Extension of the modified target nucleic acid 33 with a damage-intolerant polymerase begins at the 3' end of the gap 34 and terminates at the DNA lesion U. After denaturation, the resulting asymmetric target nucleic acid 36 is shown in Figure 3B. The asymmetric target nucleic acid 36 includes a strand of the symmetric adapter 32 along with a DNA lesion at one end. At the other end, the asymmetric target nucleic acid 36 includes an asymmetric adapter 37, e.g., a portion of the symmetric adapter sequence located between the gap and the DNA lesion. Figure 3C also shows an exemplary embodiment in which the asymmetric target nucleic acid 36 is further modified to include another adaptor. The two-domain primer 38 includes one domain 39 that anneals to the asymmetric adaptor 37 and a second domain that includes a different adaptor 40. In this exemplary embodiment, a block (*) is included to reduce extension initiated at the 3' end of the primer 38. Extension, optionally with a damage-intolerant polymerase, shown by the dotted line in Figure 3C, begins at the 3' end of the asymmetric target nucleic acid 36 and adds the different adaptor 40, resulting in the asymmetric target nucleic acid 41, as shown in Figure 3D.

[0095] Another example of a possible structure for one embodiment in which a symmetric target nucleic acid is generated by tagmentation and then one adapter is modified to yield an asymmetric target nucleic acid is shown in Figure 4. An exemplary transposome complex 41 of two transposases and a transposon includes adapters 42 (Figure 4A). Each adapter includes a primer (P5), an index (i5), a universal anchor sequence (A14), a DNA lesion uracil (U), a transposase recognition sequence (ME), and the complement of the transposase recognition sequence (ME'). The adapter also includes an optional capture agent (B) and an optional cleavable linker (CL) attached to the 5' end of one strand and an optional blocking dideoxynucleotide (ddC) attached to the 3' end of the other strand. In some embodiments, the configuration of the capture agent-cleavable linker and blocking group is switched. Figure 4B shows the tagged and fragmented nucleic acid still complexed with the transposase. For simplicity, a depiction of the dimer is shown in Figure 4A but not in Figure 4B. Figure 4C shows the structure after transposase removal and gap filling using a DNA damage-intolerant polymerase. Figure 4D shows the top strand of Figure 4C hybridized to it using a two-domain primer 43. The two-domain primer 43 contains one domain, ME, that anneals to a complementary ME' and a second domain containing different adapter sequences, B15, i7, and P7. Figure 4E shows the results of extension of the top strand based on the two-domain primer sequence. Figure 4F shows the tagged library fragments after primer removal. Extension, indicated by the dotted line in Figure 4D, begins at the 3' end of ME' and adds a different adapter 43, resulting in the asymmetric target nucleic acid shown in Figure 4F.

[0096] The library of asymmetric target nucleic acids is exposed to conditions to remove DNA damage, optionally adding one or more adapters to one or both ends of the asymmetric target nucleic acids, which may result in further modification of one or both ends with one or more universal sequences, one or more index sequences, one or more additional UMI sequences, or a combination thereof. In one embodiment, the conditions to remove DNA damage include extension with a damage-tolerant DNA polymerase. Examples of suitable damage-tolerant DNA polymerases are listed in Table 1. The damage-tolerant DNA polymerase can be used in any type of extension reaction that reads through DNA damage, and the resulting synthesized strand no longer contains DNA damage. In one embodiment, the conditions to remove DNA damage include a repair system. DNA repair systems include enzymes and mechanisms for fixing or repairing DNA damage, including, but not limited to, excision repair systems and DNA repair systems. DNA repair systems are known in the art (Chaudhuri et al., Nature Reviews Molecular Cell Biology, 2017, 18:610-621). After use of the DNA repair system, the library of asymmetric target nucleic acids is exposed to conditions that include an extension reaction.

[0097] In one embodiment, the extension is by a method that substantially increases the number of asymmetric target nucleic acids, hi one embodiment, the method can be amplification, including but not limited to polymerase chain reaction (PCR) and rolling circle amplification (RCA).

[0098] In one embodiment, the method involves the use of a transposome complex bound to a surface, such as a bead or well surface. Typically, in such embodiments, one of the strands of the transposon contains a capture agent, such as biotin. The use of a capture agent advantageously reduces the number of steps required to generate an asymmetric target nucleic acid. For example, a capture agent and an optional cleavable linker can be attached to the 5' end of one strand (e.g., the strand of adapter 42 in Figure 4, which contains primer P5, index i5, universal anchor sequence A14, the DNA damage uracil U, and the transposase recognition sequence ME). After tagmentation using the surface-bound transposome complex, a DNA damage-intolerant polymerase, dNTPs, and a two-domain primer, such as two-domain primer 43 in Figure 4D, can be added. Upon exposure to denaturing conditions, such as heat, the complement of the transposase recognition sequence ME' is removed. For example, ME' shown in Figure 4B no longer hybridizes. The polymerase extends the 3' copy of the target nucleic acid using the ME as a template, where the ME' binds to the 3' end of the target nucleic acid and terminates at the DNA damage. Following another denaturation step, a two-domain primer 43 anneals to the ME' attached to the 3' end of the target nucleic acid. Extension is initiated at the ME' using the two-domain primer as a template, resulting in an asymmetric target nucleic acid. The asymmetric target nucleic acid can then be removed from the solid surface.

[0099] In another embodiment using a transposome complex bound to a surface, such as the surface of a bead or well, the capture agent and optional cleavable linker can be attached to the 3' end of the other strand (e.g., the strand of adapter 42 in Figure 4, which contains the complement ME' of the transposase recognition sequence). After tagmentation using the surface-bound transposome complex, a DNA damage-intolerant polymerase, dNTPs, and a two-domain primer, such as two-domain primer 43 in Figure 4D, can be added. Upon exposure to denaturing conditions, such as heat, the other strand of the transposon and the bound target nucleic acid are released into solution. The two-domain primer 43 can anneal to the ME' attached to the 3' end of the target nucleic acid. Extension is initiated at the ME' using the two-domain primer as a template, resulting in an asymmetric target nucleic acid. The asymmetric target nucleic acid can then be removed from the solid surface.

[0100] Index Array In some embodiments, it may be useful to identify the source of target nucleic acids during the sequencing process. Examples of when this would be useful will be readily apparent to those skilled in the art and include, but are not limited to, the simultaneous analysis of multiple libraries from different sources (e.g., different subjects, samples, tissues, or cell types). Identifying the source of target nucleic acids can be achieved, for example, through the use of compartmentalization by distributing subsets of target nucleic acids into multiple compartments, uniquely labeling the target nucleic acids in each compartment (typically by modifying them with adapters containing unique index sequences), and then pooling the subsets. For example, single-cell combinatorial indexing ("sci-") methods typically use split-pool labeling. Thus, in some embodiments, an index bound to each target nucleic acid is present in a specific compartment, and the presence of this index is used to indicate or identify the compartment in which a population of nuclei or cells resides at this stage of the method. The use of indexes and distribution of nucleic acids into compartments, also referred to as compartmentalization, is described herein.

[0101] As used herein, an index sequence can be any suitable sequence of nucleotides in length, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more. A four-nucleotide tag provides the possibility of multiplexing 256 samples, while a six-base tag allows for the processing of 4096 samples. In some embodiments, the index is used to label nucleic acids within a specific compartment.

[0102] As described herein, modification of an asymmetric target nucleic acid to add an index can be achieved during the generation of a symmetric target nucleic acid. For example, an index can be included in a symmetric adapter. Additional indexes can be selectively added to either end of the asymmetric target nucleic acid in a subsequent step.

[0103] Methods for modifying asymmetric target nucleic acids by adding an index include, but are not limited to, direct inclusion by a primer, extension, transposition, or ligation. Examples of extension include, but are not limited to, primer hybridization, extension using reverse transcriptase, and amplification. The nucleotide sequence added to one or both ends of the asymmetric target nucleic acid can also include one or more universal sequences and / or UMIs. A universal sequence can be used as a "landing pad" in a subsequent step to anneal a nucleotide sequence that can be used, for example, as a primer for adding another nucleotide sequence, such as another index, universal sequence, and / or UMI, to the asymmetric target nucleic acid. Thus, incorporation of an index sequence can be a process involving one, two, or more steps, using essentially any combination of extension (including hybridization, reverse transcriptase, and / or amplification), ligation, or transposition.

[0104] In some embodiments, incorporation of indexes occurs over one, two, three, or more rounds of split and pool indexing, resulting in single, dual, triple, or multiple (e.g., four or more) indexed libraries, such as indexed single-cell libraries.

[0105] The method may include multiple partitioning steps in which a population of target nucleic acids, such as isolated nuclei or cells (also referred to herein as a pool), is split into subsets. While the following is discussed with respect to isolated nuclei or cells, those skilled in the art will understand that the "split and pool" step can be applied to any population of target nucleic acids. Typically, subsets of isolated nuclei or cells, e.g., subsets present in multiple compartments, are indexed with compartment-specific indexes and then pooled. This compartmentalization of target nucleic acids can occur at any stage at which indexes are added. For example, target nucleic acids may be present in compartments when symmetric adapters and / or additional adapters are added. Thus, the method typically includes at least one "split and pool" step of obtaining pooled isolated nuclei or cells, partitioning them, and adding compartment-specific indexes. The number of "split and pool" steps may depend on the number of different indexes added to the nucleic acid fragments. Each initial subset of nuclei or cells before indexing may be unique and distinct from other subsets. After indexing, subsets can be pooled, split into subsets, indexed, and pooled again as needed until a sufficient number of indexes have been added to the target nucleic acid. This process assigns a unique index or combination of indexes to each single cell or nucleus. After indexing is complete, e.g., after addition of one, two, three, or more indexes, the isolated nuclei or cells can be lysed. In some embodiments, index addition and lysis can occur simultaneously.

[0106] The number of nuclei or cells present in a subset, and therefore in each compartment, can be at least 1. In one embodiment, the number of nuclei or cells present in a subset is 100,000,000 or less, 10,000,000 or less, 1,000,000 or less, 100,000 or less, 10,000 or less, 4,000 or less, 3,000 or less, 2,000 or less, 1,000 or less, 500 or less, or 50 or less. In one embodiment, the number of nuclei or cells present in a subset can be 1 to 1,000, 1,000 to 10,000, 10,000 to 100,000, 100,000 to 1,000,000, 1,000,000 to 10,000,000, or 10,000,000 to 100,000,000. In one embodiment, the number of nuclei or cells present in each subset is approximately equal. The number of nuclei present in a subset, and therefore in each compartment, is based in part on the desire to reduce index collisions, where a collision is the presence of two nuclei or cells with the same index combination that end up in the same compartment at this step of the method. Methods for distributing nuclei or cells into subsets are known and routine to those of skill in the art and include fluorescence-activated cell sorting (FACS) and simple dilution.

[0107] The number of compartments in the partitioning step (and subsequent indexing) can depend on the format used. For example, the number of compartments can be 2 to 96 compartments (when using a 96-well plate), 2 to 384 compartments (when using a 384-well plate), or 2 to 1536 compartments (when using a 1536-well plate). In one embodiment, the number of compartments is 5000 or more (Takara Biosciences, icell8 system). In one embodiment, multiple plates can be used. In one embodiment, each compartment can be a droplet. When the type of compartment used is a droplet or well containing two or more nuclei or cells, any number of droplets or wells can be used, such as at least 10,000, at least 100,000, at least 1,000,000, or at least 10,000,000 droplets. Isolated nuclei or cell subsets are typically indexed within compartments before pooling.

[0108] FIG. 5 shows a general block diagram of a general exemplary method for single-cell combinatorial indexing according to the present disclosure. The method includes providing isolated nuclei or cells ( FIG. 5 , block 50) and distributing the isolated nuclei or cells into multiple compartments ( FIG. 5 , block 51). Block 40 refers to DNA, and those skilled in the art will recognize that the DNA can be, for example, genomic DNA or DNA derived from RNA. In an embodiment of this method, the isolated nuclei or cells are indexed with compartment-specific indexes by adding symmetric adapters ( FIG. 5 , block 52) and then pooled ( FIG. 5 , block 53). Thus, the method typically includes at least one "split and pool" step of obtaining pooled isolated nuclei or cells, distributing them, and adding compartment-specific indexes; the number of "split and pool" steps may depend on the number of different indexes added to the target nucleic acid. If asymmetric adapters are added to the second index, the pooled isolated nuclei or cells are distributed into a second plurality of compartments ( FIG. 5 , block 53) and indexed with compartment-specific indexes by adding asymmetric adapters ( FIG. 5 , block 54). Optionally, the asymmetric target nucleic acids can then be amplified ( FIG. 5 , block 55). Amplification of the asymmetric target nucleic acids can include the addition of other useful sequences to one or both ends, including, but not limited to, index sequences, UMI sequences, and / or universal sequences, and can be combined with further split-and-pool indexing.

[0109] The resulting indexed target nucleic acids collectively provide a library of nucleic acids that can be sequenced. The term library, also referred to herein as a sequencing library, refers to a collection of modified nucleic acids that contain known universal sequences at the 3' and 5' ends.

[0110] Purpose The methods provided by the present disclosure can be easily incorporated into essentially any application, including the preparation of whole genome, transcriptome, methylation, accessible (e.g., ATAC), and conformational state (e.g., HiC) sequencing libraries. This can be particularly useful in essentially any application requiring high library conversion, such as single-cell combinatorial indexing (sci) methods, including, but not limited to, sci-WGS-seq, sci-MET-seq, sci-ATAC-seq, and sci-RNA-seq. Instead of focusing sequencing library generation on generating target nucleic acids with different universal sequences on each side (e.g., asymmetric), integrating the methods provided by the present disclosure into sequencing library generation involves more efficient generation of target nucleic acids with the same universal sequence on each side (e.g., symmetric). Upon generating symmetric fragments, the methods described herein for converting symmetric fragments to asymmetric fragments can be applied. Numerous sequencing library methods are known to those of skill in the art that can be used to construct whole genome or targeted libraries (see, e.g., "Sequencing Methods Review," available at genomics.umn.edu / downloads / sequencing-methods-review.pdf).

[0111] In some embodiments, the application is whole genome or targeted sequencing. Generally, tissues, individual cells, or individual nuclei are processed as described herein to produce symmetric target nucleic acids. (See Example 2.) In some embodiments, individual cells or individual nuclei can be processed to release nucleosomes from genomic DNA (WO 2018 / 018008). The symmetrically modified target nucleic acids can then be processed as described herein to produce asymmetrically modified target nucleic acids. For example, as shown in FIG. 6, the nucleic acids are fixed to maintain nuclear integrity, exposed to conditions that remove nucleosomes from genomic DNA to make the entire genome accessible, and then have one population of adapters inserted, for example, by tagmentation, to produce symmetric target nucleic acids. The symmetric target nucleic acids can then be converted to asymmetric target nucleic acids as described herein.

[0112] In some embodiments, the application is to probe accessible DNA, such as ATAC-seq (Assay for Transposase-Accessible Chromatin Using Sequencing) for the identification of accessible DNA. Generally, tissues, individual cells, or individual nuclei with intact nucleosomes are processed as described herein to produce symmetric target nucleic acids (see Example 2). The symmetrically modified target nucleic acids can then be processed as described herein to produce asymmetrically modified target nucleic acids. For example, as shown in Figure 7, genomic DNA containing bound nucleosomes can be tagmented to produce symmetric target nucleic acids. The symmetric target nucleic acids can then be converted to asymmetric target nucleic acids as described herein.

[0113] In some embodiments, the application is for sequencing RNA, such as mRNA. In contrast to applications where RNA is converted to DNA and DNA is used as the starting material, adapters can be added to one or both ends of the RNA molecules during processing to DNA. This provides the option of 5' and / or 3' profiling of RNA or full-length RNA profiling. For example, as shown in Figure 8, in one exemplary embodiment, mRNA molecules are subjected to reverse transcriptase in the presence of a poly-T primer containing a universal sequence and a template switch primer, resulting in double-stranded DNA containing adapters (designated CS1) at each end (Figure 8A). After exposing the resulting double-stranded DNA to transposome complexes (Figure 8B) and converting the symmetric adapters to asymmetric adapters (Figure 8C), three distinct populations can result (Figure 8D). One population (the 3' end) can arise when a transposon sequence is inserted into the double-stranded DNA and the other end of the resulting target nucleic acid contains a sequence corresponding to the original 3' end of the mRNA. A second population (denoted as RNA bodies) can result when the transposon sequence is inserted into two locations within the double-stranded DNA. A third population (denoted as 5' ends) can result when the transposon sequence is inserted into the double-stranded DNA and the other end of the resulting target nucleic acid contains a sequence corresponding to the original 5' end of the mRNA.

[0114] In some embodiments, the application is methylation sequencing. A wide range of methods that allow for the analysis of methylation or hydroxymethylation status are described in the literature (Barros-Silva et al., Genes (Basel). 2018 Sep;9(9):429) or are known to those skilled in the art. Chemical (e.g., sodium bisulfite or borate chemistry) or enzymatic methods of conversion can be used in a variety of methylation sequencing methods, including, but not limited to, BS-seq, TAB-seq, RRBS-seq, MeDip-seq, methylcap-seq, MBD-seq, Nanopore-seq, oxBS-seq, SeqCap Epi CpGiant, BSAS, WGBS, and sci-MET (WO 2018 / 226708).

[0115] In one embodiment, the application is protein analysis. Proteins can be intracellular or surface-bound, isolated, or present in a biological sample. Various methods are available to those skilled in the art. A common method often used for protein detection is to label an antibody or Fab fragment with an oligonucleotide tag, allow the antibody to affinity bind to the protein of interest, and use the oligonucleotide tag for readout or detection. The oligonucleotide tag can include an index sequence, a UMI, a universal sequence, or a combination thereof.

[0116] In some embodiments, the application is a simultaneous assay, evaluating two or more different analytes or information simultaneously. Examples of analytes include, but are not limited to, DNA, RNA, and protein. Nucleic acids can be in different states, such as epigenetic states (e.g., ATAC, meC, 5-hydroxyMe, etc.), or conformational states (e.g., HiC, 3C, chromatin states, etc.). Examples include assays that analyze DNA and RNA, DNA and / or RNA and epigenetic states (e.g., ATAC, meC, 5-hydroxyMe, etc.), and DNA and conformational states (e.g., HiC, 3C, chromatin states, etc.).

[0117] An example of a simultaneous assay is the preparation of genomic DNA for genome + chromatin conformation sequencing, referred to herein as GCC-seq. GCC-seq combines whole genome sequencing and chromatin conformation analysis, capturing chromatin interactions at a higher rate than routine Hi-C-type methods when combined with single-cell or single-nucleus and split-and-pool indexing (see Example 2). As illustrated in Figure 9, genomic DNA is processed, for example, by fixation, restriction enzyme digestion, proximity ligation, and nucleosome depletion, followed by the addition of adapters to yield symmetric target nucleic acids. Optionally, molecular capture can be used. Symmetrically modified target nucleic acids can be processed as described herein.

[0118] Preparation of fixed samples for sequencing A library of indexed target nucleic acids can be prepared for sequencing. Methods for attaching indexed target nucleic acid fragments to a substrate are known in the art. In one embodiment, the indexed fragments are enriched using multiple capture oligonucleotides specific for the indexed fragments, and the capture oligonucleotides can be immobilized on the surface of a solid substrate, such as a flow cell or beads. For example, the capture oligonucleotides can include a first member of a universal binding pair, and a second member of the binding pair is immobilized on the surface of the solid substrate. Similarly, methods for amplifying immobilized target nucleic acids include, but are not limited to, bridge amplification and binding equilibrium exclusion. Methods for immobilization and amplification prior to sequencing are described, for example, in Bignell et al. (U.S. Patent No. 8,053,192), Gunderson et al. (WO 2016 / 130704), Shen et al. (U.S. Patent No. 8,895,249), and Pipenburg et al. (U.S. Patent No. 9,309,502).

[0119] Pooled samples can be immobilized in preparation for sequencing. Sequencing can be performed as a single molecule array or can be amplified prior to sequencing. Amplification can be performed using one or more immobilized primers. The immobilized primers can be, for example, on a flat surface or in a lawn on a pool of beads. The pool of beads can be isolated in an emulsion with a single bead in each "compartment" of the emulsion. At a concentration of only one template per "compartment," only a single template is amplified on each bead.

[0120] As used herein, the term "solid-phase amplification" refers to any nucleic acid amplification reaction performed on or in association with a solid support such that all or a portion of the amplification product is immobilized on the solid support upon formation. Specifically, this term encompasses solid-phase polymerase chain reaction (solid-phase PCR) and solid-phase isothermal amplification, which are reactions similar to standard solution-phase amplification except that one or both of the forward and reverse amplification primers are immobilized on the solid support. Solid-phase PCR encompasses systems such as emulsions in which one primer is immobilized on a bead and the other is in free solution, and colonization in solid-phase gel matrices in which one primer is immobilized on a surface and the other is in free solution.

[0121] In some embodiments, the solid support comprises a patterned surface. "Patterned surface" refers to the arrangement of distinct regions within or on an exposed layer of a solid support. For example, one or more regions can be features in which one or more amplification primers are present. The features can be separated by interstitial regions in which amplification primers are absent. In some embodiments, the pattern can be an xy format of features in rows and columns. In some embodiments, the pattern can be a repetitive sequence of features and / or interstitial regions. In some embodiments, the pattern can be a random sequence of features and / or interstitial regions. Exemplary patterned surfaces that can be used in the methods and compositions described herein are described in U.S. Pat. Nos. 8,778,848, 8,778,849, and 9,079,148, and U.S. Patent Application Publication No. 2014 / 0243224.

[0122] In some embodiments, the solid support comprises an array of wells or depressions on its surface, which can be fabricated as commonly known in the art using a variety of techniques, including, but not limited to, photolithography, stamping techniques, molding techniques, and microetching techniques. As understood in the art, the technique used will depend on the composition and shape of the array substrate.

[0123] The features within the patterned surface can be wells of an array of wells (e.g., microwells or nanowells) on glass, silicon, plastic, or other suitable solid support with a patterned covalently attached gel, such as poly(N-(5-azidoacetamylpentyl)acrylamide-co-acrylamide) (PAZAM, see, e.g., U.S. Patent Application Publication Nos. 2013 / 184796, WO 2016 / 066586, and 2015 / 002813, each of which is incorporated by reference in its entirety). This process creates a gel pad used for sequencing, which can be stable over many cycles of sequencing operations. Covalently attaching a polymer to the wells is useful for maintaining the gel in the structured features throughout the life of the structured substrate during various applications. However, in many embodiments, the gel need not be covalently attached to the wells. For example, in some conditions, silane-free acrylamide (SFA, see, e.g., U.S. Pat. No. 8,563,477) that is not covalently bonded to any part of the structured matrix can be used as the gel material.

[0124] In certain other embodiments, structured substrates can be fabricated by patterning a solid support material with wells (e.g., microwells or nanocells), coating the patterned support with a gel material (e.g., PAZAM, SFA, or a chemically modified variant thereof), such as an azido-SFA version, and polishing the gel-coated support, e.g., by chemical or mechanical polishing, thereby retaining the gel within the wells but removing or inactivating substantially all of the gel from the interstitial regions on the surface of the structured substrate between the wells. Primer nucleic acids can be attached to the gel material. A solution of modified target nucleic acids can then be contacted with the polished substrate, such that individual modified target nucleic acids are seeded into individual wells through interaction with primers attached to the gel material, but because the gel material is absent or inactive, the target nucleic acids do not occupy the interstitial regions. Amplification of the modified target nucleic acids will be confined to the wells because the absence or inactivity of gel within the interstitial regions prevents outward migration of growing nucleic acid colonies. The process is conveniently manufacturable and scalable, utilizing conventional micro- or nanofabrication methods.

[0125] Although the present disclosure encompasses "solid-phase" amplification methods in which only one amplification primer is immobilized (the other primer is typically in free solution), in one embodiment, a solid support is provided with both immobilized forward and reverse primers. In practice, since the amplification process requires an excess of primers to maintain amplification, there will be "multiple" identical forward primers and / or "multiple" identical reverse primers immobilized on the solid support. References herein to forward and reverse primers should be construed as encompassing "multiple" such primers, unless the context dictates otherwise.

[0126] As will be understood by those skilled in the art, any given amplification reaction requires at least one type of forward primer and at least one type of reverse primer specific to the template to be amplified. However, in certain embodiments, the forward and reverse primers may contain template-specific portions of the same sequence and may have the exact same nucleotide sequence and structure (including any non-nucleotide modifications). In other words, solid-phase amplification can be performed using only one type of primer, and such single-primer methods are encompassed within the scope of the present disclosure. Other embodiments may use forward and reverse primers that contain the same template-specific sequence but differ in some other structural features. For example, one type of primer may contain a non-nucleotide modification that is not present in the other.

[0127] Primers for solid-phase amplification are preferably immobilized to a solid support by a single-point covalent bond at or near the 5' end of the primer, leaving the template-specific portion of the primer free for annealing to its cognate template and the 3' hydroxyl group free for primer extension. Any suitable covalent attachment means known in the art can be used for this purpose. The attachment chemistry selected will depend on the nature of the solid support and any derivatization or functionalization applied to it. The primer itself may contain a moiety, which may be a non-nucleotide chemical modification, to facilitate attachment. In certain embodiments, the primer may contain a sulfur-containing nucleophile, such as a phosphorothioate or thiophosphate, at the 5' end. In the case of solid-supported polyacrylamide hydrogels, this nucleophile binds to bromoacetamide groups present in the hydrogel. A more specific means of attaching primers and templates to a solid support is via a 5' phosphorothioate bond to a hydrogel composed of polymerized acrylamide and N-(5-bromoacetamidoylpentyl)acrylamide (BRAPA), as described in WO 05 / 065814.

[0128] Certain embodiments of the present disclosure may utilize solid supports comprising an inert substrate or matrix (e.g., glass slides, polymeric beads, etc.) that has been "functionalized" by the application of a layer or coating of an intermediate material containing reactive groups that allow for covalent attachment to biomolecules, such as polynucleotides. Examples of such supports include, but are not limited to, polyacrylamide hydrogels supported on an inert substrate such as glass. In such embodiments, the biomolecule (e.g., polynucleotide) may be covalently attached directly to the intermediate material (e.g., hydrogel), or the intermediate material may itself be noncovalently attached to the substrate or matrix (e.g., glass substrate). The term "covalent attachment to a solid support" should be interpreted accordingly to encompass this type of arrangement.

[0129] The pooled samples may be amplified on beads, each bead containing a forward and reverse amplification primer. In one embodiment, a library of modified target nucleic acids is used to prepare clustered arrays of nucleic acid colonies by solid-phase amplification, more specifically solid-phase isothermal amplification, similar to those described in U.S. Patent Application Publication No. 2005 / 0100900, U.S. Patent No. 7,115,400, WO 00 / 18957, and WO 98 / 44151. The terms "cluster" and "colony" are used interchangeably herein and refer to distinct sites on a solid support containing a plurality of identical immobilized nucleic acid strands and a plurality of identical immobilized complementary nucleic acid strands. The term "clustered array" refers to an array formed from such clusters or colonies. In this context, the term "array" should not be understood as requiring an ordered arrangement of the clusters.

[0130] The terms "solid phase" or "surface" are used to refer to either a planar array in which the primers are attached to a flat surface, such as a glass, silica, or plastic microscope slide, or similar flow cell device, or beads, to which one or two primers are attached and the beads are amplified, or an array of beads on a surface after the beads have been amplified.

[0131] Clustered arrays can be prepared using a process of thermal cycling, such as that described in WO 98 / 44151, or a process in which the temperature is held constant and cycles of extension and denaturation are performed using changes in reagents. Such isothermal amplification methods are described in WO 02 / 46456 and U.S. Patent Application Publication No. 2008 / 0009420.

[0132] It will be understood that any of the amplification methods described herein or generally known in the art can be used with universal or target-specific primers to amplify immobilized DNA fragments. Suitable methods for amplification include, but are not limited to, polymerase chain reaction (PCR), strand displacement amplification (SDA), transcription-mediated amplification (TMA), and nucleic acid sequence-based amplification (NASBA), as described in U.S. Pat. No. 8,003,354. The above amplification methods can be used to amplify one or more nucleic acids of interest. For example, PCR, including multiplex PCR, SDA, TMA, NASBA, etc., can be used to amplify immobilized DNA fragments. In some embodiments, primers specifically directed to polynucleotides of interest are included in the amplification reaction.

[0133] Other suitable methods for amplifying polynucleotides can include oligonucleotide extension and ligation, rolling circle amplification (RCA) (Lizardi et al., Nat. Genet. 19:225-232 (1998)), and oligonucleotide ligation assay (OLA) techniques (see generally U.S. Pat. Nos. 7,582,420, 5,185,243, 5,679,524, and 5,573,907; European Patent Nos. 0 320 308 (B1); 0 336 731 (B1); 0 439 182 (B1); WO 90 / 01069; WO 89 / 12696; and WO 89 / 09835). It will be understood that these amplification methods can be designed to amplify immobilized DNA fragments. For example, in some embodiments, the amplification method can include a ligation probe amplification or oligonucleotide ligation assay (OLA) reaction containing primers specifically directed to the nucleic acid of interest. In some embodiments, the amplification method can include a primer extension ligation reaction containing primers specifically directed to the nucleic acid of interest. Non-limiting examples of primer extension and ligation primers that can be specifically designed to amplify the nucleic acid of interest include the primers used in the GoldenGate assay (Illumina, Inc., San Diego, CA), as exemplified by U.S. Patent Nos. 7,582,420 and 7,611,869.

[0134] DNA nanoballs can also be used in combination with the methods, systems, and compositions described herein.Methods for making and using DNA nanoblocks for genome sequencing can be found, for example, in U.S. Patents and Publications such as U.S. Patent No. 7,910,354, U.S. Patent Application Publication Nos. 2009 / 0264299, 2009 / 0011943, 2009 / 0005252, 2009 / 0155781, and 2009 / 0118488, and can be found, for example, in Drmanac et al. (2010, Science 327(5961):78-81). Briefly, after generation of the asymmetric target nucleic acid, the asymmetric target nucleic acid is circularized and amplified by rolling circle amplification (Lizardi et al., 1998. Nat. Genet. 19:225-232; U.S. Patent Application Publication No. 2007 / 0099208(A1)). The extended chain-like structure of the amplicon promotes coiling, thereby creating compact DNA nanoballs. The DNA nanoballs can be captured on a substrate, preferably forming an ordered or patterned array such that the distance between each nanoball is maintained, thereby enabling sequencing of individual DNA nanoballs. In some embodiments, such as those used by Complete Genomics (Mountain View, Calif.), sequential rounds of adapter addition, amplification, and digestion are performed prior to circularization to generate head-to-tail constructs with several target nucleic acids separated by adapter sequences.

[0135] Exemplary isothermal amplification methods that can be used in the methods of the present disclosure include, but are not limited to, Multiple Displacement Amplification (MDA), exemplified by, for example, Dean et al., Proc. Natl. Acad. Sci. USA 99:5261-66 (2002), or isothermal strand displacement nucleic acid amplification, exemplified by, for example, U.S. Pat. No. 6,214,587. Other non-PCR-based methods that can be used in the present disclosure include strand displacement amplification (SDA), as described, for example, in Walker et al., Molecular Methods for Virus Detection, Academic Press Inc., 1995; U.S. Patent Nos. 5,455,166 and 5,130,238; and Walker et al., Nucl. Acids Res. 20:1691-96 (1992); or hyperbranched strand displacement amplification, as described, for example, in Lage et al., Genome Res. 13:294-307 (2003). Isothermal amplification methods can be used, for example, with strand-displacing Phi29 polymerase or Bst DNA polymerase large fragment, 5'->3' exo for random-primed amplification of genomic DNA. The use of these polymerases takes advantage of their high processivity and strand displacement activity. Due to their high processivity, the polymerases can generate fragments of 10-20 kb in length. As noted above, polymerases with low processivity and polymerases with strand displacement activity, such as Klenow polymerase, can be used to generate smaller fragments under isothermal conditions. Further description of amplification reactions, conditions, and components is provided in detail in the disclosure of U.S. Patent No. 7,670,810.

[0136] In some embodiments, isothermal amplification can be performed using kinetic exclusion amplification (KEA), also known as exclusion amplification (ExAmp). The nucleic acid libraries of the present disclosure can be generated using a method comprising reacting amplification reagents to generate a plurality of amplification sites, each containing a substantially clonal population of amplicons, from individual target nucleic acids seeded at the sites. In some embodiments, the amplification reaction proceeds until a sufficient number of amplicons are generated to fill each amplification site to capacity. In this manner, filling a previously seeded site to capacity inhibits target nucleic acids from landing and amplifying at that site, thereby generating a clonal population of amplicons at that site. In some embodiments, apparent clonality can be achieved even if the amplification site is not filled to capacity before a second target nucleic acid arrives at the site. Under some conditions, amplification of a first target nucleic acid can proceed to a point where a sufficient number of copies are generated to effectively exceed or overwhelm the generation of copies from a second target nucleic acid transported to that site. For example, in an embodiment using a bridge amplification process on circular features less than 500 nm in diameter, it was determined that after 14 cycles of exponential amplification for a first target nucleic acid, contamination from a second target nucleic acid at the same site produces an insufficient number of contaminating amplicons to adversely affect sequence synthesis analysis on an Illumina sequencing platform.

[0137] In some embodiments, amplification sites in an array can be, but are not necessarily, completely clonal. Rather, in some applications, individual amplification sites may be populated primarily with amplicons from a first asymmetric target nucleic acid and may also have low levels of contaminating amplicons from a second asymmetric target nucleic acid. An array can have one or more amplification sites with low levels of contaminating amplicons, as long as the level of contamination does not have an unacceptable effect on subsequent use of the array. For example, if an array is used in a detection application, an acceptable level of contamination is one that does not unacceptably affect the signal-to-noise ratio or resolution of the detection technique. Thus, apparent clonality generally relates to the particular use or application of an array produced by the methods described herein. Exemplary levels of contamination that are acceptable in individual amplification sites for a particular application include, but are not limited to, up to 0.1%, 0.5%, 1%, 5%, 10%, or 25% contaminating amplicons. An array can include one or more amplification sites with these exemplary levels of contaminating amplicons. For example, up to 5%, 10%, 25%, 50%, 75%, or even 100% of the amplification sites within an array may contain contaminating amplicons. It will be understood that in an array or other collection of sites, at least 50%, 75%, 80%, 85%, 90%, 95%, or 99% or more of the sites may be clonal or appear clonal.

[0138] In some embodiments, binding equilibrium exclusion can occur when a process occurs at a rate fast enough to effectively exclude another event or process from occurring. Take the example of creating a nucleic acid array in which array sites are randomly seeded with asymmetric target nucleic acids from a solution, and copies of the asymmetric target nucleic acids are generated in an amplification process to fill each seeded site to capacity. According to the binding equilibrium exclusion method of the present disclosure, the seeding and amplification processes can proceed simultaneously under conditions in which the amplification rate exceeds the seeding rate. Thus, the relatively fast rate at which copies are generated at a site seeded by a first target nucleic acid effectively excludes a second nucleic acid from seeding that site for amplification. The binding equilibrium exclusion method can be performed as described in detail in the disclosure of U.S. Patent Application Publication No. 2013 / 0338042.

[0139] Binding equilibrium exclusion can take advantage of the relatively slow rate at which amplification is initiated (e.g., the relatively slow rate at which the first copy of the asymmetric target nucleic acid is generated) versus the relatively fast rate at which subsequent copies of the asymmetric target nucleic acid (or of the first copy of the asymmetric target nucleic acid) are generated. In the example of the previous paragraph, binding equilibrium exclusion occurs due to the relatively fast rate at which amplification occurs to fill sites with asymmetric target nucleic acid seeded copies versus the relatively slow rate at which asymmetric target nucleic acid seeded copies are generated. In another exemplary embodiment, binding equilibrium exclusion can occur due to the delayed formation of the first copy of the asymmetric target nucleic acid seeded at a site (e.g., delayed activation or slow activation) versus the relatively fast rate at which subsequent copies are generated to fill sites. In this example, several different asymmetric target nucleic acids may be seeded at individual sites (e.g., several asymmetric target nucleic acids may be present at each site prior to amplification). However, first copy formation of any given asymmetric target nucleic acid may be activated randomly, resulting in a relatively slow average rate of first copy formation compared to the rate at which subsequent copies are generated. In this case, although an individual site may be seeded with several different asymmetric target nucleic acids, binding equilibrium exclusion allows amplification of only one of them. More specifically, when a first asymmetric target nucleic acid is activated for amplification, the site is rapidly filled to capacity with copies of that nucleic acid, thereby preventing copies of a second asymmetric target nucleic acid from being made at that site.

[0140] In one embodiment, the method simultaneously (i) transports the asymmetric target nucleic acid to the amplification site at an average transport rate and (ii) amplifies the asymmetric target nucleic acid at the amplification site at an average amplification rate, so that the average amplification rate exceeds the average transport rate (U.S. Patent No. 9,169,513). Therefore, in such an embodiment, binding equilibrium exclusion can be achieved by using a relatively slow transport rate. For example, a low concentration of the asymmetric target nucleic acid slows the average transport rate, so a sufficiently low concentration can be selected to achieve the desired average transport rate. Alternatively or additionally, a high viscosity solution and / or the presence of a molecular crowding agent in the solution can be used to reduce the transport rate. Examples of useful molecular crowding agents include, but are not limited to, polyethylene glycol (PEG), Ficoll, dextran, or polyvinyl alcohol. Exemplary molecular crowding agents and formulations are described in U.S. Patent No. 7,399,590, which is incorporated herein by reference. Another factor that can be adjusted to achieve a desired transport rate is the average size of the target nucleic acid.

[0141] Amplification reagents can include additional components that promote amplicon formation, potentially increasing the rate of amplicon formation. One example is a recombinase. Recombinases can promote amplicon formation by enabling repeated invasion / extension. More specifically, recombinases can promote the invasion of asymmetric target nucleic acids by a polymerase and the extension of primers by the polymerase, using the asymmetric target nucleic acid as a template for amplicon formation. This process can be repeated as a chain reaction, with amplicons generated from each round of invasion / extension serving as templates in subsequent rounds. Because no denaturation cycles (e.g., by heating or chemical denaturation) are required, this process can be performed more quickly than standard PCR. Therefore, recombinase-promoted amplification can be performed isothermally. It is desirable to include ATP or other nucleotides (or, in some cases, non-hydrolyzable analogs thereof) in the recombinase-promoted amplification reagent to promote amplification. A mixture of recombinase and single-stranded binding (SSB) protein is particularly useful, as SSB can further promote amplification. Representative formulations for recombinase-facilitated amplification include those sold commercially as TwistAmp kits by TwistDx Ltd. (Cambridge, UK). Useful components and reaction conditions for recombinase-facilitated amplification reagents are described in U.S. Patent Nos. 5,223,414 and 7,399,590.

[0142] Another example of a component that can be included in an amplification reagent to promote amplicon formation and, in some cases, increase the rate of amplicon formation is helicase. Helicase can promote amplicon formation by enabling a chain reaction of amplicon formation. Because no denaturation cycles (e.g., by heating or chemical denaturation) are required, this process can be performed more rapidly than standard PCR. Therefore, helicase-promoted amplification can be performed isothermally. A mixture of helicase and single-stranded binding (SSB) protein is particularly useful, as SSB can further promote amplification. Representative formulations for helicase-promoted amplification include those commercially available as IsoAmp kits from Biohelix, Inc. (Beverly, MA). Further examples of useful formulations containing helicase proteins are described in U.S. Patent Nos. 7,399,590 and 7,829,284.

[0143] Yet another example of a component that can be included in an amplification reagent to facilitate amplicon formation, and in some cases increase the rate of amplicon formation, is an origin binding protein.

[0144] Sequencing methods Following the binding of the asymmetric target nucleic acid to the surface, the sequence of the immobilized and amplified asymmetric target nucleic acid is determined.Sequencing can be carried out using any suitable sequencing technology, and the method for determining the sequence of the immobilized and amplified asymmetric modified target nucleic acid, including strand resynthesis, is known in the art, and is described in, for example, Bignell et al. (U.S. Patent No. 8,053,192), Gunderson et al. (WO 2016 / 130704), Shen et al. (U.S. Patent No. 8,895,249) and Pipenburg et al. (U.S. Patent No. 9,309,502).

[0145] The methods described herein can be used in conjunction with various nucleic acid sequencing methods. Particularly applicable techniques are those in which nucleic acids are attached to fixed positions within an array so that their relative positions do not change, and the array is repeatedly imaged. For example, embodiments in which images are obtained in different color channels corresponding to different labels used to distinguish one nucleotide base type from another are particularly applicable. In some embodiments, the process of determining the nucleotide sequence of an asymmetric target nucleic acid can be an automated process. Preferred embodiments include sequencing-by-synthesis (SBS) techniques.

[0146] SBS technology generally involves the enzymatic extension of nascent nucleic acid chain by repeatedly adding nucleotide to template chain.In the traditional method of SBS, a single nucleotide monomer can be provided to target nucleic acid in the presence of polymerase in each delivery.However, in the method described herein, multiple types of nucleotide monomers can be provided to target nucleic acid in the presence of polymerase during delivery.

[0147] In one embodiment, the nucleotide monomer comprises a locked nucleic acid (LNA) or a bridged nucleic acid (BNA). The use of an LNA or BNA in the nucleotide monomer increases the hybridization strength between the nucleotide monomer and a sequencing primer sequence present on the immobilized asymmetrically modified target nucleic acid.

[0148] SBS can use nucleotide monomers with or without terminator moieties. Methods using nucleotide monomers without terminators include, for example, pyrosequencing and sequencing using γ-phosphate-labeled nucleotides, as described in more detail herein. In methods using nucleotide monomers without terminators, the number of nucleotides added in each cycle is generally variable and depends on the template sequence and the mode of nucleotide delivery. In SBS techniques using nucleotide monomers with terminator moieties, the terminators can be effectively irreversible under the sequencing conditions used, as in traditional Sanger sequencing using dideoxynucleotides, or they can be reversible, as in the sequencing method developed by Solexa (now Illumina, Inc.).

[0149] SBS techniques can use nucleotide monomers that have a label moiety or lack a label moiety. Therefore, incorporation events can be detected based on the properties of the label, such as the fluorescence of the label, the properties of the nucleotide monomer, such as molecular weight or charge, or by-products of nucleotide incorporation, such as the release of pyrophosphate. In embodiments in which two or more different nucleotides are present in the sequencing reagent, the different nucleotides may be distinguishable from one another, or the two or more different labels may be distinguishable under the detection technique used. For example, different nucleotides present in the sequencing reagent can have different labels, which can be distinguished using appropriate optical systems, as exemplified by the sequencing method developed by Solexa (now Illumina).

[0150] A preferred embodiment is pyrosequencing technology, which detects the release of inorganic pyrophosphate (PPi) when a specific nucleotide is incorporated into a nascent strand (Ronaghi, M., Karamohamed, S., Pettersson, B., Uhlen, M., and Nyren, P. (1996) "Real-time DNA sequencing using detection of pyrophosphate release." Analytical Biochemistry 242(1), 84-9; Ronaghi, M. (2001) "Pyrosequencing sheds light on DNA sequencing." Genome Res. 11(1), 3-11; Ronaghi, M., Uhlen, M., and Nyren, P. (1998) "A sequencing method based on real-time pyrophosphate." Science 281 (5375), 363, U.S. Patent Nos. 6,210,891, 6,258,568, and 6,274,320). In pyrosequencing, released PPi can be detected by its immediate conversion to adenosine triphosphate (ATP) by ATP sulfurase, and the level of generated ATP is detected via luciferase-generated photons. The nucleic acid to be sequenced can be attached to features in an array, and the array can be imaged to capture chemiluminescent signals generated by nucleotide incorporation into the features of the array. Images can be obtained after treating the array with a specific nucleotide type (e.g., T, C, or G). Images obtained after the addition of each nucleotide type differ in terms of which features in the array are detected. These differences in the images reflect the different sequence content of the features on the array. However, the relative positions of each feature remain unchanged in the image. Images can be stored, processed, and analyzed using the methods described herein.For example, images obtained after treating the array with each different nucleotide type can be processed in the same manner as exemplified herein for images obtained from different detection channels for reversible terminator-based sequencing methods.

[0151] In another exemplary type of SBS, cycle sequencing is achieved by stepwise addition of reversible terminator nucleotides containing cleavable or photobleachable dye labels, as described, for example, in WO 04 / 018497 and U.S. Pat. No. 7,057,026. This approach has been commercialized by Solexa (now Illumina) and is also described in WO 91 / 06678 and WO 07 / 123,744. The availability of fluorescently labeled terminators, both of whose termini can be reversed and from which the fluorescent labels are cleaved, facilitates efficient cyclic reversible termination (CRT) sequencing. Polymerases can also be co-engineered to efficiently incorporate and extend from these modified nucleotides.

[0152] In some reversible terminator-based sequencing embodiments, the label does not substantially inhibit extension under SBS reaction conditions. However, the detection label may be removable, for example, by cleavage or degradation. Images can be captured after incorporation of the label into arrayed nucleic acid features. In certain embodiments, each cycle involves simultaneous delivery of four different nucleotide types to the array, with each nucleotide type bearing a spectrally distinct label. Four images can then be obtained, each using a detection channel selective for one of the four different labels. Alternatively, different nucleotide types can be added sequentially, with images of the array being obtained between each addition step. In such embodiments, each image shows nucleic acid features incorporating a particular type of nucleotide. Because the sequence content of each feature varies, different features are present or absent in different images. However, the relative positions of the features remain unchanged within the image. Images obtained from such reversible terminator-SBS methods can be stored, processed, and analyzed as described herein. Following the image capture step, the label can be removed, and the reversible terminator moiety can be removed for subsequent cycles of nucleotide addition and detection. Removing the label after detection in a particular cycle and before the subsequent cycle has the advantage of reducing background signal and crosstalk between cycles. Examples of useful labeling and removal methods are described herein.

[0153] In certain embodiments, some or all of the nucleotide monomers may contain reversible terminators. In such embodiments, the reversible terminator / cleavable fluorophore may comprise a fluorophore attached to the ribose moiety via a 3' ester bond (Metzker, Genome Res. 15:1767-1776 (2005)). Another approach has separated the terminator chemistry from the fluorescent label (Ruparel et al., Proc Natl Acad Sci USA 102:5932-7 (2005)). Ruparel et al. describe the development of a reversible terminator that uses a small 3' allyl group to block extension but can be easily unblocked by brief treatment with a palladium catalyst. The fluorophore was attached to the group via a photocleavable linker that can be easily cleaved by 30 seconds of exposure to long-wavelength UV light. Thus, either disulfide reduction or photocleavage can be used as the cleavable linker. Another approach to reversible termination is the use of a natural terminator following the placement of a bulky dye on the dNTP. The presence of a charged bulky dye on the dNTP can act as an effective terminator through steric and / or electrostatic hindrance. The presence of one incorporation event prevents further binding unless the dye is removed. Cleavage of the dye removes the fluorophore, effectively reversing the terminus. Examples of modified nucleotides are also described in U.S. Patent Nos. 7,427,673 and 7,057,026.

[0154] Additional exemplary SBS systems and methods that can be utilized with the methods and systems described herein are described in U.S. Patent Application Publication Nos. 2007 / 0166705, 2006 / 0188901, 2006 / 0240439, 2006 / 0281109, 2012 / 0270305, and 2013 / 0260372, U.S. Patent No. 7,057,026, and WO 05 / 065814, U.S. Patent Application Publication No. 2005 / 0100900, and WO 06 / 064199 and WO 07 / 010,251.

[0155] Some embodiments may employ detection of four different nucleotides using fewer than four different labels. For example, SBS may be performed using the methods and systems described in the incorporated document, U.S. Patent Publication No. 2013 / 0079232. As a first example, pairs of nucleotide types may be detected at the same wavelength but may be distinguished based on differences in the intensity of one member of the pair compared to the other, or based on changes to one member of the pair (e.g., via chemical, photochemical, or physical modification) that result in the appearance or disappearance of a distinct signal compared to the signal detected for the other member of the pair. As a second example, three of the four different nucleotide types may be detected under certain conditions, while the fourth nucleotide type may have no detectable label under those conditions or be minimally detected under those conditions (e.g., minimal detection due to background fluorescence, etc.). Incorporation of the first three nucleotide types into a nucleic acid may be determined based on the presence of their corresponding signals, and incorporation of the fourth nucleotide type into a nucleic acid may be determined based on the absence or minimal detection of any signal. As a third example, one nucleotide type can include a label that is detected in two different channels, while the other nucleotide type is detected in no more than one of the channels. The three exemplary configurations above are not considered mutually exclusive and can be used in various combinations.An exemplary embodiment combining all three examples is a fluorescence-based SBS method that uses a first nucleotide type that is detected in a first channel (e.g., dATP having a label that is detected in the first channel when excited by a first excitation wavelength), a second nucleotide type that is detected in a second channel (e.g., dCTP having a label that is detected in the second channel when excited by a second excitation wavelength), a third nucleotide type that is detected in both the first and second channels (e.g., dTTP having at least one label that is detected in both channels when excited by the first and / or second excitation wavelength), and an unlabeled fourth nucleotide type that is not detected or is minimally detected in either channel (e.g., unlabeled dGTP).

[0156] Furthermore, as described in U.S. Patent Application Publication No. 2013 / 0079232, sequencing data can be obtained using a single channel. In such so-called single-dye sequencing methods, a first nucleotide type is labeled but the label is removed after the first image is generated, and a second nucleotide type is labeled only after the first image is generated. A third nucleotide type retains its label in both the first and second images, and a fourth nucleotide type remains unlabeled in both images.

[0157] Some embodiments may use sequencing by ligation techniques. Such techniques use DNA ligase to incorporate oligonucleotides and identify their incorporation. The oligonucleotides typically have different labels that correlate with the identity of specific nucleotides in the sequence to which the oligonucleotides hybridize. As with other SBS methods, images can be obtained after treating an array of nucleic acid sequences with labeled sequencing reagents. Each image shows nucleic acid features that incorporate a specific type of label. Because the sequence content of each feature varies, different features may or may not be present in different images, but the relative positions of the features remain constant within the image. Images obtained from ligation-based sequencing methods can be stored, processed, and analyzed as described herein. Exemplary SBS systems and methods that can be utilized with the methods and systems described herein are described in U.S. Patent Nos. 6,969,488, 6,172,218, and 6,306,597.

[0158] Some embodiments can use nanopore sequencing (Deamer, DW & Akeson, M. "Nanopores and nucleic acids: prospects for ultrarapid sequencing." Trends Biotechnol. 18, 147-151 (2000); Deamer, D. and D. Branton, "Characterization of nucleic acids by nanopore analysis." Acc. Chem. Res. 35:817-825 (2002); Li, J., M. Gershow, D. Stein, E. Brandin, and J. A. Golovchenko, "DNA molecules and configurations in a solid-state nanopore microscope." Nat. Mater. 2:611-615 (2003)). In such embodiments, an asymmetric target nucleic acid passes through a nanopore. The nanopore can be a synthetic pore or a biological membrane protein, such as α-hemolysin. As an asymmetric target nucleic acid passes through a nanopore, each base pair can be identified by measuring the fluctuations in the pore's electrical conductance. (U.S. Pat. No. 7,001,792; Soni, GV & Meller, "A. Progress toward ultrafast DNA sequencing using solid-state nanopores." Clin. Chem. 53, 1996-2001 (2007); Healy, K. "Nanopore-based single-molecule DNA analysis." Nanomed. 2, 459-481 (2007); Cockroft, SL, Chu, J., Amorin, M. & Ghadiri, MR "A single-molecule nanopore device detects DNA polymerase activity with single-nucleotide resolution." J. Am Chem. Soc. 130, 818-820 (2008). Data obtained from nanopore sequencing can be stored, processed, and analyzed as described herein.In particular, the data can be processed as an image according to the exemplary processing of optical and other images described herein.

[0159] Some embodiments can use methods involving real-time monitoring of DNA polymerase activity. Nucleotide incorporation can be detected via fluorescence resonance energy transfer (FRET) interactions between a fluorophore-containing polymerase and a γ-phosphate-labeled nucleotide, as described, for example, in U.S. Patent Nos. 7,329,492 and 7,211,414, or nucleotide incorporation can be detected using zero-mode waveguides, as described, for example, in U.S. Patent No. 7,315,019, and fluorescent nucleotide analogs and engineered polymerases, as described, for example, in U.S. Patent No. 7,405,281 and U.S. Patent Application Publication No. 2008 / 0108082. Illumination can be restricted to a zeptoliter-scale volume around the surface-tethered polymerase so that incorporation of fluorescently labeled nucleotides can be observed with low background (Levene, MJ et al., "Zero-mode waveguides for single-molecule analysis at high concentrations." Science, 299, 682-686 (2003); Lundquist, PM et al., "Parallel confocal detection of single molecules in real time." Opt. Lett. 33, 1026-1028 (2008); Korlach, J. et al., "Selective aluminum passivation for targeted immobilization of single DNA polymerase molecules in zero-mode waveguide nanostructures." Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008)). Images obtained from such methods can be stored, processed, and analyzed as described herein.

[0160] Some SBS embodiments involve the detection of protons released upon incorporation of a nucleotide into an extension product. For example, sequencing based on the detection of released protons can use commercially available electrical detectors and related technology from Ion Torrent (Guilford, CT, a Life Technologies subsidiary), or the sequencing methods and systems described in U.S. Patent Application Publication Nos. 2009 / 0026082, 2009 / 0127589, 2010 / 0137143, and 2010 / 0282617. The methods described herein for amplifying target nucleic acids using equilibrium exclusion can be easily adapted to substrates used to detect protons. More specifically, the methods described herein can be used to generate clonal populations of amplicons used to detect protons.

[0161] The SBS method described above can be advantageously performed in a multiplex format, allowing multiple different asymmetric target nucleic acids to be manipulated simultaneously. In certain embodiments, the different asymmetric target nucleic acids can be processed in a common reaction vessel or on the surface of a specific substrate. This allows for convenient delivery of sequencing reagents, removal of unreacted reagents, and multiplexed detection of incorporation events. In embodiments using surface-bound target nucleic acids, the asymmetric target nucleic acids can be in an array format. In an array format, the asymmetric target nucleic acids can typically be bound to a surface in a spatially distinguishable manner. The asymmetric target nucleic acids can be bound by direct covalent binding, attachment to beads or other particles, or binding to a polymerase or other molecule attached to a surface. Arrays can contain a single copy of the asymmetric target nucleic acid at each site (also called a feature), or multiple copies with the same sequence can be present at each site or feature. Multiple copies can be generated by amplification methods such as bridge amplification or emulsion PCR, as described in further detail herein.

[0162] The methods described herein can be used to fabricate, for example, at least about 10 features / cm 2 , 100 features / cm 2, 500 features / cm 2 , 1,000 features / cm 2 , 5,000 features / cm 2 , 10,000 features / cm 2 , 50,000 features / cm 2 , 100,000 features / cm 2 , 1,000,000 features / cm 2 , 5,000,000 features / cm 2 Arrays having features of any of a variety of densities, including 1000 sq. ft., ... or more, can be used.

[0163] The advantage of the method described herein is that it allows for multiple cm 2The present disclosure provides an integrated system capable of preparing and detecting nucleic acids using techniques known in the art, such as those exemplified herein. Accordingly, the integrated system of the present disclosure can include fluidic components capable of delivering amplification and / or sequencing reagents to one or more immobilized asymmetric target nucleic acids, including components such as pumps, valves, reservoirs, and fluid lines. A flow cell can be configured and / or used in the integrated system for detecting target nucleic acids. Exemplary flow cells are described, for example, in U.S. Pat. Nos. 8,241,573 and 8,951,781. As exemplified for the flow cell, one or more of the fluidic components of the integrated system can be used for amplification and detection methods. Taking the nucleic acid sequencing embodiment as an example, one or more of the fluidic components of the integrated system can be used for delivering sequencing reagents in the amplification methods described herein and the sequencing methods exemplified above. Alternatively, the integrated system can include separate fluidic systems for performing the amplification method and the detection method. Examples of integrated sequencing systems that can generate amplified nucleic acids and determine the sequence of the nucleic acids include, but are not limited to, the MiSeq™ platform (Illumina, Inc., San Diego, CA) and the device described in U.S. Pat. No. 8,951,781.

[0164] composition During the implementation of the methods provided by the present disclosure, several compositions may be produced. For example, a composition may be produced that includes a transposome complex and a damage-intolerant DNA polymerase. The transposome may include a transposase bound to a transposon sequence that includes an adapter. The adapter may include one or more DNA lesions, one or more universal sequences, one or more index sequences, one or more UMIs, or a combination thereof. The composition may further include a target nucleic acid. Optionally, the composition may include a damage-tolerant DNA polymerase.

[0165] In another embodiment, the composition may comprise a plurality of single-stranded modified target nucleic acids, primers, and a damage-intolerant DNA polymerase. For example, the target nucleic acid may comprise, from 5' to 3', a first adaptor, the target nucleic acid, and the complement of the first adaptor. The first adaptor may comprise one or more DNA damages, one or more universal sequences, one or more index sequences, one or more UMIs, or a combination thereof. In one embodiment, the universal sequence may comprise a transposase recognition site. The primer may comprise, from 5' to 3', a second adaptor and a nucleotide sequence that anneals to the complement of the first adaptor. The second adaptor may comprise one or more universal sequences, one or more index sequences, one or more UMIs, or a combination thereof. The primer may optionally comprise a blocked 3' end and may optionally comprise at least one modified nucleotide. In one embodiment, the primer anneals to the single-stranded modified target nucleic acid.

[0166] In another embodiment, the composition comprises a transposome complex. The transposome complex includes, but is not limited to, a transposase and a transposon. In one embodiment, the transposon comprises an adapter. The adapter can include, for example, a first strand having, from 5' to 3', at least one universal sequence, at least one index sequence, at least one UMI, or a combination thereof, a DNA lesion, and a transposase recognition sequence. In one embodiment, the transposase recognition sequence includes a mosaic element. The adapter can include, for example, a second strand having nucleotides complementary to at least a portion of the transposase recognition sequence. In one embodiment, the first strand also includes a capture agent at its 5' end, or the second strand also includes a capture agent at its 3' end. In one embodiment, a cleavable linker is positioned between the capture agent and the 5' end of the first strand. In one embodiment, the cleavable linker is positioned between the capture agent and the 3' end of the second strand. In one embodiment, the composition further includes a solid surface, and the transposase complex is attached to the solid surface. In another embodiment, the composition further comprises a solid surface, wherein the transposon is not associated with a transposase, and the transposon is bound to the solid surface.

[0167] kit The present disclosure also provides a kit for carrying out one or more aspects of the methods provided herein. The kit can be used to generate a library of target nucleic acids. In one embodiment, the kit can be used to generate a library of symmetric target nucleic acids. The kit can include a transposome complex and a damage-intolerant DNA polymerase in separate containers. The transposome can include a transposase bound to a transposon sequence, and the transposon sequence includes an adapter and DNA damage. In one embodiment, the kit can be used to convert a symmetric library into an asymmetric library. In this embodiment, the kit can further include a primer. In one embodiment, the primer includes a nucleotide sequence that anneals 5' to 3' to the second adapter and the complement of the first adapter.

[0168] The components of the kit may be present in suitable packaging materials in amounts sufficient to generate at least one library. Optionally, other reagents, such as buffers (either prepared or present in their components, one or more of the components may be premixed, or all of the components may be separate), are also included. Instructions for use of the packaged components are also typically included.

[0169] As used herein, the phrase "packaging material" refers to one or more physical structures used to house the contents of the kit. The packaging material is preferably constructed by known methods to provide a sterile, contaminant-free environment. The packaging material may have a label indicating that the components can be used to generate a sequencing library. In addition, the packaging material includes instructions indicating how the materials in the kit are to be used to practice one or more aspects of the methods provided herein. As used herein, the term "package" refers to a solid matrix or material, such as glass, plastic, paper, or foil, capable of holding one or more components of the kit within certain limits. "Instructions for use" typically include specific language describing at least one assay parameter, such as reagent concentrations or the relative amounts of reagents and sample to be mixed, the duration of the reagent / sample mixture, temperature, buffer conditions, etc.

[0170] The present invention is defined in the claims. However, the following provides a non-exhaustive list of non-limiting exemplary aspects. Any one or more of the features of these aspects may be combined with any one or more features of any other example, embodiment, or aspect described herein.

[0171] Exemplary Embodiments Embodiment 1 is a method for generating a sequencing library, comprising: providing a plurality of symmetrically modified target nucleic acids comprising a first adaptor sequence at each end, wherein the first adaptor sequence comprises a DNA lesion; extending the modified target nucleic acid with a damage-intolerant polymerase to generate a plurality of asymmetrically modified target nucleic acids comprising a first adapter sequence at the 5' end of each strand and a complement of a portion of the first adapter at the 3' end of each strand.

[0172] Embodiment 2 is the method of embodiment 1, wherein the plurality of symmetrically modified target nucleic acids are double-stranded, each strand comprising a first adapter sequence comprising a 5' to 3' DNA damage, the target nucleic acid, a gap comprising at least one nucleotide, and the complement of the first adapter sequence without the DNA damage.

[0173] Aspect 3 is the method of aspect 1 or 2, wherein the extension starts at a gap.

[0174] Aspect 4 is annealing a primer to a plurality of asymmetrically modified target nucleic acids, the primer comprising a second adapter sequence from 5' to 3' and an annealing domain, the annealing domain comprising a nucleotide sequence that anneals to a complement of a portion of the first adapter of the plurality of asymmetrically modified target nucleic acids; The method of any of aspects 2 or 3, further comprising extending the 3' end of the annealed asymmetrically modified target nucleic acid with a damage-intolerant polymerase, wherein the extension results in a plurality of asymmetrically modified target nucleic acids comprising, from 5' to 3', (i) the first adaptor, (ii) the target nucleic acid, (iii) a complement of a portion of the first adaptor, and (iv) a complement of the second adaptor.

[0175] Aspect 5 is the method according to any one of aspects 1 to 4, wherein the extension of the 3' end of the annealed asymmetrically modified target nucleic acid is repeated at least three times.

[0176] Aspect 6 is the method of any one of Aspects 1 to 5, wherein the DNA damage comprises at least one of an abasic site, a modified base, a mismatch, a single-strand break, or a crosslinked nucleotide.

[0177] Embodiment 7 is the method of any one of embodiments 1 to 6, wherein the DNA damage comprises at least one uracil.

[0178] Embodiment 8 is the method of any one of embodiments 1 to 7, wherein the annealing domain of the primer comprises at least one modified nucleotide that increases the melting temperature compared to the corresponding native DNA nucleotide.

[0179] Aspect 9 is the method of any one of aspects 1 to 8, wherein the modified nucleotide comprises a locked nucleic acid, a PNA, or an RNA.

[0180] Aspect 10 is the method according to any one of aspects 1 to 9, wherein the 3' end of the primer is blocked.

[0181] Example 11 is the method of any one of Examples 1 to 10, wherein the first adaptor comprises one or more universal sequences, one or more index sequences, one or more universal molecular identifiers, or a combination thereof.

[0182] Example 12 is the method of any one of Examples 1 to 11, wherein at least one of the one or more universal sequences, the one or more index sequences, and the one or more universal molecular identifiers is located in an adaptor between the DNA lesion and the end of the adaptor distal to the target nucleic acid.

[0183] Example 13 is the method of any one of Examples 1 to 12, wherein the second adaptor comprises one or more universal sequences, one or more index sequences, one or more universal molecular identifiers, or a combination thereof.

[0184] Example 14 is the method of any one of Examples 1 to 13, wherein the one or more universal sequences, one or more index sequences, and one or more universal molecular identifiers of the first adaptor are unique compared to the one or more universal sequences, one or more index sequences, and one or more universal molecular identifiers of the second adaptor.

[0185] Example 15 is the method of any one of Examples 1 to 14, wherein one or more index sequences of the first adaptor are compartment-specific.

[0186] Example 16 is the method of any one of Examples 1 to 15, wherein one or more index sequences of the second adaptor are compartment-specific.

[0187] Embodiment 17 The method of any one of embodiments 1 to 16, wherein the first adapter comprises a transposase recognition site.

[0188] Example 18 is the method of any one of Examples 1 to 17, wherein the target nucleic acid is derived from nucleic acid derived from a single cell.

[0189] Aspect 19 is the method of any one of Aspects 1 to 18, wherein the target nucleic acid is derived from nucleic acids derived from a plurality of cells.

[0190] Aspect 20 is the method of any one of Aspects 1 to 19, wherein the target nucleic acid from a single cell or multiple cells comprises RNA.

[0191] Aspect 21 is the method according to any one of Aspects 1 to 20, wherein the RNA comprises mRNA.

[0192] Aspect 22 is the method of any one of Aspects 1 to 21, wherein the target nucleic acid from a single cell or multiple cells comprises DNA.

[0193] Example 23 is the method of any one of Examples 1 to 22, wherein the DNA comprises whole cell genomic DNA.

[0194] Example 24 is the method of any one of Examples 1 to 23, wherein the whole cell genomic DNA comprises nucleosomes.

[0195] Example 25 is the method of any one of Examples 1 to 247, wherein the target nucleic acid is derived from nucleic acid derived from cell-free DNA.

[0196] Embodiment 26 is the method of any one of embodiments 1 to 25, wherein the method comprises combinatorial indexing.

[0197] Example 27 is the method of any one of Examples 1 to 26, further comprising amplifying the asymmetrically modified target nucleic acid, wherein the amplification comprises a second primer and a damage-tolerant polymerase, and wherein the second primer comprises a nucleotide sequence that anneals to the first adapter sequence or its complement.

[0198] Example 28 is the method of any one of Examples 1 to 27, wherein the second primer further comprises one or more universal sequences, one or more index sequences, one or more universal molecular identifiers, or a combination thereof.

[0199] Example 29 is the method of any one of Examples 1 to 28, wherein the one or more universal sequences, one or more index sequences, and one or more universal molecular identifiers of the second primer are unique compared to the one or more universal sequences, one or more index sequences, and one or more universal molecular identifiers of the first adaptor and the second adaptor.

[0200] Example 30 is the method of any one of Examples 1 to 29, wherein a subset of the plurality of asymmetrically modified target nucleic acids is present in a plurality of compartments, and either (i) the first adaptor comprises a first compartment-specific index, (ii) the second adaptor comprises a second compartment-specific index, or both (i) and (ii).

[0201] Embodiment 31 The method of any one of embodiments 1 to 30, further comprising combining asymmetrically modified target nucleic acids from different compartments to generate pooled indexed asymmetrically modified target nucleic acids.

[0202] Aspect 32 is 32. The method of any one of aspects 1-31, further comprising distributing a subset of the pooled indexed asymmetrically modified target nucleic acids into a second plurality of compartments and modifying the indexed asymmetrically modified target nucleic acids, wherein the modification comprises adding an additional compartment-specific index sequence to the indexed asymmetrically modified target nucleic acids present in each subset to result in indexed DNA nucleic acids, and wherein the modification comprises ligation or extension.

[0203] Example 33 is the method of any one of Examples 1 to 32, wherein the compartment comprises a well or a droplet.

[0204] Example 34 is the method of any one of Examples 1 to 33, wherein providing comprises contacting the plurality of DNA fragments with first adaptors under conditions to ligate the first adaptors to both ends of the DNA fragments.

[0205] Example 35 is the method according to any one of Examples 1 to 34, wherein the DNA fragment is double-stranded and blunt-ended.

[0206] Example 36 is the method according to any one of Examples 1 to 35, wherein the first adaptor is a double-stranded DNA oligonucleotide.

[0207] Embodiment 37 is the method of any one of embodiments 1 to 36, wherein one 3' end of the first extension oligonucleotide is blocked.

[0208] Example 38 is the method of any one of Examples 1 to 37, wherein the DNA fragment is double-stranded and comprises a single-stranded region at one or both 3' ends.

[0209] Example 39 is the method of any one of Examples 1 to 38, wherein the first adaptor is a double-stranded DNA oligonucleotide comprising a single-stranded region at one end, the single-stranded region being capable of annealing to a single-stranded region present on the DNA fragment.

[0210] Example 40 is the method of any one of Examples 1 to 38, wherein the adaptor is a forked adaptor.

[0211] Embodiment 41 is the method of any one of embodiments 1 to 40, wherein providing comprises contacting the DNA with a transposome complex, the transposome complex comprising a transposase and a first adapter, and wherein the contacting occurs under conditions suitable for ligation of the first adapter to the DNA to generate a symmetrically modified target nucleic acid. In one embodiment, the transposome complex is the transposome complex of any one of embodiments 67 to 71.

[0212] Example 42 is the method of any one of Examples 1 to 41, wherein the symmetrically modified target nucleic acid produced comprises a gap of at least one nucleotide in one strand between the ligated first adaptor and the target nucleic acid.

[0213] Example 43 is the method of any one of Examples 1 to 42, wherein the DNA is present in multiple compartments, and the first adaptor in each compartment comprises a compartment-specific index.

[0214] Example 44 is the method of any one of Examples 1 to 43, further comprising combining single-stranded modified target nucleic acids from different compartments to generate pooled symmetrically modified target nucleic acids, and distributing the symmetrically modified target nucleic acids into a second plurality of compartments.

[0215] Example 45 is the method of any one of Examples 1 to 44, wherein the method further comprises fragmenting the whole cell genomic DNA.

[0216] Example 46 is the method of any one of Examples 1 to 45, wherein the fragmenting comprises digestion of the total cellular genomic DNA with a restriction endonuclease.

[0217] Example 47 is the method of any one of Examples 1 to 46, wherein the fragmented DNA is subjected to proximity ligation to attach chimeric target nucleic acids.

[0218] Example 48 is the method of any one of Examples 1 to 47, wherein cytosine residues of the adapter are replaced with 5-methylcytosine.

[0219] Example 49 is the method of any one of Examples 1 to 48, wherein the symmetric or asymmetric target nucleic acid is subjected to a chemical or enzymatic methylation conversion.

[0220] Example 50 is the method of any one of Examples 1 to 49, wherein providing comprises fixing isolated nuclei, subjecting the isolated nuclei to conditions that dissociate nucleosomes from genomic DNA, fragmenting the genomic DNA, subjecting the fragments to proximity ligation that joins a chimeric target nucleic acid, and contacting the ligated fragments with a transposome complex, wherein the transposome complex comprises a transposase and a first adaptor, and wherein the contacting occurs under conditions suitable for ligation of the first adaptor to the DNA to generate a symmetrically modified target nucleic acid.

[0221] Embodiment 51 is the method of any one of embodiments 1 to 50, wherein the fragmentation comprises digestion with a restriction endonuclease.

[0222] Aspect 52 is providing a surface comprising a plurality of amplification sites, providing an amplification site comprising at least two populations of linked single-stranded capture oligonucleotides having free 3' ends; 52. The method of any one of aspects 1 to 51, further comprising contacting the surface comprising the amplification sites with a plurality of asymmetrically modified target nucleic acids under conditions suitable to generate a plurality of amplification sites, each of the amplification sites comprising a clonal population of amplicons from an individual asymmetrically modified target nucleic acid.

[0223] Embodiment 53 is a composition comprising a transposome complex and a DNA polymerase, wherein the transposome comprises a transposase bound to a transposon sequence, the transposon sequence comprises an adapter and DNA damage, and the DNA polymerase is a damage-intolerant polymerase.

[0224] Example 54 is the composition of example 53, wherein the adaptor comprises one or more universal sequences, one or more index sequences, one or more UMIs, or a combination thereof.

[0225] Embodiment 55 is the composition of embodiment 53 or 54, further comprising a damage-tolerant DNA polymerase.

[0226] Aspect 56 is a first adaptor comprising a 5' to 3' DNA lesion, a target nucleic acid, and a plurality of modified target nucleic acids comprising a complement of the first adaptor; a primer comprising a second adaptor from 5' to 3' and an annealing domain, wherein the annealing domain comprises a nucleotide sequence that anneals to the complement of the first adaptor; and a damage-intolerant DNA polymerase.

[0227] Embodiment 57 is the composition according to embodiment 56, wherein the primer comprises at least one modified nucleotide which increases the melting temperature compared to a corresponding natural DNA nucleotide.

[0228] Embodiment 58 is the composition of embodiment 56 or 57, wherein the primer is annealed to the target nucleic acid.

[0229] Embodiment 59 is the composition of any one of embodiments 56 to 58, wherein the 3' end of the primer is blocked.

[0230] Embodiment 60 is the composition of any one of embodiments 56 to 59, wherein the first adapter comprises a transposase recognition site.

[0231] Embodiment 61 is a kit comprising a transposome complex and a DNA polymerase in separate containers, and instructions for use, wherein the transposome comprises a transposase bound to a transposon sequence, the transposon sequence comprises a first adapter and a DNA damage, and the DNA polymerase is a damage-intolerant polymerase.

[0232] Embodiment 62 is the kit of embodiment 61, further comprising a second DNA polymerase, wherein the second DNA polymerase is a damage-tolerant polymerase.

[0233] Embodiment 63 is the kit of embodiment 61 or 62, further comprising a primer, the primer comprising a second adaptor and an annealing domain from 5' to 3', the annealing domain comprising a nucleotide sequence that anneals to the complement of the first adaptor.

[0234] Embodiment 64 is the kit according to any one of embodiments 61 to 63, wherein the 3' end of the primer is blocked.

[0235] Embodiment 65 is the kit of any one of embodiments 61 to 64, wherein the first adaptor comprises one or more universal sequences, one or more index sequences, one or more UMIs, or a combination thereof.

[0236] Embodiment 66 is the kit according to any one of embodiments 61 to 65, wherein the second adapter primer further comprises one or more universal sequences, one or more index sequences, one or more UMIs, or a combination thereof.

[0237] Embodiment 67 is a transposome complex comprising a transposase and a transposon comprising a nucleic acid comprising, on a first strand, from 5' to 3', at least one universal sequence, at least one index sequence, at least one UMI, or a combination thereof, DNA damage, or a transposase recognition sequence, and, on a second strand, an adapter comprising nucleotides complementary to at least a portion of the transposase recognition sequence.

[0238] Embodiment 68 is the transposome complex of embodiment 67, wherein the first strand further comprises a capture agent at the 5' end of the first strand.

[0239] Embodiment 69 is the transposome complex of embodiment 67 or 68, wherein the first strand further comprises a cleavable linker positioned between the capture agent and the 5' end.

[0240] Embodiment 70 is the transposome complex of any one of embodiments 67 to 69, wherein the second strand further comprises a capture agent at the 3' end of the second strand.

[0241] Embodiment 71 is any one of the transposome complexes according to embodiment 70, wherein the second strand further comprises a cleavable linker positioned between the capture agent and the 3' end. [Example]

[0242] The present disclosure is illustrated by the following examples, it being understood that the particular examples, materials, amounts, and procedures are to be interpreted broadly in accordance with the scope and spirit of the present disclosure as set forth herein.

[0243] Example 1 Proof of concept for conversion of symmetric target nucleic acids into asymmetric target nucleic acid fragments.

[0244] Experimental approach for generating symmetric target nucleic acids by tagmentation and converting them to asymmetric target nucleic acids: A sequencing library was prepared by tagmentation of DNA using a transposome complex with a single transposon to generate target nucleic acids with the same adapter at each end, which were then exposed to conditions to modify one of the adapters, resulting in an asymmetric target nucleic acid.

[0245] Cell / Nuclear Protocol.

[0246] The kit includes: a 96-well indexed TSM plate, a 384-well indexed PCR plate, 5x Tagmentation Buffer TB1, ExTB (500ul in 1.7ml screw cap tubes, LNA+TX100), Post Tagmentation Wash Buffer (10ml in 15ml conical tubes), Resuspension buffer (RSB) (10ml in 15ml conical tubes), and 0.5% SDS (500ul in 1.7ml screw cap tubes).

[0247] User prepared: Q5 2x Master Mix (NEB, M0492L), Q5U 2x Master Mix (NEB, M0597L), 80% EtOH, and AMPure XP beads (Beckman Coulter, A63880).

[0248] Equipment and consumable plastics: cell counter (ThermoFisher Countess II FL Automated Cell Counter, AMQAF1000), Countess cell counting chamber slides (ThermoFisher, PN C10228), temperature-controlled plate centrifuge, temperature-controlled benchtop centrifuge, bioanalyzer (Agilent, PN G2939BA), Agilent High Sensitivity DNA Kit (5067-4626), 96-well plates (Eppendorf twin.tec PCR Plate 96 LoBind, skirted, PN 0030129512), 384-well plates (Eppendorf twin.tec PCR Plate 384 LoBind, skirted, PN0030129547), disposable reagent reservoirs (VWR, PN89094-658) or equivalent, magnetic stand for bead collection, thermal cycler for 96- and 384-well plates, plate shaker, and Falcon 15 mL collection tubes (ThermoFisher, PN14-959-53A or SARSTEDT, PN62.554.205).

[0249] Reagents for nuclei preparation: Pierce™ 16% Formaldehyde (w / v), methanol-free (ThermoFisher, PN28906), TryPLE (Fisher Scientific, PN12-604-039), PBS buffer (Sigma, PN#806552-1L), Pierce protease inhibitor mini-tablets, EDTA-free (PN A32955), and trypan blue solution (ThermoFisher, PN15250061).

[0250] Recommended buffers for cell lines: Lysis buffer is 10 mM HEPES, 10 mM NaCl, 3 mM MgCl, 0.1% Igepal, 0.1% Tween®, and protease inhibitors. NIB buffer is 10 mM Tris (pH 7.5), 10 mM NaCl, 3 mM MgCl, 0.1% Tween, and protease inhibitors. 10x Xlink buffer is 5 M NaCl, 1 M Tris HCl (pH 7.5), 1 M MgCl, and 100 ng / uL BSA.

[0251] Nuclei preparation and nucleosome depletion. Cells were cultured at 1 × 10 in a T25 flask (PN) the day before. 6 Cells were seeded in flasks at 100°C and were subconfluent at the time of harvest. Cells were washed in flasks with 5 mL of ice-cold PBS, trypsinized with TrypLE (1 mL, 5 min at 37°C), harvested by spinning at 500 rcf for 3 min at 4°C, washed with 1 mL of ice-cold PBS, and proceeded to nuclear isolation.

[0252] Nuclei isolation. Cells were spun down at 500 rcf for 3 minutes at 4°C, resuspended in 1 mL of lysis buffer, and incubated on ice for 10 minutes. Cells were spun down at 500 rcf for 3 minutes at 4°C, and resuspended in 300 uL of lysis buffer. Nuclei were counted using a 1:5 dilution (2 uL sample + 8 uL lysis buffer + 10 uL trypan blue solution). 1 x 10 6 was aliquoted for fixation.

[0253] Nuclei fixation. The volume was increased to 5 mL of lysis buffer and 246 μL of 16% formaldehyde from a freshly opened ampoule was added (total 0.75% formaldehyde; a range of 0.5% to 0.75% is acceptable). Incubate for 10 minutes at room temperature with gentle shaking, pellet by centrifugation at 500 rcf for 3 minutes at 4°C, wash with 1 mL of ice-cold NIB, spin at 500 rcf for 3 minutes at 4°C, and wash with 200 μL of 1x Xlink buffer (ice-cold). During the wash, nuclei were transferred to a 1.5 mL tube for better pelleting and spun at 500 rcf for 3 minutes at 4°C.

[0254] Nucleosome depletion (for whole genome sequencing): The pellet was resuspended in 760 μL of 1x Xlink buffer with 40 μL of 1% SDS and incubated at 37°C for 20 minutes with shaking (400 rpm). (0.05% final SDS) spun at 500 rcf for 3 minutes at 4°C, washed with 200 μL of 1x NIB, spun at 500 rcf for 3 minutes at 4°C, and resuspended in 50-100 μL of NIB. To a 2 μL sample, 8 μL of NIB and 10 μL of trypan blue were added, and 10 μL was loaded onto a cell counter. Nuclei were concentrated or diluted to 500 nuclei / μL, as needed, for pipetting out.

[0255] Protocol for plate-based combinatorial indexing workflow (Figure 10).

[0256] Tagmentation. Mix nuclei with buffer: 350 ul (approximately 100K) nuclei, 500 ul of 5x tagmentation buffer (TB1), and 1350 ul of HO. Add 20 ul to each well of a 96TSM plate and incubate at 55°C for 15 minutes on a thermal cycler. Add 100 ul of 200 mM EDTA to a 15 mL collection tube and pool the nuclei from the 96-well plate into the 15 mL collection tube on ice (total: 25 ul x 96 + 100 ul = 2.5 mL). Pellet the nuclei at 4°C and 500 rcf and suspend the nuclei in 500 ul of wash buffer. Determine the concentration of nuclei by removing a 2 uL sample, adding 8 uL of NIB and 10 uL of trypan blue, and loading 10 uL into a cell counter. Then, dilute the nuclei to nuclei / uL and load 4 uL into each well of the plate.

[0257] Extension: Add reagents in the following order: 1 ul of 0.5% SDS, heat to 55°C for 10 minutes, add 2 ul of ExTB, and 7 ul of 2x Q5 Master Mix (NEB) for a total of 14 ul. Mix wells and run the program on the thermocycler: 1. 72°C for 10 minutes, 2. 98°C for 30 seconds, 3. 98°C for 10 seconds, 4. 59°C for 20 seconds, 5. 72°C for 10 seconds, 6. Repeat steps 3-5 for a total of 10 cycles, 7. 72°C for 2 minutes, and 8. Hold temperature at 10°C.

[0258] Indexed PCR: Transfer 1 ul of PCR primer from the extension well from the 384-well PCR plate to the nuclei plate. Add 15 ul 2x NEB Q5U and run the PCR program on a thermal cycler: 1. 98°C for 30 seconds, 2. 98°C for 10 seconds, 3. 55°C for 20 seconds, 4. 72°C for 30 seconds, 5. Repeat steps 2-4 for a total of 20 cycles, 6. 72°C for 2 minutes, and 7. 10°C hold temperature. Libraries are typically amplified between 12-14 cycles.

[0259] Library cleanup: Pool 10ul per well, total 3840ul into a 15ml collection tube (PN), concentrate through a Qiagen PCR cleanup column (PN), elute into 50ul, add 50ul Ampure XP beads, wash twice with 100ul 80% EtOH, elute into 20ul RSB, and quantify with the Bioanalyzer DNA HS kit (PN).

[0260] AA → AB (symmetric vs. asymmetric) protocol for genomic DNA (gDNA).

[0261] Tnp assembly: Add 5ul 10x annealing buffer, 5ul SBS12-U-ME (Mosaic Elements) 100uM, 5ul ME' 100uM, and 35ul H2O for a total of 50ul. Run on a thermocycler: 95°C for 1 minute, 80°C for 30 seconds, ramp down to 20°C by 1°C per cycle, 20°C for 1 hour, 10°C hold temperature.

[0262] TSM assembly: Add 79ul SDB buffer, 1ul Tn5 200uM, and 20ul Tnp for a total of 100ul from Tnp assembly. Incubate overnight at 37°C and dilute 4x in SDB buffer to 500uM TSM.

[0263] Tagmentation of gDNA: Add 4 ul of 20 ng of gDNA, 5 ul of 2x TD buffer (tagmentation buffer), and 1 ul of TSM from the TSM assembly for a total of 10 ul and incubate at 55°C for 10 minutes.

[0264] AA to AB conversion: Add 1 ul of 1% SDS, incubate at 55°C for 10 minutes, add 2 ul of 10% Triton®-X100 mixed with 1 uM LNA-ME_A14 oligo, add 2 ul of 2x NPM master mix (Illumina) for a total of 15 ul. Run on a thermocycler: 1. 72°C for 10 minutes, 2. 98°C for 30 seconds, 3. 98°C for 10 seconds, 4. 59°C for 20 seconds, 5. 72°C for 10 seconds, 6. Repeat steps 3-5 for a total of 10 cycles, 7. 72°C for 2 minutes, and 8. 10°C hold temperature.

[0265] PCR: Add 1 μl of 25 μM SBS12, 1 μl of 25 μM A14, 8 μl of H2O, and 25 μl of 2x NEB Q5U Master Mix for a total of 50 μl. Run on a thermocycler: 1. 98°C for 30 seconds, 2. 98°C for 10 seconds, 3. 55°C for 20 seconds, 4. 72°C for 30 seconds, 5. Repeat steps 2-4 for a total of 20 cycles, 6. 72°C for 2 minutes, and 7. 10°C hold temperature. Libraries are typically amplified for 12-14 cycles. Libraries can be checked by loading 5 μl of PCR product onto a 1.2% Lonza agrose gel and resolving the product at 180V for 15 minutes.

[0266] Effect of DNA lesion size.

[0267] Proof-of-concept data for the AA → AB approach to gDNA. Three different TSMs containing different numbers of uracils were tested as the DNA lesion between the ME and the index: U, UU, or UUU. The first extension was repeated 10 times. All TSMs were functional and generated libraries with varying efficiencies, with a single U being the most efficient (Figure 11). The AB system was compared to a control, where SBS12-ME TSM was mixed with A14-ME-loaded TSM. qPCR demonstrated that the AA → AB system increased template yield by approximately 4-fold compared to the standard AB system. Titration of LNA-ME concentration (data not shown here) indicated that 100 nM was efficient for the second extension.

[0268] Effect of modified nucleotides on extension to add adapters.

[0269] The data show that the standard A14-ME oligo (no locked nucleic acid (LNA) in the primer) performed poorly in the AA-to-AB conversion. An AB system containing SBS12-ME and A14-ME TSM was compared as a control (Figure 12). Instead of LNA-A14, an oligo made with regular bases was applied in the second extension. The final library yield from PCR was significantly reduced and shows a broad smear compared to Figure 11.

[0270] LNA-ME enhances the second extension. Increasing the number of cycles of LNA-ME extension improves yield, reaching nearly the theoretical maximum at 10 cycles (Figure 13). Compared to the nearly complete library conversion using modified A14-ME oligos with LNA modifications, the difference in poor library generation with standard A14-ME oligos (no LNA) was surprising and unexpected, a significant advantage. Additionally, the two-fold yield difference between the AA→AB and AB systems indicates that nearly complete maximum conversion was obtained.

[0271] Effect of annealing temperature.

[0272] Titration of LNA-ME annealing temperature in the nuclear ATAC bulk assay. The AB system containing SBS12 and A14ME was used as a control. Genomic DNA in the same number of nuclei was transposed by TSM. The second extension AA→AB workflow was performed at different annealing temperatures. Approximately 59.5°C showed optimal efficiency, enhancing amplifiable template by approximately 5-fold compared to the AB control according to qPCR (Figure 14).

[0273] Example 2 Improved single-cell combinatorial indexing A major challenge in single-cell omics is the efficient conversion of genomic features for each cell into a sequencing library. Herein, we describe an adapter-switching strategy for a single-cell combinatorial indexing workflow (sci) that is generalizable to multiple assays and does not require custom sequencing chemistry. In this technology, symmetric stranded sci (s3) provides one to two orders of magnitude improvement in the reads obtained per cell for various features, including chromatin accessibility (s3-ATAC), whole-genome sequencing (s3-WGS), and genome and chromatin conformation (s3-GCC).

[0274] Main Single-cell genomics assays are rapidly becoming a powerful platform for interrogating complex biological systems across the entire spectrum of life science disciplines. Platforms for capturing various properties at the single-cell level typically suffer from a trade-off between cellular throughput and the depth of information that can be obtained per cell. We demonstrate the use of transposase-based library construction to assess various genomic properties in a high-throughput manner. 2 Leveraging Single-Cell Combinatorial Indexing (SCI) 1We have described a workflow utilizing the tagmentation method. The tagmentation reaction itself is highly efficient, but viable sequencing library molecules are generated only if different adapters in the form of forward or reverse primary sequences are incorporated at each end of the molecule. During the tagmentation reaction, there is an equal probability of incorporating each of the two sequences, resulting in half of the molecules being forward-forward or reverse-reverse adapter combinations, reducing the theoretical yield to 50%. To counter this efficiency, the use of a larger complement of adapter species is required. 3 , incorporation of T7 promoter sequences to pass through an RNA intermediate 4~6 , or targeting 7 or random priming 8 Several strategies have been developed, including the incorporation of a second adapter using a transposase-adapter complex. Herein, we present an alternative strategy that utilizes adapter exchange to generate library molecules tagged with both forward and reverse adapters on both the top and bottom strands. Furthermore, this format allows for the use of DNA index sequences embedded within the transposase-adapter complex, enabling single-cell combinatorial indexing (sci) applications in which two rounds of indexing are performed, once at the transposition stage and a second time at the PCR stage. 1,9,10 .

[0275] This technology, symmetric strand SCI (S3), leverages the efficiency of single-adapter transposition to incorporate a forward primer sequence in addition to a universal mosaic end sequence and compartment-specific DNA barcode. The adapter is designed so that a uracil base is present immediately after the transposase recognition sequence (mosaic end) on the top strand of the resulting product, which is covalently incorporated during the tagmentation reaction. Polymerase extension using a uracil-intolerant enzyme results in a copy of the mosaic end sequence on the bottom strand without extension to the DNA barcode or forward primer sequence. Subsequent denaturation and addition of a mosaic end-locked nucleic acid (LNA) template containing the reverse primer sequence along with a uracil-resistant polymerase allows extension of the library molecule to incorporate additional sequences. To ensure maximum efficiency, the template oligonucleotide can be blocked from extension to prevent its action as a primer, allowing multiple linear extension reactions to be performed (Figure 7). An additional advantage of the S3 platform is that the adapter sequence is designed so that standard sequencing recipes can be used instead of the custom workflow and primers required for SCI technology. Single-cell chromatin accessibility libraries (s3-ATAC), with a 16-fold improvement in median passage reads per cell, compared to previous SCI-seq / sci-DNA-seq 11 Single-cell whole genome sequencing (s3-WGS), a 126-fold improvement over previous combinatorial index Hi-C methods. 12 We demonstrate this workflow for generating a novel technique that captures both genomic sequence and chromatin conformation information (s3-GCC) in single cells with a higher percentage of chromatin interaction signals than conventional methods.

[0276] Prior to the novel components of this technique, we first sought to establish an S3 technique for assessing chromatin accessibility, as it requires minimal pre-processing of nuclei. In S3-ATAC, nuclei are isolated and then tagmented as in traditional sci-ATAC-seq, but instead, single-end indexed transposomes are used and then run through the adapter-switching S3 workflow (Figure 7). To ensure a true single-cell library free of genomic contamination from other nuclei and with minimal barcode collisions, we performed a mixed-species experiment, also known as a "barnyard" test, on primary frozen human cortical tissue and frozen whole mouse brain tissue (Figure 15). To more accurately capture the rate of cross-cell contamination, we chose to perform this test on primary tissue samples rather than an idealized cell line setting. We further designed the experiment to assess the level of crosstalk at both points of potential introduction by mixing nuclei from two samples, pre- and post-tagmentation, at the tagmentation and PCR stages. Furthermore, we generated pure-species libraries by leveraging the inherent sample multiplexing capabilities of the single-cell combinatorial indexing workflow. In total, we generated 1,366 human and 1,054 mouse single-cell ATAC-seq profiles with a median of 30,886 and 26,530 unique, high-mapping-quality reads (hereafter referred to as "pass reads") aligned to chromosomes 1-22, 23 (human), X, and Y per cell in human and mouse, respectively. Notably, the libraries were highly complex, with a median of 69.05% of reads assigned to cells as unique, indicating that additional sequencing depth could be obtained beyond and significantly increase the coverage of currently sequenced cells. Upon further sequencing, predicted estimates fall within 2% of empirical data. 9Using established methods for projecting unique reads per cell, we found that our libraries reach median values ​​of 128,144 and 174,858 passing reads per cell at 95% library saturation for human cortex and mouse whole brain samples, respectively. We then compared the current depth and projections of our mouse brain samples with published datasets of equivalent tissues and, where available, their library projections. We found that our libraries represent an order of magnitude improvement over any other library or self-reported library projections (Figure 16).

[0277] To confirm that the improvement was not due to index duplication or genomic crosstalk, we verified the purity of the samples by assessing the number of unique reads corresponding to the combined human-mouse reference genome. Before any processing, i.e., before tagmentation, under experimental conditions in which nuclei were mixed, a collision rate of 5.12% was observed (Figure 17, 2 × 2.56% detected human-mouse collisions), well within acceptable levels. As expected, zero collisions were observed under experimental conditions after tagmentation, suggesting that the collisions observed in the pre-tagmentation experiment were due to double sampling as opposed to crosstalk or surrounding chromatin. We also confirmed that the read gains were indeed capturing biological signal that was not due to excessive background by assessing transcription start site (TSS) enrichment at median values ​​of 2.77 and 3.93, respectively, and fraction of reads in called peak (FRIP) at 14.60% and 19.40% for human and mouse samples, respectively, with both metrics being comparable to other platforms for matched tissue types. Next, we sought to use sufficient signal to identify the cell types present within the samples. For each species, we used the called peaks on the aggregate data to construct a count matrix, followed by the topic modeling tool cisTopic. 13We then build a dimensionality reduction using UMAP 14 We visualized the data using and finally performed graph-based clustering at the topic level. We found clear separation of cell types in both the visualization space and the identified clusters, with distinct signals in cell type-specific genes within each cluster for both human cortex and mouse whole brain samples (Figures 18-21).

[0278] Second, the improvement in data quality generated by s3-ATAC is expected to be significant for high-throughput, low-pass single-cell genome sequencing. 9 We reasoned that our method should be translatable to other single-cell combinatorial indexing workflows, including our sci-DNA-seq method previously reported for s3. In addition to using the s3 workflow (Figure 7), we also investigated other improvements to the nucleosome depletion component of the technique used to obtain uniform coverage. First, we deployed s3-WGS (Figure 6) on a control lymphoblastoid cell line (GM12878) and found that an optimized version of detergent-based nucleosome depletion (×SDS) provided the highest uniformity and read count, with 92 cells in that condition demonstrating a median of 6,584,602 reads passed per cell, translating to a median genome capture rate of 37.12%. We also confirmed that coverage was uniform and comparable to other single-cell genome sequencing techniques by assessing the median absolute deviation (MAD), which fell within 0.1–0.3 (median 0.18). Using this optimized protocol, we next deployed s3-WGS to sequence two cell lines derived from primary pancreatic ductal adenocarcinoma (PDAC) tumors after minimal passages.

[0279] PDAC is a devastating form of cancer that typically presents at an advanced stage, making early detection and testing key to tumor progression. Because PDAC testing suffers from low cancer cell fractions in biopsy samples, we used continuously regenerating cell lines (CRCs) derived at low passages from purified tumors. This method allows for multiple modalities of characterization and perturbation while preserving much of the heterogeneity present in tumor samples, as evidenced by karyotyping. 15 We targeted two lines (termed PDAC-1 and PDAC-2) harboring two distinct subclonal missense mutations (p.G12D and p.G12C) in the oncogene KRAS and significant genomic instability as measured by G-banding-based and spectral karyotyping. For these lines, we obtained 709 and 267 single-cell libraries with median expected pass-through reads of 2,096,207 and 1,445,381 for PDAC-1 and PDAC-2, respectively (Figures 22-24). While lower than the initial GM12878 control sample, it significantly exceeds the coverage achieved by conventional methods. The MAD scores for the two lines (Figure 24) were higher than those of the relatively normal karyotype of GM12878, with median values ​​of 0.28 and 0.32, respectively; however, this is expected given the extensive copy number alterations present in the samples. We validated this expectation with paired whole-exome sequencing and copy number calling from PDAC-1 primary tumors, normal blood, and CRC lines, revealing strong evidence of characteristic genomic instability. Next, we performed single-cell copy number profiling to identify highly altered genomic landscapes within each of the two lines. Based on limited karyotyping and whole-exome data, we confirmed similar cell-by-cell patterns of multimegabase copy number abnormalities. Using inferred copy number profiles within genomic windows, we performed hierarchical and K-means clustering for three samples from GM12878 and two PDAC lines, revealing multiple clonal genomic arrangements.

[0280] With single-cell resolution, we were able to assess the occurrence of copy number aberrations in known PDAC-associated oncogenes and tumor suppressors. As an example of interpatient differences, cluster 7, exclusively populated by PDAC-2 samples, showed unique amplification of genomic regions including TGFβR2 and PBRM1, regions associated with cell proliferation and previously associated with a high tumor cell fraction in PDAC patients. PDAC-1 samples revealed heterogeneous amplification of a genomic region including the oncogene MYC (absolute copy number 2.26 ± 2.36). Furthermore, we found focal amplification of a genomic region overlapping with the oncogene KRAS, which is known to occur in >90% of PDAC cases. Cluster 1 contained the fewest cells with KRAS amplification (23.3%, 40 / 172 cells), while cluster 5 contained the highest frequency of KRAS copy number gain (82.6%, 138 / 167 cells). We validated this heterogeneous copy number aberration by utilizing whole-exome data for genotyping and digital droplet PCR and found that 53% of KRAS alleles sampled from PDAC-1 CRC lines displayed mutant KRAS alleles associated with overexpression.

[0281] Duplications and deletions are not the only forms of genomic rearrangements that can induce a competitive advantage in cancer cell proliferation. Genomic inversions are difficult to assess through standard karyotyping and chromosome painting methods, while chromosomal translocations are difficult to uncover with whole-genome amplification methods because only reads capturing the breakpoints provide supporting evidence. To address both of these limitations, we utilized s3-WGS technology with an additional preprocessing workflow: restriction digestion after fixation and nucleosome depletion, followed by religation following s3 library preparation (as in HiC, but without incorporating biotinylated bases). We reasoned that this additional processing would yield a portion of the reads spanning chimeric ligation junctions, representing distal chromatin contact points, while the remaining reads would serve as whole-genome sequencing data, enabling both genomic and chromatin conformation analysis (s3-GCC) (Figure 9). We performed s3-GCC on the same two PDAC cell lines as in the s3-WGS experiments (Figures 25-28), generating 22 and 93 cell profiles with comparable predicted median reads per cell for PDAC-1 and PDAC-2, 1,034,014 and 1,245,266, respectively. Copy number calling was then performed, and the results were compared with the s3-WGS libraries, revealing similar patterns of interspersed profiles from each method within the cell line panel. To obtain an initial measure of chromatin conformation signal, we assessed the proportion of interchromosomal read pairs, and both s3-GCC preparations contained an excess over their s3-WGS counterparts, with a 68.91- and 58.91-fold increase. The proportion of reads with insert sizes greater than 1 kbp was then measured, with a median of 15.6% and 17.0%, respectively, and a mean of 16%, for each lineage, for a fold enrichment of 361-fold and 402-fold over s3-WGS, again with comparable medians. To assess the predicted total unique chromosome contacts per cell, we first assumed that the read count projection of chromatin contacts performed identically to the bulk of the data, representing standard genome sequencing reads, and then took the percentage of total pass reads.This generated predicted median contact points per cell of 20,451 for PDAC-1 and 20,611 for PDAC-2. Furthermore, we performed read count projections specifically on the portion of reads representing chromatin contacts, yielding similar values ​​of 244,728 and 245,560. We then demonstrated the ability to generate chromatin contact maps using contact points obtained from relatively shallow sequencing depths, with aggregate profiles exhibiting distinct topological patterns. We separated single cells by their distal contact information via scHiCluster and observed three distinct clusters. Notably, even at this low sequencing depth, we could reliably distinguish the sparse contact profiles of the cell lines. To assess unique translocation and inversion events across the sampled cells, we examined the differences between the aggregated contact maps of clusters 0 (occupied exclusively by PDAC-1) and 1. We found that single-cell contact data replicated the chromosome arm-scale translocations reported from spectral karyotyping (SKY) data in the case of the t(3;14)(q24-26;q21-24) translocation, which was not specifically identified for PDAC-1 samples. We also found an increased frequency of interchromosomal contacts between the TGFβR2 and PBRM1 regions of chromosome 3 seen in s3-WGS data, suggesting aberrant genomic compartmentalization of copy number gains toward chromosomes 2 and 4.

[0282] In summary, the s3 workflow represents a significant improvement over conventional sci platforms in terms of the number of pass reads obtained per cell, without sacrificing signal enrichment in the case of s3-ATAC or coverage uniformity in the case of s3-WGS. We also introduce another variant of the combinatorial indexing workflow, namely s3-GCC, to obtain both genome sequencing and chromatin conformation, improving the number of chromatin contacts obtained per cell compared to sci-HiC. We demonstrate the utility of these approaches by evaluating two patient-derived tumor cell lines with dramatic chromatin instability. We uncover patterns of focal amplification of disease-associated genes, revealing widespread heterogeneity at a throughput not achievable with standard karyotyping. Furthermore, we highlight the collaborative analysis of protocols to reveal the effects of copy number aberrations that disrupt chromatin compartmentalization. Furthermore, the s3 workflow possesses the same inherent throughput potential of standard single-cell combinatorial indexing. We also expect this platform to be compatible with other transposase-based technologies, including sci-MET. 10 One possible drawback of the S3 platform is the need to use a full set of unique transposome complexes, as opposed to using a series of 8 forward and 12 reverse complexes (corresponding to the rows and columns of a 96-well plate), increasing the number of oligos required for the workflow. However, these costs ultimately pay off as proportionally fewer oligos are required per experiment. Finally, unlike the SCI workflow, the S3 platform does not require custom sequencing primers or custom sequencing recipes, eliminating one of the major hurdles labs can face while implementing these technologies.

[0283] Reference to Example 2 1. Cusanovich,D.A.et al.Multiplex single-cell profiling of chromatin accessibility by combinatorial cellular indexing.Science(80-.)。348,910-914(2015)。 2. Adey,A.et al.Rapid,low-input,low-bias construction of shotgun fragment libraries by high-density in vitro transposition.Genome Biol.11,R119(2010)。 3. Tan,L.,Xing,D.,Chang,C.H.,Li,H.& Xie,X.S.Three-dimensional genome structures of single diploid human cells.Science(80-.)。361,924-928(2018)。 4. Sos,B.C.et al.Characterization of chromatin accessibility with a transposome hypersensitive sites sequencing(THS-seq) assay.Genome Biol.17,20(2016)。 5. Yin,Y.et al.High-Throughput Single-Cell Sequencing with Linear Amplification.Mol.Cell 76,676-690.e10(2019)。 6. Chen,C.et al.Single-cell whole-genome analyses by Linear Amplification via Transposon Insertion(LIANTI).Science(80-.)。356,189-194(2017)。 7. Adey,A.& Shendure,J.Ultra-low-input,tagmentation-based whole-genome bisulfite sequencing.Genome Res.22,1139-1143(2012)。 8. Mulqueen,R.M.et al.Highly scalable generation of DNA methylation profiles in single cells.Nat.Biotechnol.36,428-431(2018)。 9. Vitak,S.A.et al.Sequencing thousands of single-cell genomes with combinatorial indexing.Nat.Methods 14,302-308(2017)。 10. Mulqueen,R.M.et al.Highly scalable generation of DNA methylation profiles in single cells.Nat.Biotechnol.36,428-431(2018)。 11. Vitak,S.A.et al.Sequencing thousands of single-cell genomes with combinatorial indexing.Nat.Methods 14,(2017)。 12. Ramani,V.et al.Massively multiplex single-cell Hi-C.Nat.Methods 14,263-266(2017)。 13. Bravo Gonzalez-Blas,C.et al.cisTopic:cis-regulatory topic modeling on single-cell ATAC-seq data.Nat.Methods 16,397-400(2019)。 14. Becht, E. et al. Dimensionality reduction for visualizing single-cell data using UMAP.Nat.Biotechnol.37,38-44(2018). 15. Lindenburger,K.et al.AB024.S024.Drug responses of patient-derived cell lines in vitro that match drug responses of patient PDAc tumors in situ.Ann.Pancreat.Cancer 1,AB024-AB024(2018).

[0284] method Generation of s3-ATAC library Prior to sample handling, composite transposases were obtained from Illumina Inc. 96 unique indexed transposases were loaded with one of each adapter, diluted to 2.5 uM, and stored at -20°C. Fifty milliliters of nuclear isolation buffer (NIB-HEPES) was freshly prepared with a final concentration of 10 mM HEPES-KOH (Fisher Scientific BP310-500 and Sigma-Aldrich 1050121000, respectively), pH 7.2, 10 mM NaCl (Fisher Scientific S271-3), 3 mM MgCl2 (Fisher Scientific AC223210010), 0.1% (v / v) IGEPAL CA-630 (Sigma-Aldrich I3021), and 0.1% (v / v) Tween (Sigma-Aldrich P-7949), and diluted with PCR-grade ultrapure distilled water (Thermo Fisher Scientific 10977015). After dilution, two tablets of Pierce™ Protease Inhibitor Mini Tablets, EDTA-Free (Thermo Fisher A32955) were dissolved and suspended to prevent protease degradation during nuclei isolation.

[0285] For s3-ATAC tissue handling, primary samples of C57 / B6 mouse whole brain and human cortex were extracted, flash-frozen in a liquid nitrogen bath, and then stored at -80°C. A bench dissection stage was set up prior to nuclear extraction. Petri dishes were placed on dry ice, and a new, sterilized razor blade was pre-cooled by embedding it in dry ice. A 7 mL Dounce homogenizer was filled with 2 mL of NIB-HEPES buffer and kept on wet ice. The Dounce homogenizer pestle was cooled by placing a 15 mL tube of ice-cold 70% (v / v) ethanol (Decon Laboratories Inc. 2701) on ice. Immediately before use, the pestle was rinsed with chilled distilled water. For tissue dissociation, mouse and human brain samples were treated similarly. The still-frozen tissue block was placed on a clean, pre-cooled Petri dish and coarsely chopped with a razor blade. Approximately 1 mg of minced tissue was then transferred to chilled NIB-HEPES buffer in a Dounce homogenizer. The suspended sample was allowed 5 minutes to equilibrate to the salt concentration change before Dounce. It was then homogenized with 5 strokes of a loose (A) pestle, incubated for an additional 5 minutes, and then homogenized with 5-10 strokes of a tight (B) pestle. The sample was then filtered through a 35 μm cell strainer (Corning 352235) while being transferred to a 15 mL conical tube. The nuclei were kept on ice until ready to proceed. The nuclei were pelleted by centrifugation at 400 rcf for 10 minutes at 4°C. The supernatant was removed, and the pellet was resuspended in 1 mL of NIB-HEPES buffer. This process was repeated for a second wash, and the nuclei were again kept on ice until ready to proceed. A 10 μL aliquot of the suspended nuclei was diluted with 90 μL of NIB-HEPES (1:10 dilution) and quantified using either a hemocytometer or a BioRad TC-20 automated cell counter according to the manufacturer's recommended protocol. The stock nuclei suspension was then diluted to a concentration of 1400 nuclei / μL.

[0286] Tagmentation plates were prepared by combining 420 μL of the 1400 nuclei / μL solution with 540 μL of 2×TD buffer (Nextera XT kit, Illumina Inc.). From this mixture, 8 μL (approximately 5000 nuclei total) was pipetted into each well of a 96-well plate according to the well schema. Then, 1 μL of 25 μM unique indexed transposase was pipetted into each well. Tagmentation was performed at 55°C for 10 minutes on a 300 rcf Eppendorf ThermoMixer. After this incubation, the plate temperature was reduced by a brief incubation on ice to stop the reaction. Pools of tagged nuclei were mixed according to the experimental schema, and 2 μL of 5 mg / mL DAPI (Thermo Fisher Scientific D1306) was added.

[0287] Nuclei were then flow-sorted through a Sony SH800 to remove debris and obtain accurate counts per well prior to PCR. A 96-well plate was prepared with 9 μL of 1×TD buffer (diluted with ultrapure water) and then held in a sample chamber at 4°C. Fluorescent nuclei were then flow-sorted for single nuclei, gating by size, internal complexity, and DAPI fluorescence. Immediately after sorting was completed, the plate was sealed and spun down at 500 rcf for 5 minutes at 4°C to ensure the nuclei were within the buffer.

[0288] Nucleosomes and remaining transposase were then denatured by adding 1 μL of 0.1% SDS (approximately 0.01% FC) per well. Subsequently, 4 μL of NPM (Nextera XT kit, Illumina Inc.) was added per well to gap-fill the tagged genomic DNA and incubated at 72°C for 10 minutes. 1.5 μL of 1 μM A14-LNA-ME oligo was then added to provide the template for adapter switching. Polymerase-based adapter switching was then performed under the following conditions: initial denaturation at 98°C for 30 seconds, 10 cycles of 98°C for 10 seconds, 59°C for 20 seconds, and 72°C for 10 seconds. The plate was then held at 10°C. After adapter switching, the SDS dwell time was quenched by adding 1% (v / v) Triton-X100 dissolved in ultrapure HO (Sigma 93426). At this point, some plates were stored at -20°C for several weeks, while other plates were processed immediately.

[0289] The following was then mixed per well for PCR: 16.5 μl of sample, 2.5 μL of 10 μM indexed i7 primer, 2.5 μL of 10 μM indexed i5 primer, 3 μL of ultrapure water, and 25 μL of NEBNext Q5U 2X Master Mix (New England Biolabs M0597S), and 0.5 μL of 100X SYBR Green I (Thermo Scientific S7563) for a 50 μL reaction per well. Real-time PCR was performed on a BioRad CFX with SYBR fluorescence measured after each cycle: 98°C for 30 seconds, 16–18 cycles of 98°C for 10 seconds, 55°C for 20 seconds, 72°C for 30 seconds, fluorescence read, 72°C for 10 seconds. After the fluorescence passed exponential growth and began to refract, the sample was held at 72°C for an additional 30 seconds and then stored at 4°C.

[0290] The amplified library was then cleaned up by pooling 25 μL per well into a 15 mL conical tube and cleaned up via a Qiaquick PCR purification column according to the manufacturer's protocol (Qiagen 28106). The pooled sample was eluted in 50 μL of 10 mM Tris-HCl, pH 8.0 (Life Technologies AM9855). The library molecules then underwent size selection via SPRI selection beads (Mag-Bind® TotalPure NGS Omega Biotek M1378-01). After vortexing at room temperature, 50 μL of fully suspended SPRI beads were mixed with 50 μL of the library (1× cleanup) and incubated at room temperature for 5 minutes. The reaction was then clarified once by placing it on a magnetic rack, and the supernatant was removed. The remaining pellet was rinsed twice with 100 μL of fresh 80% ethanol. After pipetting off the ethanol, the tube was spun down and returned to the magnetic rack to remove any residual ethanol. The beads were then removed from the magnetic rack, resuspended in 31 μL of 10 mM Tris-HCl, pH 8.0, and incubated at room temperature for 5 minutes. The tubes were then placed back on the magnetic rack, cleared, and the entire supernatant transferred to a clean tube. DNA was then quantified using the Qubit dsDNA High Sensitivity Assay according to the manufacturer's instructions (Thermo Fisher Q32851). The libraries were then diluted to 2 ng / μL and run on an Agilent Tapestation 4150 D5000 tape (Agilent 5067-5592). Library molecule concentrations ranging from 100 to 1000 bp were used, up to a final library dilution of 1 nM. The diluted libraries were then sequenced using a high- or medium-capacity 150 bp sequencing kit on a Nextseq 500 system according to the manufacturer's recommendations (Illumina Inc.).

[0291] Generation of s3-WGS libraries Prior to processing, the following buffers were prepared: 50 mL of NIB HEPES buffer as described above, and 50 mL of a Tris-based NIB (NIB Tris) variant containing final concentrations of 10 mM Tris HCl pH 7.4 (Life Technologies AM9855), 10 mM NaCl, 3 mM MgCl, 0.1% (v / v) IGEPAL CA-630, and 0.1% (v / v) Tween, diluted with PCR-grade ultrapure distilled water. After dilution, two Pierce™ protease inhibitor mini-tablets (EDTA-free) were dissolved and suspended to prevent protease degradation during nuclei isolation.

[0292] s3-WGS library preparation was performed for cell lines as follows. For patient-derived CRC cell lines, cells were seeded at a density of 1 x 106 onto a T25 flask prior to treatment. Cells were washed twice with ice-cold 1x PBS (VWR75800-986) and then trypsinized with 5 mL of 1x TrypLE (Thermo Fisher 12604039) at 37°C for 15 minutes. Suspension cells were then collected and pelleted at 300 rcf for 5 minutes at 4°C. For the suspension-growing cell line (GM12878), cells were pipetted from the growth medium and pelleted at 300 rcf for 5 minutes at 4°C.

[0293] Following the initial pellet, the cells were washed twice with 1 mL of ice-cold NIB HEPES. After the second wash, the pellet was resuspended in 300 μL of NIB HEPES. Nuclei were aliquoted and quantified as described above, and then 1 million nuclei aliquots were generated based on the quantification. The aliquots were pelleted by centrifugation at 300 rcf for 5 minutes at 4°C and resuspended in 5 mL of NIB HEPES. 246 μL of 16% (w / v) formaldehyde (Thermo Fisher 28906) was then added to the nuclei suspension (fc0.75% formaldehyde) to gently fix the nuclei. Nuclei were fixed by incubating in the formaldehyde solution for 10 minutes on an orbital shaker set at 50 rpm. The suspension was then pelleted at 500 rcf for 4 minutes at 4°C, and the supernatant was aspirated. The pellet was then resuspended in 1 mL of NIB Tris buffer to quench any remaining formaldehyde. The nuclei were pelleted again at 500 rcf for 4 minutes at 4°C, and the supernatant was aspirated. The pellet was washed once with 500 μL 1x NE Buffer 2.1 (NEB B7202S) and then resuspended in 760 μL 1x NE Buffer 2.1. 40 μL of 1% SDS (v / v) was added, and the sample was incubated on a ThermoMixer set at 300 rcf for 20 minutes at 37°C. The nucleosome-depleted nuclei were then pelleted at 500 rcf for 5 minutes at 4°C and then resuspended in 50 μL of NIB Tris. A 5 μL nuclei aliquot was taken, diluted 1:10 with NIB Tris, and quantified as described above. Nuclei were diluted to 500 nuclei / uL by adding NIB Tris based on quantification. Depending on the experimental setup, 420 uL of nuclei at 500 nuclei / uL were mixed with 540 uL 2x TD buffer. Following this, nuclei were tagmented, stained, flow-sorted, gap-filled genomic DNA, and adapter switching were performed as described for the s3-ATAC protocol. Library amplification was performed by PCR as described above, with a low total number of cycles (13–15) to allow for a high number of initial capture events per library. Libraries were then cleaned up, size-selected, and sequenced as previously described.

[0294] Generate s3-GCC library The same cultured cell line samples were sampled and processed from the same pool of fixed, nucleosome-depleted nuclei as described for generating s3-WGS libraries. Following nuclei quantification, the remaining nuclear suspension (approximately 2-3 million nuclei per sample) was pooled into each of the samples. Nuclei were pelleted at 500 rcf for 5 minutes at 4°C and resuspended in 90 μL of 1x Cutsmart buffer (NEB B7204S). 10 μL of 10 U / μL AluI restriction enzyme (NEB R0137S) was added to each sample. Samples were then digested for 2 hours at 300 rpm and 37°C on a ThermoMixer. After digestion, nuclear fragments underwent proximity ligation. Nuclei were pelleted at 500 rcf for 5 minutes at 4°C and resuspended in 100 μL of ligation reaction buffer. The ligation buffer was a mixture of 1x T4 DNA ligase buffer + ATP (NEB M0202S), 0.01% Triton X-100, 0.5 mM DTT (Sigma D0632), and 200 U of T4 DNA ligase diluted in ultrapure water to a final concentration. Ligation was performed at 16°C for 14 hours (overnight). After this incubation, nuclei were pelleted at 500 rcf for 5 minutes at 4°C and resuspended in 100 μL of NIB HEPES buffer. Nuclei aliquots were quantified as previously described, then diluted, aliquoted, tagmented, pooled, DAPI stained, and flow-sorted. Genomic DNA was gap-filled and adapter-switched as described for the s3-ATAC protocol. Library amplification occurred at the same rate as for s3-WGS libraries (13–15 cycles), after which libraries were pooled, cleaned up, and sequenced as described above.

[0295] Example 3 Library preparation by combined tagmentation and indexing In the following examples, methods and systems are presented for preparing dual-indexed paired-end libraries from nucleic acid samples using tagmentation and indexing steps, where a first index sequence is added via tagmentation and a second index sequence is added via hybridization and extension.

[0296] This example uses an immobilized transposome complex having a transposon with a first strand of 5'-primer-index-adapter-uracil-transposase recognition domain, e.g., 5'-P5-i5-A14-U-ME-3', and a second strand that is the complement of the transposase recognition sequence, e.g., 5'-ME'-3'. An exemplary first strand of the transposon is SEQ ID NO: 1, and an exemplary second strand of the transposon is the complement of nucleotides 53-71 of SEQ ID NO: 1. The transposome complex is immobilized on beads via biotin attached to the 5' end of the first strand via a cleavable linker. An oligonucleotide containing the second index sequence has a sequence of 5'-primer-index-adapter-transposase recognition sequence, e.g., 5'-P7-i7-B15-ME-3'. An exemplary oligonucleotide containing the second index sequence is SEQ ID NO: 2. The second index sequence is optionally blocked at the 3' end, for example, using a dideoxy or locked nucleic acid. Exemplary sequences for P5, i5, P7, i7, ME, A14, and B15 are SEQ ID NOs: 3-9, respectively.

[0297] The nucleic acid in solution is added to each well of a 96-well plate at the desired length. A suspension of bead-bound transposomes carrying the transposon sequence described above is added to each well, and the plate is incubated under conditions suitable for fragmenting the transposase and inserting the transposon sequence. The transposase enzyme is removed, for example, by adding SDS or heating. A uracil-intolerant polymerase (e.g., proofreading polymerase or Phusion) is added to fill the gap between the ends of the nucleic acid fragment and the second transposon sequence. The uracil-intolerant polymerase stops when it reaches the uracil inserted between A14 of the inserted first transposon sequence and the ME sequence. The tagmented nucleic acid is denatured, and a second indexing oligonucleotide is hybridized to the tagmented nucleic acid by enzymatic extension. The doubly-indexed nucleic acid can be used for sequencing, including but not limited to amplification, or further processed.

[0298] The complete disclosures of all patents, patent applications, and publications, as well as electronically available materials cited herein (e.g., nucleotide sequence submissions in GenBank and RefSeq, amino acid sequence submissions in SwissProt, PIR, PRF, PDB, and translations from annotated coding regions in GenBank and RefSeq), are incorporated by reference in their entirety. Supplementary materials referenced in publications (such as supplementary tables, figures, materials and methods, and / or experimental data) are likewise incorporated by reference in their entirety. In the event of a conflict between the disclosure of this application and the disclosure of a document incorporated by reference herein, the disclosure of this application shall control. The foregoing detailed description and examples are provided for clarity of understanding only. No unnecessary limitations should be understood therefrom. The disclosure is not limited to the exact details shown and described, since variations obvious to those skilled in the art are included in the disclosure defined by the claims.

[0299] Unless otherwise indicated, all numbers expressing quantities of ingredients, molecular weights, and the like used in the specification and claims should be understood to be modified in all instances by the term "about." Accordingly, unless otherwise indicated, the numerical parameters set forth in the specification and claims are approximations that may vary depending upon the desired properties sought to be obtained by the present disclosure. At the very least, and not as an attempt to limit the scope of the claims to the doctrine of equivalents, each numerical parameter should, at the very least, be construed in light of the number of reported significant digits and by applying ordinary rounding techniques.

[0300] Notwithstanding that the numerical ranges and parameters setting forth the broad scope of the present disclosure are approximations, the numerical values ​​set forth in the specific examples are reported as precisely as possible, however, all numerical values ​​inherently contain ranges necessarily resulting from the standard deviations found in their respective testing measurements.

[0301] All headings are for the convenience of the reader and should not be used to limit the meaning of the text that follows the heading, unless specifically stated. The present invention provides, for example, the following items. (Item 1) 1. A method for generating a sequencing library, comprising: providing a plurality of symmetrically modified target nucleic acids comprising a first adaptor sequence at each end, wherein the first adaptor sequence comprises a DNA lesion; extending the modified target nucleic acid with a damage-intolerant polymerase to generate a plurality of asymmetrically modified target nucleic acids comprising the first adapter sequence at the 5' end of each strand and the complement of a portion of the first adapter at the 3' end of each strand. (Item 2) 2. The method of claim 1, wherein the plurality of symmetrically modified target nucleic acids are double-stranded, each strand comprising the first adapter sequence comprising the DNA damage 5' to 3', the target nucleic acid, a gap comprising at least one nucleotide, and the complement of a portion of the first adapter sequence that does not comprise the DNA damage. (Item 3) 2. The method of claim 1, wherein the extension begins at the gap. (Item 4) annealing a primer to the plurality of asymmetrically modified target nucleic acids, the primer comprising a second adapter sequence and an annealing domain from 5' to 3', the annealing domain comprising a nucleotide sequence that anneals to the complement of the portion of the first adapter of the plurality of asymmetrically modified target nucleic acids; extending the 3' end of the annealed asymmetrically modified target nucleic acid with a damage-intolerant polymerase; 3. The method of claim 2, wherein the extension results in a plurality of asymmetrically modified target nucleic acids comprising, from 5' to 3', (i) the first adaptor, (ii) the target nucleic acid, (iii) the complement of the portion of the first adaptor, and (iv) the complement of the second adaptor. (Item 5) 5. The method of claim 4, wherein the extension of the 3' end of the annealed asymmetrically modified target nucleic acid is repeated at least three times. (Item 6) 2. The method of claim 1, wherein the DNA damage comprises at least one of an abasic site, a modified base, a mismatch, a single-strand break, or a crosslinked nucleotide. (Item 7) 2. The method of claim 1, wherein the DNA damage comprises at least one uracil. (Item 8) 5. The method of claim 4, wherein the annealing domain of the primer comprises at least one modified nucleotide that increases the melting temperature compared to the corresponding natural DNA nucleotide. (Item 9) 9. The method of claim 8, wherein the modified nucleotide comprises a locked nucleic acid, PNA, or RNA. (Item 10) 5. The method of claim 4, wherein the 3' end of the primer is blocked. (Item 11) Item 10. The method of item 1, wherein the first adapter comprises one or more universal sequences, one or more index sequences, one or more universal molecular identifiers, or a combination thereof. (Item 12) At least one of the one or more universal sequences, one or more index sequences, and one or more universal molecular identifiers is located distal to the DNA lesion and the target nucleic acid. Item 12. The method of item 11, wherein the adapter is located between the ends of the adapter. (Item 13) 5. The method of claim 4, wherein the second adapter comprises one or more universal sequences, one or more index sequences, one or more universal molecular identifiers, or a combination thereof. (Item 14) 14. The method of claim 13, wherein the one or more universal sequences, one or more index sequences, and one or more universal molecular identifiers of the first adaptor are unique compared to the one or more universal sequences, one or more index sequences, and one or more universal molecular identifiers of the second adaptor. (Item 15) Item 12. The method of item 11, wherein the one or more index sequences of the first adapter are compartment-specific. (Item 16) Item 14. The method of item 13, wherein the one or more index sequences of the second adapter are compartment-specific. (Item 17) Item 10. The method of item 1, wherein the first adapter comprises a transposase recognition site. (Item 18) 2. The method of claim 1, wherein the target nucleic acid is derived from nucleic acid derived from a single cell. (Item 19) 2. The method of claim 1, wherein the target nucleic acid is derived from nucleic acids derived from multiple cells. (Item 20) 20. The method of claim 18 or 19, wherein the target nucleic acid from a single cell or multiple cells comprises RNA. (Item 21) 21. The method of claim 20, wherein the RNA comprises mRNA. (Item 22) 20. The method of claim 18 or 19, wherein the target nucleic acid derived from a single cell or multiple cells comprises DNA. (Item 23) 23. The method of claim 22, wherein the DNA comprises whole cell genomic DNA. (Item 24) 24. The method of claim 23, wherein the whole cell genomic DNA comprises nucleosomes. (Item 25) 2. The method of claim 1, wherein the target nucleic acid is derived from cell-free DNA. (Item 26) 17. The method according to any one of items 11 to 16, wherein the method comprises combinatorial indexing. (Item 27) further comprising amplifying the asymmetrically modified target nucleic acid; the amplification comprises a second primer and a damage-tolerant polymerase; 2. The method of claim 1, wherein the second primer comprises a nucleotide sequence that anneals to the first adapter sequence or its complement. (Item 28) 28. The method of claim 27, wherein the second primer further comprises one or more universal sequences, one or more index sequences, one or more universal molecular identifiers, or a combination thereof. (Item 29) The one or more universal sequences, one or more index sequences of the second primer 29. The method of claim 28, wherein the sequence, and the one or more universal molecular identifiers are unique compared to the one or more universal sequences, one or more index sequences, and one or more universal molecular identifiers of the first adaptor and the second adaptor. (Item 30) 2. The method of claim 1, wherein a subset of the plurality of asymmetrically modified target nucleic acids is present within a plurality of compartments, and wherein (i) the first adapter comprises a first compartment-specific index, (ii) the second adapter comprises a second compartment-specific index, or both (i) and (ii). (Item 31) 31. The method of claim 30, further comprising combining the asymmetrically modified target nucleic acids from different compartments to generate pooled indexed asymmetrically modified target nucleic acids. (Item 32) further comprising distributing a subset of the pooled indexed asymmetrically modified target nucleic acids into a second plurality of compartments and modifying the indexed asymmetrically modified target nucleic acids; the modification comprises adding an additional compartment-specific index sequence to the indexed asymmetrically modified target nucleic acids present in each subset to provide indexed DNA nucleic acids; 32. The method of claim 31, wherein the modification comprises ligation or extension. (Item 33) 33. The method according to any one of items 30 to 32, wherein the compartment comprises a well or a droplet. (Item 34) 2. The method of claim 1, wherein the providing comprises contacting a plurality of DNA fragments with the first adaptors under conditions that ligate the first adaptors to both ends of the DNA fragments. (Item 35) 35. The method of claim 34, wherein the DNA fragment is double-stranded and blunt-ended. (Item 36) 36. The method of claim 34 or 35, wherein the first adaptor is a double-stranded DNA oligonucleotide. (Item 37) 36. The method according to item 34 or 35, wherein one 3' end of the first adaptor is blocked. (Item 38) 35. The method of claim 34, wherein the DNA fragment is double-stranded and contains a single-stranded region at one or both 3' ends. (Item 39) 39. The method of claim 34 or 38, wherein the first adaptor is a double-stranded DNA oligonucleotide comprising a single-stranded region at one end, the single-stranded region being capable of annealing to the single-stranded region present on the DNA fragment. (Item 40) 39. The method of claim 34, 35, or 38, wherein the adapter is a forked adapter. (Item 41) 2. The method of claim 1, wherein the providing comprises contacting DNA with a transposome complex, the transposome complex comprising a transposase and the first adapter, and the contacting occurs under conditions suitable for ligation of the first adapter to the DNA to generate the symmetrically modified target nucleic acid. (Item 42) Item 4. The symmetrically modified target nucleic acid produced contains a gap of at least one nucleotide in one strand between the ligated first adapter and the target nucleic acid. The method described in 1. (Item 43) 43. The method of claim 41 or 42, wherein the DNA is present in multiple compartments, and the first adapter in each compartment comprises a compartment-specific index. (Item 44) 43. The method of claim 42, further comprising combining single-stranded modified target nucleic acids from different compartments to generate pooled symmetrically modified target nucleic acids and distributing the symmetrically modified target nucleic acids into a second plurality of compartments. (Item 45) 24. The method of claim 23, wherein the method further comprises fragmenting the whole cell genomic DNA. (Item 46) 46. ​​The method of claim 45, wherein the fragmenting comprises digestion of the whole cell genomic DNA with a restriction endonuclease. (Item 47) 47. The method of claim 45 or 46, wherein the fragmented DNA is subjected to proximity ligation to attach chimeric target nucleic acids. (Item 48) 5. The method according to item 1 or 4, wherein the cytosine residues of the adapter are replaced with 5-methylcytosine. (Item 49) 49. The method of claim 48, wherein the symmetric or asymmetric target nucleic acid is subjected to chemical or enzymatic methylation conversion. (Item 50) 2. The method of claim 1, wherein said providing comprises fixing isolated nuclei, subjecting said isolated nuclei to conditions that dissociate nucleosomes from genomic DNA, fragmenting said genomic DNA, subjecting said fragments to proximity ligation that joins a chimeric target nucleic acid, and contacting said ligated fragments with a transposome complex, wherein said transposome complex comprises a transposase and said first adapter, and said contacting occurs under conditions suitable for ligation of said first adapter to said DNA to produce said symmetrically modified target nucleic acid. (Item 51) 51. The method of claim 50, wherein the fragmentation comprises digestion with a restriction endonuclease. (Item 52) providing a surface comprising a plurality of amplification sites, providing the amplification site comprising at least two populations of linked single-stranded capture oligonucleotides having free 3' ends; 5. The method of claim 4, further comprising contacting a surface comprising the amplification sites with the plurality of asymmetrically modified target nucleic acids under conditions suitable to generate a plurality of amplification sites each comprising a clonal population of amplicons from an individual asymmetrically modified target nucleic acid. (Item 53) A composition comprising a transposome complex and a DNA polymerase, the transposome comprises a transposase bound to a transposon sequence; the transposon sequence comprises an adapter and a DNA lesion; The composition, wherein the DNA polymerase is a damage-intolerant polymerase. (Item 54) 54. The composition of item 53, wherein the adaptor comprises one or more universal sequences, one or more index sequences, one or more UMIs, or a combination thereof. (Item 55) 55. The composition of item 53 or 54, further comprising a damage-tolerant DNA polymerase. (Item 56) 1. A composition comprising: a plurality of modified target nucleic acids comprising a first adaptor comprising a 5' to 3' DNA lesion, a target nucleic acid, and a complement of the first adaptor; a primer comprising a second adaptor from 5' to 3' and an annealing domain, wherein the annealing domain comprises a nucleotide sequence that anneals to the complement of the first adaptor; A composition comprising: a damage-intolerant DNA polymerase. (Item 57) 57. The composition of item 56, wherein the primer comprises at least one modified nucleotide that increases the melting temperature compared to the corresponding natural DNA nucleotide. (Item 58) 58. The composition of claim 56 or 57, wherein the primer is annealed to a target nucleic acid. (Item 59) 57. The composition of claim 56, wherein the 3' end of the primer is blocked. (Item 60) 57. The composition of item 56, wherein the first adapter comprises a transposase recognition site. (Item 61) a transposome complex and a DNA polymerase in separate containers, the transposome comprises a transposase bound to a transposon sequence; the transposon sequence comprises a first adapter and a DNA lesion; a transposome complex and a DNA polymerase, wherein the DNA polymerase is a damage-intolerant polymerase; A kit including instructions for use. (Item 62) further comprising a second DNA polymerase in a separate container; 62. The kit of item 61, wherein the second DNA polymerase is a damage-tolerant polymerase. (Item 63) further comprising a primer, 63. The kit of claim 61 or 62, wherein the primer comprises a second adapter and an annealing domain from 5' to 3', the annealing domain comprising a nucleotide sequence that anneals to the complement of the first adapter. (Item 64) Item 64. The kit of item 63, wherein the 3' end of the primer is blocked. (Item 65) 62. The kit of item 61, wherein the first adaptor comprises one or more universal sequences, one or more index sequences, one or more UMIs, or a combination thereof. (Item 66) 64. The kit of item 63, wherein the second adapter primer further comprises one or more universal sequences, one or more index sequences, one or more UMIs, or a combination thereof. (Item 67) A transposome complex comprising: Transposase and a transposon comprising a nucleic acid comprising, on a first strand, from 5' to 3', at least one universal sequence, at least one index sequence, at least one UMI, or a combination thereof, a DNA lesion, and a transposase recognition sequence, and, on a second strand, an adapter comprising a nucleotide complementary to at least a portion of the transposase recognition sequence; and a transposome complex comprising: (Item 68) 68. The transposome complex of item 67, wherein the first strand further comprises a capture agent at the 5' end of the first strand. (Item 69) 69. The transposome complex of item 68, wherein the first strand further comprises a cleavable linker located between the capture agent and the 5' end. (Item 70) 68. The transposome complex of item 67, wherein the second strand further comprises a capture agent at the 3' end of the second strand. (Item 71) 71. The transposome complex of item 70, wherein the second strand further comprises a cleavable linker positioned between the capture agent and the 3' end.

Claims

1. 1. A method for generating a sequencing library, the method comprising: providing a plurality of double-stranded symmetrically modified target nucleic acids comprising a first adapter sequence at each end; wherein the first adapter sequence comprises a DNA lesion; wherein each strand of the symmetrically modified target nucleic acid comprises, from 5' to 3', the first adapter sequence comprising the DNA lesion, the target nucleic acid, a gap comprising at least one nucleotide, and the complement of a portion of the first adapter sequence that does not comprise the DNA lesion. To provide and extending the modified target nucleic acid with a damage-intolerant polymerase to generate a plurality of asymmetrically modified target nucleic acids comprising the first adapter sequence at the 5'-end of each strand and the complement of a portion of the first adapter at the 3'-end of each strand; annealing a primer to the plurality of asymmetrically modified target nucleic acids, the primer comprising a second adapter sequence and an annealing domain from 5' to 3', the annealing domain comprising a nucleotide sequence that anneals to the complement of the portion of the first adapter of the plurality of asymmetrically modified target nucleic acids; wherein the annealing domain comprises at least one modified nucleotide that increases the melting temperature compared to the corresponding natural DNA nucleotide. Annealing, and extending the 3' end of the annealed asymmetrically modified target nucleic acid with a damage-intolerant polymerase using the primer as a template; wherein the extending results in a plurality of asymmetrically modified target nucleic acids comprising, from 5' to 3': (i) the first adaptor, (ii) the target nucleic acid, (iii) the complement of the portion of the first adaptor, and (iv) the complement of the second adaptor.

2. The method of claim 1 , wherein the extension begins at the gap.

3. The method of claim 1 , wherein the extension of the 3′ end of the annealed asymmetrically modified target nucleic acid is repeated at least three times.

4. 2. The method of claim 1, wherein the DNA damage comprises at least one of an abasic site, a modified base, a mismatch, a single-strand break, or a cross-linked nucleotide.

5. 2. The method of claim 1, wherein the DNA damage comprises at least one uracil.

6. 10. The method of claim 1, wherein the modified nucleotide comprises a locked nucleic acid, a PNA, or an RNA.

7. The method of claim 1 , wherein the 3′ end of the primer is blocked.

8. 10. The method of claim 1, wherein the first adaptor comprises one or more universal sequences, one or more index sequences, one or more universal molecular identifiers, or a combination thereof.

9. At least one of the one or more universal sequences, one or more index sequences, and one or more universal molecular identifiers is located distal to the DNA lesion and the target nucleic acid. The method of claim 8 , located in the adapter between the ends of the adapter.

10. 10. The method of claim 1, wherein the second adaptor comprises one or more universal sequences, one or more index sequences, one or more universal molecular identifiers, or a combination thereof.

11. 11. The method of claim 10, wherein the one or more universal sequences, one or more index sequences, and one or more universal molecular identifiers of the first adaptor are unique compared to the one or more universal sequences, one or more index sequences, and one or more universal molecular identifiers of the second adaptor.

12. 9. The method of claim 8, wherein the one or more index sequences of the first adaptor are compartment-specific.

13. 11. The method of claim 10, wherein the one or more index sequences of the second adaptor are compartment-specific.

14. 2. The method of claim 1, wherein the first adapter comprises a transposase recognition site.

15. 10. The method of claim 1, wherein the target nucleic acid is derived from nucleic acid derived from a single cell, and the nucleic acid comprises RNA or DNA.

16. 10. The method of claim 1, wherein the target nucleic acid is derived from nucleic acids derived from a plurality of cells, and the nucleic acids comprise RNA or DNA.

17. 17. The method of claim 15 or 16, wherein the RNA comprises mRNA.

18. 17. The method of claim 15 or 16, wherein the DNA comprises whole cell genomic DNA.

19. 19. The method of claim 18, wherein the whole cell genomic DNA comprises nucleosomes.

20. 10. The method of claim 1, wherein the target nucleic acid is derived from nucleic acid derived from cell-free DNA.

21. The method of any one of claims 8 to 13, wherein the method comprises combinatorial indexing.

22. further comprising amplifying the asymmetrically modified target nucleic acid; the amplification comprises a second primer and a damage-tolerant polymerase; The method of claim 1 , wherein the second primer comprises a nucleotide sequence that anneals to the first adapter sequence or its complement.

23. 23. The method of claim 22, wherein the second primer further comprises one or more universal sequences, one or more index sequences, one or more universal molecular identifiers, or a combination thereof.

24. The one or more universal sequences, one or more index sequences of the second primer 24. The method of claim 23, wherein the sequence, and one or more universal molecular identifiers are unique compared to the one or more universal sequences, one or more index sequences, and one or more universal molecular identifiers of the first adaptor and the second adaptor.

25. 2. The method of claim 1, wherein a subset of the plurality of asymmetrically modified target nucleic acids is present within a plurality of compartments, and wherein either (i) the first adapter comprises a first compartment-specific index, (ii) the second adapter comprises a second compartment-specific index, or both (i) and (ii).

26. 26. The method of claim 25, further comprising combining the asymmetrically modified target nucleic acids from different compartments to generate pooled indexed asymmetrically modified target nucleic acids.

27. further comprising distributing a subset of the pooled indexed asymmetrically modified target nucleic acids into a second plurality of compartments and modifying the indexed asymmetrically modified target nucleic acids; the modification comprises adding an additional compartment-specific index sequence to the indexed asymmetrically modified target nucleic acids present in each subset to provide indexed DNA nucleic acids; 27. The method of claim 26, wherein the modification comprises ligation or extension.

28. A method described in any one of claims 25 to 27, wherein subsets of the multiple asymmetrically modified target nucleic acids are present in multiple compartments, and the compartments comprise wells or droplets.

29. 2. The method of claim 1, wherein said providing comprises contacting a plurality of DNA fragments with said first adaptors under conditions to ligate said first adaptors to both ends of said DNA fragments.

30. 30. The method of claim 29, wherein the DNA fragment is double-stranded and blunt-ended.

31. 31. The method of claim 29 or 30, wherein the first adaptor is a double-stranded DNA oligonucleotide.

32. 31. The method of claim 29 or 30, wherein one 3' end of the first adaptor is blocked.

33. 30. The method of claim 29, wherein the DNA fragment is double-stranded and contains a single-stranded region at one or both 3' ends.

34. 34. The method of claim 29 or 33, wherein the first adaptor is a double-stranded DNA oligonucleotide comprising a single-stranded region at one end, wherein the single-stranded region is capable of annealing to the single-stranded region present on the DNA fragment.

35. 34. The method of any one of claims 29, 30, or 33, wherein the adapter is a forked adapter.

36. 2. The method of claim 1, wherein said providing comprises contacting DNA with a transposome complex, said transposome complex comprising a transposase and said first adapter, and said contacting occurs under conditions suitable for ligation of said first adapter to said DNA to generate said symmetrically modified target nucleic acid.

37. 37. The method of Claim 36, wherein the symmetrically modified target nucleic acid produced comprises a gap of at least one nucleotide in one strand between the ligated first adaptor and the target nucleic acid.

38. 38. The method of claim 36 or 37, wherein the DNA is present in multiple compartments and the first adapter in each compartment comprises a compartment-specific index.

39. 38. The method of claim 37, further comprising combining single-stranded modified target nucleic acids from different compartments to generate pooled symmetrically modified target nucleic acids, and distributing the symmetrically modified target nucleic acids into a second plurality of compartments.

40. 20. The method of claim 18, wherein the method further comprises fragmenting the whole cell genomic DNA.

41. 41. The method of claim 40, wherein said fragmenting comprises digestion of said total cellular genomic DNA with a restriction endonuclease.

42. 42. The method of claim 40 or 41, wherein the fragmented DNA is subjected to proximity ligation to attach a chimeric target nucleic acid.

43. 2. The method of claim 1, wherein the cytosine residues of the adapter are replaced with 5-methylcytosine.

44. 44. The method of claim 43, wherein the symmetric or asymmetric target nucleic acid is subjected to a chemical or enzymatic methylation conversion.

45. 2. The method of claim 1, wherein said providing comprises fixing isolated nuclei, subjecting the isolated nuclei to conditions that dissociate nucleosomes from genomic DNA, fragmenting the genomic DNA, subjecting the fragments to proximity ligation that joins a chimeric target nucleic acid, and contacting the ligated fragments with a transposome complex, wherein the transposome complex comprises a transposase and the first adapter, and wherein said contacting occurs under conditions suitable for ligation of the first adapter to the DNA to generate the symmetrically modified target nucleic acid.

46. 46. ​​The method of claim 45, wherein said fragmenting comprises digestion with a restriction endonuclease.

47. providing a surface comprising a plurality of amplification sites, providing the amplification site comprising at least two populations of linked single-stranded capture oligonucleotides having free 3' ends; 2. The method of claim 1, further comprising contacting a surface comprising the amplification sites with the plurality of asymmetrically modified target nucleic acids under conditions suitable to generate a plurality of amplification sites each comprising a clonal population of amplicons from an individual asymmetrically modified target nucleic acid.

48. 1. A composition comprising: a plurality of modified target nucleic acids comprising, from 5' to 3', a first adaptor comprising a DNA lesion, a target nucleic acid, a gap comprising at least one nucleotide, and a complement of a portion of the first adaptor sequence that does not comprise the DNA lesion; a primer comprising, from 5' to 3', a second adapter and an annealing domain, wherein the annealing domain comprises a nucleotide sequence that anneals to the complement of the portion of the first adapter, and wherein the annealing domain comprises at least one modified nucleotide that increases the melting temperature compared to a corresponding natural DNA nucleotide; A composition comprising: a damage-intolerant DNA polymerase;

49. 49. The composition of claim 48, wherein the primer is annealed to a target nucleic acid.

50. 49. The composition of claim 48, wherein the 3' end of the primer is blocked.

51. 49. The composition of claim 48, wherein the first adapter comprises a transposase recognition site.

52. The composition of claim 7, wherein the 3' end of the primer comprises a dideoxynucleotide.

Citation Information

Patent Citations

  • Transposon end compositions and methods for modifying nucleic acids

    JP2012506704A

  • Compositions including a double stranded nucleic acid molecule and a stem-loop oligonucleotide

    US10208337B2

  • Methods for asymmetric DNA library generation and optionally integrated duplex sequencing

    WO2020043803A1