Normalization Reagents and Methods
Novel polynucleotide adaptors with complementary and non-complementary regions and linkers address inefficiencies in traditional gDNA library normalization, ensuring accurate and efficient sequencing by minimizing errors and costs.
Patent Information
- Application Number
- JP2025547462
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-06
- Filing Date
- 2024-02-16
- Publication Date
- 2026-02-20
AI Technical Summary
Traditional methods for normalizing genomic DNA (gDNA) libraries are inefficient, requiring additional handling, reagents, time, labor, and cost, and are often inaccurate due to manual qPCR and multiple liquid transfer steps, which are exacerbated in high-throughput environments.
The use of novel polynucleotide adaptors with complementary and non-complementary regions connected by linkers, such as PEG units, to normalize gDNA libraries efficiently and accurately, allowing for balanced sequence loading across samples.
This approach achieves high-efficiency ligation with minimal error, reducing transformation inefficiencies and maintaining accuracy, thereby improving the quality of genomic sequencing.
Smart Images

Figure 2026506068000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 485,676, filed February 17, 2023, U.S. Provisional Patent Application No. 63 / 511,086, filed June 29, 2023, and U.S. Provisional Patent Application No. 63 / 550,574, filed February 6, 2024, each of which is incorporated by reference in its entirety.
[0002] Sequence Listing This application contains a Sequence Listing that has been submitted electronically in XML format under the file name "00415-0056-00304 Sequence Listing.xml," which is incorporated by reference in its entirety. The Sequence Listing was created on February 16, 2024, and is 103,870 bytes in size. [Background technology]
[0003] High-fidelity, low-cost identification of genomic variants in complex nucleic acid samples plays a central role in biotechnology, medicine, and basic biomedical research. To achieve balanced sequence loading in genomic DNA (gDNA) libraries, the number of nucleic acid molecules must be normalized across all samples. Traditional normalization methods are inefficient because they require additional handling, reagents, time, labor, and cost to measure individual gDNA concentrations and subsequently perform dilution steps. Furthermore, such traditional methods tend to be inaccurate because they require manual quantitative polymerase chain reaction (qPCR) followed by multiple liquid transfer steps for dilution. These inefficiencies and inaccuracies of normalization are exacerbated in high-throughput environments.
[0004] Alternative known normalization methods either require significant modifications to the library preparation workflow (e.g., the use of protected adapters followed by enzymatic processing) or are limited to transposase-based methods, which increase error and imprecision and reduce transformation efficiency. Such modifications to the library preparation workflow increase costs and reduce accuracy.
[0005] Novel reagents and methods are needed to generate and normalize gDNA libraries with greater efficiency while maintaining accuracy during sequencing. Summary of the Invention
[0006] Provided herein are compositions and methods for library size normalization.
[0007] (a) a first polynucleotide adaptor comprising a first strand comprising a first terminal adaptor region, a first non-complementary region, and a first yoke region; and a second strand comprising a second terminal adaptor region, a second non-complementary region, and a second yoke region, wherein the first yoke region and the second yoke region are complementary, and the first non-complementary region and the second non-complementary region are not complementary; and (b) a second polynucleotide adaptor comprising a third terminal adaptor region, a Provided herein is a composition comprising a third strand comprising three non-complementary regions and a third yoke region; a fourth strand comprising a fourth terminal adaptor region, a fourth non-complementary region, and a fourth yoke region, wherein the third yoke region and the fourth yoke region are complementary, and the third non-complementary region and the fourth non-complementary region are not complementary, wherein the first polynucleotide adaptor and the second polynucleotide adaptor are connected via a linker. Further provided herein is a composition in which a linker is attached to the first terminal adaptor region and the third terminal adaptor region. Further provided herein is a composition in which a linker is attached to the second terminal adaptor region and the fourth terminal adaptor region. Further provided herein is a composition in which the linker comprises an alkene, alkyne, triazine, triazole, oxime, sulfide, amide, or cyclooctane. Further provided herein is a composition in which the linker comprises a PEG unit. Further provided herein is a composition wherein the linker comprises 5 to 50 PEG units. Further provided herein is a composition wherein the linker comprises Sp18. Further provided herein is a composition wherein the linker comprises 1 to 10 Sp18. Further provided herein is a composition wherein one or more of the first polynucleotide adaptor and the second polynucleotide adaptor comprise at least one barcode. Further provided herein is a composition wherein the at least one barcode comprises one or more of a cell index, a sample index, or a unique molecular identifier (UMI).Further provided herein are compositions wherein one or more of the first polynucleotide adaptor and the second polynucleotide adaptor comprise at least one alpha thiophosphatase. Further provided herein are compositions wherein the 5' end of at least one of the strands comprises thymidine. Further provided herein are compositions wherein a linker connects the 5' ends of one or more strands. Further provided herein are compositions wherein a linker connects the 3' ends of one or more strands. Further provided herein are compositions wherein the linker comprises a polynucleotide. Further provided herein are compositions wherein the linker comprises at least one cleavable base. Further provided herein are compositions wherein the linker comprises uracil. Further provided herein are compositions wherein the linker comprises at least one restriction endonuclease site. Further provided herein are compositions wherein the linker comprises a double-stranded region. Further provided herein are compositions wherein the linker comprises at least one splint polynucleotide. Further provided herein are compositions in which at least one splint polynucleotide is at least partially complementary to a portion of a first polynucleotide adaptor or a second polynucleotide adaptor. Further provided herein are compositions in which a linker comprises at least two splint polynucleotides, wherein the at least two splint polynucleotides at least partially overlap with each other. Further provided herein are compositions in which a linker comprises at least two splint polynucleotides, wherein the at least two splint polynucleotides at least partially overlap with a portion of a first polynucleotide adaptor or a second polynucleotide adaptor. Further provided herein are compositions in which the first polynucleotide adaptor and the second polynucleotide adaptor are covalently linked to the linker. Further provided herein are compositions in which the first polynucleotide adaptor and the second polynucleotide adaptor are non-covalently linked to the linker. Further provided herein are compositions in which one or more of the first polynucleotide adaptor and the second polynucleotide adaptor comprise at least one barcode.
[0008] Provided herein are adaptor-ligated sample polynucleotides comprising the compositions provided herein, wherein the compositions further comprise at least one sample polynucleotide. Further provided herein are adaptor-ligated sample polynucleotides, wherein at least one sample polynucleotide is attached to the 3'-end of a first strand and the 5'-end of a second strand. Further provided herein are adaptor-ligated sample polynucleotides, wherein at least one sample polynucleotide is attached to the 3'-end of a third strand and the 5'-end of a fourth strand. Further provided herein are adaptor-ligated sample polynucleotides, wherein at least one sample polynucleotide is attached to the 3'-end of a first strand and the 3'-end of a second strand. Further provided herein are adaptor-ligated sample polynucleotides, wherein at least one sample polynucleotide is attached to the 5'-end of a third strand and the 5'-end of a fourth ... the adaptor-ligated sample polynucleotides comprise genomic DNA. Further provided herein are adaptor-ligated sample polynucleotides, wherein the adaptor-ligated sample polynucleotides comprise cDNA. Further provided herein are adaptor-ligated sample polynucleotides, wherein the adaptor-ligated sample polynucleotides comprise cDNA.
[0009] Provided herein is a library of polynucleotides, the library comprising a plurality of adaptor-ligated sample polynucleotides disclosed herein. Provided herein are a plurality of sequencing libraries described herein. Further provided herein are libraries, each library obtained from a different sample. Further provided herein are libraries, in which 70% or less of the sample polynucleotides are present within one standard deviation of the average sample polynucleotide amount. Further provided herein are libraries, in which 50% or less of the sample polynucleotides are present within one standard deviation of the average sample polynucleotide amount. Further provided herein are libraries, in which 25% or less of the sample polynucleotides are present within one standard deviation of the average sample polynucleotide amount.
[0010] Provided herein are methods for preparing a library, the method comprising providing a plurality of sample polynucleotides and ligating at least one composition provided herein to at least one sample polynucleotide. Further provided herein are methods wherein the sample polynucleotides comprise genomic DNA. Further provided herein are methods wherein the sample polynucleotides comprise cDNA. Further provided herein are methods wherein the molar ratio of the at least one composition to the plurality of sample polynucleotides is 1:5 or less. Further provided herein are methods wherein the molar ratio of the at least one composition to the plurality of sample polynucleotides is 1:2 or less. Further provided herein are methods wherein the molar ratio of the at least one composition to the plurality of sample polynucleotides is 1:1 or less. Further provided herein are methods wherein the ligation occurs with at least 25% efficiency. Further provided herein are methods wherein the ligation occurs with at least 50% efficiency. Further provided herein are methods wherein the ligation occurs with at least 75% efficiency. Further provided herein are methods further comprising cleaving the linker. Further provided herein is a method, wherein the step of cleaving the linker comprises contacting the conjugate with an enzyme. Further provided herein is a method, wherein the step of cleaving the enzyme comprises a USER or a site-specific restriction endonuclease.
[0011] Provided herein is a conjugate comprising a first strand, the first strand comprising a first terminal adaptor region, a first non-complementary region, and a first yoke region, and a second strand, the second strand comprising a second terminal adaptor region, a second non-complementary region, and a second yoke region, wherein the first strand and the second strand are connected via a linker. Further provided herein is a conjugate wherein the linker is attached to the 5' end of the first strand and the 5' end of the second strand. Further provided herein is a conjugate wherein the linker is attached to the 3' end of the first strand and the 3' end of the second strand.
[0012] Provided herein are methods for producing the compositions or conjugates provided herein, comprising contacting a first strand with a second strand, wherein the first strand comprises a first reactive group and the second strand comprises a second reactive group, and reaction of the first reactive group with the second reactive group produces the conjugate. Further provided herein are methods wherein the first strand further comprises a first linker. Further provided herein are methods wherein the first strand further comprises a second linker. Further provided herein are methods wherein the first linker comprises a first reactive group. Further provided herein are methods wherein the second linker comprises a second reactive group. Further provided herein are methods wherein the first reactive group is the first reactive group or the second reactive group.
[0013] Provided herein are methods for producing the compositions or conjugates provided herein, comprising: (a) attaching nucleotide monomers to a growing strand on a surface; (b) repeating step (a) to produce a first strand; (c) attaching a linker comprising a reactive group to the end of the first strand; and (d) repeating step (a) to produce a third strand. Also provided herein are methods further comprising cleaving the composition from the solid support. Also provided herein are methods comprising hybridizing a second strand to the first strand. Also provided herein are methods comprising hybridizing a fourth strand to the third strand. [Brief explanation of the drawings]
[0014] [Figure 1A] FIG. 1A shows a fully assembled adapter comprising a 5′-5′ sequence annealed to a reverse complementary sequence to create two adapters covalently linked by a linker comprising a hexaethylene glycol spacer 18 (S18) group, according to an embodiment of the present disclosure. [Figure 1B]FIG. 1B is a diagram showing adapters that can be conjugated to another adapter or fragment thereof, according to an embodiment of the present disclosure, where each adapter includes a primer binding region (stripes), a non-complementary region (solid white), and a yoke region (dots). [Figure 2A] FIG. 2A is a diagram illustrating a 5′ linking polynucleotide that can then be assembled with a reverse complement sequence to create an adapter, according to an embodiment of the disclosure. [Figure 2B] FIG. 2B is a diagram illustrating a 3′ linking polynucleotide that can then be assembled with a reverse complement sequence to create an adapter, according to an embodiment of the disclosure. [Figure 2C] FIG. 2C shows an exemplary method for the synthesis of hybrid cyclic adaptors using "click" chemistry, according to an embodiment of the present disclosure. [Figure 3A] FIG. 3A is a line graph showing simulated ligation of gDNA as a function of adapter dosage for a standard adapter, according to an embodiment of the present disclosure, where the solid line shows relative gDNA conversion as indicated by the amount of adapter used (gDNA amount out) as the amount of adapter dosage (gDNA amount in) increases, and the dotted line represents the point where the adapter and gDNA amounts are equal. [Figure 3B] FIG. 3B is a line graph showing simulated ligation of gDNA as a function of adaptor dosage for hybrid circular adaptors provided herein, according to embodiments of the present disclosure, where the solid line shows relative gDNA conversion as indicated by the amount of adaptor used (gDNA amount out) as the amount of adaptor dosage (gDNA amount in) increases, and the dotted line represents the point where the amounts of adaptor and gDNA are equal. [Figure 4A] FIG. 4A is a bar graph showing read representation as a function of gDNA to adapter ratio, according to an embodiment of the present disclosure; as shown in the graph, as the mass dosage of gDNA increased relative to the dosage of adapter, the read representation remained between 0.6 and 1.3. [Figure 4B]FIG. 4B is a scatter plot graph showing read representation at seven different gDNA mass amounts between 1 and 100 ng, according to an embodiment of the present disclosure; as shown in the graph, read representation remained between 0.6 and 1.3 across the various amounts. [Figure 5] FIG. 5 is a panel of two bioanalyzer read values assessing cleavage of a cleavable base from an exemplary adapter conjugate (e.g., a hybrid cyclic adapter) according to an embodiment of the present disclosure. [Figure 6A] FIG. 6A shows an exemplary adapter having a double-stranded linkage, according to an embodiment of the disclosure. [Figure 6B] FIG. 6B shows an exemplary adapter with an overhang ligation, according to an embodiment of the present disclosure. [Figure 7A] FIG. 7A is a diagram illustrating barcoded adapter pairs matched by adapter (A or B), according to an embodiment of the present disclosure. [Figure 7B] FIG. 7B shows a barcoded adapter pair that is mismatched (crosstalking) with both adapters (A and B) according to an embodiment of the present disclosure. [Figure 7C] FIG. 7C is a bar graph showing the relative abundance of reads with various barcoded adapter pairs, according to an embodiment of the present disclosure. [Figure 8A] FIG. 8A is a panel of six figures showing various exemplary adapters according to embodiments of the present disclosure. [Figure 8B] FIG. 8B is a bar graph showing the percentage of linear competitive barcodes for adapters modified to have overhangs of various lengths and to have various numbers of spacers, according to embodiments of the present disclosure. [Figure 8C] FIG. 8C is a bar graph showing pass-through filter-normalized reads across various mass (ng) doses of gDNA samples, according to an embodiment of the present disclosure. [Figure 9A] FIG. 9A illustrates an exemplary in-line barcode adapter according to an embodiment of the present disclosure. [Figure 9B]FIG. 9B is a bar graph showing pass-through filter-normalized reads across 60 barcoded adapters, each with a unique in-line barcode, according to an embodiment of the present disclosure. [Figure 9C] FIG. 9C is a bar graph showing conversion rates (as measured by normalized matches) during ligation across 12 unique barcoded adapters, according to an embodiment of the present disclosure. [Figure 10A] FIG. 10A is a histogram showing 1,024 possible barcoded adapters, with the left cluster containing 992 mismatched (crosstalk) barcodes and the right cluster containing 32 matched barcodes, according to an embodiment of the present disclosure. [Figure 10B] FIG. 10B is a bar graph showing the amount of crosstalk after pooled ligation in an initial test of three groups of barcoded adapters without subsequent additional steps (none), a heat denaturation step (heat quench), or a bead purification step (beads); as shown in the graph, all three groups exhibited relatively high amounts of crosstalk before any additional steps. [Figure 10C] FIG. 10C is a bar graph showing the amount of crosstalk in the group of FIG. 10B after the additional steps of heat denaturation or bead purification according to an embodiment of the present disclosure; as shown in the graph, the amount of crosstalk was dramatically reduced after heat denaturation and bead purification. [Figure 11A] FIG. 11A is a bar graph showing pass-through filter-normalized reads across six different mass doses of human gDNA samples sequenced after ligation with 12 exemplary adapters, according to an embodiment of the present disclosure. [Figure 11B] FIG. 11B is a bar graph showing pass-through filter-normalized reads across six different mass doses of corn gDNA samples sequenced after ligation with 12 exemplary adapters, according to an embodiment of the present disclosure. [Figure 12] FIG. 12 is a bar graph showing fraction chimeras across six different mass doses of human and maize gDNA samples sequenced after ligation with 12 exemplary adapters, according to an embodiment of the present disclosure. [Figure 13] FIG. 13 is a bar graph showing average target coverage after target enrichment across six different mass doses of human and corn gDNA samples sequenced after ligation with 12 exemplary adapters, according to an embodiment of the present disclosure. [Figure 14] FIG. 14 is a bar graph showing uniformity after target enrichment (multiple 80 base penalty) across six different mass doses of human and maize gDNA samples sequenced after ligation with 12 exemplary adapters, according to an embodiment of the present disclosure. [Figure 15] FIG. 15 is a genome coverage plot showing genome coverage after target enrichment across six different mass doses of human gDNA samples sequenced after ligation with 12 exemplary adapters, according to an embodiment of the present disclosure. [Figure 16] FIG. 16 is a genome coverage plot showing genome coverage after target enrichment across six different mass doses of maize gDNA samples sequenced after ligation with 12 exemplary adapters, according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0015] Provided herein are adaptor conjugates and compositions comprising the same. Further provided herein are compositions and methods for generating synthetic polynucleotide libraries. Further provided herein are compositions and methods for polynucleotide sample normalization.
[0016] definition The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of any embodiments.
[0017] Throughout this disclosure, numerical characteristics are presented in range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of any embodiment. Accordingly, the description of a range should be considered to have specifically disclosed all possible subranges and individual numerical values within that range, to the tenth of the unit of the lower limit, unless the context clearly dictates otherwise. For example, description of a range such as 1 to 6 should be considered to have specifically disclosed subranges such as 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, etc., as well as individual values within that range, e.g., 1.1, 2, 2.3, 5, and 5.9. This applies regardless of the breadth of the range. The upper and lower limits of these intervening ranges may independently be included in the smaller ranges and are also encompassed within the invention, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding one or both of those included limits are also included in the invention, unless the context clearly dictates otherwise.
[0018] As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly dictates otherwise. Furthermore, it will be understood that the terms "comprises" and / or "comprising," as used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0019] Unless otherwise specified or clear from the context, as used herein, the term "about" in reference to a number or range of numbers is intended to mean the specified number and number + / - 10%, or 10% below the recited lower limit and 10% above the recited upper limit for the recited values for a range.
[0020] As used herein, the terms "polynucleotide," "oligonucleotide," "oligo," "oligonucleic acid," and "nucleic acid molecule" are used interchangeably. These terms refer to nucleic acid molecules (e.g., DNA, RNA) that can be single-stranded, double-stranded, or triple-stranded. Double-stranded and triple-stranded nucleic acid molecules need not be coextensive; for example, a double-stranded nucleic acid molecule need not be double-stranded along the entire length of both strands. Unless otherwise specified, the sequence of a nucleic acid molecule, if provided, is listed in the 5' to 3' direction. The length of a nucleic acid molecule, if provided, is given as the number of bases and can be abbreviated, for example, as nt (nucleotides), bp (base pairs), kb (kilobases), Mb (megabases), or Gb (gigabases). The methods described herein provide for the production of isolated nucleic acid molecules. The methods described herein further provide for the production of isolated and purified nucleic acid molecules. As described herein, a polynucleotide can encode a gene or gene fragment from an organism, including but not limited to prokaryotes (e.g., bacteria) and eukaryotes (e.g., mice, rabbits, humans, and non-human primates).
[0021] As used herein, the terms "synthetic" or "synthesized" refer to polynucleotides that are chemically synthesized and / or synthesized de novo.
[0022] As used herein, the term "library" refers to a library of synthetic polynucleotides described herein, which may include a plurality of polynucleotides that collectively encode one or more genes or gene fragments. In some examples, a polynucleotide library includes coding or non-coding sequences. In some examples, a polynucleotide library encodes a plurality of circular DNA (cDNA) sequences. The reference gene sequence on which the cDNA sequences are based may include introns, but the cDNA sequences exclude introns. In some examples, a polynucleotide library includes one or more polynucleotides, each of which encodes a sequence of multiple exons. Each polynucleotide in a library described herein may encode a different (e.g., non-identical) sequence. In some examples, each polynucleotide in a library described herein includes at least one portion that is complementary to the sequence of another polynucleotide in the library. The polynucleotide libraries described herein can contain at least 10, 20, 50, 100, 200, 500, 1,000, 2,000, 5,000, 10,000, 20,000, 30,000, 50,000, 100,000, 200,000, 500,000, 1,000,000, or more than 1,000,000 polynucleotides. The polynucleotide libraries described herein can have fewer than 10, 20, 50, 100, 200, 500, 10,000, 20,000, 30,000, 50,000, 100,000, 200,000, 500,000, or 1,000,000 polynucleotides. The polynucleotide libraries described herein can contain 10 to 500, 20 to 1000, 50 to 2000, 100 to 5000, 500 to 10,000, 1,000 to 5,000, 10,000 to 50,000, 100,000 to 50,000, or 50,000 to 1,000,000 polynucleotides. The polynucleotide libraries described herein can contain greater than about 370,000, 400,000, 500,000, or 500,000 different polynucleotides.
[0023] As used herein, the terms "preselected sequence," "predefined sequence," and "predetermined sequence" are used interchangeably. These terms refer to the sequence of a polymer that is known and selected prior to the synthesis or assembly of the polymer. In particular, various aspects of the invention are described herein with respect to the preparation of nucleic acid molecules, where the sequence of the polynucleotide is known and selected prior to the synthesis or assembly of the nucleic acid molecule.
[0024] As used herein, the terms "electrophile group," "electrophile," and the like refer to an atom or group of atoms that can accept an electron pair to form a covalent bond. The term "electrophile" includes, but is not limited to, halide-containing compounds, carbonyl-containing compounds, and epoxide-containing compounds. Common electrophiles include halides (e.g., thiophosgene, glycerin dichlorohydrin, phthaloyl chloride, succinyl chloride, chloroacetyl chloride, chlorosuccinyl chloride, etc.), ketones (e.g., chloroacetone, bromoacetone, etc.), aldehydes (e.g., glyoxal, etc.), isocyanates (e.g., hexamethylene diisocyanate, tolylene diisocyanate, meta-xylylene diisocyanate, cyclohexylmethane-4,4-diisocyanate, etc.), and derivatives of these compounds.
[0025] As used herein, the terms "nucleophilic group," "nucleophile," and the like refer to an atom or group of atoms having a pair of electrons capable of forming a covalent bond. This type of group can be an ionizable group that reacts as an anionic group. The term "nucleophilic group" includes, but is not limited to, hydroxyl, primary amine, secondary amine, tertiary amine, and thiol.
[0026] Polynucleotide Adapter Conjugates Conjugation can occur by linking two polynucleotide molecules (e.g., two adapters) by hybridization, chemical coupling of different groups, etc., or enzymatically. In some examples, the adapters are conjugated using solid-phase synthesis. In some examples, the adapters are conjugated via a nucleotide coupling process. In some examples, a linker is added during one or more steps of the nucleotide coupling process to generate the conjugate. In some examples, a reverse amidite (e.g., a phosphoramidite at the 3' position) is used. For example, a series of nucleotide monomers are added in a 3' to 5' direction to form a first polynucleotide, a linker is added, and then a series of reverse amidites are added to form a 5' to 3' second polynucleotide.
[0027] In some examples, the universal adapter is conjugated. As shown in FIG. 1B, in some examples, the universal adapter disclosed herein can include a universal polynucleotide adapter 100 including a first strand 101a and a second strand 101b. In some examples, the first strand 101a includes a first primer binding region 102a, a first non-complementary region 103a, and a first yoke region 104a. In some examples, the second strand 101b includes a second primer binding region 102b, a second non-complementary region 103b, and a second yoke region 104b. In some examples, the primer (e.g., 102a / 102b) binding region enables PCR amplification of the universal polynucleotide adapter 100. In some examples, the primer (e.g., 102a / 102b) binding region enables PCR amplification of the polynucleotide adapter 100 and simultaneous addition of one or more barcodes to the polynucleotide adapter. In some examples, the first yoke region 104a is complementary to the second yoke region 104b. In some examples, the first non-complementary region 103a is not complementary to the second non-complementary region 103b. In some examples, the universal polynucleotide adaptor 100 is a Y-shaped adaptor or a forked adaptor. In some examples, one or more yoke regions contain nucleobase analogs that increase the Tm between the first yoke region and the second yoke region. The primer binding region described herein can be in the form of a terminal adaptor region of the polynucleotide. In some examples, the universal adaptor contains one index sequence. In some examples, the universal adaptor contains one unique molecular identifier (UMI).
[0028] Conjugation (e.g., of two adapters) can occur by reacting a nucleophilic reactive group of one adapter with an electrophilic reactive group of another adapter. In some examples, a first adapter and a second adapter are linked by reacting a nucleophilic reactive moiety on the first adapter with an electrophilic reactive moiety on the second adapter. In some examples, the first adapter comprises a linker, the linker comprises a reactive group, and the second adapter comprises a complementary reactive group. Non-limiting examples of nucleophilic reactive groups include amino, thiol, and hydroxyl. Non-limiting examples of electrophilic reactive groups include carboxyl, acyl chloride, anhydride, ester, succinimide ester, alkyl halide, sulfonate ester, maleimide, haloacetyl, and isocyanate. In some examples, the adapter comprises a reactive group. In some examples, the adapter comprises a linker. In some examples, the linker comprises a reactive group. In some examples, the linker connects the 5' end of the first polynucleotide adaptor (or fragment thereof) to the 5' end of the second polynucleotide adaptor (or fragment thereof). In some examples, the linker connects the 3' end of the first polynucleotide adaptor (or fragment thereof) to the 3' end of the second polynucleotide adaptor (or fragment thereof). In some examples, the linker comprises a polynucleotide. In some examples, the first polynucleotide adaptor and the second polynucleotide adaptor are covalently attached to the linker. In some examples, the first polynucleotide adaptor and the second polynucleotide adaptor are non-covalently attached to the linker.
[0029] Provided herein are conjugates comprising one or more polynucleotide strands. In some examples, the conjugate comprises a first strand. In some examples, the conjugate comprises a second strand. In some examples, the conjugate comprises a first strand and a second strand. In some examples, the first strand comprises one or more of a first terminal adaptor region, a first non-complementary region, and a first yoke region. In some examples, the second strand comprises one or more of a second terminal adaptor region, a second non-complementary region, and a second yoke region. In some examples, the first strand and the second strand are connected via a linker. In some examples, the linker is attached to the 5' end of the first strand and the 5' end of the second strand. In some examples, the linker is attached to the 3' end of the first strand and the 3' end of the second strand.
[0030] Compositions comprising ligated adaptors (adapter conjugates) are provided herein. In some examples, the compositions comprise a first polynucleotide adaptor and a second polynucleotide adaptor. In some examples, the first polynucleotide adaptor comprises one or more first strands. In some examples, the first polynucleotide adaptor comprises one or more second strands. In some examples, the first polynucleotide adaptor comprises a first strand and a second strand. In some examples, the first strand comprises one or more of a first non-complementary region and a first yoke region. In some examples, the second strand comprises one or more of a second non-complementary region and a second yoke region. In some examples, the first yoke region and the second yoke region are complementary, and the first non-complementary region and the second non-complementary region are not complementary. In some examples, the second polynucleotide adaptor comprises one or more third strands. In some examples, the second polynucleotide adaptor comprises one or more fourth strands. In some examples, the second polynucleotide adapter comprises a third strand and a fourth strand. In some examples, the third strand comprises one or more of a third non-complementary region and a third yoke region. In some examples, the fourth strand comprises one or more of a fourth non-complementary region and a fourth yoke region. In some examples, the third yoke region and the fourth yoke region are complementary, and the third non-complementary region and the fourth non-complementary region are not complementary. In some examples, the first polynucleotide adapter and the second polynucleotide adapter are connected via a linker. In some examples, the linker is attached to the first terminal adapter region and the third terminal adapter region. In some examples, the linker is attached to the second terminal adapter region and the fourth terminal adapter region. In some examples, the adapter comprises at least one barcode. In some cases, the barcode comprises one or more of a cell index, a sample index, or a unique molecular identifier (UMI). In some examples, the adaptor comprises at least one exonuclease resistant base. In some examples, the at least one exonuclease resistant base comprises an alpha thiophosphate.In some examples, the 5' of at least one strand comprises a thymidine. In some examples, the composition comprises two or more adaptor conjugates.
[0031] In some embodiments, conjugation can be achieved through the use of organosilanes, such as aminosilanes treated with glutaraldehyde, carbonyldiimidazole (CDI) activation of silanol groups, or dendrimers. Various dendrimers are known in the art, including poly(amidoamine) (PAMAM) dendrimers synthesized by a divergent method starting from ammonia or ethylenediamine initiator core reagents; a subclass of PAMAM dendrimers based on a tris-aminoethylene-imine core; radial layered poly(amidoamine-organosilicon) dendrimers (PAMAMOS), which are inverted unimolecular micelles consisting of a hydrophilic nucleophilic polyamidoamine (PAMAM) interior and a hydrophobic organosilicon (OS) exterior; poly(propyleneimine) (PPI) dendrimers, which are generally polyalkylamines with primary amines as terminal groups, but the dendrimer interior consists of multiple tertiary tris-propionamines; and poly(propyleneamine) (POP) dendrimers. These include AM) dendrimers, diaminobutane (DAB) dendrimers, amphiphilic dendrimers, micellar dendrimers, which are unimolecular micelles of water-soluble hyperbranched polyphenylenes, polylysine dendrimers, and dendrimers based on poly-benzyl ether hyperbranched backbones.
[0032] In some embodiments, the conjugation can be carried out by olefin metathesis. In some embodiments, the first adaptor, the second adaptor, and / or the linker (L) comprise an alkene or alkyne moiety that can undergo metathesis. In some embodiments, a suitable catalyst (e.g., copper, ruthenium) is used to accelerate the metathesis reaction. Suitable methods for carrying out olefin metathesis reactions are described in the art. For example, see Schafmeister et al., J Am Chem Soc, 2000, Vol. 122, Pages 5891-5892; Walensky et al., Science, 2004, Vol. 305, Pages 1466-1470; and Blackwell et al., Angew Chem Int Ed., 1998, Vol. 37, Pages 3281-3284.
[0033] In some embodiments, conjugation can be carried out using click chemistry. The "click reaction" is broad in scope, easy to perform, uses only readily available reagents, and is insensitive to oxygen and water. In some embodiments, the click reaction is a cycloaddition reaction between an alkynyl group and an azide group to form a triazolyl group. In some embodiments, the click reaction uses a copper or ruthenium catalyst. Suitable methods for performing click reactions are described in the art. For example, see Kolb et al., Drug Discovery Today, 2003, Vol. 8, Pages 1128-1137; Kolb et al., Angew Chem Int Ed., 2001, Vol. 40, Pages 2004-2021; Rostovtsev et al., Angew. Chem. Int. Ed. 41:2596 (2002); Tornoe et al., J Org Chem., 2002, Vol. 67, Pages 3057-3064; Manetsch et al., J Am Chem Soc., 2004, Vol. 126, Pages 12809-12818; Lewis et al., Angew. Chem. Int. Ed. 41:1053 (2002); Speers et al., J Am Chem Soc., 2003, Vol. 125, Pages See, e.g., Chan et al., Organic Letters, 2004, Vol. 6, Pages 2853-2855; Zhang et al., J Am Chem Soc., 2005, Vol. 127, Pages 15998-15999; and Waser et al., J Am Chem Soc., 2005, Vol. 127, Pages 8294-8295. In some instances, the click reaction is performed "copper-free" using strained olefins such as trans-cyclooctene. Indirect conjugation via high-affinity specific binding partners (e.g., streptavidin / biotin, avidin / biotin, or lectin / carbohydrate) is also contemplated.
[0034] The adaptors can be connected via a linker, L, for example, in the formula adaptor-Z-adaptor. In some embodiments, Z is a linking group. In some embodiments, Z is a bifunctional linker and contains only two reactive groups prior to conjugation to the adaptor. In embodiments in which both adaptors have electrophilic reactive groups, Z contains two identical or two different nucleophilic groups (e.g., amine, hydroxyl, or thiol) prior to conjugation to the adaptor. In embodiments in which both the first adaptor and the second adaptor have nucleophilic reactive groups, Z contains two identical or two different electrophilic groups (e.g., carboxyl groups, activated forms of carboxyl groups, or compounds with leaving groups) prior to conjugation to the adaptor. In embodiments in which the first adaptor contains a nucleophilic reactive group and the second adaptor contains an electrophilic reactive group, Z contains one nucleophilic reactive group and one electrophilic group prior to conjugation to the adaptor.
[0035] A linker can be any molecule having at least one reactive group (prior to conjugation to the adapters) that can react with each of the adapters. In some embodiments, the linker has only reactive groups and is bifunctional. Prior to conjugation to the adapters, the linker can be represented by the formula ALB, where L and B are independently nucleophilic or electrophilic reactive groups. In some embodiments, A and B are both nucleophilic reactive groups. In some embodiments, A and B are both electrophilic reactive groups. In some embodiments, A is a nucleophilic reactive group and B is an electrophilic reactive group. In some embodiments, A is an electrophilic reactive group and B is a nucleophilic reactive group.
[0036] In some embodiments, A and B may contain alkene and / or alkyne functional groups suitable for olefin metathesis reactions. In some embodiments, A and B contain moieties suitable for click chemistry (e.g., alkene, alkyne, nitrile, or azide moieties). Other non-limiting examples of reactive groups (A and B) include pyridyldithiols, aryl azides, diazirines, carbodiimides, and hydrazides.
[0037] In some embodiments, the linker is hydrophobic.Hydrophobic linkers or linking groups are known in the art.See, for example, Bioconjugate Techniques, GT Hermanson (Academic Press, San Diego, CA, 1996), which is incorporated herein by reference in its entirety.Suitable hydrophobic linking groups known in the art include, for example, 8-hydroxyoctanoic acid and 8-mercaptooctanoic acid.Before conjugation, the hydrophobic linker comprises at least two reactive groups (A and B), as shown in the formula A-(hydrophobic linking group)-B, as described herein.
[0038] In some embodiments, the hydrophobic linker comprises either a maleimide group or an iodoacetyl group and either a carboxylic acid or an activated carboxylic acid (e.g., an NHS ester) as the reactive group. In these embodiments, the maleimide group or iodoacetyl group can be coupled to a thiol moiety on a first adaptor, and the carboxylic acid or activated carboxyl group can be coupled to an amine on a second adaptor, with or without the use of a coupling agent. For example, the carboxylic acid can be coupled to a free amine using any coupling agent known to those skilled in the art, such as DCC, DIC, HATU, HBTU, TBTU, and other activating agents described herein. In some embodiments, the hydrophilic linking group comprises an aliphatic chain of 2 to 100 methylene groups, and A and B are carboxyl groups or derivatives thereof (e.g., succinic acid). In some embodiments, L is iodoacetic acid.
[0039] In some examples, prior to conjugation to the adaptor, the hydrophilic linking group comprises at least two reactive groups (A and B), as described herein and shown below: A-(hydrophilic linking group)-B. In certain embodiments, the linking group comprises polyethylene glycol (PEG). In some examples, a PEG "unit" comprises -(OCH2CH2)-. In certain embodiments, the PEG has a molecular weight of about 100 daltons to about 10,000 daltons, e.g., about 500 daltons to about 5000 daltons. In some embodiments, the PEG has a molecular weight of about 10,000 daltons to about 40,000 daltons. In some examples, the PEG-containing linker comprises an Sp3, Sp9, Sp12, or Sp18 (hexaethylene glycol spacer 18, comprising 18 PEG units) group. In some examples, the PEG-containing linker has a molecular weight of about 1, 2, 3, 4, 5, 6, 8, In some examples, the PEG-containing linker comprises 10, 12, 14, 16, 18, 20, 22, 24, 30, 40, or about 50 PEG units. In some examples, the PEG-containing linker comprises 1-1000, 1-500, 1-250, 1-200, 1-150, 1-100, 1-80, 1-75, 1-60, 1-50, 1-40, 1-30, 1-25, 1-20, 5-1000, 5-500, 5-250, 5-200, 5-80, 5-75, 5-60, 5-50, 5-40, 5-30, 5-25, 5-20, 10-1000, 10-500, 10-250, 10 Includes units of ~200, 10~150, 10~100, 10~80, 10~75, 10~60, 10~50, 10~40, 10~30, 10~25, 10~20, 12~50, 12~40, 12~30, 12~20, 15~50, 15~30, 15~25, 15~20, 17~19, 17~25, 17~35, 18~50, 18~30, 18~25, 20~30, 20~50, 30~50, or 50~100. In some examples, the PEG-containing linker comprises 1 or less, 2 or less, 3 or less, 4 or less, 5 or less, 6 or less, 8 or less, 10 or less, 12 or less, 14 or less, 16 or less, 18 or less, 20 or less, 22 or less, 24 or less, 30 or less, 40 or less, or 50 or less PEG units. In some examples, the PEG-containing linker comprises at least 1, 2, 3, 4, 5, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 30, 40, or at least 50 PEG units.
[0040] In some embodiments, the hydrophilic linking group comprises either a maleimide group or an iodoacetyl group and either a carboxylic acid or an activated carboxylic acid (e.g., an NHS ester) as the reactive group. In these embodiments, the maleimide group or iodoacetyl group can be coupled to a thiol moiety on an adaptor, and the carboxylic acid or activated carboxylic acid can be coupled to an amine on another adaptor or linker, with or without the use of a coupling reagent. The carboxylic acid can be coupled to the amine using any suitable coupling agent known to those skilled in the art, such as DCC, D1C, HATU, HBTU, TBTU, and other activating agents described herein. In some embodiments, the linking group is maleimide-polymer(0.1-2.5 kDa)-COOH, iodoacetyl-polymer(0.1-2.5 kDa)-COOH, maleimide-polymer(0.1-2.5 kDa)-NHS, or iodoacetyl-polymer(0.1-2.5 kDa)-NHS.
[0041] In some embodiments, the linker comprises an amino acid, a dipeptide, a tripeptide, or a polypeptide, wherein the amino acid, dipeptide, tripeptide, or polypeptide comprises at least two activating groups described herein. In some embodiments, the linker comprises a moiety selected from the group consisting of amino, ether, thioether, maleimide, disulfide, amide, ester, thioester, alkene, cycloalkene, alkyne, triazole, carbamate, carbonate, cathepsin B cleavable, and hydrazone.
[0042] In some embodiments, the linker comprises a chain of atoms from 1 to about 60, 1 to about 30, 10 to 20, 2 to 10, 2 to 5, or 5 to 10 atoms long. In some embodiments, all of the chain atoms are carbon atoms. In some embodiments, the chain atoms in the backbone of the linker are selected from the group consisting of C, O, N, and S, with the chain atoms and linker being selected according to their expected solubility (hydrophilicity), in some instances to provide a more soluble conjugate. In some embodiments, L provides a functional group that is subject to cleavage by an enzyme or other catalyst or to hydrolysis conditions found in a target tissue, organ, or cell. In some embodiments, the length of L is long enough to reduce the possibility of steric hindrance.
[0043] In some examples, a suitable polymer backbone has the formula X-polymer-LY, where the polymer is poly(ethylene glycol), X is a functional group that does not react with azide groups, and Y is a suitable leaving group. Examples of suitable functional groups include, but are not limited to, hydroxyl, protected hydroxyl, acetal, alkenyl, amine, aminooxy, protected amine, protected hydrazide, protected thiol, carboxylic acid, protected carboxylic acid, maleimide, dithiopyridine, and vinylpyridine, and ketone. Examples of suitable leaving groups include, but are not limited to, chloride, bromide, iodide, mesylate, tresylate, and tosylate.
[0044] Linkers can have a wide range of molecular weights or lengths. Linkers with larger or smaller molecular weights can be used to provide the desired spatial relationship or conformation between the adapter and the linked entity (e.g., the second adapter). Linkers with longer or shorter molecular lengths can also be used to provide the desired space or flexibility between the adapter and the linked entity.
[0045] In some embodiments, the linker comprises a water-soluble bifunctional linker having a dumbbell structure comprising: a) an azide, alkyne, hydrazine, hydrazide, hydroxylamine, or carbonyl-containing moiety on at least a first end of the polymer backbone; and b) at least a second functional group on a second end of the polymer backbone. The second functional group may be the same as or different from the first functional group. In some embodiments, the second functional group does not react with the first functional group. In some embodiments, the water-soluble compound comprises at least one arm of a branched molecular structure. For example, the branched molecular structure may be dendritic.
[0046] In exemplary embodiments, the polymer is linked to the adaptor via a linker. For example, the linker may comprise one or two amino acids attached at one end to the polymer (such as an albumin-binding moiety) and at the other end to any available position on the polypeptide backbone. Further exemplary linkers include hydrophilic linkers, such as chemical moieties containing at least five non-hydrogen atoms, 30-50% of which are either N or O.
[0047] In some embodiments, the adaptors are linked by a polypeptide linker, which in some embodiments is one or more (e.g., 1, 1-2, 1-3, 1-4, 1-5, 1-6, 1-7, 1-8, 1-9, 1-10, 1-11, 1-12) amino acids in length or longer.
[0048] Generally, carbon electrophiles are susceptible to attack by complementary nucleophiles, including carbon nucleophiles, which donate an electron pair to the carbon electrophile to form a new bond between the nucleophile and the carbon electrophile. Non-limiting examples of carbon nucleophiles include, but are not limited to, alkyl, alkenyl, aryl, and alkynyl Grignards, organolithium, organozinc, alkyl-alkenyl, aryl-, and alkynyl-tin reagents (organostannanes), and alkyl-, alkenyl-, aryl-, and alkynyl-borane reagents (organoboranes and organoborates). These carbon nucleophiles have the advantage of being kinetically stable in water or polar organic solvents. Other non-limiting examples of carbon nucleophiles include phosphorus hydrides, enol, and enolate reagents. These carbon nucleophiles have the advantage of being relatively easy to generate from precursors well known to those skilled in the art of synthetic organic chemistry. Carbon nucleophiles, when used in conjunction with carbon electrophiles, generate new carbon-carbon bonds between the carbon nucleophile and the carbon electrophile. Non-limiting examples of non-carbon nucleophiles suitable for coupling to carbon electrophiles include, but are not limited to, primary and secondary amines, thiols, thiolates, and thioethers, alcohols, alkoxides, azides, semicarbazides, and the like. These non-carbon nucleophiles, when used in conjunction with carbon electrophiles, typically generate heteroatom linkages (CXC), where X is a heteroatom, including, but not limited to, oxygen, sulfur, or nitrogen.
[0049] In some cases, the polymers used herein terminate at one end with hydroxy or methoxy, i.e., X is H or CH ("methoxy PEG"). Alternatively, the polymer may terminate with a reactive group, thereby forming a bifunctional polymer. Exemplary reactive groups may include reactive groups commonly used to react with functional groups found on the 20 common amino acids (including, but not limited to, maleimide groups, activated carbonates (including, but not limited to, p-nitrophenyl esters), activated esters (including, but not limited to, N-hydroxysuccinimide, p-nitrophenyl esters), and aldehydes), as well as functional groups that are inert toward the 20 common amino acids but specifically react with complementary functional groups (including, but not limited to, azide groups, alkyne groups). Note that the other end of the polymer, designated Y in the above formula, is directly or indirectly attached to an adaptor. Alternatively, an alkyne group on the polymer can react with an azide group present on the adaptor. In some embodiments, strong nucleophiles (including, but not limited to, hydrazine, hydrazide, hydroxylamine, semicarbazide) can react with aldehyde or ketone groups present on the adapter to form hydrazones, oximes, or semicarbazones, if applicable, which can be further reduced by treatment with an appropriate reducing agent. Alternatively, strong nucleophiles can be incorporated into the adapter and used to preferentially react with ketone or aldehyde groups present in the water-soluble polymer.
[0050] The activated ester of a carboxylic acid can be, for example, N-hydroxysuccinimide (NHS), tosylate (Tos), mesylate, triflate, carbodiimide, or hexafluorophosphate. In some embodiments, the carbodiimide is 1,3-dicyclohexylcarbodiimide (DCC), 1,1'-carbonyldiimidazole (CDI), 1-ethyl-3-(3-dimethylaminopropyl)carbodiimide hydrochloride (EDC), or 1,3-diisopropylcarbodiimide (DICD). In some embodiments, the hexafluorophosphate is hexafluorophosphate benzotriazol-1-yl-oxy-tris(dimethylamino)phosphatase. phosphonium hexafluorophosphate (BOP), benzotriazol-1-yl-oxytripyrrolidinophosphonium hexafluorophosphate (PyBOP), 2-(1H-7-azabenzotriazol-1-yl)-1,1,3,3-tetramethyl)uranium hexafluorophosphate (HATU), and o-benzotriazole-N,N,N',N'-tetramethyl-uronium-hexafluorophosphate (HBTU).
[0051] Any molecular weight of the polymer can be used as practically desired, including, but not limited to, from about 0.1 Daltons (Da) to 2,500 Da or more. The molecular weight of the polymer can range widely, including, but not limited to, from about 100 Da to about 5,000 Da or more. In some examples, the polymer is 50-5,000 Da, 50-3,000 Da, 50-2,500 Da, 100-2,500 Da, 250-2,500 Da, 250-5,000 Da, or 500-5,000 Da. Branched chain polymers include, but are not limited to, polymer molecules in which each chain has a molecular weight ranging from 0.1-5 kDa, 0.1-4 kDa, 0.1-3 kDa, 0.1-2.5 kDa, or 0.1-1.5 kDa.
[0052] The polymers may include azide- and acetylene-containing polymer derivatives containing a water-soluble polymer backbone with an average molecular weight of about 800 Da to about 100,000 Da. The polymer backbone of the water-soluble polymer may be polyethylene glycol. However, it should be understood that a wide variety of water-soluble polymers may also be used, including, but not limited to, poly(ethylene) glycol and other related polymers, including poly(dextran) and polypropylene glycol, and the use of the term PEG or poly(ethylene glycol) is intended to encompass and include all such molecules. The term PEG includes, but is not limited to, poly(ethylene glycol) in any of its forms, including bifunctional PEG, multi-armed PEG, derivatized PEG, forked PEG, branched PEG, pendant PEG (i.e., PEG or related polymers with one or more functional groups pendant to the polymer backbone), or PEG with degradable linkages therein.
[0053] In addition to these forms of polymers, polymers can also be prepared with weak or degradable bonds in the backbone. For example, polymers can be prepared with ester bonds in the polymer backbone that undergo hydrolysis. This hydrolysis results in the cleavage of the polymer into lower molecular weight fragments, as shown below: -polymer-CO2-polymer-+HOapolymer-CO2H+HO-polymer-.
[0054] The linker may comprise a polymer, such as one comprising a water-soluble backbone. In some embodiments, the water-soluble polymer backbone comprises from 2 to about 300 termini. Examples of suitable polymers include, but are not limited to, other poly(alkylene glycols), such as poly(propylene glycol) ("PPG"), their copolymers (including, but not limited to, copolymers of ethylene glycol and propylene glycol), their terpolymers, mixtures thereof, and the like. The molecular weight of each chain of the polymer backbone can vary, but typically ranges from about 800 Da to about 100,000 Da, often from about 6,000 Da to about 80,000 Da. The molecular weight of each chain of the polymer backbone can be from about 100 Da to about 100,000 Da, and can be 100,000 Da, 95,000 Da, 90,000 Da, 85,000 Da, 80,000 Da, 75,000 Da, 70,000 Da, 65,000 Da, 60,000 Da, 55,000 Da, 50,000 Da, 45,000 Da, 40,000 Da, 35,000 Da, 30,000 Da, 25,000 Da, 30,000 Da, 45,000 Da, 40,000 Da, 55,000 Da, 50,000 Da, 45,000 Da, 40,000 Da, 55,000 Da, 50,000 Da, 55,000 Da, 50,000 Da, 65,000 Da, 60,000 Da, 65,000 Da, 60,000 Da, 75,000 Da, 70,000 Da, 85,000 Da, 80,000 Da, 80,000 Da, 95,000 Da, 90,000 Da, 95,000 Da, 90,000 Da, 10 ... Examples of molecular weights include, but are not limited to, 00 Da, 20,000 Da, 15,000 Da, 10,000 Da, 9,000 Da, 8,000 Da, 7,000 Da, 6,000 Da, 5,000 Da, 4,000 Da, 3,000 Da, 2,000 Da, 1,000 Da, 900 Da, 800 Da, 700 Da, 600 Da, 500 Da, 400 Da, 300 Da, 200 Da, and 100 Da. In some embodiments, the molecular weight of each chain of the polymer backbone is from about 100 Da to about 50,000 Da. In some embodiments, the molecular weight of each chain of the polymer backbone is from about 100 Da to about 40,000 Da. In some embodiments, the molecular weight of each chain of the polymer backbone is from about 1,000 Da to about 40,000 Da. In some embodiments, the molecular weight of each chain of the polymer backbone is from about 5,000 Da to about 40,000 Da. In some embodiments, the molecular weight of each chain of the polymer backbone is from about 10,000 Da to about 40,000 Da.
[0055] In some examples, the adaptor is linked via a water-soluble polymer via the methods described herein. In some embodiments, the method includes contacting an adaptor comprising a reactive amino acid side chain with a linker. In some examples, the conjugate is synthesized by reacting a functional group present on the adaptor with a reactive group present on the linker. In some examples, the adaptor conjugate is synthesized by reacting a functional group present on the linker with a reactive group present on the adaptor.
[0056] In some embodiments, the linker has a molecular weight of 0.1 kDa to 5 kDa. In some embodiments, the linker has a molecular weight of 0.1 kDa to 2.5 kDa. In some embodiments, the linker or polymer is linear, branched, multimeric, or dendrimeric. In some embodiments, the linker or polymer is a bifunctional or multifunctional linker or a bifunctional or multifunctional polymer.
[0057] In other embodiments, the polymer is a water-soluble polymer. In other embodiments, the water-soluble polymer is polyethylene glycol (PEG). In some embodiments, the PEG has a molecular weight between 0.1 kDa and 10 kDa. In other embodiments, the PEG has a molecular weight between 0.1 kDa and 5 kDa. In other embodiments, the PEG has a molecular weight between 0.1 kDa and 4 kDa. In other embodiments, the PEG has a molecular weight between 0.1 kDa and 3 kDa. In other embodiments, the PEG has a molecular weight between 0.1 kDa and 2 kDa. In other embodiments, the PEG has a molecular weight between 0.1 kDa and 2.5 kDa. In some embodiments, the polyethylene glycol molecule has a molecular weight between about 0.1 kDa and about 10 kDa. In some embodiments, the poly(ethylene glycol) molecule has a molecular weight between 0.1 kDa and 50 kDa. In some embodiments, the polyethylene glycol has a molecular weight of 0.1 kDa to 2.5 kDa, or 0.2 to 2.2 kDa, or 0.5 kDa to 2 kDa. For example, the molecular weight of the poly(ethylene glycol) polymer is, in some instances, about 0.5 kDa, or about 1 kDa, or about 2 kDa, or about 2.5 kDa. For example, the molecular weight of the poly(ethylene glycol) polymer is, in some instances, 0.1 kDa, 0.5 kDa, 1 kDa, or 2.5 kDa. In some embodiments, the poly(ethylene glycol) molecule is a branched PEG. In some embodiments, the poly(ethylene glycol) molecule is a branched 1K PEG. In some embodiments, the poly(ethylene glycol) molecule is a branched 2.5K PEG. In some embodiments, the poly(ethylene glycol) molecule is a branched 5K PEG. In some embodiments, the poly(ethylene glycol) molecule is a linear PEG. In some embodiments, the polyethylene glycol molecule is a linear 2.5K PEG. In some embodiments, the polyethylene glycol molecule is a linear 10K PEG. In some embodiments, the poly(ethylene glycol) molecule is a linear 2K PEG. In some embodiments, the poly(ethylene glycol) molecule is a linear 0.5K PEG.In some embodiments, the molecular weight of the polyethylene glycol polymer is an average molecular weight. In certain embodiments, the average molecular weight is a number average molecular weight (Mn). The average molecular weight can be determined or measured using GPC or SEC, SDS / PAGE analysis, RP-HPLC, mass spectrometry, or capillary electrophoresis.
[0058] The linker may contain a cleavable base. In some cases, the cleavable base may be removed to separate one or more adapters or portions thereof from each other. In some examples, the cleavable base includes a base that can be enzymatically removed. In some examples, the cleavable base is uracil. In some examples, the cleavable base is removable by USER. The linkers described herein may contain a nucleotide analog recognized by a specific enzyme. In some examples, the supporting linker includes a nucleotide analog. In some examples, the supporting linker includes deoxyuridine or 8-oxo-deoxyguanosine recognized by a specific glycosylase (e.g., uracil deoxyglycosylase followed by endonuclease VIII and 8-oxoguanine DNA glycosylase, respectively). In some embodiments, cleavage by glycosylases and / or endonucleases may require a double-stranded DNA substrate. In some embodiments, the support linker comprises a base analogue cleavable by endonuclease III, including, but not limited to, urea, thymine glycol, methyltartonyl urea, alloxan, uracil glycol, 6-hydroxy-5,6-dihydrocytosine, 5-hydroxyhydantoin, 5-hydroxycytosine, trans-1-carbamoyl-2,-oxo-4,5-dihydroxyimidazolidine, 5,6-dihydrouracil, 5-hydroxycytosine, 5-hydroxyuracil, 5-hydroxy-6-hydrouracil, 5-hydroxy-6-hydrothymine, 5,6-dihydrothymine. In some embodiments, the support linker comprises a base analogue cleavable by formamidopyrimidine DNA glycosylase, including, but not limited to, 7,8-dihydro-8-oxoguanine, 7,8-dihydro-8-oxoinosine, 7,8-dihydro-8-oxoadenine, 7,8-dihydro-8-oxonebularine, 4,6-diamino-5-formamidopyrimidine, 2,6-diamino-4-hydroxy-5-formamidopyrimidine, 2,6-diamino-4-hydroxy-5-N-methylformamidopyrimidine, 5-hydroxycytosine, 5-hydroxyuracil.In some embodiments, the supported linker comprises a base analog cleavable by hNeil 1, including, but not limited to, guanidinohydantoin, spiroiminodihydantoin, 5-hydroxyuracil, and thymine glycol. In some embodiments, the supported linker comprises a base analog cleavable by thymine DNA glycosylase, including, but not limited to, 5-formylcytosine and 5-carboxycytosine. In some embodiments, the supported linker comprises a base analog cleavable by thymine DNA glycosylase, including, but not limited to, 3-methyladenine, 3-methylguanine, 7-methylguanine, 7-(2-chloroethyl)-guanine, 7-(2-hydroxyethyl)-guanine, 7-(2-ethoxyethyl)-guanine, 1,2-bis-(7-guanyl)ethane, 1,N-methyl-2-methyl-3-methyl-4-methyl-5-methyl-6-methyl-7-methyl-8-methyl-9-methyl-10-methyl-11-methyl-12-methyl-13-methyl-14-methyl-15-methyl-16-methyl-17-methyl-18-methyl-19 ... 6 -Ethenoadenine, 1,N 2 -Ethenoguanine, N 2 ,3-Ethenoguanine, N 2 The linker may comprise a base analogue cleavable by human alkyladenine DNA glycosylase, including 3-ethanoguanine, 5-formyluracil, 5-hydroxymethyluracil, and hypoxanthine. In some embodiments, the supported linker comprises a 5-methylcytosine cleavable by 5-methylcytosine DNA glycosylase. In some examples, the linker comprises 1, 2, 3, 4, 5, 6, 8, 10, or more than 10 cleavable bases.
[0059] A linker comprising a polynucleotide may contain multiple cleavable bases. In some examples, the multiple cleavable bases comprise restriction endonuclease (RE) sites. The linker can be cleaved by treatment with an appropriate endonuclease that recognizes the RE site. In some examples, the RE site is unique to the linker portion of the adapter conjugate. In some examples, treatment with a restriction enzyme does not cleave other portions of the conjugate, such as the adapter or sample nucleic acid (if ligated). In some cases, the RE site is present in both the linker and other portions of the adapter / sample nucleic acid, but these other portions are blocked from cleavage (e.g., methylated).
[0060] The linker may be non-covalently attached. In some cases, complementary overlapping regions of one or both adapters and / or the linker are used for attachment. In some examples, the linker comprises a double-stranded region. In some examples, the double-stranded region comprises at least partially complementary strands of nucleic acid. Such linkers can be removed / cleaved by denaturation. In some examples, the linker is formed by overlapping at least a portion of a first adapter with a portion of a second adapter. In some examples, one or more splint polynucleotides are used to generate the linker. In some examples, at least one splint polynucleotide is at least partially complementary to a portion of the first polynucleotide adapter or the second polynucleotide adapter. In some examples, the linker comprises at least two splint polynucleotides, where the at least two splint polynucleotides at least partially overlap with each other. In some examples, the linker comprises at least two splint polynucleotides, where the at least two splint polynucleotides at least partially overlap with a portion of the first polynucleotide adapter or the second polynucleotide adapter.
[0061] Provided herein are adaptor-ligated samples. In some examples, the adaptor-ligated samples comprise one or more adaptors and a sample polynucleotide. In some examples, at least one sample polynucleotide is attached to the 3' end of a first strand and the 5' end of a second strand. In some examples, at least one sample polynucleotide is attached to the 3' end of a third strand and the 5' end of a fourth strand. In some examples, the sample polynucleotide comprises genomic DNA. In some examples, the sample polynucleotide comprises cDNA. In some examples, the sample polynucleotide comprises cDNA. Further provided herein are libraries of adaptor-ligated samples. In some examples, the library comprises a plurality of adaptor-ligated samples. In some examples, each library is obtained from a different sample. Further provided herein are sequencing libraries comprising a plurality of libraries.
[0062] In some embodiments, sequencing libraries can contain or be derived from different amounts of sample doses (sample polynucleotides). In some examples, the use of adapter conjugates described herein results in normalization of sample dose. In some examples, 10, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, or 95% or less of the sample polynucleotides are within one standard deviation of the average sample polynucleotide amount. In some examples, 10-95%, 10-90%, 10-75%, 10-60, 10-50, 20-95, 20-97, 40-95, 50-95, 75-99, 75-90%, or 50-99% or less of the sample polynucleotides are within one standard deviation of the average sample polynucleotide amount.
[0063] After the adapters are generated, they may be resuspended. In some embodiments, the adapters may be resuspended in a buffer. In some embodiments, the buffer comprises Tris-HCl and EDTA·Na2 ("TE buffer"). In some embodiments, the TE buffer is at a 1× concentration. In some embodiments, the adapters are resuspended in 1× TE buffer to a concentration of about 100 μM and mixed to reach a final working concentration of about 10 μM.
[0064] Adapter systems and sample normalization Provided herein are adapter systems and methods of using same that allow for library size normalization without further steps such as further dilution or enzymatic processing. In some embodiments, the incorporation of in-line barcodes into the adapter system allows for increased multiplexing capacity of the barcode system for high-throughput library preparation.
[0065] Further provided herein are methods for sample normalization using the adapters described herein. In some examples, the adapters comprise adapter conjugates. In some examples, multiple samples are processed for sequencing. In some examples, at least some of the multiple samples are derived from different sources. In some examples, at least some of the multiple samples contain different amounts of sample nucleic acid. In some examples, the sample nucleic acid is subjected to one or more steps of shearing / fragmentation (mechanically or using amplification), end repair, tailing, ligation to an adapter conjugate described herein, enrichment / capture, amplification, and sequencing. In some examples, the method includes one or more steps of (a) providing multiple sample polynucleotides, and (b) ligating at least one composition (e.g., an adapter conjugate) described herein to at least one sample polynucleotide. In some examples, the sample polynucleotides comprise genomic DNA. In some examples, the sample polynucleotides comprise cDNA. In some examples, the molar ratio of adapter conjugate to the plurality of sample polynucleotides is 1:10, 1:5, 1:2, 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 10:1, 15:1, 20:1, or 50:1 or less. In some examples, the molar ratio of adapter conjugate to the plurality of sample polynucleotides is 1:10 to 10:1, 1:10 to 5:1, 1:10 to 3:1, 1:10 to 2:1, 1:10 to 1:1, 1:5 to 10:1, 1:5 to 5:1, 1:5 to 2:1, 1:5 to 1:1, 1:3 to 1:5, and 1:3 to 1:1 or less. In some examples, ligation occurs with an efficiency of at least 75%, 80%, 85%, 90%, 92%, 95%, 97%, 98%, or at least 99%. In some examples, ligation occurs at an efficiency of 75-99%, 80-99%, 85-99%, 90-99%, 90-95%, or 95-99%. In some examples, the multiple sample polynucleotides represent a sequencing library. In some examples, the use of adapter conjugates results in a normalized display of signal during sequencing for two or more samples.
[0066] Multiplexing gDNA on a sequencer can improve efficiency and accuracy if the sequencing load is balanced equally across all samples in terms of the number of molecules.For example, some samples may contain relatively high concentrations of gDNA, while other samples may contain relatively low concentrations of gDNA.Normalizing sequencing load can improve the yield and accuracy of sequencing results and reduce sequencing errors.
[0067] In conventional methods, a large excess of adapters is used for library construction. Because conventional methods use two independent ligation products per molecule, conversion cannot be easily controlled by modifying the adapter concentration. In a high-throughput setting, the inability to control conversion by modifying the adapter concentration can pose challenges, requiring qPCR and dilution for each sample.
[0068] This disclosure describes systems and methods that include hybrid circular adapters (FIG. 1A), as opposed to conventional adapters (FIG. 1B). Hybrid circular adapters are generally compatible with standard sequencing workflows without substantial modifications to the workflow. By ligating two adapters, two independent ligation events are replaced with a slower intermolecular ligation and a faster intramolecular ligation. In some embodiments, steric interactions of the adapters can be reduced by a flexible linker. Unlike conventional adapters, ligation conversion can be controlled by the number of adapter molecules to normalize sample volume.
[0069] A comparison of the ligation conversion between a conventional adapter (FIG. 3A) and the adapter system of the present disclosure (FIG. 3B) is shown in the Appendix. Further examples, features, and functionality of hybrid cyclic adapters are described in U.S. Provisional Patent Application No. 63 / 511,086.
[0070] In some embodiments suitable for high-throughput environments, more than about 1024 fragmented gDNA samples can be ligated to equal concentrations of hybrid circular adapters to generate an adapter-ligated polynucleotide library for each sample. The concentration of the hybrid circular adapters can be adjusted based on the minimum sample concentration. In some examples, the samples are enriched (e.g., exome enriched), amplified, and / or subjected to next-generation sequencing. The degree of signal normalization (e.g., counts) can be measured.
[0071] Hybridization and capture A polynucleotide library can be designed to contain polynucleotide sequences that are identical to or complementary to (target, hybridize with) one or more variants. In some examples, at least some of the polynucleotide molecules are each configured to hybridize to a genomic region containing at least two variants. In some examples, at least some of the polynucleotide molecules are each configured to hybridize to a genomic region containing at least one, two, three, four, five, six, or more than six variants. In some examples, at least some of the polynucleotides are each configured to hybridize to a genomic region containing one to four variants. In some examples, at least some of the polynucleotides are each configured to hybridize to a genomic region containing one to two or three variants. In some examples, at least 50% of the polynucleotides are each configured to hybridize to a genomic region containing at least two variants. In some examples, at least 50% of the polynucleotides are each configured to hybridize to a genomic region containing at least one, two, three, four, five, six, or more than six variants. In some examples, at least 50% of the polynucleotides are each configured to hybridize to a genomic region comprising one to four variants. In some examples, at least 50% of the polynucleotides are each configured to hybridize to a genomic region comprising one to two or three variants. In some examples, at least 25% of the polynucleotides are each configured to hybridize to a genomic region comprising at least two variants. In some examples, at least 25% of the polynucleotides are each configured to hybridize to a genomic region comprising at least one, two, three, four, five, six, or more than six variants. In some examples, at least 25% of the polynucleotides are each configured to hybridize to a genomic region comprising one to four variants. In some examples, at least 25% of the polynucleotides are each configured to hybridize to a genomic region comprising one to two or three variants.In some examples, at least 5% of the polynucleotides are each configured to hybridize to a genomic region comprising at least two variants. In some examples, at least 5% of the polynucleotides are each configured to hybridize to a genomic region comprising at least one, two, three, four, five, six, or more than six variants. In some examples, at least 5% of the polynucleotides are each configured to hybridize to a genomic region comprising one to four variants. In some examples, at least 5% of the polynucleotides are each configured to hybridize to a genomic region comprising one to two or three variants.
[0072] The polynucleotide library can be configured to bind to a large number of variants. In some examples, the polynucleotide library is collectively configured to bind to a genomic region that contains about 50, 100, 200, 500, 800, 1000, 2000, 5000, 8000, 10,000, 20,000, 50,000, 80,000, 100,000, 250,000, 500,000, 750,000, 1 million, 1.5 million, 2 million, 2.5 million, 3 million, 3.5 million, 4 million, 4.5 million, or about 5 million variants. In some examples, the polynucleotide libraries are collectively configured to bind to genomic regions that include at least 50, 100, 200, 500, 800, 1000, 2000, 5000, 8000, 10,000, 20,000, 50,000, 80,000, 100,000, 250,000, 500,000, 750,000, 1 million, 1.5 million, 2 million, 2.5 million, 3 million, 3.5 million, 4 million, 4.5 million, or at least 5 million variants. In some examples, the polynucleotide libraries are collectively configured to bind to genomic regions comprising 100-1,000, 50-100, 50-500, 50-5,000, 50-10,000, 100,000-5 million, 250,000-3 million, 500,000-2 million, 750,000-4 million, 1 million-5 million, 1 million-3 million, 1 million-4 million, or 4 million-6 million variants.
[0073] A polynucleotide library for identifying variants can be optimized. In some examples, the library is homogeneous (each unique polynucleotide is equally represented). In some examples, the library is not homogeneous. In some examples, the polynucleotides are represented in an amount within at least about 1.5-fold of the average representation in the polynucleotide library. In some examples, the polynucleotides are represented in an amount within at least about 2-fold of the average representation in the polynucleotide library. In some examples, the polynucleotides are represented in an amount within at least about 1.2-fold of the average representation in the polynucleotide library. In some examples, the polynucleotides are represented in an amount within at least about 1.7-fold of the average representation in the polynucleotide library. In some examples, at least 80% of the polynucleotides are represented in an amount within at least about 1.5-fold of the average representation in the polynucleotide library. In some examples, at least 80% of the polynucleotides are represented in an amount within at least about 2-fold of the average representation in the polynucleotide library. In some examples, at least 80% of the polynucleotides are represented in an amount within at least about 1.7-fold of the average representation in the polynucleotide library. In some examples, at least 80% of the polynucleotides are represented in an amount within at least about 2-fold of the average representation in the polynucleotide library. In some examples, at least 90% of the polynucleotides are represented in an amount within at least about 1.5-fold of the average representation in the polynucleotide library. In some examples, at least 90% of the polynucleotides are represented in an amount within at least about 2-fold of the average representation in the polynucleotide library. In some examples, at least 80% of the polynucleotides are represented in an amount within at least about 1.7-fold of the average representation in the polynucleotide library. In some examples, at least 90% of the polynucleotides are represented in an amount within at least about 2-fold of the average representation in the polynucleotide library. In some examples, at least 95% of the polynucleotides are represented in an amount within at least about 1.5-fold of the average representation in the polynucleotide library.In some examples, at least 95% of the polynucleotides are represented in an amount within at least about 2-fold of the average representation in the polynucleotide library. In some examples, at least 95% of the polynucleotides are represented in an amount within at least about 1.7-fold of the average representation in the polynucleotide library. In some examples, at least 95% of the polynucleotides are represented in an amount within at least about 2-fold of the average representation in the polynucleotide library. In some examples, the polynucleotide library contains at least some polynucleotides, each of which comprises an overlapping region with another polynucleotide in the library. In some examples, at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or at least 90% of the polynucleotides each comprise an overlapping region with another polynucleotide in the library. In some examples, about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or about 90% of the polynucleotides each comprise an overlapping region with another polynucleotide in the library. In some examples, 10% to 90%, 10% to 80%, 10% to 75%, 25% to 50%, 25% to 90%, 50% to 90%, 1% to 35%, or 80% to 99% of the polynucleotides comprise regions of overlap with other polynucleotides in the library. In some examples, the abundance of at least some of the polynucleotides in the library is 5, 10, 20, 25, 50, 75, 100, 150, 200, 250, 300, 400, 500, or 600 times higher than the average representation of the polynucleotide library. In some examples, the abundance of at least 1% of the polynucleotides in the library is 5, 10, 20, 25, 50, 75, 100, 150, 200, 250, 300, 400, 500, or 600 times higher than the average representation of the polynucleotide library. In some examples, the abundance of at least 2% of the polynucleotides in the library is 5, 10, 20, 25, 50, 75, 100, 150, 200, 250, 300, 400, 500, or 600 times higher than the average representation of the polynucleotide library.In some examples, the abundance of at least 5% of the polynucleotides in the library is 5, 10, 20, 25, 50, 75, 100, 150, 200, 250, 300, 400, 500, or 600 times higher than the average representation of the polynucleotide library. In some examples, the abundance of 5% or less of the polynucleotides in the library is 5, 10, 20, 25, 50, 75, 100, 150, 200, 250, 300, 400, 500, or 600 times higher than the average representation of the polynucleotide library. In some examples, the abundance of 10% or less of the polynucleotides in the library is 5, 10, 20, 25, 50, 75, 100, 150, 200, 250, 300, 400, 500, or 600 times higher than the average representation of the polynucleotide library. In some examples, the abundance of at least 1% to 10% of the polynucleotides in the library is 5, 10, 20, 25, 50, 75, 100, 150, 200, 250, 300, 400, 500, or 600 times higher than the average representation in the polynucleotide library. In some examples, the abundance of at least 1% to 20% of the polynucleotides in the library is 5, 10, 20, 25, 50, 75, 100, 150, 200, 250, 300, 400, 500, or 600 times higher than the average representation in the polynucleotide library. In some examples, the relative abundance of the polynucleotide library is adjusted based on high or low GC content.
[0074] A polynucleotide library for identifying mutants can collectively target a desired number of bases (bait territory). In some examples, the polynucleotide library comprises a bait territory of at least 5 million, 10 million, 15 million, 20 million, 25 million, 30 million, 40 million, 50 million, 60 million, 70 million, 80 million, 90 million, or at least 100 million bases. In some examples, the polynucleotide library comprises a bait territory of about 5 million, 10 million, 15 million, 20 million, 25 million, 30 million, 40 million, 50 million, 60 million, 70 million, 80 million, 90 million, or about 100 million bases. In some examples, the polynucleotide library comprises a bait territory of 5 million, 10 million, 15 million, 20 million, 25 million, 30 million, 40 million, 50 million, 60 million, 70 million, 80 million, 90 million, or 100 million bases or less.
[0075] Systems and methods for generating polynucleotide libraries, such as variant-targeting polynucleotide libraries, are provided herein. In some examples, the system includes generating an in silico polynucleotide library including a sequence. In some examples, the system generates the nucleic acid standard described herein. In some examples, the system for generating a polynucleotide library includes at least one processor and instructions executable by the at least one processor to perform one or more of the following operations: (a) receiving as input, (b) generating a polynucleotide library by saturating at least one target region with one or more polynucleotides, and (c) generating one or more outputs including the sequence of the polynucleotide library. In some examples, the administration includes a nucleic acid reference sequence, one or more of the at least one target region, the nucleic acid reference sequence, and one or more variables. In some examples, the administered nucleic acid reference sequence includes a genome. In some examples, the administered nucleic acid reference sequence includes an mRNA. In some examples, at least one target region includes at least one exon. In some examples, at least one target region includes a variant. In some examples, variants include single nucleotide variants (SNVs), insertions / deletions (indels), or structural variants (SVs).
[0076] The variables controlling the sequences generated by the system can be adjusted for a particular application or target region. In some examples, one or more variables independently include polynucleotide length, offset, number of probes, overlap, overhang, target region merging, and tiling depth. In some examples, the relationship between polynucleotide size and target is used to generate a library. In some examples, the target region is smaller than the polynucleotide length. In some examples, when the target region is smaller than the polynucleotide length, polynucleotides are generated with a 1, 2, 3, 4, 5, or 6 base offset. In some examples, the target region is larger than the polynucleotide length. In some examples, when the target region is larger than the polynucleotide length, polynucleotides are generated to ensure even coverage of the entire target region.
[0077] Unique Molecular Identifier (UMI) Provided herein are adapters containing unique molecular identifiers (UMIs). In some examples, the adapters include universal adapters. In some examples, the adapters include a Y annealing region (which anneals to form a yoke), one or more Y step non-annealing regions, a first index region, a second index region, a first UMI (index) region, a second UMI (index) region, and one or more regions outside the index. In some examples, the adapters are ligated to sample polynucleotides to form adapter-ligated polynucleotides. After denaturation, upper and lower strand ligation products are formed. In some examples, each strand is labeled with a different UMI. After amplification with forward and backward primers, upper and lower strand PCR products are generated. In some examples, the adapter-ligated polynucleotides generated using universal adapters are further amplified using barcoded primers. In some examples, the adapters described herein include an "in-line" UMI, where at least one of the 5' or 3' UMI is not complementary to the other corresponding strand of the adapter. In some examples, the adaptors described herein include a "duplex" UMI, where at least one of the 5' or 3' UMI is complementary to the other corresponding strand of the adaptor.
[0078] Adapter-ligated libraries containing unique molecular identifiers can be used to distinguish "true" mutations from polynucleotide sample libraries from artifacts generated during sequencing library preparation (e.g., PCR errors, sequencing errors, or other incorrect base calls). In some examples, a workflow is used to analyze a library of adapter-ligated sample polynucleotides. Each adapter-ligated sample polynucleotide contains two distinct UMIs, represented by letters (A-F; for simplicity, six barcode combinations are shown), which are attached to the sample polynucleotides. After sequencing, forward and reverse read pairs from the sequencing are sorted into read pair groups. Next, the read pairs are grouped by barcode and barcode position. Single-stranded consensus sequences are then generated from each group of barcode-grouped read pairs. Errors from D-C and F-E are identified, while errors in A-B remain. Finally, double-stranded consensus sequences are generated by comparing each set of single-stranded consensus sequences. Errors in A-B can be identified, and true mutations E-F can be confirmed. In some examples, the error comprises a substitution, deletion, or insertion. In some examples, the error is present in the sample polynucleotide portion of the adaptor-ligated polynucleotide. In some examples, the error is present in a barcode configured to identify the sample origin (e.g., index) or to uniquely identify the sample polynucleotide. In some cases, the error is present in the UMI. In some examples, the error is present in the sample index. The compositions and methods described herein are used, in some examples, to identify such errors.
[0079] Described herein are UMI sets, where the UMI sets have defined properties. In some examples, the UMI sets include a plurality of different polynucleotides with unique sequences. In some examples, the UMI sets are 8, 12, 16, 20, 24, 30, 32, 36, 39, 48, or 64 unique sequences. In some examples, the sequences of the UMI sets differ by a Hamming distance of 1, 2, 3, 4, or 5 or less. In some examples, the sequences of the UMI sets differ by a Hamming distance of at least 1, 2, 3, 4, or 5. In some examples, the sequences of the UMI sets differ by a Hamming distance of at least 2. In some examples, the sequences of the UMI sets differ by a Hamming distance of at least 1.
[0080] UMIs can be any length depending on the desired application. In some examples, the UMIs are 15, 12, 10, 8, 7, 6, 5, 4 or less, or 3 or less bases in length. In some examples, the UMIs are about 15, 12, 10, 8, 7, 6, 5, 4, or about 3 bases in length. In some examples, the UMIs are about 3-12, 3-10, 3-8, 4-12, 4-10, 4-8, 6-12, or 8-12 bases in length. The UMIs in a set can include multiple lengths. In some examples, 10, 20, 25, 30, 40, 50, 60, or 70 percent of the UMIs in a set are a first length, and 90, 80, 75, 70, 60, 50, 40, or 30 percent are a second length. In some examples, the first length is 3-5 bases, and the second length is 3-5 bases. In some instances, the UMI comprises 5 or 6 bases in length.
[0081] After adding UMI-containing adaptors to sample polynucleotides, at least some of the sample polynucleotides can be uniquely labeled.In some examples, at least 30%, 50%, 75%, 80%, 90%, 95%, or at least 98% of the sample polynucleotides are linked to adaptors containing UMI.In some examples, at least 1%, 2%, 5%, 10%, 15%, 20%, 30%, 50%, 75%, 80%, 90%, 95%, or at least 98% of the sample polynucleotides are labeled with unique UMI sequences.In some examples, 1%, 2%, 5%, 10%, 15%, 20%, 30%, 50%, 75%, 80%, 90%, 95%, or at most 98% of the sample polynucleotides are labeled with unique UMI sequences. In some examples, at least 1%, 2%, 5%, 10%, 15%, 20%, 30%, 50%, 75%, 80%, 90%, 95%, or at least 98% of the sample polynucleotides are uniquely identifiable after labeling with a UMI.
[0082] In some embodiments, the UMI comprises one or more of the following polynucleotide sequences: AAGGA (SEQ ID NO: 1), ACAAC (SEQ ID NO: 2), ATACG (SEQ ID NO: 3), CACTG (SEQ ID NO: 4), CATGA (SEQ ID NO: 5), CGATA (SEQ ID NO: 6), CGTGT (SEQ ID NO: 7), GCCAT (SEQ ID NO: 8), GCTGT (SEQ ID NO: 9), GTCAC (SEQ ID NO: 10), GTCGT (SEQ ID NO: 11), TACGA (SEQ ID NO: 12), TCCTA (SEQ ID NO: 13), TCGTG (SEQ ID NO: 14), TGTCG (SEQ ID NO: 15), TTGGC (SEQ ID NO: 16), AACAC (SEQ ID NO: 17), AATGC (SEQ ID NO: 18), ACTAG (SEQ ID NO: 19), AGCAT (SEQ ID NO: 20), AGTAC (SEQ ID NO: 21), ATCTC (SEQ ID NO: 22), CAGAC (SEQ ID NO: 23), CAGTA (SEQ ID NO: 24), CGAAT (SEQ ID NO: 25), CGGTT (SEQ ID NO: 26), CTTGG (SEQ ID NO: 27), GCATA (SEQ ID NO: 28), GCTAA (SEQ ID NO: 29), GTGAG (SEQ ID NO: 30), GTGTC (SEQ ID NO: 31), and TGTGC (SEQ ID NO: 32). In some embodiments, the UMI comprises two or more of the following polynucleotide sequences: AAGGA (SEQ ID NO: 1), ACAAC (SEQ ID NO: 2), ATACG (SEQ ID NO: 3), CACTG (SEQ ID NO: 4), CATGA (SEQ ID NO: 5), CGATA (SEQ ID NO: 6), CGTGT (SEQ ID NO: 7), GCCAT (SEQ ID NO: 8), GCTGT (SEQ ID NO: 9), GTCAC (SEQ ID NO: 10), GTCGT (SEQ ID NO: 11), TACGA (SEQ ID NO: 12), TCCTA (SEQ ID NO: 13), TCGTG (SEQ ID NO: 14), TGTCG (SEQ ID NO: 15), TTGGC (SEQ ID NO: 16), AACAC (SEQ ID NO: 17), AATGC (SEQ ID NO: 18), ACTAG (SEQ ID NO: 19), AGCAT (SEQ ID NO: 20), AGTAC (SEQ ID NO: 21), ATCTC (SEQ ID NO: 22), CAGAC (SEQ ID NO: 23), CAGTA (SEQ ID NO: 24), CGAAT (SEQ ID NO: 25), CGGTT (SEQ ID NO: 26), CTTGG (SEQ ID NO: 27), GCATA (SEQ ID NO: 28), GCTAA (SEQ ID NO: 29), GTGAG (SEQ ID NO: 30), GTGTC (SEQ ID NO: 31), and TGTGC (SEQ ID NO: 32).In some embodiments, the UMI comprises five or more of the following polynucleotide sequences: AAGGA (SEQ ID NO: 1), ACAAC (SEQ ID NO: 2), ATACG (SEQ ID NO: 3), CACTG (SEQ ID NO: 4), CATGA (SEQ ID NO: 5), CGATA (SEQ ID NO: 6), CGTGT (SEQ ID NO: 7), GCCAT (SEQ ID NO: 8), GCTGT (SEQ ID NO: 9), GTCAC (SEQ ID NO: 10), GTCGT (SEQ ID NO: 11), TACGA (SEQ ID NO: 12), TCCTA (SEQ ID NO: 13), TCGTG (SEQ ID NO: 14), TGTCG (SEQ ID NO: 15), TTGGC (SEQ ID NO: 16), AACAC (SEQ ID NO: 17), AATGC (SEQ ID NO: 18), ACTAG (SEQ ID NO: 19), AGCAT (SEQ ID NO: 20), AGTAC (SEQ ID NO: 21), ATCTC (SEQ ID NO: 22), CAGAC (SEQ ID NO: 23), CAGTA (SEQ ID NO: 24), CGAAT (SEQ ID NO: 25), CGGTT (SEQ ID NO: 26), CTTGG (SEQ ID NO: 27), GCATA (SEQ ID NO: 28), GCTAA (SEQ ID NO: 29), GTGAG (SEQ ID NO: 30), GTGTC (SEQ ID NO: 31), and TGTGC (SEQ ID NO: 32). In some embodiments, the UMI comprises 10 or more of the following polynucleotide sequences: AAGGA (SEQ ID NO: 1), ACAAC (SEQ ID NO: 2), ATACG (SEQ ID NO: 3), CACTG (SEQ ID NO: 4), CATGA (SEQ ID NO: 5), CGATA (SEQ ID NO: 6), CGTGT (SEQ ID NO: 7), GCCAT (SEQ ID NO: 8), GCTGT (SEQ ID NO: 9), GTCAC (SEQ ID NO: 10), GTCGT (SEQ ID NO: 11), TACGA (SEQ ID NO: 12), TCCTA (SEQ ID NO: 13), TCGTG (SEQ ID NO: 14), TGTCG (SEQ ID NO: 15), TTGGC (SEQ ID NO: 16), AACAC (SEQ ID NO: 17), AATGC (SEQ ID NO: 18), ACTAG (SEQ ID NO: 19), AGCAT (SEQ ID NO: 20), AGTAC (SEQ ID NO: 21), ATCTC (SEQ ID NO: 22), CAGAC (SEQ ID NO: 23), CAGTA (SEQ ID NO: 24), CGAAT (SEQ ID NO: 25), CGGTT (SEQ ID NO: 26), CTTGG (SEQ ID NO: 27), GCATA (SEQ ID NO: 28), GCTAA (SEQ ID NO: 29), GTGAG (SEQ ID NO: 30), GTGTC (SEQ ID NO: 31), and TGTGC (SEQ ID NO: 32).
[0083] UMIs can be represented in a preselected proportion in the library of UMIs. In some examples, at least 90% of the UMIs are present in a proportion of 1-5%. In some examples, at least 90% of the UMIs are present in a proportion of 0.5%, 1%, 1.5%, 2%, 2.5%, 3%, 3.5%, 4%, 4.5%, 5%, 5.5%, 6%, 7%, or 8%. In some examples, at least 90% of the UMIs are present in a proportion of 0.5-8%, 1-7%, 1.5-7%, 2-7%, 2.5-6%, 3-8%, 3-6%, 1-5%, 0.5-5.5%, 1-4%, 1-6%, or 1-8%.
[0084] Any amount of sample polynucleotide (e.g., input DNA or other nucleic acid) can be ligated to the adaptors described herein. In some examples, the amount of sample polynucleotide is about 1, 5, 8, 10, 15, 20, 25, 30, 50, 75, or about 100 ng. In some examples, the amount of sample polynucleotide is 1, 5, 8, 10, 15, 20, 25, 30, 50, 75, or 100 ng or less. In some examples, the amount of sample polynucleotide is at least 1, 5, 8, 10, 15, 20, 25, 30, 50, 75, or at least 100 ng. In some examples, the amount of sample polynucleotide is 1-10 ng, 1-100 ng, 3-10 ng, 5-100 ng, 5-75 ng, 5-50 ng, 10-100 ng, 10-50 ng, 25-100 ng, or 25-75 ng.
[0085] A method for generating an adapter containing a UMI is provided herein. A first method for adapter synthesis involves synthesizing an upper strand of an adapter containing at least one UMI and a complementary lower strand. After annealing the upper and lower adapter strands, an adapter containing the adapter structure is formed. A second method for adapter synthesis involves synthesizing an upper strand without a UMI and a lower strand containing a complementary region and a UMI. After annealing, PCR is used to generate a complementary UMI on the upper strand, and terminal transferase adds a T to the 3' end of the upper strand to generate the adapter. A third synthesis method involves synthesizing an upper strand without a UMI, a lower strand containing a UMI, a restriction site, and a 5' overhang. After annealing, the upper strand is extended by PCR, and a restriction endonuclease is used to cleave a portion of the 3' upper strand and the 5' lower strand to generate the adapter. In a fourth method of adapter synthesis, two complementary strands each containing a UMI, a restriction site, and an overhang portion (3' top strand, 5' bottom strand) are synthesized, annealed, and cleaved with a restriction enzyme to generate an adapter. There can be multiple UMIs per adapter. In some examples, an adapter contains 1, 2, 3, 4, 5, or more UMIs. In some examples, an adapter comprises a first UMI and a second UMI. In some examples, the first UMI and the second UMI are complementary. In some examples, an adapter comprises a first UMI and a second UMI. In some examples, the first UMI and the second UMI are not complementary. In some examples, adapters are combined into a library of adapters. In some examples, adapters in the library contain UMIs. In some examples, adapters in the library contain unique combinations of first UMIs and second UMIs.
[0086] Universal Adapter Universal adapters are provided herein. In some examples, the universal adapter comprises one or more unique molecular identifiers. In some examples, the universal adapter disclosed herein can comprise a universal polynucleotide adapter comprising a first strand and a second strand. In some examples, the first strand comprises a first primer binding region, a first non-complementary region, and a first yoke region. In some examples, the second strand comprises a second primer binding region, a second non-complementary region, and a second yoke region. In some examples, the primer binding region enables PCR amplification of the polynucleotide adapter. In some examples, the primer binding region enables PCR amplification of the polynucleotide adapter and simultaneous addition of one or more barcodes to the polynucleotide adapter. In some examples, the first yoke region is complementary to the second yoke region. In some examples, the first non-complementary region is not complementary to the second non-complementary region. In some examples, the universal adapter is a Y-shaped or forked adapter. In some examples, one or more yoke regions comprise a nucleobase analogue that increases the Tm between the first yoke region and the second yoke region. The primer binding region described herein can be in the form of a terminal adapter region of a polynucleotide. In some examples, the universal adapter comprises an index sequence. In some examples, the universal adapter comprises a unique molecular identifier. In some examples, the universal adapter is configured for use with a barcoded primer, and after ligation, the barcoded primer is added via PCR.
[0087] Universal adapters comprise polynucleotide sequences that may be shortened compared to the polynucleotide sequences of typical barcoded adapters (e.g., full-length "Y adapters"). For example, universal adapters may comprise polynucleotide sequences between 20 and 45 bases in length. In some examples, universal adapters may comprise polynucleotide sequences between 25 and 40 bases in length. In some examples, universal adapters may comprise polynucleotide sequences between 30 and 35 bases in length. In some examples, universal adapters may comprise polynucleotide sequences that are 50 bases or less in length, 45 bases or less in length, 40 bases or less in length, 35 bases or less in length, 30 bases or less in length, or 25 bases or less in length. In some examples, universal adapters may comprise polynucleotide sequences that are approximately 25, 27, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, or 60 bases in length. In some examples, universal adapters may comprise polynucleotide sequences that are approximately 60 base pairs in length. In some examples, the universal adaptor may comprise a polynucleotide sequence that is about 58 base pairs in length. In some examples, the universal adaptor may comprise a polynucleotide sequence that is about 52 base pairs in length. In some examples, the universal adaptor may comprise a polynucleotide sequence that is about 33 base pairs in length.
[0088] The universal adaptor may be modified to facilitate ligation with the sample polynucleotide. For example, the 5' end is phosphorylated. In some instances, the universal adaptor comprises one or more non-natural nucleobase linkages, such as phosphorothioate linkages. For example, the universal adaptor comprises a phosphorothioate linkage between the 3' end base and the base adjacent to the 3' end base. In some instances, the sample polynucleotide comprises nucleic acid from various sources, such as DNA or RNA of human, bacterial, plant, animal, fungal, or viral origin. In some instances, the adaptor-linked sample polynucleotide comprises a sample polynucleotide (e.g., a sample nucleic acid) having adaptor-universal adaptors ligated to both the 5' and 3' ends of the sample polynucleotide to form an adaptor-linked polynucleotide. The double-stranded sample polynucleotide comprises both a first strand (forward) and a second strand (reverse).
[0089] A universal adaptor can contain any number of different nucleobases (DNA, RNA, etc.), nucleobase analogs, or non-nucleobase linkers or spacers. For example, a universal adaptor can contain one or more nucleobase analogs or linkers that allow hybridization (T) between the two polynucleotide strands of the universal adaptor. m In some instances, the nucleobase analog is present in the yoke region of the universal adaptor. Nucleobase analogs and other groups include, but are not limited to, locked nucleic acids (LNAs), bicyclic nucleic acids (BNAs), C5-modified pyrimidine bases, 2'-O-methyl substituted RNAs, peptide nucleic acids (PNAs), glycol nucleic acids (GNAs), threose nucleic acids (TNAs), xenon nucleic acids (XNAs), morpholino backbone-modified bases, minor groove binders (MGBs), spermine, G-clamps, or anthraquinone (Uaq) caps.
[0090] Universal adapters are used to identify the desired hybridization T mDepending on the type, any number of nucleobase analogs (such as LNA or BNA) may be included. For example, a universal adapter includes 1 to 20 nucleobase analogs. In some examples, a universal adapter includes 1 to 8 nucleobase analogs. In some examples, an adapter includes at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or at least 12 nucleobase analogs. In some examples, an adapter includes about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or about 16 nucleobase analogs. In some examples, the number of similar nucleobases is expressed as a percentage of the total bases in the adapter. For example, an adapter includes at least 1%, 2%, 5%, 10%, 12%, 18%, 24%, 30%, or more than 30% nucleobase analogs. In some examples, the universal adapters described herein include a methylated nucleobase, such as a methylated cytosine.
[0091] Polynucleotide barcodes A barcode sequence (or index sequence) is a defined short polynucleotide sequence that can be used as a sample-specific label during PCR amplification. A polynucleotide primer may contain a barcode sequence (hereinafter "barcode"). An adapter, as described herein, may contain one or more barcodes. In some examples, the adapter contains at least one index barcode and at least one unique molecular identifier (UMI) barcode. Barcodes can be attached to universal adapters, for example, using PCR and barcoded primers to generate barcoded adapter-ligated sample polynucleotides. Primer binding sites, such as universal primer binding sites, facilitate simultaneous amplification of all members or subpopulations of members of a barcoded primer library. In some examples, the primer binding site comprises a primer binding region that binds to a flow cell or other solid support during next-generation sequencing. In some examples, the barcoded primer comprises a P5 or P7 linkage region, as shown in Table 1 below.
[0092] [Table 1]
[0093] In some examples, the primer binding site is configured to bind to a universal adapter sequence and facilitate amplification and generation of a barcoded adapter. In some examples, the barcoded primer is 60 bases or less in length. In some examples, the barcoded primer is 55 bases or less in length. In some examples, the barcoded primer is 50-60 bases in length. In some examples, the barcoded primer is about 60 bases in length. In some examples, the barcodes described herein comprise a methylated nucleobase, such as a methylated cytosine.
[0094] The number of unique barcode sequences available in a barcode set (a collection of unique barcodes or barcode combinations configured to be used together to uniquely define a sample) can depend on the length. In some examples, the Hamming distance is defined by the number of base differences between any two barcodes. In some examples, the Levenshtein distance is defined by the number of changes (insertions, substitutions, or deletions) required to change one barcode into another. In some examples, the barcode sets described herein comprise a Levenshtein distance of at least 2, 3, 4, 5, 6, 7, or at least 8. In some examples, the barcode sets described herein comprise a Hamming distance of at least 2, 3, 4, 5, 6, 7, or at least 8.
[0095] Barcodes may be erroneously associated with a different sample than the one to which they were assigned. In some instances, inaccurate barcodes result from PCR errors (e.g., substitutions) during library amplification. In some instances, entire barcodes "hop" or are transferred from one polynucleotide sample to another. In some instances, such transfers result from cross-contamination of free adapters or primers during the library generation workflow. In some instances, barcode sets are selected to minimize "barcode hopping." In some instances, the barcode hopping (for a single barcode) for the barcode sets described herein is 7%, 5%, 4%, 3%, 2%, 1%, 0.5%, or less than 0.1%. In some instances, the barcode hopping (for a single barcode) for the barcode sets described herein is 0.1-6%, 0.1-5%, 0.2-5%, 0.5-5%, 1-7%, 1-5%, or 0.5-7%. In some examples, the barcode hopping (for two barcodes) for a barcode set described herein is 0.7%, 0.5%, 0.4%, 0.3%, 0.2%, 0.1%, 0.05%, or 0.1% or less. In some examples, the barcode hopping (for two barcodes) for a barcode set described herein is 0.01-0.6%, 0.01-0.5%, 0.02-0.5%, 0.05-0.5%, 0.1-0.7%, 0.1-0.5%, or 0.05-0.7%.
[0096] The barcoded primer comprises one or more barcodes. In some examples, the barcode is added to the universal adapter by a PCR reaction. The barcode is a nucleic acid sequence that allows the barcode to identify some feature of the associated polynucleotide. In some examples, the barcode comprises an index sequence. In some examples, the index sequence allows for identification of the sample or unique source of the nucleic acid being sequenced. The barcode or combination of barcodes, in some examples, identifies a specific patient. In some examples, the barcode or combination of barcodes identifies a specific sample from a patient among other samples from the same patient. After sequencing, the barcode (or barcode region) provides an index for identifying a feature associated with the coding region or sample source. Barcodes can be designed with a length suitable for allowing a sufficient degree of identification, for example, at least about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55 or more bases in length. Multiple barcodes, such as about 2, 3, 4, 5, 6, 7, 8, 9, 10 or more, can be used on the same molecule, optionally separated by non-barcode sequences. In some examples, barcodes are placed on the 5' and 3' ends of the sample polynucleotide. In some examples, each barcode in the plurality of barcodes differs from every other barcode in the plurality at at least three base positions, such as at least about 3, 4, 5, 6, 7, 8, 9, 10, or more. The use of barcodes allows for the pooling and simultaneous processing of multiple libraries for downstream applications, such as sequencing (multiplexing). In some examples, at least 4, 8, 16, 32, 48, 64, 128, or 512 or more barcoded libraries are used.In some examples, at least 400, 500, 800, 1000, 2000, 5000, 10,000, 12,000, 15,000, 18,000, 20,000, or at least 25,000 barcodes are used.
[0097] The barcoded primer or adapter may include a unique molecular identifier (UMI). In some examples, such a UMI uniquely tags every nucleic acid in a sample. In some examples, at least 60%, 70%, 80%, 90%, 95%, or more than 95% of the nucleic acids in a sample are tagged with a UMI. In some examples, at least 85%, 90%, 95%, 97%, or at least 99% of the nucleic acids in a sample are tagged with a unique barcode or UMI. The barcoded primer, in some examples, includes an index sequence and one or more UMIs. The UMI allows for internal measurement of initial sample concentration or stoichiometry prior to downstream sample processing (e.g., PCR or enrichment steps) that may introduce bias. In some examples, the UMI includes one or more barcode sequences. In some examples, each strand of the adapter-ligated sample polynucleotide (forward vs. reverse) has one or more unique barcodes. Such barcodes are optionally used to uniquely tag each strand of the sample polynucleotide. In some examples, the barcoded primers comprise an index barcode and a UMI barcode. In some examples, after amplification with at least two barcoded primers, the resulting amplicon comprises two index sequences and two UMIs. In some examples, after amplification with at least two barcoded primers, the resulting amplicon comprises an index barcode and one UMI barcode. In some examples, each strand of the universal adapter sample polynucleotide duplex is tagged with a unique barcode, such as a UMI or index barcode.
[0098] The barcoded primers in the library contain a region complementary to the primer binding region on the universal adapter. For example, the universal adapter binding region is complementary to the primer region of the universal adapter, and the universal adapter binding region is complementary to the primer region of the universal adapter. This configuration facilitates extension of the universal adapter during PCR, allowing the barcoded primer to bind. In some instances, the Tm between the primer and the primer binding region is 40-65°C. In some instances, the Tm between the primer and the primer binding region is 42-63°C. In some instances, the Tm between the primer and the primer binding region is 50-60°C. In some instances, the Tm between the primer and the primer binding region is 53-62°C. In some instances, the Tm between the primer and the primer binding region is 54-58°C. In some instances, the Tm between the primer and the primer binding region is 40-57°C. In some instances, the Tm between the primer and the primer binding region is 40-50°C. In some examples, the Tm between the primer and the primer binding region is about 40, 45, 47, 50, 52, 53, 55, 57, 59, 61, or 62°C.
[0099] Hybridization Blockers Blockers can contain any number of different nucleobases (e.g., DNA, RNA), nucleobase analogs (non-standard), or non-nucleobase linkers or spacers. In some examples, blockers include universal blockers. Such blockers may, in some examples, be described as a "set," where a set includes two or more blockers configured to prevent undesired interactions with the same adapter sequence. In some examples, a universal blocker prevents adapter-adapter interactions regardless of the one or more barcodes present on at least one of the adapters. For example, a blocker includes one or more nucleobase analogs or other groups that enhance hybridization (Tm) between the blocker and the adapter. In some examples, a blocker includes one or more nucleobases (e.g., "universal" bases) that decrease hybridization (Tm) between the blocker and the adapter. In some examples, blockers described herein include both one or more nucleobases that increase hybridization (Tm) between the blocker and the adapter and one or more nucleobases that decrease hybridization (Tm) between the blocker and the adapter.
[0100] Described herein are hybridization blockers that include one or more regions (e.g., adapters) that enhance binding to a target sequence and one or more regions (e.g., adapters) that decrease binding to a target sequence. In some examples, each region is timed for a given desired level of off-bait activity during target enrichment applications. In some examples, each region can be modified with either a single type of chemical modification / moiety or multiple types to increase or decrease the overall affinity of the molecule for the target sequence. In some examples, the melting temperatures of all individual members of a blocker set are maintained above a specific temperature (e.g., with the addition of moieties such as LNAs and / or BNAs). In some examples, a given set of blockers improves off-bait performance regardless of index length, index sequence, and the number of adapter indexes present in the hybridization.
[0101] Blockers may contain moieties that increase and / or decrease affinity for target sequencing, such as adapters. In some instances, such specific regions can be thermodynamically tuned to specific melting temperatures to avoid or increase affinity for specific target sequences. This combination of modifications is designed to, in some instances, help increase the affinity of the blocker molecule for specific and unique adapter sequences and decrease the affinity of the blocker molecule for repetitive adapter sequences (e.g., the Y stem annealing portion of the adapter). In some instances, blockers contain a moiety that decreases binding of the blocker to the Y stem region of the adapter. In some instances, blockers contain a moiety that decreases binding of the blocker to the Y stem region of the adapter and a moiety that increases binding of the blocker to non-Y stem regions of the adapter.
[0102] Blockers (e.g., universal blockers) and adapters can form several distinct populations during hybridization. Population "A," in some instances, contains blockers correctly bound to the non-index region of the adapter. In population "B," a region of the blocker binds to the "yoke" region of the adapter, while the remainder of the blocker does not bind to the adjacent region of the adapter. In population "C," two blockers dimerize non-productively. In population "D," blockers do not bind to any other nucleic acid. In some instances, as the number of DNA modifications that reduce affinity at the Y stem annealing region of the blocker increases, populations "A" and "D" become dominant, either with the desired effect or with minimal effect. In some instances, as the number of DNA modifications that reduce affinity at the Y stem annealing region of the blocker decreases, populations "B" and "C" become dominant, either with the undesirable effect of daisy chaining or annealing to other adapters ("B"), or by isolating blockers that cannot function properly ("C").
[0103] The index in both single-index adapter and dual-index adapter designs can be partially or fully covered by a universal blocker extended with DNA modifications specifically designed to cover the adapter index bases. In some examples, such modifications include moieties that reduce annealing to the index, such as universal bases. In some examples, the index of a dual-index adapter is partially covered (or overlaps) by one or more blockers. In some examples, the index of a dual-index adapter is completely covered by one or more blockers. In some examples, the index of a single-index adapter is partially covered by one or more blockers. In some examples, the index of a single-index adapter is completely covered by one or more blockers. In some examples, the blocker overlaps the index sequence by at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, or more than 20 bases. In some examples, the blocker overlaps with the index sequence by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20 or less bases, or by 25 or less bases. In some examples, the blocker overlaps with the index sequence by about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20 or about 30 bases. In some examples, the blocker overlaps with the index sequence by 1-5, 1-3, 2-5, 2-8, 2-10, 3-6, 3-10, 4-10, 4-15, 1-4, or 5-7 bases. In some examples, the region of the blocker that overlaps with the index sequence comprises at least one 2-deoxyinosine or 5-nitroindole nucleobase.
[0104] One or two blockers can overlap with the index sequence present on the adapter. In some examples, the combined one or two blockers overlap at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20 bases, or more than 20 bases of the index sequence. In some examples, the combined one or two blockers overlap 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20 or less bases of the index sequence. In some examples, the combined one or two blockers overlap about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20 or about 20 bases of the index sequence. In some examples, one or two of the combined blockers overlap the index sequence by 1 to 5, 1 to 3, 2 to 5, 2 to 8, 2 to 10, 3 to 6, 3 to 10, 4 to 10, 4 to 15, 1 to 4, or 5 to 7 bases. In some examples, the region of the blocker that overlaps with the index sequence comprises at least one 2-deoxyinosine or 5-nitroindole nucleobase.
[0105] In the first configuration, the length of the adapter index overhang can be varied. When designed from a single side, the adapter index overhang can be modified to cover 0 to n adapter index bases from either side of the index. This allows the ability to design such adapter blockers for both single-indexed and dual-indexed adapter systems.
[0106] In the second configuration, the adapter index bases are covered from both sides. When the adapter index bases are covered from both sides, the length of each blocker's coverage region can be selected so that a single pair of blockers can interact with a range of adapter index lengths while still covering a significant portion of the total number of index bases. As an example, take two blockers designed with a 3-bp overhang that covers the adapter index. In the context of 6-bp, 8-bp, or 10-bp adapter index lengths, these blockers leave 0-bp, 2-bp, or 4-bp exposed, respectively, during hybridization.
[0107] In the third configuration, modified nucleobases are selected to cover the index adapter bases. Current commercially available examples of these modifications include degenerate bases (e.g., mixed bases of A, T, C, and G), 2'-deoxylinosine, and 5-nitroindole.
[0108] In the fourth configuration, a blocker with an adapter index overhang binds to either the sense (i.e., "top") strand or the antisense (i.e., "bottom") strand of a next-generation sequencing library.
[0109] In the fifth configuration, the blocker is further expanded to cover, in addition to standard adapter index bases of defined length and composition, other polynucleotide sequences (e.g., poly-A tails added in previous biochemical steps to facilitate ligation or other methods for introducing defined adapter sequences, unique molecular identifiers for bioinformatics assignment after sequencing, etc.). These types of sequences can be placed at multiple positions on the adapter, in which case the most widely utilized case (e.g., unique molecular index next to the genomic insert) is presented. Other positions of the unique molecular identifier (e.g., next to the adapter index base) can also be addressed with a similar approach.
[0110] In a sixth configuration, all of the aforementioned configurations are utilized in various combinations to meet target performance metrics for off-bait performance during target enrichment under specified conditions.
[0111] Blockers may contain moieties such as nucleobase analogs. Nucleobase analogs and other groups include, but are not limited to, locked nucleic acids (LNAs), bicyclic nucleic acids (BNAs), C5-modified pyrimidine bases, 2'-O-methyl-substituted RNA, peptide nucleic acids (PNAs), glycol nucleic acids (GNAs), threose nucleic acids (TNAs), inosine, 2'-deoxyinosine, 3-nitropyrrole, 5-nitroindole, xenon nucleic acids (XNAs), morpholino backbone-modified bases, minor groove binders (MGBs), spermine, G-clamps, or anthraquinone (Uaq) caps. In some examples, the nucleobase analogs contain universal bases, and the nucleobases have a lower Tm for binding to their cognate nucleobases. In some examples, the universal bases include 5-nitroindole or 2'-deoxyinosine. In examples, the blockers contain a spacer element connecting two polynucleotide strands. In some examples, the blockers contain one or more nucleobase analogs. In some examples, such nucleobase analogs are added to control the Tm of the blocker. The blocker can contain any number of nucleobase analogs (such as LNA or BNA) depending on the desired hybridization Tm. For example, the blocker contains 20 to 40 nucleobase analogs. In some examples, the blocker contains 8 to 16 nucleobase analogs. In some examples, the blocker contains at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or at least 12 nucleobase analogs. In some examples, the blocker contains about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or about 16 nucleobase analogs. In some examples, the number of analogous nucleobases is expressed as a percentage of the total bases in the blocker. For example, a blocker comprises at least 1%, 2%, 5%, 10%, 12%, 18%, 24%, 30%, or more than 30% nucleobase analogs. In some examples, a blocker comprising a nucleobase analog increases the Tm by about 2°C to about 8°C for each nucleobase analog.In some examples, the Tm is increased by at least or about 1°C, 2°C, 3°C, 4°C, 5°C, 6°C, 7°C, 8°C, 9°C, 10°C, 12°C, 14°C, or 16°C for each nucleobase analog. Such blockers are, in some examples, configured to bind to the top or "sense" strand of the adapter. Blockers are optionally configured to bind to the bottom or "antisense" strand of the adapter. In some examples, the set of blockers includes sequences configured to bind to both the top and bottom strands of the adapter. In some examples, additional blockers are configured to the complementary, reverse complementary, forward complementary, or reverse complementary strand of the adapter sequence. In some examples, a set of blockers targeting the top strand (binding to the top) or the bottom strand (or both) is designed and tested, followed by optimization, such as replacing top blockers with bottom blockers or replacing bottom blockers with top blockers. In some examples, the blockers are configured to fully or partially overlap with bases of an index or barcode on the adapter. In some examples, the set of blockers includes at least one blocker that overlaps with the adapter index sequence. In some examples, the set of blockers includes at least one blocker that overlaps with the adapter index sequence and at least one blocker that does not overlap with the adapter sequence. In some examples, the set of blockers includes at least one blocker that does not overlap with the yoke region sequence. In some examples, the set of blockers includes at least one blocker that does not overlap with the yoke region sequence and at least one blocker that overlaps with the yoke region sequence. In some examples, the set of blockers includes 2, 3, 4, 5, 6, 7, 8, 9, 10, or more than 10 blockers.
[0112] The blocker can be any length, depending on the size or hybridization Tm of the adapter. For example, the blocker is 20-50 bases long. In some examples, the blocker is 25-45, 30-40, 20-40, or 30-50 bases long. In some examples, the blocker is 25-35 bases long. In some examples, the blocker is at least 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or at least 35 bases long. In some examples, the blocker is 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 bases or less. In some examples, the blocker is about 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or about 35 bases long. In some examples, the blocker is about 50 bases long. The set of blockers targeting adapter-tagged genomic library fragments includes, in some instances, blockers of multiple lengths. In some instances, two blockers are tethered together with a linker. Various linkers are well known in the art, and in some instances, include alkyl, polyether, amine, amide, or other chemical groups. In some instances, the linker includes individual linker units connected to each other (or attached to the blocker polynucleotides) via a backbone such as a phosphate, thiophosphate, amide, or other backbone. In an exemplary configuration, the linker spans the index region between a first blocker targeting the 5' end of the adapter sequence and a second blocker targeting the 3' end of the adapter sequence. In some instances, a capping group is added to the 5' or 3' end of the blocker to prevent downstream amplification. Capping groups variously include polyethers, polyalcohols, alkanes, or other non-hybridizing groups that prevent amplification. In some instances, such groups are connected via a phosphate, thiophosphate, amide, or other backbone. In some instances, more than one blocker is used. In some instances, at least four non-identical blockers are used.In some examples, the first blocker spans a first 3' end of the adapter sequence, the second blocker spans a first 5' end of the adapter sequence, the third blocker spans a second 3' end of the adapter sequence, and the fourth blocker spans a second 5' end of the adapter sequence. In some examples, the first blocker is at least 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or at least 35 bases in length. In some examples, the second blocker is at least 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or at least 35 bases in length. In some examples, the third blocker is at least 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or at least 35 bases in length. In some examples, the fourth blocker is at least 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or at least 35 bases in length. In some examples, the first blocker, second blocker, third blocker, or fourth blocker comprises a nucleobase analog. In some examples, the nucleobase analog is LNA.
[0113] The design of a blocker can be influenced by the desired hybridization Tm to the adapter sequence. In some examples, a non-standard nucleic acid (e.g., a locked nucleic acid, a bridged nucleic acid, or other non-standard nucleic acid or analog) is inserted into the blocker to increase or decrease the Tm of the blocker. In some examples, the Tm of a blocker is calculated using a tool specific for calculating the Tm of polynucleotides containing non-standard amino acids. In some examples, the Tm is calculated using the Exiqon™ online prediction tool. In some examples, the blocker Tm described herein is calculated in silico. In some examples, the blocker Tm is calculated in silico and correlated with experimental in vitro conditions. Without being bound by theory, an experimentally determined Tm can be further influenced by experimental parameters such as salt concentration, temperature, the presence of additives, or other factors. In some examples, the Tm described herein is an in silico-determined Tm used to design or optimize blocker performance. In some examples, the Tm value is predicted, estimated, or determined from a melting curve analysis experiment. In some examples, the blocker has a Tm of 70°C to 99°C. In some examples, the blocker has a Tm of 75°C to 90°C. In some examples, the blocker has a Tm of at least 85°C. In some examples, the blocker has a Tm of at least 70, 72, 75, 77, 80, 82, 85, 88, 90, or at least 92°C. In some examples, the blocker has a Tm of about 70, 72, 75, 77, 80, 82, 85, 88, 90, 92, or about 95°C. In some examples, the blocker has a Tm of 78°C to 90°C. In some examples, the blocker has a Tm of 79°C to 90°C. In some examples, the blocker has a Tm of 80°C to 90°C. In some examples, the blocker has a Tm of 81°C to 90°C. In some examples, the blocker has a Tm of 82°C to 90°C. In some examples, the blockers have a Tm between 83° C. and 90° C. In some examples, the blockers have a Tm between 84° C. and 90° C. In some examples, the set of blockers has an average Tm between 78° C. and 90° C.In some examples, the set of blockers has an average Tm of 80° C. to 90° C. In some examples, the set of blockers has an average Tm of at least 80° C. In some examples, the set of blockers has an average Tm of at least 81° C. In some examples, the set of blockers has an average Tm of at least 82° C. In some examples, the set of blockers has an average Tm of at least 83° C. In some examples, the set of blockers has an average Tm of at least 84° C. In some examples, the set of blockers has an average Tm of at least 86° C. Blocker Tms are, in some examples, modified as a result of other components described herein, such as the use of fast hybridization buffers and / or hybridization enhancers.
[0114] The molar ratio of blocker to adaptor target can affect the rate of off-bait (and subsequent off-target) during hybridization. The more efficient the blocker is in binding to the target adaptor, the less blocker is required. The blockers described herein, in some examples, achieve sequencing results with 20% or less off-target reads (blocker target) at a molar ratio of less than 20:1. In some examples, 20% or less off-target reads are achieved at a molar ratio (blocker target) of less than 10:1. In some examples, 20% or less off-target reads are achieved at a molar ratio (blocker target) of less than 5:1. In some examples, 20% or less off-target reads are achieved at a molar ratio (blocker target) of less than 2:1. In some examples, 20% or less off-target reads are achieved at a molar ratio (blocker target) of less than 1.5:1. In some examples, 20% or less off-target reads are achieved at a molar ratio (blocker target) of less than 1.2:1. In some instances, 20% or less off-target reads are achieved at a molar ratio (blocker to target) of less than 1.05:1.
[0115] Universal blockers can be used with panel libraries of various sizes. In some embodiments, the panel library comprises at least or about 0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, 1.0, 2.0, 4.0, 8.0, 10.0, 12.0, 14.0, 16.0, 18.0, 20.0, 22.0, 24.0, 26.0, 28.0, 30.0, 40.0, 50.0, 60.0, or more than 60.0 megabases (Mb).
[0116] The blockers described herein can improve on-target performance. In some embodiments, the on-target performance is improved by at least or about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or more than 95%. In some embodiments, the on-target performance is improved by at least or about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or more than 95% for various refractive index designs. In some embodiments, on-target performance is improved by at least or about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or more than 95% for various panel sizes.
[0117] De novo synthesis of small polynucleotide populations for amplification reactions. Described herein are methods for synthesizing polynucleotides from a surface, e.g., a plate. In some examples, the polynucleotide library comprises a sample polynucleotide library. In some examples, polynucleotides are synthesized on clusters of loci for polynucleotide extension, released, and then subjected to an amplification reaction, e.g., PCR. In some examples, the silicon plate comprises multiple clusters. Each cluster has multiple loci. Polynucleotides are de novo synthesized on the plate from the clusters. The polynucleotides are cleaved and removed from the plate to form a population of released polynucleotides. The population of released polynucleotides is then amplified to form a library of amplified polynucleotides.
[0118] The present invention provides a method in which amplification of polynucleotides synthesized on clusters provides enhanced control over polynucleotide expression compared to amplification of polynucleotides over the entire surface of a structure without such a clustered configuration. In some examples, amplification of polynucleotides synthesized from a surface with a clustered configuration of loci for polynucleotide extension provides for overcoming the negative effects on display caused by the repeated synthesis of large polynucleotide populations. Exemplary negative effects on display caused by the repeated synthesis of large polynucleotide populations include, but are not limited to, amplification bias caused by high / low GC content, repetitive sequences, trailing adenines, secondary structures, affinity for target sequence binding, or modified nucleotides in the polynucleotide sequence.
[0119] In contrast to the amplification of polynucleotides across the entire plate without clustering, cluster amplification can result in a more strict distribution around the average.For example, when 100,000 reads are randomly sampled, an average of 8 reads per sequence will result in a library with a distribution of about 1.5X from the average.In some cases, single cluster amplification will result in a maximum of about 1.5X, 1.6X, 1.7X, 1.8X, 1.9X, or 2.0X from the average.In some cases, single cluster amplification will result in at least about 1.0X, 1.2X, 1.3X, 1.5X, 1.6X, 1.7X, 1.8X, 1.9X, or 2.0X from the average.
[0120] The cluster amplification method described herein can result in a polynucleotide library that requires less sequencing for equivalent sequence representation compared to plate-wide amplification.In some instances, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% less sequencing is required.In some instances, up to 10%, up to 20%, up to 30%, up to 40%, up to 50%, up to 60%, up to 70%, up to 80%, up to 90%, or up to 95% less sequencing is required.In some instances, 30% less sequencing is required after cluster amplification compared to plate-wide amplification.In some instances, the sequencing of polynucleotides is verified by high-throughput sequencing, such as next-generation sequencing. Sequencing of the sequencing library can be performed using any suitable sequencing technology, including, but not limited to, single molecule real-time (SMRT) sequencing, polony sequencing, ligation sequencing, reversible terminator sequencing, proton detection sequencing, ion semiconductor sequencing, nanopore sequencing, electron sequencing, pyrosequencing, Maxam-Gilbert sequencing, chain termination (e.g., Sanger) sequencing, +S sequencing, or sequencing by synthesis. The number of times a single nucleotide or polynucleotide is identified or "read" is defined as sequencing depth or read depth. In some cases, read depth is referred to as fold coverage, for example, 55-fold (or 55X) coverage, optionally describing the percentage of bases.
[0121] In some instances, amplification from a clustered configuration results in fewer dropouts or sequences that are not detected after sequencing of the amplified products compared to amplification across a plate. Dropouts can be AT and / or GC dropouts. In some instances, the number of dropouts is at most about 1%, 2%, 3%, 4%, or 5% of the polynucleotide population. In some instances, the number of dropouts is zero.
[0122] The clusters described herein comprise a collection of distinct, non-overlapping loci for polynucleotide synthesis. Clusters can comprise approximately 50-1,000, 75-900, 100-800, 125-700, 150-600, 200-500, or 300-400 loci. In some examples, each cluster comprises 121 loci. In some examples, each cluster comprises approximately 50-500, 50-200, or 100-150 loci. In some examples, each cluster comprises at least approximately 50, 100, 150, 200, 500, 1,000, or more loci. In some examples, a single plate comprises 100, 500, 10,000, 20,000, 30,000, 50,000, 100,000, 500,000, 700,000, 1,000,000, or more loci. The loci can be spots, wells, microwells, channels, or posts. In some examples, each cluster has a distinct feature of at least 1X, 2X, 3X, 4X, 5X, 6X, 7X, 8X, 9X, 10X, or more redundancy that supports the extension of polynucleotides having identical sequences.
[0123] Generation of polynucleotide libraries with controlled stoichiometry of sequence content In some instances, a polynucleotide library (such as a sample polynucleotide set for variant detection) is synthesized with a specific distribution of desired polynucleotide sequences. In some instances, tailoring the polynucleotide library for enrichment of specific desired sequences results in improved downstream application results.
[0124] One or more specific sequences can be selected based on their evaluation in downstream applications.In some examples, the evaluation is the binding affinity to the target sequence for amplification, enrichment, or detection, stability, melting temperature, biological activity, ability to assemble into larger fragments, or other properties of polynucleotides.In some examples, the evaluation is empirical or predicted from previous experiments and / or computer algorithms.An exemplary application includes increasing the sequences in the probe library corresponding to the region of the genome target that has less than the average read depth.
[0125] The selected sequences in a polynucleotide library may represent at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or more than 95% of the sequences. In some examples, the selected sequences in a polynucleotide library represent at most 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or up to 100% of the sequences. In some cases, the selected sequences are within the range of about 5-95%, 10-90%, 30-80%, 40-75%, or 50-70% of the sequences.
[0126] Polynucleotide libraries can be adjusted for the frequency of each selected sequence. In some examples, polynucleotide libraries favor a greater number of selected sequences. For example, libraries are designed with an increased polynucleotide frequency of selected sequences ranging from about 40% to about 90%. In some examples, polynucleotide libraries contain a smaller number of selected sequences. For example, libraries are designed with an increased polynucleotide frequency of selected sequences ranging from about 10% to about 60%. Libraries can be designed to favor higher and lower frequencies of selected sequences. In some examples, libraries favor uniform sequence representation. For example, polynucleotide frequencies are uniform with respect to selected sequence frequency, ranging from about 10% to about 90%. In some examples, libraries contain polynucleotides with a selected sequence frequency of about 10% to about 95% of the sequences.
[0127] The generation of a polynucleotide library having a specific selected sequence frequency can be achieved by combining at least two polynucleotide libraries having different selected sequence frequency contents. In some instances, at least 2, 3, 4, 5, 6, 7, 10, or more than 10 polynucleotide libraries can be combined to generate a population of polynucleotides having a specific selected sequence frequency. In some instances, 2, 3, 4, 5, 6, 7, or 10 or fewer polynucleotide libraries can be combined to generate a population of non-identical polynucleotides having a specific selected sequence frequency.
[0128] In some examples, the frequency of selected sequences is adjusted by synthesizing fewer or more polynucleotides per cluster. For example, at least 25, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or more than 1000 non-identical polynucleotides are synthesized on a single cluster. In some cases, no more than about 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 non-identical polynucleotides are synthesized on a single cluster. In some examples, 50 to 500 non-identical polynucleotides are synthesized on a single cluster. In some examples, 100 to 200 non-identical polynucleotides are synthesized on a single cluster. In some examples, about 100, about 120, about 125, about 130, about 150, about 175, or about 200 non-identical polynucleotides are synthesized on a single cluster.
[0129] In some cases, the frequency of selected sequences is adjusted by synthesizing non-identical polynucleotides of various lengths.For example, the length of each non-identical polynucleotide that is synthesized can be at least or at least about 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 150, 200, 300, 400, 500, 2000 nucleotides or more.The length of the non-identical polynucleotide that is synthesized can be at most or at most about 2000, 500, 400, 300, 200, 150, 100, 50, 45, 35, 30, 25, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10 nucleotides or less. The length of each of the non-identical polynucleotides to be synthesized can be in the range of 10 to 2000, 10 to 500, 9 to 400, 11 to 300, 12 to 200, 13 to 150, 14 to 100, 15 to 50, 16 to 45, 17 to 40, 18 to 35, and 19 to 25.
[0130] kit Kits containing a library of polynucleotides are provided herein. In some examples, the kits include one or more reference standards (controls), which include a sample polynucleotide set and a background set, instructions for use of the kit contents, and packaging for holding and describing the kit contents. In some examples, the kits include at least two standards selected from sample polynucleotides having a VAF of 0%, 0.1%, 0.25%, 0.5%, 1%, or 2% relative to the wild-type genomic sequence. In some examples, the kits include five standards, each having a VAF of 0%, 0.1%, 0.25%, 0.5%, 1%, or 2% relative to the wild-type genomic sequence. In some examples, the kits include instructions for use of the reference standards with one or more sequencing instruments or other instruments configured to measure genomic variants. In some examples, the reference standards are packaged in a buffer solution. In some examples, the reference standards are packaged in a tube. In some examples, the reference standards are not packaged in a plasma-like format. In some examples, the reference standards include 500 ng to 5 micrograms of total DNA.
[0131] Next-generation sequencing The downstream application of polynucleotide library (for example, sample polynucleotide set or reference standard) can include next-generation sequencing.For example, enrichment of target sequence by controlled stoichiometry polynucleotide probe library leads to more efficient sequencing.The performance of polynucleotide library to capture or hybridize to target can be defined by several different metrics that describe efficiency, precision and accuracy. For example, Picard metrics may include variables such as HS library size (the number of unique molecules in the library corresponding to the target region is calculated from read pairs), average target coverage (the percentage of bases that reach a particular coverage level), depth of coverage (the number of reads containing a given nucleotide), fold enrichment (sequence reads that uniquely map to the target / reads that map to all samples multiplied by the total sample length / target length), percent off-bait bases (the percentage of bases that do not correspond to a base in the probe / bait), percent off-target (the percentage of bases that do not correspond to the base of interest), available bases on the target, AT or GC dropout rate, fold 80 base penalty (the fold overcoverage required to bring 80 percent of the non-zero targets up to the average coverage level), percent zero-coverage targets, PF reads (the number of reads that pass the quality filter), percent selected bases (the sum of on-bait and near-bait bases divided by the total aligned bases), percent overlap, or other variables that meet specifications.
[0132] Read depth (sequencing depth, or sampling) represents the total number of times a sequenced nucleic acid fragment ("read") is obtained for a sequence. Theoretical read depth is defined as the expected number of times the same nucleotide would be read if the reads were perfectly distributed across an idealized genome. Read depth is expressed as a function of percent coverage (or coverage width). For example, 10 million reads for a perfectly distributed 1 million-base genome would theoretically result in a 10x read depth of 100% of the sequence. In practice, a larger number of reads (higher theoretical read depth, or oversampling) may be required to achieve the desired read depth for a percentage of the target sequence. Enriching target sequences with a controlled stoichiometry probe library increases the efficiency of downstream sequencing because fewer total reads are required to obtain results with acceptable read numbers over a desired percentage (%) of the target sequence. For example, in some instances, a 55x theoretical read depth of the target sequence results in at least 30x coverage of at least 90% of the sequence. In some examples, a theoretical read depth of 55x or less of a target sequence results in a read depth of at least 30x for at least 80% of the sequences. In some examples, a theoretical read depth of 55x or less of a target sequence results in a read depth of at least 30x for at least 95% of the sequences. In some examples, a theoretical read depth of 55x or less of a target sequence results in a read depth of at least 10x for at least 98% of the sequences. In some examples, a theoretical read depth of 55x or less of a target sequence results in a read depth of at least 20x for at least 98% of the sequences. In some examples, a theoretical read depth of 55x or less of a target sequence results in a read depth of at least 5x for at least 98% of the sequences. Increasing the concentration of the probe during hybridization with the target can lead to increased read depth. In some examples, the concentration of the probe is increased by at least 1.5x, 2.0x, 2.5x, 3x, 3.5x, 4x, 5x, or more than 5x.In some examples, increasing the probe concentration results in at least a 1000% increase in read depth, or a 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 500%, 750%, 1000% or greater increase. In some examples, increasing the probe concentration by 3 times results in a 1000% increase in read depth. In some examples, sequencing is performed to achieve a theoretical read depth of at least 30X, 50X, 100X, 150X, 200X, 250X, 300X, 500X, or at least 1000X. In some examples, sequencing is performed to achieve a theoretical read depth of about 30X, 50X, 100X, 150X, 200X, 250X, 300X, 500X, or about 1000X. In some examples, sequencing is performed to achieve a theoretical read depth of 30X or less, 50X or less, 100X or less, 150X or less, 200X or less, 250X or less, 300X or less, 500X or less, or 1000X or less. In some examples, sequencing is performed to achieve an actual read depth of at least 30X, 50X, 100X, 150X, 200X, 250X, 300X, 500X, or at least 1000X. In some examples, sequencing is performed to achieve an actual read depth of 30X or less, 50X or less, 100X or less, 150X or less, 200X or less, 250X or less, 300X or less, 500X or less, or 1000X or less. In some examples, sequencing is performed to achieve an actual read depth of about 30X, 50X, 100X, 150X, 200X, 250X, 300X, 500X, or about 1000X.
[0133] The on-target rate represents the percentage of sequencing reads that correspond to the desired target sequence. In some examples, the controlled stoichiometric polynucleotide probe library results in an on-target rate of at least 30%, or at least 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, or at least 90%. Increasing the concentration of polynucleotide probes in contact with target nucleic acids results in an increase in on-target rate. In some examples, the probe concentration increases by at least 1.5x, 2x, 2.5x, 3x, 3.5x, 4x, 5x, or more than 5x. In some examples, increasing the probe concentration results in an increase in on-target binding of at least 20%, or 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, or at least 500%. In some instances, a three-fold increase in probe concentration results in a 20% increase in on-target rate.
[0134] Coverage uniformity is sometimes calculated as the read depth as a function of target sequence identity. Higher coverage uniformity results in fewer sequencing reads being required to achieve the desired read depth. For example, the characteristics of the target sequence can affect read depth, such as high or low GC or AT content, repetitive sequences, trailing adenines, secondary structure, affinity for target sequence binding (for amplification, enrichment, or detection), stability, melting temperature, biological activity, ability to assemble into larger fragments, sequences containing modified nucleotides or nucleotide analogs, or any other characteristics of polynucleotides. Enriching target sequences with a controlled stoichiometric polynucleotide probe library results in higher coverage uniformity after sequencing. In some examples, 95% of the sequences have a read depth within 1x of the average library read depth, or within about 0.05, 0.1, 0.2, 0.5, 0.7, 1, 1.2, 1.5, 1.7, or 2x of the average library read depth. In some examples, 80%, 85%, 90%, 95%, 97%, or 99% of the sequences have a read depth that is within 1x of the average.
[0135] Enrichment of target nucleic acid molecules with polynucleotide probe libraries The probe libraries described herein can be used to enrich target polynucleotides present in a population of sample polynucleotides for various downstream applications. In some examples, a sample is obtained from one or more sources, and a population of sample polynucleotides is isolated. The sample is obtained from a biological source (non-limiting examples include saliva, blood, tissue, skin, or a completely synthetic source). Multiple polynucleotides obtained from the sample are fragmented, end-repaired, and adenylated to form double-stranded sample nucleic acid fragments. In some examples, end-repair is achieved by treatment with one or more enzymes, such as T4 DNA polymerase, Klenow enzyme, and T4 polynucleotide kinase, in an appropriate buffer. Nucleotide overhangs to facilitate ligation to adapters are added, optionally with a 3'-5' exo-minus Klenow fragment and dATP.
[0136] Adapters (e.g., universal adapters) may be ligated to both ends of sample polynucleotide fragments using a ligase such as T4 ligase to create a library of adapter-tagged polynucleotide strands, and the adapter-tagged polynucleotide library is amplified using primers such as universal primers. In some examples, the adapters are Y-shaped adapters containing one or more primer binding sites, one or more graft regions, and one or more index (or barcode) regions. In some examples, one or more index regions are present on each strand of the tire adapter. In some examples, the graft regions are complementary to the flow cell surface, facilitating next-generation sequencing of the sample library. In some examples, the Y-shaped adapters contain partially complementary sequences. In some examples, the Y-shaped adapters contain a single thymidine overhang that hybridizes to the overhanging adenine of the double-stranded adapter-tagged polynucleotide strands. The Y-shaped adapters may contain modified nucleic acids that are resistant to cleavage. For example, a phosphorothioate backbone is used to attach an overhanging thymidine to the 3' end of the adapter. When using universal primers, library amplification is performed to add barcoded primers to adapters. The library of double-stranded adapter-tagged polynucleotide strands is contacted with polynucleotide probes to form hybrid pairs. These pairs are separated from unhybridized fragments and isolated from the probes to produce an enriched library. The enriched library can then be sequenced.
[0137] The library of double-stranded sample nucleic acid fragments is then denatured in the presence of an adapter blocker. The adapter blocker minimizes off-target hybridization of probes to adapter sequences (instead of target sequences) present on the adapter-tagged polynucleotide strands and / or prevents intermolecular hybridization of adapters (e.g., "daisy chaining"). Denaturation is sometimes performed at 96°C, or about 85, 87, 90, 92, 95, 97, 98, or about 99°C. The polynucleotide targeting library (probe library) is denatured in a hybridization solution, in some instances at 96°C, about 85, 87, 90, 92, 95, 97, 98, or 99°C. The denatured adapter-tagged polynucleotide library and hybridization solution are incubated for an appropriate amount of time and at an appropriate temperature to allow the probes to hybridize with their complementary target sequences. In some examples, a suitable hybridization temperature is about 45-80°C, or at least 45, 50, 55, 60, 65, 70, 75, 80, 85, or 90°C. In some examples, the hybridization temperature is 70°C. In some examples, a suitable hybridization time is 16 hours, or at least 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, or more than 22 hours, or about 12-20 hours. A binding buffer is then added to the hybridized adaptor-tagged polynucleotide probes, and a solid support containing a capture moiety is used to selectively bind the hybridized adaptor-tagged polynucleotide probes. The solid support is washed with a buffer to remove unbound polynucleotides, and then an elution buffer is added to release the enriched tagged polynucleotide fragments from the solid support. In some examples, the solid support is washed twice, or one, two, three, four, five, or six times. The enriched library of adaptor-tagged polynucleotide fragments is amplified and the enriched library is sequenced.
[0138] A plurality of nucleic acids (e.g., genomic sequences) are obtained from a sample, fragmented, optionally end-repaired, and adenylated. Adapters are ligated to both ends of the polynucleotide fragments to generate a library of adapter-tagged polynucleotide strands, and the adapter-tagged polynucleotide library is amplified. The adapter-tagged polynucleotide library is then denatured in the presence of an adapter blocker at elevated temperatures, preferably 96°C. A polynucleotide targeting library (probe library) is denatured in a hybridization solution at elevated temperatures, preferably about 90-99°C, and mixed with the denatured tagged polynucleotide library in the hybridization solution at about 45-80°C for about 10-24 hours. A binding buffer is then added to the hybridized tagged polynucleotide probes, and a solid support containing a capture moiety is used to selectively bind the hybridized adapter-tagged polynucleotide probes. The solid support is washed with buffer one or more times, preferably about two and five times, to remove unbound polynucleotides, followed by the addition of an elution buffer to release the enriched adapter-tagged polynucleotide fragments from the solid support. The enriched library of adaptor-tagged polynucleotide fragments is amplified, and then the library is sequenced. Alternative variables such as incubation time, temperature, reaction volume / concentration, number of washes, or other variables consistent with this specification may also be used in this method.
[0139] In either case, detection or quantitative analysis of the oligonucleotides can be achieved by sequencing. Subunits or the entire synthesized oligonucleotide can be detected by any suitable method known in the art, including, for example, the sequencing methods described herein, such as Illumina sequencing by synthesis, PacBio nanopore sequencing, or BGI / MGI nanopore sequencing, for complete sequencing of all oligonucleotides.
[0140] Sequencing can be achieved by the classical Sanger sequencing method, which is well known in the art.Sequencing can also be achieved using high-throughput systems, some of which allow the detection of sequenced nucleotide immediately after its incorporation into growing chain or after incorporation, for example, detection of sequence in real time or substantially real time.In some cases, high-throughput sequencing generates at least 1,000, at least 5,000, at least 10,000, at least 20,000, at least 30,000, at least 40,000, at least 50,000, at least 100,000 or at least 500,000 sequence reads per hour, and each read is at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 120 or at least 150 bases / read.
[0141] In either example, detection or quantitative analysis of the oligonucleotides can be achieved by sequencing. Subunits or the entire synthesized oligonucleotide can be detected by any suitable method known in the art, including the sequencing methods described herein, such as complete sequencing of all oligonucleotides by Illumina sequencing by synthesis, PacBio nanopore sequencing, or BGI / MGI nanopore sequencing.
[0142] In some instances, high-throughput sequencing involves the use of technology available from the ABI Solid Systems. This genetic analysis platform allows for massively parallel sequencing of clonally amplified DNA fragments linked to beads. The sequencing method is based on sequential ligation with dye-labeled oligonucleotides.
[0143] Next-generation sequencing can include ion semiconductor sequencing. Ion semiconductor sequencing takes advantage of the fact that ions can be released when a nucleotide is incorporated into a DNA strand. To perform ion semiconductor sequencing, a high-density array of microfabricated wells can be formed. Each well can hold a single DNA template. An ion-sensitive layer can be located beneath the well, and an ion sensor can be located beneath the ion-sensitive layer. When a nucleotide is added to the DNA, H+ can be released, which can be measured as a change in pH. The H+ ions can be converted to a voltage and recorded by the semiconductor sensor. The array chip can be flooded sequentially with nucleotides one after the other. No vacuum cleaner, light, or camera is required. In some cases, nucleic acids are sequenced using the IONPROTON™ sequencer. In other cases, the IONPGM™ sequencer is used. The Ion Torrent Personal Genome Machine (PGM) can generate 10 million reads in two hours.
[0144] In some instances, high-throughput sequencing involves the use of technology such as the Single Molecule Sequencing by Synthesis (SMSS) method available from Helicos BioSciences Corporation (Cambridge, Massachusetts). SMSS is unique because it allows sequencing of the entire human genome at up to 24 cycles. Finally, SMSS is powerful because, like MW technology, it does not require a pre-amplification step before hybridization. In fact, SMSS does not require any amplification.
[0145] In some cases, high-throughput sequencing involves the use of technology available from 454 Lifesciences, Inc. (Branford, Connecticut), such as the Pico Titer Plate device, which contains a fiber optic plate that transmits the chemiluminescent signal generated by the sequencing reaction and is recorded by a CCD camera within the instrument. The use of this fiber optic allows for the detection of at least 20 million base pairs in 4.5 hours.
[0146] A method using bead amplification followed by fiber optic detection is described in Marguiles et al., "Genome sequencing in microfabricated high-density picolitre reactors," Nature, 2005, Vol. 437, Pages 376-380.
[0147] In some examples, high-throughput sequencing is performed using Clonal Single Molecule Array (Solexa, Inc.) or sequencing by synthesis (SBS) using reversible terminator chemistry. See, for example, Constants, A., The Scientist, 2003. Vol. 17, Issue 13, Page 36. High-throughput sequencing of oligonucleotides can be achieved using any suitable sequencing method known in the art, such as those commercially available from Pacific Biosciences, Complete Genomics, Genia Technologies, Halcyon Molecular, Oxford Nanopore Technologies, etc. Overall, such systems involve sequencing a target oligonucleotide molecule having multiple bases by measuring the time-dependent addition of bases via a polymerization reaction on the oligonucleotide molecule, i.e., the activity of a nucleic acid polymerase on the template oligonucleotide molecule to be sequenced is tracked in real time. The catalytic activity of the nucleic acid polymerase at each step in the sequence of base addition can then be used to determine which bases are incorporated into the growing complementary strand of the target oligonucleotide, thereby deducing the sequence. A polymerase on the target oligonucleotide molecule complex is provided in a position suitable for moving along the target oligonucleotide molecule and extending the oligonucleotide primer at the active site. Multiple labeled types of nucleotide analogs are provided in proximity to the active site, with each distinct type of nucleotide analog being complementary to a different nucleotide in the target oligonucleotide sequence. The growing oligonucleotide chain is extended by using the polymerase to add nucleotide analogs to the oligonucleotide chain at the active site, with the added nucleotide analogs being complementary to nucleotides of the target oligonucleotide at the active site. The nucleotide analogs added to the oligonucleotide primer as a result of the polymerization step are identified.The steps of providing a labeled nucleotide analog, polymerizing the growing oligonucleotide chain, and identifying the added nucleotide analog are repeated so that the oligonucleotide chain is further extended and the sequence of the target oligonucleotide is determined.
[0148] Next-generation sequencing technologies include Pacific Biosciences' real-time (SMRT™) technology. In SMRT, each of the four DNA bases can be conjugated to one of four different fluorescent dyes. These dyes can be phosphoconjugated. A single DNA polymerase can be immobilized with a single molecule of template single-stranded DNA at the bottom of a zero-mode waveguide (ZMW). The ZMW can be a confining structure that allows observation of the incorporation of a single nucleotide by the DNA polymerase against a background of fluorescent nucleotides that can rapidly diffuse out of the ZMW (in microseconds). Incorporation of the nucleotide into the growing strand can take several milliseconds. During this time, the fluorescent label can be excited to generate a fluorescent signal, and the fluorescent tag can be cleaved. The ZMW can be illuminated from below. Attenuated light from the excitation beam can penetrate the bottom 20–30 nm of each ZMW. Microscopes with detection limits of 20 zeptoliters (10 in. liters) can be created. The small detection volume can provide a 1000-fold improvement in background noise reduction. Detection of the corresponding fluorescence of the dye can indicate which base has been incorporated. This process can be repeated.
[0149] In some cases, next-generation sequencing is nanopore sequencing. See, for example, Soni et al., Clin Chem., 2007, Vol. 53, Pages 1996-2001. Nanopores can be small, with diameters on the order of about 1 nanometer. When a nanopore is immersed in a conductive fluid and a potential is applied across it, a small current can be generated due to the conduction of ions through the nanopore. The amount of current that flows can be sensitive to the size of the nanopore. When a DNA molecule passes through the nanopore, each nucleotide on the DNA molecule can block the nanopore to a different extent. Therefore, changes in the current passing through the nanopore as the DNA molecule passes through the nanopore can represent a readout of the DNA sequence. Nanopore sequencing technology is available from Oxford Nanopore Technologies, e.g., the GridlON system. A single nanopore can be inserted into the polymer membrane across the top of a microwell. Each microwell can have an individual sensing electrode. The microwells can be fabricated into array chips with 100,000 or more microwells per chip (e.g., greater than 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, or 1,000,000). Instruments (or nodes) can be used to analyze the chip. Data can be analyzed in real time. One or more instruments can be operated at a time. The nanopore can be a protein nanopore, e.g., protein alpha hemolysin, a heptameric protein pore. The nanopore can be a solid-state nanopore, e.g., fabricated with a nanometer-sized pore formed in a synthetic membrane (e.g., SiNx or SiO2). The nanopore can be a hybrid pore (e.g., incorporation of a protein pore into a solid membrane). The nanopore can be a nanopore with an integrated sensor (e.g., a tunneling electrode detector, a capacitance detector, or a graphene-based nanogap or edge state detector). See, for example, Garaj et al., Nature, 2010, Vol. 467, Pages 190-193.Nanopores can be functionalized to analyze specific types of molecules (e.g., DNA, RNA, or proteins). Nanopore sequencing can include "strand sequencing," in which intact DNA polymers can be passed through a protein nanopore while the DNA is sequenced in real time as it moves through the pore. An enzyme can separate the strands of double-stranded DNA and feed the strands through the nanopore. The DNA can have a hairpin at one end, allowing the system to read both strands. In some cases, nanopore sequencing is "exonuclease sequencing," in which individual nucleotides can be cleaved from the DNA strand by a processive exonuclease and the nucleotides can pass through the protein nanopore. The nucleotides can be transiently bound to a molecule (e.g., cyclodextran) in the pore. A characteristic disruption of the current can be used to identify the base.
[0150] Nanopore sequencing technology from GENIA can be used. Engineered protein pores can be embedded in lipid bilayer membranes. "Active control" techniques can be used to enable efficient nanopore-membrane assembly and control of DNA translocation through the channel. In some cases, the nanopore sequencing technology is from NABsys. Genomic DNA can be fragmented into strands with an average length of approximately 100 kb. The 100 kb fragments can be single-stranded and subsequently hybridized with 6-mer probes. The genomic fragments with the probes can be driven through a nanopore, which can create a current versus time trace. The current trace can provide the location of the probe on each genomic fragment. The genomic fragments can be aligned to create a probe map of the genome. This process can be performed in parallel for a library of probes. A genome-length probe map for each probe can be generated. Errors can be fixed with a process called "moving window sequencing by hybridization (mwSBH)." In some cases, the nanopore sequencing technology is from IBM / Roche. Nanopore-sized openings can be created in microchips using electron beams. Electric fields can be used to pull or thread DNA through the nanopore. A DNA transistor device within the nanopore can involve modifying alternating nanometer-sized layers of metal and dielectric. Discrete charges in the DNA backbone can be trapped by the electric field inside the DNA nanopore. By turning the gate voltage off and on, the DNA sequence can be read.
[0151] Next-generation sequencing can include DNA nanoball sequencing. See, for example, Complete Genomics and Drmanac et al., Science, 2010, Vol. 327, pp. 78-81. DNA can be isolated, fragmented, and size-selected. For example, DNA can be fragmented (e.g., by sonication) to an average length of approximately 500 bp. Adapters (Adl) can be attached to the ends of the fragments. The adapters can be used to hybridize to anchors for sequencing reactions. DNA with adapters attached to each end can be PCR-amplified. The adapter sequence can be modified so that complementary single-stranded ends can ligate to each other to form circular DNA. DNA can be methylated to protect it from cleavage by type IIS restriction enzymes used in subsequent steps. The adapter (e.g., the correct adapter) can have a restriction recognition site, or the restriction recognition site can remain unmethylated. The unmethylated restriction recognition site in the adapter can be recognized by a restriction enzyme (e.g., Acul), and the DNA can be cleaved 13 bp to the right of the right adapter by Acul to form a linear double-stranded DNA. A second-round right and left adapter (Ad2) can be ligated to either end of the linear DNA, and all DNA bound by both adapters can be PCR amplified (e.g., by PCR). The Ad2 sequence can be modified to allow them to ligate together to form circular DNA. The DNA can be methylated, but the restriction enzyme recognition site can remain unmethylated on the left Ad1 adapter. A restriction enzyme (e.g., Acul) can be applied, cleaving the DNA 13 bp to the left of Ad1 to form a linear DNA fragment. A third-round left and right adapter (Ad3) can be ligated to the left and right flanks of the linear DNA, and the resulting fragments can be PCR amplified. The adapters can be modified to allow them to ligate together to form circular DNA. A type III restriction enzyme (eg, EcoP15) can be added, which can cut DNA 26 bp to the left for Ad3 and 26 bp to the right for Ad2.This cleavage removes large fragments of DNA and can re-linearize the DNA. Fourth-round right and left adapters (Ad4) can be ligated to the DNA, the DNA amplified (e.g., by PCR), and modified so that they can recombine with each other and form the completed circular DNA template.
[0152] Rolling circle replication (e.g., using Phi 29 DNA polymerase) can be used to amplify small fragments of DNA. Four adapter sequences can contain palindromic sequences that can hybridize, allowing a single strand to fold back on itself to form a DNA nanoball (DNB™), which can be approximately 200–300 nanometers in diameter on average. The DNA nanoballs can be attached (e.g., by adsorption) to a microarray (sequencing flow cell). The flow cell can be a silicon wafer coated with silicon dioxide, titanium, hexamethyldisilazane (HMDS), and a photoresist material. Sequencing can be performed by non-linked sequencing by ligating fluorescent probes to DNA. The fluorescent color at the interrogation position can be visualized with a high-resolution camera. The identity of the nucleotide sequence between the adapter sequences can be determined.
[0153] A population of polynucleotides can be enriched prior to adapter ligation. In one example, a plurality of polynucleotides are obtained from a sample, fragmented, optionally end-repaired, and denatured at high temperatures, preferably 90-99°C. A polynucleotide targeting library (probe library) is denatured in a hybridization solution at high temperatures, preferably about 90-99°C, and mixed with the denatured tagged polynucleotide library in the hybridization solution at about 45-80°C for about 10-24 hours. A binding buffer is then added to the hybridized tagged polynucleotide probes, and a solid support containing a capture moiety is used to selectively bind the hybridized adapter-tagged polynucleotide probes. The solid support is washed with buffer one or more times, preferably about two and five times, to remove unbound polynucleotides, followed by the addition of an elution buffer to release the enriched adapter-tagged polynucleotide fragments from the solid support. The enriched polynucleotide fragments are then polyadenylated, adapters are ligated to both ends of the polynucleotide fragments to generate a library of adapter-tagged polynucleotide strands, and the adapter-tagged polynucleotide library is amplified. The adapter-tagged polynucleotide library is then sequenced.
[0154] Undesired sequences can also be filtered from a plurality of polynucleotides by hybridizing to the undesired fragments using a targeting library. For example, a plurality of polynucleotides are obtained from a sample, fragmented, optionally end-repaired, and adenylated. Adapters are ligated to both ends of the polynucleotide fragments to generate a library of adapter-tagged polynucleotide strands, and the adapter-tagged polynucleotide library is then amplified. Alternatively, the adenylation and adapter ligation steps are performed after enrichment of the sample polynucleotides. The adapter-tagged polynucleotide library is then denatured at high temperatures, preferably 90-99°C, in the presence of an adapter blocker. A polynucleotide filtering library (probe library) designed to remove undesired nonspecific sequences is denatured in a hybridization solution at high temperatures, preferably about 90-99°C, and mixed with the denatured tagged polynucleotide library in the hybridization solution at about 45-80°C for about 10-24 hours. A solid support containing a capture moiety is then used to selectively bind the hybridized adapter-tagged polynucleotide probes. The solid support is washed with buffer one or more times, preferably about 1-5 times, to elute unbound adaptor-tagged polynucleotide fragments. The enriched library of unbound adaptor-tagged polynucleotide fragments is amplified, and the amplified library is then sequenced.
[0155] Highly parallel novel nucleic acid molecule synthesis Described herein is a platform approach that utilizes miniaturization, parallelization, and vertical integration of end-to-end processes from polynucleotide synthesis to gene assembly in silicon nanowells to create a revolutionary synthesis platform. The device described herein uses the same footprint as a 96-well plate, enabling the silicon synthesis platform to increase throughput by 100-1,000-fold compared to traditional synthesis methods, producing up to approximately 1,000,000 polynucleotides in a single, highly parallelized run. In some examples, a single silicon plate described herein provides for the synthesis of approximately 6,100 non-identical polynucleotides. In some examples, each non-identical polynucleotide is located within a cluster. A cluster can contain 0-500 non-identical polynucleotides.
[0156] The methods described herein provide for the synthesis of a library of polynucleotides, each encoding a predetermined variant of at least one predetermined reference nucleic acid sequence. In some cases, the predetermined reference sequence is a protein-encoding nucleic acid sequence, and the variant library includes sequences encoding at least a single codon variation, such that multiple different variants of a single residue in the subsequent protein encoded by the synthesized nucleic acid are generated by standard translation processes. Synthesized specific modifications in nucleic acid sequences can be introduced by incorporating nucleotide changes into overlapping or blunt-ended polynucleotide primers. Alternatively, a population of polynucleotides can collectively encode a long nucleic acid (e.g., a gene) and its variants. In this configuration, the population of polynucleotides can be hybridized and subjected to standard molecular biology techniques to form the long nucleic acid (e.g., a gene) and its variants. When the long nucleic acid (e.g., a gene) and its variants are expressed in a cell, a mutant protein library is generated. Similarly, methods are provided herein for synthesizing a variant library encoding RNA sequences (e.g., miRNA, shRNA, and mRNA) or DNA sequences (e.g., enhancer, promoter, UTR, and terminator regions). Also provided herein are downstream applications of variants selected from libraries synthesized using the methods described herein, including the identification of variant nucleic acid or protein sequences with enhanced biologically relevant functions, such as biochemical affinities, enzymatic activities, altered cellular activities, and the treatment or prevention of disease states.
[0157] substrate Provided herein is a substrate comprising a plurality of clusters, each cluster comprising a plurality of loci supporting polynucleotide binding and synthesis. As used herein, the term "locus" refers to a structurally distinct region extending from a surface that supports a polynucleotide encoding a single predetermined sequence. In some examples, a locus is on a two-dimensional surface, such as a substantially flat surface. In some examples, a locus refers to a discrete elevated or depressed site on a surface, such as a well, microwell, channel, or post. In some examples, the surface of the locus comprises a material that is actively functionalized to bind at least one nucleotide for polynucleotide synthesis, or preferably a population of identical nucleotides for the synthesis of a population of polynucleotides. In some examples, a polynucleotide refers to a population of polynucleotides encoding the same nucleic acid sequence. In some examples, the surface of the device comprises one or more surfaces of a substrate.
[0158] Provided herein are structures that can include a surface that supports the synthesis of multiple polynucleotides with different predetermined sequences at addressable locations on a common support. In some examples, the device can accommodate 2,000, 5,000, 10,000, 20,000, 30,000, 50,000, 75,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000, 1, Support for the synthesis of 200,000, 1,400,000, 1,600,000, 1,800,000, 2,000,000, 2,500,000, 3,000,000, 3,500,000, 4,000,000, 4,500,000, 5,000,000, 10,000,000 or more non-identical polynucleotides is provided. In some examples, the device may include a plurality of different sequences encoding 2,000, 5,000, 10,000, 20,000, 30,000, 50,000, 75,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000, In some instances, the present invention provides support for the synthesis of more than 1,000, 1,200,000, 1,400,000, 1,600,000, 1,800,000, 2,000,000, 2,500,000, 3,000,000, 3,500,000, 4,000,000, 4,500,000, 5,000,000, 10,000,000 or more polynucleotides. In some instances, at least some of the polynucleotides have identical sequences or are configured to be synthesized with identical sequences.
[0159] surface material Provided herein are devices including surfaces that are modified to support polynucleotide synthesis at predetermined locations, resulting in low error rates, low dropout rates, high yields, and high oligo expression. In some examples, the surfaces of the polynucleotide synthesis apparatus provided herein are fabricated from a variety of materials that can be modified to support de novo polynucleotide synthesis reactions. In some cases, the device is sufficiently conductive, e.g., capable of forming a uniform electric field across all or a portion of the device. The devices described herein may include flexible materials. Exemplary flexible materials include, but are not limited to, modified nylon, unmodified nylon, nitrocellulose, and polypropylene. The devices described herein may include rigid materials. Exemplary rigid materials include, but are not limited to, glass, fused silica, silicon, silicon dioxide, silicon nitride, plastics (e.g., polytetrafluoroethylene, polypropylene, polystyrene, polycarbonate, and blends thereof), and metals (e.g., gold, platinum). The devices disclosed herein can be fabricated from materials including silicon, polystyrene, agarose, dextran, cellulose polymers, polyacrylamide, polydimethylsiloxane (PDMS), glass, or any combination thereof. In some cases, the devices disclosed herein are manufactured using a combination of the materials listed herein or any other suitable materials known in the art.
[0160] Surface Architecture Devices with raised and / or depressed features are provided herein. One advantage of having such features is an increased surface area for supporting polynucleotide synthesis. In some examples, devices with raised and / or depressed features are referred to as three-dimensional substrates. In some examples, the three-dimensional device includes one or more channels. In some examples, one or more locations include channels. In some examples, the channels are accessible for reagent deposition via a deposition device, such as a polynucleotide synthesizer. In some examples, reagents and / or fluids are collected in a larger well that is fluidically connected to one or more channels. For example, the device includes multiple channels corresponding to multiple loci having a cluster, and the multiple channels are fluidically connected to one well of the cluster. In some methods, a library of polynucleotides is synthesized at multiple loci of the cluster.
[0161] surface modification In various examples, surface modification is employed to chemically and / or physically alter a surface by an additive or subtractive process to alter one or more chemical and / or physical properties of the device surface or selected sites or regions of the device surface. For example, surface modification includes, but is not limited to, (1) changing the wetting properties of the surface, (2) functionalizing the surface, e.g., providing, modifying, or substituting surface functional groups, (3) defunctionalizing the surface, e.g., removing surface functional groups, (4) otherwise altering the chemical composition of the surface, e.g., through etching, (5) increasing or decreasing surface roughness, (6) providing a coating on the surface, e.g., a coating that exhibits wetting properties different from those of the surface, and / or (7) depositing particulates on the surface.
[0162] Polynucleotide Synthesis The disclosed method for polynucleotide synthesis may include a process involving phosphoramidite chemistry. In some examples, polynucleotide synthesis involves combining a base with a phosphoramidite. Polynucleotide synthesis may involve combining a base by deposition of a phosphoramidite under coupling conditions, where the same base is optionally deposited multiple times with the phosphoramidite, e.g., double-coupled. Polynucleotide synthesis may include capping of unreacted sites. In some examples, capping is optional. Polynucleotide synthesis may also include oxidation, one oxidation step, or multiple oxidation steps. Polynucleotide synthesis may include deblocking, detritylation, and sulfurization. In some examples, polynucleotide synthesis includes either oxidation or sulfurization. In some examples, between one or each step during the polynucleotide synthesis reaction, the device is washed, for example, using tetrazole or acetonitrile. The time frame for any one step in the phosphoramidite synthesis method can be less than about 2 minutes, 1 minute, 50 seconds, 40 seconds, 30 seconds, 20 seconds, and 10 seconds.
[0163] Large polynucleotide libraries with low error rates The average error rate of polynucleotides synthesized in libraries using the provided systems and methods can be less than 1 in 1000, less than 1 in 1250, less than 1 in 1500, less than 1 in 2000, less than 1 in 3000, or even less. In some examples, the average error rate of polynucleotides synthesized in libraries using the provided systems and methods is 1 / 500, 1 / 600, 1 / 700, 1 / 800, 1 / 900, 1 / 1000, 1 / 1100, 1 / 1200, 1 / 1250, 1 / 1300, 1 / 1400, 1 / 1500, 1 / 1600, 1 / 1700, 1 / 1800, 1 / 1900, 1 / 2000, 1 / 3000, or less. In some examples, the average error rate of polynucleotides synthesized in libraries using the provided tire systems and methods is less than 1 / 1000.
[0164] In some examples, the aggregation error rate of polynucleotides synthesized in a library using the provided systems and methods is 1 / 500, 1 / 600, 1 / 700, 1 / 800, 1 / 900, 1 / 1000, 1 / 1100, 1 / 1200, 1 / 1250, 1 / 1300, 1 / 1400, 1 / 1500, 1 / 1600, 1 / 1700, 1 / 1800, 1 / 1900, 1 / 2000, 1 / 3000, or less, compared to a predetermined sequence. In some examples, the total error rate of polynucleotides synthesized in a library using the provided systems and methods is less than 1 / 500, 1 / 600, 1 / 700, 1 / 800, 1 / 900, or 1 / 1000. In some examples, the total error rate of polynucleotides synthesized in a library using the provided systems and methods is less than 1 / 1000.
[0165] In some examples, error correction enzymes can be used in the polynucleotides synthesized in the library using the provided system and available methods.In some examples, the total error rate of the polynucleotides with error correction compared with the predetermined sequence is 1 / 500, 1 / 600, 1 / 700, 1 / 800, 1 / 900, 1 / 1000, 1 / 1100, 1 / 1200, 1 / 1300, 1 / 1400, 1 / 1500, 1 / 1600, 1 / 1700, 1 / 1800, 1 / 1900, 1 / 2000, 1 / 3000 or less.In some examples, the total error rate of the polynucleotides with error correction synthesized in the library using the provided system and methods can be 1 / 500, 1 / 600, 1 / 700, 1 / 800, 1 / 900 or less than 1 / 1000. In some examples, the total error rate with error correction for polynucleotides synthesized in libraries using the systems and methods provided can be less than 1 in 1000.
[0166] The error rate can limit the value of gene synthesis for producing libraries of gene variants. With an error rate of 1 / 300, approximately 0.7% of clones in a 1500 base pair gene are correct. Because most errors from polynucleotide synthesis result in frameshift mutations, more than 99% of the clones in such a library do not produce full-length proteins. Reducing the error rate by 75% increases the percentage of correct clones by 40-fold. The disclosed methods and compositions enable rapid de novo synthesis of large polynucleotide and gene libraries with lower error rates than commonly observed gene synthesis methods due to both improved synthesis quality and the applicability of error correction methods, which are enabled in a massively parallel and time-efficient manner. Thus, libraries may be designed to provide a range of 1 / 300, 1 / 400, 1 / 500, 1 / 600, 1 / 700, 1 / 800, 1 / 900, 1 / 1000, 1 / 1250, 1 / 1500, 1 / 2000, 1 / 2500, 1 / 3000, 1 / 4000, 1 / 5000, 1 / 6000, 1 / 7000, 1 / 80 ...1000, 1 / 1250, 1 / 1500, 1 / 2000, 1 / 2500, 1 / 3000, 1 / 4000, 1 / 5000, 1 / 6000, 1 / 7000, 1 / 8000, 1 / 1000, 1 / 1250, 1 / 1500, 1 / 2000, 1 / 2500, 1 / 3000, 1 / 4000, 1 / 5000, 1 / 6000, 1 / 7000, 1 / 8000, 1 / 1000, 1 / 1000, 1 / 1250, 1 / 1500, 1 / 2000, 1 / 2500, 1 / 3000, 1 / 4000, 1 / 5000, 1 / 6000, 1 / 7000, 1 / 8 1 / 9000, 1 / 10000, 1 / 12000, 1 / 15000, 1 / 20000, 1 / 25000, 1 / 30000, 1 / 40000, 1 / 50000, 1 / 60000, 1 / 70000, 1 / 80000, 1 / 90000, 1 / 100000, 1 / 125000, 1 / 150000, 1 / 200000, 1 / 300000, 1 / 400000, 1 / 500000, 1 / 600000, 1 / 700000, 1 / 800000, 1 / 900000, 1 / 1000000 or less.The disclosed methods and compositions further relate to large synthetic polynucleotide and gene libraries that have a low error rate associated with at least 30%, 40%, 50%, 60%, 70%, 75%, 80%, 85%, 90%, 93%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.8%, 99.9%, 99.95%, 99.98%, 99.99% or more of the polynucleotides or genes in at least a subset of the library that are associated with error-free sequences compared to predetermined / preselected sequences. In some examples, at least 30%, 40%, 50%, 60%, 70%, 75%, 80%, 85%, 90%, 93%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.8%, 99.9%, 99.95%, 99.98%, 99.99% or more of the polynucleotides or genes in an isolated volume in a library have the same sequence. In some examples, at least 30%, 40%, 50%, 60%, 70%, 75%, 80%, 85%, 90%, 93%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.8%, 99.9%, 99.95%, 99.98%, 99.99% or more of any related polynucleotides or genes with greater than 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8% or more similarity or identity have the same sequence. In some examples, the error rate associated with a particular locus on a polynucleotide or gene is optimized.Thus, a given locus or multiple selected loci of one or more polynucleotides or genes as part of a larger library may have a nucleotide sequence of 1 / 300, 1 / 400, 1 / 500, 1 / 600, 1 / 700, 1 / 800, 1 / 900, 1 / 1000, 1 / 1250, 1 / 1500, 1 / 2000, 1 / 2500, 1 / 3000, 1 / 4000, 1 / 5000, 1 / 6000, 1 / 7000, 1 / 8000, 1 / 9000, 1 / 10000, 1 / 10000, 1 / 2000, 1 / 2500, 1 / 3000, 1 / 4000, 1 / 5000, 1 / 6000, 1 / 7000, 1 / 8000, 1 / 9000, 1 / 10000, 1 / 20 ...10000, 1 / 2000, 1 / 2500, 1 / 3000, 1 / 4000, 1 / 5000 The error rate may be less than 12,000, 1 / 15,000, 1 / 20,000, 1 / 25,000, 1 / 30,000, 1 / 40,000, 1 / 50,000, 1 / 60,000, 1 / 70,000, 1 / 80,000, 1 / 90,000, 1 / 100,000, 1 / 125,000, 1 / 150,000, 1 / 200,000, 1 / 300,000, 1 / 400,000, 1 / 500,000, 1 / 600,000, 1 / 700,000, 1 / 800,000, 1 / 900,000, or 1 / 1,000,000. In various examples, such error optimization loci may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 80 It may include 0, 900, 1000, 1500, 2000, 2500, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 30000, 50000, 75000, 100000, 500000, 1000000, 2000000, 3000000 or more loci. The error optimized loci may be distributed over at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 30000, 75000, 100000, 500000, 1000000, 2000000, 3000000 or more polynucleotides or genes.
[0167] Error rates can be achieved with or without error correction. Error rates can be achieved across the library or across more than 80%, 85%, 90%, 93%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.8%, 99.9%, 99.95%, 99.98%, 99.99% or more of the library.
[0168] The present disclosure is further illustrated by the following non-limiting sections.
[0169] Item 1. (a) a first polynucleotide adaptor, a first strand, the first strand comprising a first terminal adaptor region, a first non-complementary region, and a first yoke region; a first polynucleotide adaptor comprising a second strand, the second strand comprising a second terminal adaptor region, a second non-complementary region, and a second yoke region, wherein the first yoke region and the second yoke region are complementary and the first non-complementary region and the second non-complementary region are not complementary; (b) a second polynucleotide adaptor, a third strand, the third strand comprising a third terminal adaptor region, a third non-complementary region, and a third yoke region; a second polynucleotide adaptor comprising a fourth strand, the fourth strand comprising a fourth terminal adaptor region, a fourth non-complementary region, and a fourth yoke region, wherein the third yoke region and the fourth yoke region are complementary, and the third non-complementary region and the fourth non-complementary region are not complementary, The composition, wherein the first polynucleotide adaptor and the second polynucleotide adaptor are connected via a linker.
[0170] Item 2. The composition of item 1, wherein the linker is attached to the first terminal adaptor region and the third terminal adaptor region.
[0171] Item 3. The composition of item 1, wherein the linker is attached to the second terminal adaptor region and the fourth terminal adaptor region.
[0172] Item 4. The composition of item 1, wherein the linker connects the 5' ends of one or more strands.
[0173] Item 5. The composition of item 1, wherein the linker connects the 3' ends of one or more strands.
[0174] Item 6. The composition of any one of items 1 to 3, wherein the linker comprises an alkene, alkyne, triazine, triazole, oxime, sulfide, amide, or cyclooctane.
[0175] Item 7. The composition of any one of items 1 to 4, wherein the linker comprises a PEG unit.
[0176] Item 8. The composition of item 7, wherein the linker comprises 5 to 50 PEG units.
[0177] Item 9. The composition of any one of items 1 to 8, wherein the linker comprises spacer 18 (Sp18).
[0178] Item 10. The composition of any one of items 1 to 8, wherein the linker comprises 1 to 20 SpI8.
[0179] Item 11. The composition of any one of items 1 to 13, wherein the linker comprises a polynucleotide.
[0180] Item 12. The composition of item 11, wherein the linker comprises at least one cleavable base.
[0181] Item 13. The composition of item 12, wherein the cleavable base is a uracil base.
[0182] Item 14. The composition of item 12, wherein the linker comprises at least one restriction endonuclease site.
[0183] Item 15. The composition of item 11, wherein the linker comprises a duplex region.
[0184] Item 16. The composition of item 11, wherein the linker comprises at least one splint polynucleotide.
[0185] Item 17. The composition of item 1, wherein at least one splint polynucleotide is at least partially complementary to a portion of the first polynucleotide adaptor or the second polynucleotide adaptor.
[0186] Item 18. The composition of item 17, wherein the linker comprises at least two splint polynucleotides, and the at least two splint polynucleotides at least partially overlap each other.
[0187] Item 19. The composition of item 16, wherein the linker comprises at least two splint polynucleotides, and the at least two splint polynucleotides at least partially overlap with a portion of the first polynucleotide adaptor or the second polynucleotide adaptor.
[0188] Item 20. The composition of any one of items 1 to 19, wherein the first polynucleotide adaptor and the second polynucleotide adaptor are covalently attached to a linker.
[0189] Item 21. The composition of any one of items 1 to 19, wherein the first polynucleotide adaptor and the second polynucleotide adaptor are non-covalently attached to the linker.
[0190] Item 22. The composition of any one of items 1 to 21, wherein one or more of the first polynucleotide adaptor and the second polynucleotide adaptor comprises at least one barcode.
[0191] Item 23. The composition of item 15, wherein at least one barcode comprises one or more of a cell index, a sample index, or a unique molecular identifier (UMI).
[0192] Item 24. The composition of any one of items 1 to 23, wherein one or more of the first polynucleotide adaptor and the second polynucleotide adaptor comprises at least one alpha thiophosphate.
[0193] Item 25. The composition of any one of items 1 to 24, wherein the 5' end of at least one of the strands comprises a thymidine.
[0194] Item 26. An adaptor-ligated sample polynucleotide comprising the composition of any one of items 1 through 14, wherein the composition further comprises at least one sample polynucleotide.
[0195] Item 27. The adaptor-ligated sample polynucleotide of Item 26, wherein at least one sample polynucleotide is attached to a 3' end of a first strand and a 5' end of a second strand.
[0196] Item 28. The adaptor-ligated sample polynucleotide of item 26 or 27, wherein at least one sample polynucleotide is attached to the 3' end of the third strand and the 5' end of the fourth strand.
[0197] Item 29. The adaptor-ligated sample polynucleotide of Item 26, wherein at least one sample polynucleotide is attached to a 3'-end of a first strand and a 3'-end of a second strand.
[0198] Item 30. The adaptor-ligated sample polynucleotide of item 26 or 29, wherein at least one sample polynucleotide is attached to a 5'-end of the third strand and a 5'-end of the fourth strand.
[0199] Item 31. The adaptor-ligated sample polynucleotide according to any one of Items 26 to 30, wherein the sample polynucleotide comprises genomic DNA.
[0200] Item 32. The adaptor-ligated sample polynucleotide of any one of Items 26 to 30, wherein the sample polynucleotide comprises cDNA.
[0201] Item 33. A library of polynucleotides comprising a plurality of adaptor-ligated sample polynucleotides of any one of items 26 to 32.
[0202] Item 34. A plurality of sequencing libraries comprising one or more libraries of item 33.
[0203] Item 35. The plurality of sequencing libraries of Item 34, wherein each library is obtained from a different sample.
[0204] Item 36. The library of any one of Items 33 to 35, wherein 70% or less of the sample polynucleotides are present within one standard deviation of the mean sample polynucleotide abundance.
[0205] Item 37. The library of any one of Items 33 to 35, wherein 50% or less of the sample polynucleotides are present within one standard deviation of the average sample polynucleotide amount.
[0206] Item 38. The library of any one of Items 33 to 35, wherein 25% or less of the sample polynucleotides are present within one standard deviation of the mean sample polynucleotide abundance.
[0207] Item 39. A method for library preparation, comprising: (a) providing a plurality of sample polynucleotides; (b) ligating at least one composition described in any one of items 1 to 10 to at least one sample polynucleotide.
[0208] Item 40. The method of item 39, wherein the plurality of sample polynucleotides comprises genomic DNA (gDNA).
[0209] Item 41. The method of item 29, wherein the plurality of sample polynucleotides comprises circular DNA (cDNA).
[0210] Item 42. The method of any one of Items 39 to 41, wherein the molar ratio of the at least one composition to the plurality of sample polynucleotides is 1:5 or less.
[0211] Item 43. The method of any one of Items 39 to 41, wherein the molar ratio of the at least one composition to the plurality of sample polynucleotides is 1:2 or less.
[0212] Item 44. The method of any one of Items 39 to 41, wherein the molar ratio of the at least one composition to the plurality of sample polynucleotides is 1:1 or less.
[0213] Item 45. The method of any one of items 39 to 44, wherein the ligation occurs with an efficiency of at least 25%.
[0214] Item 46. The method of any one of items 39 to 44, wherein the ligation occurs with an efficiency of at least 50%.
[0215] Item 47. The method of any one of items 39 to 44, wherein the ligation occurs with an efficiency of at least 75%.
[0216] Item 48. The method according to any one of Items 39 to 47, further comprising cleaving the linker.
[0217] Item 49. The method of any one of items 39 to 47, wherein the step of cleaving the linker comprises contacting the conjugate with an enzyme.
[0218] Item 50. The method of item 49, wherein the enzymatic cleavage comprises a USER or site-specific restriction endonuclease.
[0219] Item 51. A conjugate comprising: a first strand comprising a first terminal adaptor region, a first non-complementary region, and a first yoke region; a second strand comprising a second terminal adaptor region, a second non-complementary region, and a second yoke region; The first strand and the second strand are connected via a linker, conjugate.
[0220] Item 52. The conjugate of Item 51, wherein the linker is attached to the 5' end of the first strand and the 5' end of the second strand.
[0221] Item 53. The conjugate of Item 51, wherein the linker is attached to the 3' end of the first strand and the 3' end of the second strand.
[0222] Item 54. A method for producing the composition according to any one of items 1 to 25 or the conjugate according to any one of items 51 to 53, comprising: 1. A method comprising contacting a first strand with a second strand, wherein the first strand comprises a first reactive group and the second strand comprises a second reactive group, and wherein reaction of the first reactive group with the second reactive group produces a conjugate.
[0223] Item 55. The method of item 54, wherein the first strand further comprises a first linker.
[0224] Item 56. The method of items 54 or 55, wherein the second strand further comprises a second linker.
[0225] Item 57. The method of items 55 or 56, wherein the first linker comprises a first reactive group.
[0226] Item 58. The method of items 56 or 57, wherein the second linker comprises a second reactive group.
[0227] Item 59. A method for producing the composition according to any one of items 1 to 25 or the conjugate according to any one of items 51 to 53, comprising: (a) attaching nucleotide monomers to a growing chain on a surface; (b) repeating step (a) to produce the first strand described in any one of items 1 to 25; (c) attaching a linker comprising a reactive group to the end of the first strand; (d) repeating step (a) to produce the third strand according to any one of items 1 to 25.
[0228] Item 60. The method of item 59, further comprising cleaving the composition from the solid support.
[0229] Item 61. The method of items 59 or 60, comprising hybridizing the second strand to the first strand.
[0230] Item 62. The method of any one of Items 59 to 61, comprising hybridizing the fourth strand to the third strand. [Example]
[0231] The following examples are given for the purpose of illustrating various embodiments of the present disclosure and are not intended to be limiting in any way.
[0232] Example 1: Synthesis of hybrid cyclic adaptors
[0233] This example demonstrates the synthesis of an exemplary adapter conjugate described herein.
[0234] As shown in Figures 2A-2C, the adapter strands were synthesized by the click reaction of two stubby-Y adapters ("stub adapters") having the sequences shown in Table 2 below. Both 500 stub adapters were reacted in the presence of a catalyzed Cu(I) molecule to form an exemplary 5'-linked adapter dimer.
[0235] [Table 2]
[0236] In another embodiment, the adapter strand was synthesized directly using solid-phase polynucleotide synthesis. Phosphoramidites were added to generate the adapter polynucleotide sequence, which was then reacted with the linker. Phosphoramidite synthesis was then continued in the 5' to 3' direction using reverse amidites to complete the sequence. The adapter-linker-adapter conjugates were cleaved from the surface and hybridized to the 500-stub adapter and 700-stub adapter shown in Table 3A below to generate compositions having the sequences shown in Table 3B below.
[0237] [Table 3A]
[0238] [Table 3B]
[0239] Example 2: High-throughput sample balancing with hybrid circular adaptors
[0240] This example demonstrates the ability of the exemplary adapter conjugates (hybrid circular adapters) described herein to normalize gDNA sample amounts during high-throughput sequencing to improve efficiency without reducing accuracy.
[0241] In some instances, gDNA libraries multiplexed on a sequencer benefit from balancing the sequencing load evenly across all samples in terms of number of molecules (i.e., some samples have very high concentrations of gDNA, while others have much lower concentrations). Generally, a large excess of adapters is used for library construction. Because two independent ligation products are used per molecule, conversion cannot be easily controlled by modifying adapter concentration. In some instances, in high-throughput settings, this leads to challenges when qPCR / dilution is performed sample by sample. Unbalanced sample load can also reduce the accuracy of sequencing results. Existing methods for balancing samples generally use substantial modifications to the library preparation workflow for A-tailing (Normalase) or are limited to transposase-based approaches that limit conversion (Nextera Flex, SeqWell).
[0242] Hybrid circular adapters are generally compatible with standard sequencing workflows without substantial workflow modifications. By linking two adapters, two independent ligation events are replaced with a slower intermolecular ligation and a faster intramolecular ligation. Furthermore, steric interactions can be reduced by flexible linkers. Unlike standard adapters, ligation conversion can be controlled by the number of adapter molecules to normalize sample volume.
[0243] Hybrid circular adapters are generated using the general method described in Example 1. In a high-throughput format, 1024 fragmented gDNA samples are ligated to equal concentrations of hybrid circular adapters adjusted to the minimum sample concentration to generate an adapter-ligated polynucleotide library for each sample. The samples are then optionally enriched (e.g., exome enriched), amplified, and subjected to next-generation sequencing. The degree of signal normalization (i.e., counts) is measured.
[0244] Example 3: Hybrid Circular Adapters with Cleavable Bases
[0245] This example demonstrates the synthesis of exemplary adapter conjugates (hybrid circular adapters) with cleavable bases described herein and their ability to normalize gDNA sample amounts during high-throughput sequencing to improve efficiency without reducing accuracy.
[0246] Using the general method of Example 1, make modified adapters with the cleavable bases and sequences shown in Table 4 below.
[0247] [Table 4]
[0248] These adapters are then tested for sample normalization at 100 mM concentration using the general method of Example 2.
[0249] Example 4: Hybrid circular adaptors with double-stranded ligations
[0250] This example demonstrates the synthesis of exemplary adapter conjugates (hybrid circular adapters) with temperature-sensitive double-stranded ligation described herein and their ability to normalize gDNA sample amounts during high-throughput sequencing to improve efficiency without reducing accuracy.
[0251] Using the general method of Example 1, make modified adapters with the double-stranded ligation and sequences shown in Table 5 below.
[0252] [Table 5]
[0253] These adapters are then tested for sample normalization at 100 mM concentration using the general method of Example 2.
[0254] Example 5: Hybrid Annular Adapter with Overhanging Link
[0255] This example demonstrates the synthesis of exemplary adapter conjugates (hybrid circular adapters) with temperature-sensitive overhang binding described herein and their ability to normalize gDNA sample amounts during high-throughput sequencing to improve efficiency without reducing accuracy.
[0256] Using the general method of Example 1, make modified adapters with the double-stranded ligation and sequences shown in Table 6 below.
[0257] [Table 6]
[0258] These adapters are then tested for sample normalization at 100 mM concentration using the general method of Example 2.
[0259] Example 6: Normalization study using hybrid circular adapters
[0260] This example demonstrates the ability of an exemplary adapter conjugate (hybrid circular adapter) used to accurately sequence gDNA samples despite variations in sample mass.
[0261] Human gDNA molecules were fragmented by mechanical shearing (COVARIS® platform) to obtain DNA molecule fragments with an average size of approximately 200 base pairs. The DNA fragments were aliquoted into various masses ranging from 1 ng to 200 ng for fragment end repair and dA tailing (i.e., the addition of deoxyadenosine monophosphate (dAMP) nucleotides to the 3' ends of the DNA fragments).
[0262] Hybrid circular adapters were generated according to the methods described herein (e.g., Example 1 above) and ligated to DNA fragment molecules in a gDNA sample. The reaction was cleaned up, and the DNA fragment molecules were amplified by PCR to create a sequencing-ready gDNA library. The library was then sequenced by next-generation sequencing (NEXTSEQ™ 550), and library size was measured after demultiplexing. Results were reported as levels of read representation, with increased read representation indicating greater confidence in the accuracy of the results.
[0263] As shown in Figure 4A and Figure 4B, the level of read representation remained relatively uniform with increasing mass dosing of the gDNA library, with significant variation across seven different ratios of administered gDNA-to-adapter ratios.
[0264] Example 7: Normalization study of hybrid circular adapters using cleavable bases
[0265] This example demonstrates the ability of an exemplary adapter conjugate (hybrid circular adapter) modified with a cleavable base to improve denaturation, minimize disruption of downstream PCR steps, and be used to accurately sequence gDNA samples regardless of the varying mass of the samples.
[0266] Hybrid circular adapters with a cleavable uracil base and the sequences shown in Table 7 below were generated according to methods described herein (eg, Example 3 above).
[0267] [Table 7]
[0268] Adapters were then prepared and tested for sample normalization according to the method described above in Example 6, except that after ligation and before PCR, the adapters were cleaved with USER Enzyme, and cleavage was assessed and optimized based on the bioanalyzer readout as shown in Figure 5.
[0269] Example 8: Hybrid circular adaptor with duplex and overhang ligation
[0270] This example demonstrates the ability of exemplary adapter conjugates modified to include double-stranded ligations and / or overhang ligations (hybrid circular adapters) to be used to improve denaturation and accurately sequence gDNA samples despite sample mass variation.
[0271] Hybrid circular adapters having duplex or overhang ligations and the sequences shown in Table 8 below were generated according to methods described herein (see, e.g., Examples 4 and 5 above). The structures of the duplex and / or overhang ligations are shown in Figures 6A and 6B.
[0272] [Table 8]
[0273] Adapters were blended and pooled at a final working concentration ratio of 10 μM and tested for sample normalization according to the method described in Example 6 above.
[0274] Example 9: Determining adapter flexibility for intramolecular connections
[0275] This example demonstrates the ability of flexible adapter conjugates (hybrid circular adapters) to be used to accurately sequence gDNA samples.
[0276] Flexible hybrid circular adapters minimize steric hindrance during the ligation process as adapters interact with library molecules, thereby accelerating the intramolecular ligation step for sequencing. However, it is important that flexible adapters are ligated one-to-one to library molecules, as ligation to multiple library molecules hinders accuracy.
[0277] To validate individual hybrid circular adapter molecule-library molecule pairs on a one-to-one basis, hybrid circular adapters bearing internal (in-line) barcodes ("barcoded adapters") and having the sequences shown in Table 9A below were generated according to methods described herein (e.g., Example 1 above). An assay was used to assess in-line barcode crosstalk by quantifying the extent of converted library molecules ligated to a single adapter (e.g., see the in-line barcode crosstalk diagrams in Figures 7A and 7B). The barcoded adapters were then annealed and pooled for ligation with library molecules having the sequences shown in Table 9B below.
[0278] [Table 9A]
[0279] [Table 9B]
[0280] The ligated library molecules were then subjected to next-generation sequencing, and the paired-end read data was analyzed to identify inline barcodes and analyze the amount of output crossmatches. As shown in Figure 7C, an inline barcode crossmatch indicates that a library molecule was ligated to a distinct circular adapter.
[0281] Example 10: Hybrid Circular Adapters with Increased Flexible Linkages
[0282] This example demonstrates the effect of various factors (overhang length, base content, melting temperature, linker type, linker length, and linker amount) on adapter performance with respect to increasing ligation rates, reducing crosstalk, and sample normalization.
[0283] Based on the results described in Example 9 above, two adapters (SEQ ID NO:51 and SEQ ID NO:53) were selected for testing. These adapters were then modified to create the adapters shown in Table 10 below. Iterations of adapter overhang length, base content, melting temperature, linker type, linker length, and linker amount were evaluated. See, e.g., Figure 8A. Evaluation criteria included determining the adapter that resulted in the most favorable (e.g., strongest) intramolecular connection, as measured using the assay described in Example 9 above. As shown in Figure 8B, different iterations of the adapter exhibited different amounts of crosstalk.
[0284] Adapters that promote minimal amounts of crosstalk were selected and tested for normalization according to the methods described herein. As shown in Figure 8C, normalization and similar conversion of library molecules was observed within a 10-fold range of 20-200 ng total input mass into library preparation and ligation with selected hybrid cyclic adapters.
[0285] [Table 10-1]
[0286] [Table 10-2]
[0287] Example 11: Hybrid Circular Adapter with In-Line Barcode
[0288] This example demonstrates the ability of an exemplary adapter conjugate (hybrid circular adapter) to incorporate an in-line barcode for use in sample normalization during next-generation sequencing without affecting accuracy.
[0289] The 60 inline barcode sequences were designed to have intermediate GC content, a Hamming distance of at least 2 from each other, and a length of 6-7 bases. As shown in Figure 9B, the accuracy of the sequencing results (measured via read representation) varied across the six barcodes. Of these 16 inline barcode sequences, a set of 12 sequences was selected and incorporated into each of three different adapters, as shown in Table 11 below. A diagram of the adapter structure with the inline barcodes is shown in Figure 9A. The adapters were then blended in a single pool at a final concentration of 10 µM, ligated with gDNA samples, and tested for sample normalization according to methods described herein (e.g., Example 6 above). In testing the selected barcodes, each sample had a fixed mass dose to assess barcode uniformity.
[0290] [Table 11-1]
[0291] [Table 11-2]
[0292] [Table 11-3]
[0293] As shown in Figure 9C, molecular conversion rates were similar across the 12 selected barcodes.
[0294] Example 12: In-line barcoding for PCR artifact removal
[0295] This example demonstrates the use of the exemplary in-line barcodes described herein to assess the presence of artifacts during the PCR step of next-generation sequencing.
[0296] Initial screening of 32 in-line barcode sequences was performed using a Y-adapter structure, in which individual barcodes were ligated to DNA library samples in individual isolated wells. After pooling, solid-phase immobilization (SPRI), and PCR amplification to sequencing, a single barcode was predicted for every sample or read sequenced. Any molecules determined to contain mismatched (crosstalking) in-line barcodes may be due to artifacts arising from processing steps downstream of ligation. Such artifacts can be removed with silica to improve the accuracy of downstream analyses.
[0297] All possible barcodes (1,024) were divided into matched barcodes (32) and mismatched (crosstalked) barcodes (992), as shown in the histogram in Figure 10A. Initial testing of pooled ligation reactions resulted in high crosstalk barcodes, as shown in Figure 10B. After individual bead purification or a post-ligation heat denaturation step, crosstalk barcodes were reduced or eliminated, as shown in Figure 10C. Therefore, inline barcodes were used to identify and eliminate artifacts during the PCR process.
[0298] Example 13: High-throughput library preparation workflow using adapters
[0299] This example demonstrates the ability of the exemplary barcoded adapters described herein to pool samples immediately after normalization and handle multiple samples sequenced in parallel.
[0300] The library workflow described herein was performed using the barcoded adapters described in Example 12 above. In the workflow, human and maize gDNA samples were enzymatically fragmented, normalized, and tagged with adapters. The mass dose of gDNA samples was varied over a range of 40 ng to 200 ng. The fragmented gDNA was end-repaired and dA-tailed for ligation with barcoded adapters. Ligation reactions were heat-denatured and pooled together for SPRI cleanup at a 0.8X ratio. The eluted pool was subjected to a 12-cycle PCR reaction and then subjected to next-generation sequencing.
[0301] As shown in Figure 11A (human sample) and Figure 11B (corn sample), the pass filter normalized matches for both human and corn gDNA samples were relatively uniform across the range of mass doses. As shown in Figure 12, fractional chimerism remained relatively uniform for both species, with corn samples consistently showing greater fractional chimerism than human samples.
[0302] Average target coverage, uniformity, and post-target enrichment genome coverage (TE) were captured for both human samples (800 kb panel) and corn samples (1.25 Mb panel), with read depth downsampled to 150x per library in the pool. The overlap rate for the human samples was approximately 15%, while the overlap rate for the corn samples was approximately 32%. As shown in Figures 13 and 14, the average target coverage and uniformity after TE were relatively uniform across the range of mass doses. Figure 15 shows genome coverage for the human KDM5C locus, and Figure 16 shows genome coverage for the corn Chr5 locus. Both graphs show one replicate for each mass dose from the 12-plex.
[0303] While preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will occur to those skilled in the art without departing from the invention. It is understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. It is intended that the following claims define the scope of the invention, and that methods and structures within the scope of these claims and their equivalents be covered thereby.
Claims
1. (a) a first polynucleotide adaptor, a first strand comprising a first terminal adaptor region, a first non-complementary region, and a first yoke region; a second strand comprising a second terminal adaptor region, a second non-complementary region, and a second yoke region; a first polynucleotide adaptor, wherein the first yoke region and the second yoke region are complementary, and the first non-complementary region and the second non-complementary region are non-complementary; (b) a second polynucleotide adaptor, a third strand comprising a third terminal adaptor region, a third non-complementary region, and a third yoke region; a fourth strand comprising a fourth terminal adaptor region, a fourth non-complementary region, and a fourth yoke region; a second polynucleotide adaptor, wherein the third yoke region and the fourth yoke region are complementary, and the third non-complementary region and the fourth non-complementary region are not complementary, The composition, wherein the first polynucleotide adaptor and the second polynucleotide adaptor are connected via a linker.
2. The composition of claim 1 , wherein the linker is attached to the first terminal adaptor region and the third terminal adaptor region.
3. The composition of claim 1 , wherein the linker is attached to the second terminal adaptor region and the fourth terminal adaptor region.
4. The linker is alkenes, alkynes, triazines, triazoles, oximes, sulfides, amides, and / or cyclooctane, and / or at least one PEG unit, and / or The composition of claim 1 , comprising at least one spacer 18 (Sp18).
5. The composition of claim 1 , wherein the linker comprises a polynucleotide.
6. the polynucleotide is at least one cleavable base, and / or a double-stranded region, and / or The composition of claim 5 , comprising at least one splint polynucleotide.
7. 7. The composition of claim 6, wherein the polynucleotide comprises at least one splint polynucleotide, and the at least one splint polynucleotide is at least partially complementary to a portion of the first polynucleotide adaptor or the second polynucleotide adaptor.
8. The composition of claim 7 , wherein the polynucleotide comprises at least two splint polynucleotides, and the at least two splint polynucleotides at least partially overlap each other.
9. 8. The composition of claim 7, wherein the polynucleotide comprises at least two splint polynucleotides, and the at least two splint polynucleotides at least partially overlap with a portion of the first polynucleotide adaptor or the second polynucleotide adaptor.
10. 10. The composition of claim 1, wherein one or more of the first polynucleotide adaptor and the second polynucleotide adaptor comprises at least one barcode.
11. 11. The composition of claim 10, wherein the at least one barcode comprises one or more of a cell index, a sample index, or a unique molecular identifier (UMI).
12. A composition according to any one of claims 1 to 11; and at least one sample polynucleotide. Adaptor-ligated sample polynucleotides.
13. 13. The adaptor-ligated sample polynucleotide of claim 12, wherein the at least one sample polynucleotide is attached to a 3'-end of a first strand and a 3'-end or a 5'-end of a second strand.
14. 13. The adaptor-ligated sample polynucleotide of claim 12, wherein the at least one sample polynucleotide is attached to the 3' or 5' end of a third strand and the 5' end of a fourth strand.
15. A method for producing a composition according to any one of claims 1 to 11, comprising the steps of: contacting a first strand with a second strand, wherein the first strand comprises a first reactive group and the second strand comprises a second reactive group, and wherein reaction of the first reactive group with the second reactive group produces a conjugate.
16. A method for producing a composition according to any one of claims 1 to 11, comprising the steps of: (a) attaching nucleotide monomers to a growing chain on a surface; (b) repeating step (a) to produce a first strand according to any one of claims 1 to 11; (c) attaching a linker comprising a reactive group to the end of the first strand; (d) repeating step (a) to produce the third strand of any one of claims 1 to 11.
17. 17. The method of claim 16, further comprising hybridizing the second strand to the first strand.
18. 17. The method of claim 16, further comprising hybridizing the fourth strand to the third strand.