Cleavable linkers for tethering polymerases to nucleotides

By using conjugates of cleavable linkers and combining proteases with esterase activity, controlled DNA synthesis is achieved, solving the problem of high error rate in phosphoramidite synthesis and improving the synthesis accuracy and length of oligonucleotides.

CN120500533APending Publication Date: 2025-08-15ANSA BIOTECHNOLOGIES INC
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202380080714.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-26
Filing Date
2023-09-25
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Existing phosphoramidite synthesis methods have low levels of depurination, branching and other types of damage when synthesizing growth oligonucleotides, resulting in high error rates, limiting the length of oligonucleotides and failing to meet the needs of high-throughput applications.

Method used

Using a conjugate containing a cleavable linker, the polymerase and nucleotide are linked by amino acid ester, and the linker is cleaved under specific conditions using an esterase-active protease to achieve step-by-step, controlled DNA synthesis, avoiding undesirable insertions and deletions.

Benefits of technology

It improves the accuracy and length of oligonucleotide synthesis, reduces the error rate, and meets the needs of high-throughput applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120500533A_ABST
    Figure CN120500533A_ABST
Patent Text Reader

Abstract

The present disclosure provides methods of polynucleotide synthesis and compounds, compositions useful in the synthesis of polynucleotides. The chemical compounds include nucleotides and analogs thereof that are linked to a polymerase via a cleavable linker comprising an enzymatically cleavable amino acid ester.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 377,121, filed September 26, 2022; the contents of which are hereby incorporated by reference in their entirety. Background Art

[0003] Most biological research and bioengineering all relate to synthetic DNA (which can include oligonucleotides), synthetic genes or even chromosomes.Synthetic DNA is customized using the phosphoramidite method.Unfortunately, phosphoramidite synthesis can not produce high-quality long oligomers.Each synthesis cycle induces low-level but detectable depurination, branching and other types of damage in the nascent oligonucleotide, which causes final sequence errors.For the short oligonucleotides used in low-throughput applications (such as PCR primers or Sanger sequencing), this error rate can be ignored, but as length and throughput increase, this error rate becomes significant.Therefore, the maximum oligonucleotide length obtainable by phosphoramidite synthesis is limited to about 200bp.

[0004] Enzymatic methods for single-stranded DNA (ssDNA) synthesis have long been sought as an alternative to phosphoramidite chemistry. TdT (TdT) is capable of polymerizing thousands of non-templated nucleotides into a DNA strand, but synthesis of a defined sequence requires the strict restriction of incorporation to one base at a time. In next-generation sequencing (NGS), this functionality is provided by nucleotides with removable 3' blocking groups ("reversible terminators"), but TdT does not readily incorporate NTPs containing functional groups that block extension.

[0005] To achieve stepwise, controlled enzymatic DNA synthesis, a TdT conjugate can be used, bound to a nucleotide via a cleavable linker. When exposed to the free 3' end of an oligonucleotide, the conjugate adds its tethered nucleotide and remains attached to the extended primer, thereby blocking further extension of other conjugates. The linker is then cleaved to release the TdT and expose the oligomer's terminus for the next extension. These two steps of "extension" and "deprotection" are repeated to synthesize a defined sequence.

[0006] However, if the cleavable linker spontaneously cleaves prior to the controlled cleavage step, it may result in unwanted insertions during oligomer synthesis due to the presence of free nucleotides and polymerase in the added conjugate solution. Furthermore, a linker that is not cleaved during the controlled cleavage step may result in one or more unwanted deletions during synthesis because new nucleotides cannot be added while the polymerase remains bound.

[0007] Therefore, there is a need for improved linkers that are highly stable under storage and oligomer synthesis reaction conditions, but are quantitatively cleaved within a short time frame while leaving only benign residues or "scars" at the bases. Summary of the Invention

[0008] In some embodiments, a conjugate is provided herein comprising a polymerase, a nucleotide, and a cleavable linker connected to the polymerase and the nucleotide, wherein the cleavable linker comprises an amino acid ester. In some embodiments, the amino acid ester is connected to an amino acid. In some embodiments, the amine group of the amino acid ester is bound to the amino acid. In some embodiments, the conjugate comprises a peptide of at least 2, at least 3, at least 4, or at least 5 amino acids bound to the amine group of the amino acid ester.

[0009] In some embodiments, the amino acid is selected from the group consisting of alanine, arginine, asparagine, aspartic acid, cysteine, glutamic acid, glutamine, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, and valine. In some embodiments, the amino acid is glycine, or the amino acid comprises glycine. In some embodiments, the amino acid is a non-naturally occurring amino acid, or the amino acid comprises a non-naturally occurring amino acid.

[0010] In some embodiments, the cleavable linker is attached to the alpha phosphate, sugar, or nucleobase of the nucleotide.

[0011] In some embodiments, the amino acid ester is represented by:

[0012]

[0013] where R 1 and R 1' are each independently selected from hydrogen and optionally substituted C 1-6 Alkyl, or optionally together with the atoms to which they are attached, form an optionally substituted C3-C7 carbocycle.

[0014] In some embodiments, the amino acid ester is represented by a compound selected from the group consisting of:

[0015]

[0016] In some embodiments, the linker comprises the following structure:

[0017]

[0018] where R 1 and R 1' are each independently selected from hydrogen and optionally substituted C 1-6alkyl, or optionally together with the atoms to which they are attached form an optionally substituted C3-C7 carbocycle; each R 2 is an optionally substituted group independently selected from the group consisting of hydrogen, C 1-6 Alkyl, phenyl, C1-C6 carbon ring and 3-7 heterocycle; each R 3 is hydrogen or optionally substituted C 1-6 and n is 1, 2, 3, 4 or 5. In some embodiments, R 3 In some embodiments, R 2 In some embodiments, R 2 Selected from the group consisting of: hydrogen, -Me, -iso-Pr, -tert-butyl, isobutyl, -CH2Ph, -CH2OH, -CH2SH, -CH2CH2SCH3, -CH2COOH, -CH2CH2COOH, -CH2CONH2, -CH2CH2CONH2, -CH2CH2, CH2CH2NH2,

[0019] In some embodiments, n is 1. In some embodiments, R 1 and R 1' Together they form an optionally substituted C3-C7 carbocycle. 1 and R 1' Together they form an optionally substituted C3 carbocycle.

[0020] In some embodiments, the linker comprises the following structure:

[0021]

[0022] In some embodiments, the conjugate comprises the following structure:

[0023] Nuc—L1—L2—L3—Pol

[0024] wherein Nuc is a nucleotide; Pol is a polymerase; L1 is the first part of the linker that connects the nucleotide to L2; and L2 is the second part of the linker represented by the following formula:

[0025]

[0026] where R 1 and R 1' are each independently selected from optionally substituted C 1-6 alkyl, halogen, or optionally together with the atoms to which they are attached form an optionally substituted C3-C7 carbocycle; each R 2 is an optionally substituted group independently selected from the group consisting of hydrogen, C 1-6Alkyl, phenyl, C1-C6 carbon ring and 3-7 heterocycle; each R 3 is hydrogen or optionally substituted C 1-6 Alkyl; n is 0, 1, 2, 3, 4 or 5; wherein * represents the point of attachment of L2 to L1; and ** represents the point of attachment of L2 to L3; wherein L 2 Is cleavable; L 3 To connect pol to L 2 connector.

[0027] In some embodiments, L 1 Selected from the group consisting of: a bond, an optionally substituted C 1-12 Alkylene chain, C4-C 20 Polyethylene glycol, optionally substituted C 2-12 Alkenylene chain and C 2-12 Alkyne chain, where L 1 1-6 methylene units are optionally and independently replaced by -O-, -N(R b )-, -N=C(H)-, -C(O)-, -S-, -S(O)-, -S(O)2-, optionally substituted phenylene or optionally substituted cyclopropylene.

[0028] In some embodiments, L 1 include:

[0029]

[0030] Each R a are independently selected from the group consisting of halogen, hydroxy, cyano, optionally substituted C 1-6 Alkyl and optionally substituted C 1-6 Alkoxy.

[0031] In some embodiments, L 2 comprising an amino acid ester selected from the group consisting of:

[0032]

[0033] In some embodiments, L 2 It is represented by:

[0034]

[0035] In some embodiments, L 1 In some embodiments, L 1 In some embodiments, the nucleobase is selected from the group consisting of:

[0036]

[0037] In some embodiments, L 1 Binds to the sugar of the nucleotide.

[0038] In some embodiments, L 1 In some embodiments, the nucleotide is a phosphate. In some embodiments, the phosphate is an alpha phosphate. In some embodiments, the nucleotide is a polyphosphate ribonucleotide or a polyphosphate deoxyribonucleotide. In some embodiments, the nucleotide is selected from the group consisting of adenine, guanine, cytosine, uracil, and thymine.

[0039] In some embodiments, the polymerase is a template-independent polymerase. In some embodiments, the polymerase is TdT.

[0040] In some embodiments, the linker is cleavable by a protease comprising esterase activity. In some embodiments, the linker is cleavable by proteinase K. In some embodiments, the linker is cleavable at the ester group on L2, leaving a compound represented by Nuc-L1-OH after said cleavage.

[0041] In some embodiments, the present invention also provides a method for synthesizing a polynucleotide, comprising: incubating a polynucleotide with the conjugate described herein. In some embodiments, the method further comprises extending the polynucleotide by adding a nucleotide bound to the conjugate to the 3'OH of the polynucleotide.

[0042] In some embodiments, the method further comprises cleaving the cleavable linker after adding the nucleotide to the precursor polynucleotide. In some embodiments, the method further comprises repeating the incubation, extension, and cleavage steps one or more times. In some embodiments, the cleavage comprises contacting the extended polynucleotide with an enzyme comprising esterase activity under conditions sufficient to cleave the linker, thereby releasing the polymerase from the extension product. In some embodiments, the enzyme is a protease comprising esterase activity.

[0043] In some embodiments, the method further comprises removing the scar remaining after cleavage of the linker to the nucleotide.

[0044] In some embodiments, scar removal is performed after polynucleotide synthesis is complete. In some embodiments, scar removal is performed after a portion of the polynucleotide is synthesized. In some embodiments, scar removal is performed during polynucleotide synthesis after cleavage of a linker and before addition of the next nucleotide.

[0045] According to some embodiments, the present invention also provides a method for synthesizing a polynucleotide, comprising: (a) incubating the nucleic acid with a first conjugate as described herein under conditions where a polymerase catalyzes the covalent addition of nucleotides of the first conjugate to the 3' hydroxyl group of the nucleic acid to produce a first extension product; (b) cleaving the cleavable bond of the linker, thereby releasing the polymerase from the extension product to unmask the 3' hydroxyl end of the first extension product; (c) incubating the extension product with a second conjugate as described herein under conditions where a polymerase catalyzes the covalent addition of nucleotides of the second conjugate to the 3' end of the first extension product to produce a second extension product; (d) repeating steps (b)-(c) multiple times on the second extension product to produce an extended nucleic acid of a defined sequence.

[0046] According to some embodiments, the present invention also provides a sequencing method, which comprises: incubating a duplex comprising a primer and a template with a composition comprising a set of conjugates as described in any one of claims 1 to 36, wherein the conjugates correspond to G, A, T (or U) and C and are distinguishably labeled; detecting which nucleotide has been added to the primer by detecting a signal from the distinguishable label; cleaving the cleavable bond of the linker, thereby releasing the polymerase from the extension product to demask the 3' hydroxyl end of the first extension product; and repeating the incubation, detection and cleavage steps to determine the sequence of the template.

[0047] In some embodiments, a modified nucleotide is also provided herein, comprising a cleavable linker, wherein the cleavable linker comprises an amino acid ester. In some embodiments, the amino acid ester is connected to an amino acid. In some embodiments, the amine group of the amino acid ester is bound to an amino acid. In some embodiments, the modified nucleotide comprises a peptide of at least 2, at least 3, at least 4, or at least 5 amino acids bound to the amine group of the amino acid ester. In some embodiments, the amino acid is selected from the group consisting of: alanine, arginine, asparagine, aspartic acid, cysteine, glutamic acid, glutamine, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, and valine. In some embodiments, the amino acid is glycine, or the amino acid includes glycine. In some embodiments, the amino acid is a non-naturally occurring amino acid, or the amino acid includes a non-naturally occurring amino acid. In some embodiments, the cleavable linker is bound to the α-phosphate, sugar, or nucleobase of the nucleotide.

[0048] In some embodiments, the amino acid ester is represented by:

[0049]

[0050] where R 1 and R 1'are each independently selected from hydrogen and optionally substituted C 1-6 Alkyl, or optionally together with the atoms to which they are attached, form an optionally substituted C3-C7 carbocycle.

[0051] In some embodiments, the amino acid ester is represented by a compound selected from the group consisting of:

[0052]

[0053] In some embodiments, the linker comprises the following structure:

[0054]

[0055] where R 1 and R 1' are each independently selected from hydrogen and optionally substituted C 1-6 alkyl, or optionally together with the atoms to which they are attached form an optionally substituted C3-C7 carbocycle; each R 2 is an optionally substituted group independently selected from the group consisting of hydrogen, C 1-6 Alkyl, phenyl, C1-C6 carbon ring and 3-7 heterocycle; each R 3 is hydrogen or optionally substituted C 1-6 and n is 1, 2, 3, 4 or 5. In some embodiments, R 3 In some embodiments, R 2 In some embodiments, R 2 Selected from the group consisting of: hydrogen, -Me, -iso-Pr, -tert-butyl, isobutyl, -CH2Ph, -CH2OH, -CH2SH, -CH2CH2SCH3, -CH2COOH, -CH2CH2COOH, -CH2CONH2, -CH2CH2CONH2, -CH2CH2, CH2CH2NH2,

[0056] In some embodiments, n is 1. In some embodiments, R 1 and R 1' Together they form an optionally substituted C3-C7 carbocycle. 1 and R 1' Together they form an optionally substituted C3 carbocycle.

[0057] In some embodiments, the linker comprises the following structure:

[0058]

[0059] In some embodiments, the conjugate comprises the following structure:

[0060] Nuc—L1—L2

[0061] wherein Nuc is a nucleotide; L1 is the first part of a linker that connects the nucleotide to L2; and L2 is the second part of a linker represented by the formula:

[0062]

[0063] where R 1 and R 1' are each independently selected from optionally substituted C 1-6 alkyl, halogen, or optionally together with the atoms to which they are attached form an optionally substituted C3-C7 carbocycle; each R 2 is an optionally substituted group independently selected from the group consisting of hydrogen, C 1-6 Alkyl, phenyl, C1-C6 carbon ring and 3-7 heterocycle; each R 3 is hydrogen or optionally substituted C 1-6 Alkyl; n is 0, 1, 2, 3, 4 or 5; wherein * represents the point of attachment of L2 to L1; and wherein L 2 is cleavable.

[0064] In some embodiments, L 1 Selected from the group consisting of: a bond, an optionally substituted C 1-12 Alkylene chain, C4-C 20 Polyethylene glycol, optionally substituted C 2-12 Alkenylene chain and C 2-12 Alkyne chain, where L 1 1-6 methylene units are optionally and independently replaced by -O-, -N(R b )-, -N=C(H)-, -C(O)-, -S-, -S(O)-, -S(O)2-, optionally substituted phenylene or optionally substituted cyclopropylene.

[0065] In some embodiments, L 1 include:

[0066]

[0067] Each R a are independently selected from the group consisting of halogen, hydroxy, cyano, optionally substituted C 1-6 Alkyl and optionally substituted C 1-6 Alkoxy.

[0068] In some embodiments, L 2 comprising an amino acid ester selected from the group consisting of:

[0069]

[0070] In some embodiments, L 2 It is represented by:

[0071]

[0072] In some embodiments, L 1 In some embodiments, L 1 Binds to the nucleobase at the oxygen or nitrogen that participates in base pairing.

[0073] In some embodiments, L 1 Binds to the sugar of the nucleotide.

[0074] In some embodiments, L 1 In some embodiments, the nucleotide is a phosphate. In some embodiments, the phosphate is an alpha phosphate. In some embodiments, the nucleotide is a polyphosphate ribonucleotide or a polyphosphate deoxyribonucleotide. In some embodiments, the nucleotide is selected from the group consisting of adenine, guanine, cytosine, uracil, and thymine.

[0075] In some embodiments, the linker is cleavable by a protease comprising esterase activity. In some embodiments, the linker is cleavable by proteinase K. In some embodiments, the linker is cleavable at the ester group on L2, leaving a compound represented by Nuc-L1-OH after said cleavage.

[0076] In some embodiments, a conjugate is provided herein comprising a polymerase, a nucleotide and a cleavable linker connected to the polymerase and the nucleotide, wherein the cleavable linker is enzymatically cleavable. In some embodiments, the cleavable linker can be cleaved by a protease. In some embodiments, a modified nucleotide is provided herein comprising a cleavable linker, wherein the cleavable linker is enzymatically cleavable. In some embodiments, the cleavable linker can be cleaved by a protease. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] The foregoing and other objects, features and advantages will be apparent from the following description of specific embodiments as illustrated in the accompanying drawings in which like reference characters refer to the same parts throughout the different views. The drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the various embodiments.

[0078] Figure 1 Two amino acid ester dTTP analogs used for oligomer synthesis and linker cleavage are shown. One is based on a hydroxypropyl scar (Linker 1) and the other is based on a smaller hydroxymethyl scar (Linker 2). Two amino acid ester dTTP analogs (Linkers 1 and 2, Figure 1; synthesized by Jena Bioscience) was linked to a cysteine-reactive cross-linker and conjugated to TdT, with the final structure shown in FIG. Figure 1 Also shown are the alcohol scarred cleavage products following cleavage of the linker ester.

[0079] Figure 2 (AC) shows the addition of conjugates to unscarred oligomers ( Figure 2 (A and B)) and oligomers of hydroxymethyl scar ( Figure 2 -C) Kinetic diagram. Figure 2 -A: Natural DNA primer exposed to a dTTP conjugate containing an ester bond for 1 second resulted in approximately 35% extension yield. Figure 2 -B: The oligo synthesis reaction proceeds to completion and the linker is cleaved to generate a primer with a hydroxymethyl scar on the last base. Figure 2 -C. Exposure of the scarred primer to a dTTP conjugate containing an ester bond for 1 second again resulted in approximately 35% extension yield.

[0080] Figure 3 Shown are the results of primer extension using TdT-dTTP conjugates based on Linker 1 or Linker 2, as measured by gel shift assay on SDS-PAGE. ssDNA primers were extended for 60 s using 1) Linker 1 conjugate, 2) Linker 2 conjugate, 3) Linker 2 conjugate (repeat), and 4) no conjugate. T / P: TdT / DNA primer complex. P: ssDNA primer.

[0081] Figure 4 Shown are primer extension products as measured by capillary electrophoresis. Extension was performed by storing the adapter 2 conjugate overnight at a specified pH or only in buffer (negative control). Extension without insertion showed a peak at approximately 58 nt. The peak indicating unwanted insertion (extension product) was at approximately 59 nt in some samples and indicated the presence of free dNTPs in the incubated conjugate.

[0082] Figure 5 The results of enzymatic synthesis of 100-mer and 200-mer dT oligomers using an adapter 2 dNTP conjugate as measured using capillary electrophoresis are shown (Part A). In Part B, a magnified view of the product distribution of the 100-mer from enzymatic synthesis (upper panel) and chemical synthesis (lower panel) as observed by capillary electrophoresis is shown.

[0083] Figure 6Shown are the results of extending oligonucleotides using TdT-dATP, TdT-dCTP, TdT-dGTP, and TdT-dTTP conjugates containing linker 6 as measured by capillary electrophoresis (Panel A), and the cleavage time course of the TdT-dTTP conjugate containing linker 6 incorporated into the oligonucleotide and cleaved by proteinase K for 30-240 seconds as measured by capillary electrophoresis (Panel B).

[0084] Figure 7 Shown are the structures of linker nucleotides comprising glycine amino acid ester (Gly-OMe-U) and ACC amino acid ester (ACC-OMe-U) and the ester-labile products of the two linkers (HOMe-U) (top), as well as a comparison of the intact (Gly-OMe-U or ACC-OMe-U) and hydrolysis (HOMe-U) products after exposure to 45°C for 60 minutes.

[0085] Figure 8A 、 8B A and 8C show a comparison of linker cleavage efficiencies of various TdT-nucleotide conjugates. The data shown are for the 60-second ProK treatment time point.

[0086] Figure 9 Shown are a series of electropherograms characterizing the cleavage rate of proteinase K (ProK) on illustrative linkers with an aminocyclopropylcarboxyethyl group and one (1×G) or two (2×G) glycines. The cleavage reaction was terminated after 15 seconds (s), 30 seconds, 60 seconds, 4 minutes (m), 8 minutes, or 16 minutes.

[0087] Figure 10 Shown are the results for each TdT-nucleotide conjugate (L2 = ACC, Gly-ACC, or 2xGly-ACC) where the conjugate was added to the primer 3.8 seconds after addition of the conjugate.

[0088] Figure 11 Graph showing the % ester hydrolysis of compounds 14-18 (expanded series linker nucleotides) after exposure to 50°C for 1 minute to 20 hours.

[0089] Figure 12 Shown are the results of oligonucleotides extended with allyl G, ACC, AiB, AC4C, AC5C, or AC6C conjugates exposed to a temperature of 50°C for 1 hour, 4 hours, or overnight, as measured by capillary electrophoresis to show the proportions of intact and hydrolysis products. DETAILED DESCRIPTION

[0090] The details of various embodiments of the present disclosure are set forth in the following description. Other features, objects, and advantages of the present disclosure will be apparent from the description and drawings, and from the claims.

[0091] definition

[0092] The term "alkyl" refers to a straight or branched fully saturated hydrocarbon chain. Exemplary alkyl groups are methyl, ethyl, propyl, isopropyl, butyl, isobutyl, and tert-butyl.

[0093] The term "haloalkyl" refers to a straight or branched chain alkyl group substituted with one or more halogen atoms.

[0094] As described herein, compounds of the present disclosure may contain "optionally substituted" parts. In general, no matter whether the term "optionally" is added in front, the term "substituted" means that one or more hydrogens of the specified part are replaced by suitable substituents. Unless otherwise indicated, the "optionally substituted" group can have a suitable substituent at each substitutable position of the group, and when more than one position in any given structure can be substituted by more than one substituent selected from a specified group, the substituent on each position can be the same or different. The combination of substituents envisioned by the present disclosure is preferably such that stable or chemically feasible compounds are formed. As used herein, the term "stable" refers to a compound that does not substantially change when subjected to conditions that allow the compound to be produced, detected, and in certain embodiments, recovered, purified, and used for one or more purposes disclosed herein.

[0095] Suitable monovalent substituents on the substitutable carbon atoms of an "optionally substituted" group are independently halogen; -(CH2) 0-4 R ° ;—(CH2) 0-4 OR ° ; —O(CH2) 0-4 R ° 、—O—(CH2) 0-4 C(O)OR ° ;—(CH2) 0-4 CH(OR ° )2;—(CH2) 0-4 SR ° ;—(CH2) 0-4 Ph, which can be R ° Substitution; —(CH2) 0-4 O(CH2) 0-1 Ph, which can be R ° substituted; —CH═CHPh, which can be replaced by R ° Substitution; —(CH2) 0-4 O(CH2) 0-1 -pyridyl, which may be R ° Substitution; —NO2; —CN; —N3; —(CH2) 0-4 N(R ° )2;—(CH2) 0-4 N(R° )C(O)R ° ;—N(R ° )C(S)R ° ;—(CH2) 0-4 N(R ° )C(O)NR ° 2;—N(R ° )C(S)NR ° 2;—(CH2) 0-4 N(R ° )C(O)OR ° ;—N(R ° )N(R ° )C(O)R ° ;—N(R ° )N(R ° )C(O)NR ° 2;—N(R ° )N(R ° )C(O)OR ° ;—(CH2) 0-4 C(O)R ° ;—C(S)R ° ;—(CH2) 0-4 C(O)OR ° ;—(CH2) 0-4 C(O)SR ° ;—(CH2) 0-4 C(O)OSiR ° 3;—(CH2) 0-4 OC(O)R ° ;—OC(O)(CH2) 0-4 SR ° 、SC(S)SR ° ;—(CH2) 0-4 SC(O)R ° ;—(CH2) 0-4 C(O)NR ° 2;—C(S)NR ° 2;—C(S)SR ° ;—SC(S)SR ° 、—(CH2) 0-4 OC(O)NR ° 2;—C(O)N(OR ° )R ° ;—C(O)C(O)R ° ;—C(O)CH2C(O)R ° ;—C(NOR ° )R ° ;—(CH2) 0-4 SSR° ;—(CH2) 0-4 S(O)2R ° ;—(CH2) 0-4 S(O)2OR ° ;—(CH2) 0-4 OS(O)2R ° ;—S(O)2NR ° 2;—(CH2) 0-4 S(O)R ° ;—N(R ° )S(O)2NR ° 2;—N(R ° )S(O)2R ° ;—N(OR ° )R ° ;—C(NH)NR ° 2;—P(O)2R ° ;—P(O)R ° 2;—OP(O)R ° 2;—OP(O)(OR ° )2;SiR ° 3;—(C 1-4 Straight or branched alkylene)O—N(R ° )2; or—(C 1-4 Straight or branched alkylene) C(O)O—N(R ° )2, where each R ° may be substituted and independently be hydrogen, C 1-6 Aliphatic group, -CH2Ph, -O(CH2) 0-1 Ph, -CH2-(5-6 membered heteroaryl ring), or a 5-6 membered saturated, partially unsaturated or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen or sulfur, or, notwithstanding the above definitions, two independent occurrences of R ° Together with the atoms therebetween they form a 3-12 membered saturated, partially unsaturated or aromatic monocyclic or bicyclic ring having 0-4 heteroatoms independently selected from nitrogen, oxygen or sulfur, which may be substituted as defined below.

[0096] R ° (or by replacing two independent occurrences of R ° Suitable monovalent substituents on the ring formed by combining the atoms between them are independently halogen, -(CH2) 0-2 R ● 、-(halogenated R ● ),—(CH2) 0-2 OH, —(CH2) 0-2 OR ● 、—(CH2) 0-2 CH(OR● )2;—O(halogenated R ● )、—CN、—N3、—(CH2) 0-2 C(O)R ● 、—(CH2) 0-2 C(O)OH, —(CH2) 0-2 C(O)OR ● 、—(CH2) 0-2 SR ● 、—(CH2) 0-2 SH, —(CH2) 0-2 NH2, —(CH2) 0-2 NHR ● 、—(CH2) 0-2 NR ● 2.—NO2,—SiR ● 3. OSiR ● 3. —C(O)SR ● 、—(C 1-4 linear or branched alkylene)C(O)OR ● or—SSR ● , where each R ● is unsubstituted or preceded by "halo" where it is substituted only by one or more halogens and is independently selected from C 1-4 Aliphatic groups, -CH2Ph, -O(CH2) 0-1 Ph, or a 5-6 membered saturated, partially unsaturated or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen or sulfur. ° Suitable divalent substituents on a saturated carbon atom of include =0 and =S.

[0097] Suitable divalent substituents on a saturated carbon atom of an "optionally substituted" group include the following: ═O, ═S, ═NNR*2, ═NNHC(O)R*, ═NNHC(O)OR*, ═NNHS(O)2R*, ═NR*, ═NOR*, —O(C(R*2)), 2-3 O—or—S(C(R*2)) 2-3 S—, wherein each independent occurrence of R* is selected from hydrogen, C 1-6 an aliphatic group, or an unsubstituted 5-6 membered saturated, partially unsaturated or aromatic ring having 0-4 heteroatoms independently selected from nitrogen, oxygen or sulfur. Suitable divalent substituents bonded to the ortho-substitutable carbon of the "optionally substituted" group include: -O(CR*2) 2-3 O—, wherein each independent occurrence of R* is selected from hydrogen, C 1-6 an aliphatic group, or an unsubstituted 5-6 membered saturated, partially unsaturated or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen or sulfur.

[0098] Suitable substituents on the aliphatic group of R* include halogen, -R ● 、-(halogenated R ● ),—OH,—OR ● 、—O(halogenated R ● ), —CN, —C(O)OH, —C(O)OR ● 、—NH2、—NHR ● 、—NR ● 2 or —NO2, where each R ● is unsubstituted or preceded by "halo" where it is substituted only by one or more halogens, and is independently C 1-4 Aliphatic groups, -CH2Ph, -O(CH2) 0-1 Ph, or a 5-6 membered saturated, partially unsaturated or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen or sulfur.

[0099] Suitable substituents on a substitutable nitrogen of an "optionally substituted" group include Each of these is independently hydrogen, and the substituted C 1-6 an aliphatic group, an unsubstituted -OPh, or an unsubstituted 5-6 membered saturated, partially unsaturated or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen or sulfur, or, notwithstanding the above definitions, two independent occurrences of Together with the atoms therebetween, they form an unsubstituted 3-12 membered saturated, partially unsaturated or aromatic monocyclic or bicyclic ring having 0-4 heteroatoms independently selected from nitrogen, oxygen or sulfur.

[0100] Suitable substituents on the aliphatic group are independently halogen, -R ● 、-(halogenated R ● ),—OH,—OR ● 、—O(halogenated R ● ), —CN, —C(O)OH, —C(O)OR ● 、—NH2、—NHR ● 、—NR ● 2 or —NO2, where each R ● is unsubstituted or preceded by "halo" where it is substituted only by one or more halogens, and is independently C 1-4 Aliphatic groups, -CH2Ph, -O(CH2) 0-1 Ph, or a 5-6 membered saturated, partially unsaturated or aryl ring having 0-4 heteroatoms independently selected from nitrogen, oxygen or sulfur.

[0101] The recitation of a list of chemical groups in any definition of a variable herein includes definitions of that variable as any single group or combination of listed groups. The recitation of an embodiment of a variable herein includes that embodiment as any single embodiment or in combination with any other embodiment or portion thereof.

[0102] In alternative embodiments, the compounds described herein may also contain one or more isotopic substitutions. For example, hydrogen may be 2 H (D or deuterium) or 3 H (T or tritium); carbon can be e.g. 13 C or 14 C; oxygen can be e.g. 18 O; nitrogen can be e.g. 15 N, etc. In other embodiments, specific isotopes (e.g., 3 H. 13 C. 14 C. 18 O or 15 N) can represent at least 1%, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or at least 99.9% of the total isotopic abundance of the element occupying a particular site in the compound.

[0103] As used herein, the terms "about" and "approximately" refer to values or compositions within an acceptable error range for a particular value or composition as determined by one of ordinary skill in the art, which will depend in part on how the value or composition is measured or determined, i.e., the limitations of the measurement system. For example, "about" or "approximately" can mean within one or more standard deviations practiced in the art. Alternatively, "about" or "approximately" can mean a range of up to 10% (i.e., ± 10%) or greater, depending on the limitations of the measurement system. For example, about 5 mg can include any number between 4.5 mg and 5.5 mg. In addition, particularly for biological systems or processes, these terms can mean up to an order of magnitude or up to 5 times the value. When a specific value or composition is provided in the present disclosure, unless otherwise stated, the meaning of "about" or "approximately" should be considered to be within the acceptable error range for that specific value or composition. In addition, where a range and / or subrange of a value is provided, the range and / or subrange can include the endpoints of the range and / or subrange.

[0104] The terms "nucleic acid," "polynucleotide," and "oligonucleotide," as well as other related terms used herein, are used interchangeably and refer to polymers of nucleotides and are not limited to any specific length. Nucleic acids include recombinant and chemically synthesized forms. Nucleic acids can be isolated. Nucleic acids include DNA molecules (e.g., cDNA or genomic DNA), RNA molecules (e.g., mRNA), analogs of DNA or RNA produced using nucleotide analogs (e.g., peptide nucleic acids (PNA) and non-naturally occurring nucleotide analogs), and chimeric forms containing DNA and RNA. Nucleic acids can be single-stranded or double-stranded. Nucleic acids comprise polymers of nucleotides, wherein the nucleotides comprise natural or non-natural bases and / or sugars. Nucleic acids comprise naturally occurring internucleoside bonds, such as phosphodiester bonds. Nucleic acids can lack phosphate groups. Nucleic acids comprise non-natural internucleoside bonds, including phosphorothioate, phosphorothioate, or peptide nucleic acid (PNA) bonds. In some embodiments, nucleic acids comprise a mixture of one type of polynucleotide or two or more different types of polynucleotides.

[0105] As used herein, the terms "operably linked" and "operably joined" or related terms refer to the juxtaposition of components. The juxtaposed components can be covalently linked together. For example, two nucleic acid components can be enzymatically linked together, wherein the bond linking the two components together comprises a phosphodiester bond. The first nucleic acid component and the second nucleic acid component can be linked together, wherein the first nucleic acid component can confer a function to the second nucleic acid component. For example, the bond between the primer binding sequence and the sequence of interest forms a nucleic acid library molecule having a portion that can be bound to a primer. In another example, a transgenic (e.g., a nucleic acid encoding a polypeptide or a nucleic acid sequence of interest) can be linked to a vector, wherein the bond allows expression or function of the transgenic sequence contained in the vector. In some embodiments, the transgenic is operably linked to a host cell regulatory sequence (e.g., a promoter sequence) that affects transgenic expression. In some embodiments, the vector comprises at least one host cell regulatory sequence, including a promoter sequence, an enhancer, a transcription and / or translation initiation sequence, a transcription and / or translation termination sequence, a polypeptide secretion signal sequence, etc. In some embodiments, the host cell regulatory sequence controls the expression level, time, and / or position of the transgenic.

[0106] The terms "linked," "joined," "attached," "appended," and variations thereof include any type of fusion, bonding, adhesion, or association between any combination of compounds or molecules that is sufficiently stable to withstand use in a particular procedure. The procedure may include, but is not limited to, nucleotide binding; nucleotide incorporation; deblocking (e.g., removal of a chain terminating moiety); washing; removal; flow; detection; imaging and / or identification. Such bonds may include, for example, covalent bonding, ionic bonding, hydrogen bonding, dipole-dipole bonding, hydrophilic bonding, hydrophobic bonding, or affinity bonding, bonds or associations involving van der Waals forces, mechanical bonding, and the like. In some embodiments, such bonds occur within a molecule, such as joining the ends of a single-stranded or double-stranded linear nucleic acid molecule together to form a circular molecule. In some embodiments, such bonds may occur between a combination of different molecules, or between a molecule and a non-molecule, including, but not limited to, a bond between a nucleic acid molecule and a solid surface; a bond between a protein and a detectable reporter moiety; a bond between a nucleotide and a detectable reporter moiety; and the like. Some examples of bonds can be found in, e.g., Hermanson, G., “Bioconjugate Techniques”, 2nd ed. (2008); Aslam, M., Dent, A., “Bioconjugation: Protein Coupling Techniques for the Biomedical Sciences”, London: Macmillan (1998); Aslam, M., Dent, A., “Bioconjugation: Protein Coupling Techniques for the Biomedical Sciences”, London: Macmillan (1998).

[0107] When used with reference to nucleic acids, the terms "extend," "extending," and "extension" and other variants refer to the incorporation of one or more nucleotides into a nucleic acid molecule. Nucleotide incorporation involves the polymerization of one or more nucleotides onto the 3' OH terminus of a nucleic acid chain (e.g., a nucleic acid primer), thereby resulting in an extension of the nucleic acid chain (e.g., an extended primer). Nucleotide incorporation can be performed with natural nucleotides and / or nucleotide analogs.

[0108] As used herein, the term "cleavable linker" or "cleavable moiety" refers to a divalent or monovalent moiety, respectively, that is capable of being separated (e.g., dissociated, split, broken, hydrolyzed, stable bonds within the moiety are destroyed) into different entities. In embodiments, the cleavable linker is cleavable (e.g., specifically cleavable) in response to an external stimulus (e.g., an enzyme, a nucleophilic / alkaline reagent, a reducing agent, photoirradiation, an electrophilic / acidic reagent, an organometallic and metallic reagent, or an oxidizing agent).

[0109] The use of term " cleavable joint " does not mean to show that need to remove whole joint.Cleavage site can be positioned at the position on the joint, and described position guarantees that the part of joint remains connected to dyestuff and / or substrate moiety after cracking.As limiting examples, cleavable joint can be electrophilic cleavable joint, nucleophilic cleavable joint, light cleavable joint, the joint (for example containing disulfide or azide joint) that can be cleaved under reducing conditions, the joint that can be cleaved under oxidizing conditions, by using the cleavable joint of safety lock joint and by the cleavable joint of elimination mechanism.Use cleavable joint that dye compound is connected to substrate moiety to ensure when needed, after detection, can remove mark, thereby avoid any interfering signal in downstream step.

[0110] In embodiments, the cleavable linker is cleaved by contacting the cleavable linker with a cleavage agent (e.g., a reducing agent). In embodiments, the cleavage agent is...

[0111] As used herein, the terms "polymerase-compatible cleavable moiety" and "polymerase-compatible cleavable linker" refer to a cleavable moiety or cleavable linker that does not interfere with the function of a polymerase (e.g., a DNA polymerase or modified DNA polymerase that incorporates the nucleotide to which the polymerase-compatible cleavable moiety is attached into the 3' end of the newly formed nucleotide chain). The methods contemplated herein for determining polymerase function are described in B. Rosenblum et al. (Nucleic Acids Res. 1997 Nov 15; 25(22): 4500-4504); and Z. Zhu et al. (Nucleic Acids Res. 1994 Aug 25; 22(16): 3418-3422), which are incorporated herein by reference in their entirety for all purposes. In embodiments, the polymerase-compatible cleavable moiety does not reduce the function of the polymerase relative to the absence of the polymerase-compatible cleavable moiety. In embodiments, the polymerase-compatible cleavable moiety does not adversely affect DNA polymerase recognition. In embodiments, the polymerase-compatible cleavable moiety does not adversely affect (eg, limit) the read length of the DNA polymerase.

[0112] Enzymatic polynucleotide synthesis

[0113] The present disclosure describes a method for enzymatic polynucleotide synthesis that uses a polymerase-nucleotide conjugate to control the repeated addition of a single nucleotide per cycle to the 3' hydroxyl terminus of a growing polynucleotide chain by a nucleotide-bound polymerase to perform polynucleotide synthesis. This control is achieved through the so-called "shielding effect." Shielding describes the following steric hindrance: while the polymerase remains attached to the added nucleotide, the 3' hydroxyl terminus that has been extended by the conjugate is prevented from being accessed by another conjugate, and the polymerase that is tethered to the nucleotide at the 3' terminus is prevented from accessing the nucleotides of the other conjugate.

[0114] A typical method for synthesizing a determined sequence using a template-independent polymerase is described in PCT Publication No. WO2017 / 223517, which is incorporated herein by reference in its entirety. The nucleic acid used as the initial substrate for extension (i.e., "starting molecule") is incubated with a first polymerase-nucleotide conjugate. Once the nucleic acid has been extended by the tethered nucleotides of the conjugate, extension no longer occurs because the conjugate realizes a termination mechanism. In the second step of the process, the joint is cleaved to release the polymerase and reverse the termination mechanism, thereby making subsequent extension possible. The extension product is then exposed to the second conjugate, and these two steps are repeated to extend the nucleic acid by determining the sequence. A synthetic method using a conjugate comprising TdT and a photocleavable joint is also described in WO2017 / 223517. As described above, other strategies can be used for connecting and cleaving joints.

[0115] An important step in this polynucleotide synthesis method is deprotection, or the removal of the tethered polymerase from the extended polynucleotide so that the 3' end is available for continued extension in the next synthesis cycle. In order to be useful for polynucleotide synthesis, the tethered polymerase is preferably removed with rapid kinetics to reduce the synthesis cycle time, while also being performed under benign conditions to prevent damage to the polynucleotide being synthesized. It is also preferred that the removal of the tethered polymerase be carried out to completion and produce a cleavage product that does not interfere with the continued extension or downstream applications of the intact DNA synthesis product. In some embodiments, the tethering also allows for the efficient conjugation of nucleotides to the polymerase, followed by the efficient positioning of the nucleotide within the active site to facilitate rapid incorporation into the free primer 3' end.

[0116] Herein, we describe optimized cleavable linker designs for tethering polymerases to nucleotides that are highly stable during storage and under oligomer synthesis reaction conditions (prior to controlled linker cleavage) and that are enzymatically cleavable to completion within a short time frame suitable for oligomer synthesis.

[0117] Cleavable linker

[0118] Provided herein is a conjugate comprising a polymerase and a nucleotide connected by a linker, wherein the linker comprises an enzymatically cleavable bond. The polymerase portion of the conjugate can extend a nucleic acid using the nucleotide to which it is connected (i.e., the polymerase can catalyze the connection of the nucleotide to which it is connected to the nucleic acid) and remains connected to the extended nucleic acid through the linker until the linker is enzymatically cleaved.

[0119] In the conjugate, the linker comprises an atom that connects the nucleotide to the polymerase. In some embodiments, the linker connects the base, sugar, or α-phosphate of the nucleotide to the polymerase. In some embodiments, the linker connects the terminal phosphate of the nucleotide to the polymerase. In some embodiments, the linker connects the C α In some embodiments, the polymerase and the nucleotide are covalently linked, and the distance between the linking atom of the nucleotide and the polymerase to which it is linked can be in the range of 4- For example, 15- or 20- The linker used should be long enough to allow the nucleotide to access the active site of the polymerase to which it is tethered. As will be described in more detail below, the polymerase of the conjugate is capable of catalyzing the addition of the nucleotide to which it is attached to the 3' end of the nucleic acid.

[0120] The linkers contemplated herein are also of sufficient length and stability to allow efficient hydrolysis by enzymatic means. The number of carbons or atoms in the linker, optionally derivatized with other functional groups, must be long enough to allow enzymatic cleavage from the nucleotide by a polymerase.

[0121] In some aspects, the cleavable linker comprises an amino acid ester. In some aspects, the amino acid ester is the cleavage site of the linker, thereby promoting release of the polymerase upon exposure to an esterase or a protease comprising esterase activity. The portion of the cleavable linker comprising an amino acid ester is referred to herein as the "L" of the linker. 2 "Part.L 2 Can be designed and optimized for enzymatic cleavage by esterases or proteases containing esterase activity, for example by modifying the chemical group attached to the alpha carbon of the amino acid ester or by including one or more amino acids adjacent to the amino acid ester as L 2 part of.

[0122] Described herein are polymerase-nucleotide conjugates comprising a cleavable linker that is highly stable and rapidly enzymatically cleaved by a protease comprising esterase activity. In some embodiments, the polymerase-nucleotide conjugate comprises a nucleotide linked to a polymerase using an enzymatically cleavable linker. In some embodiments, the polymerase-nucleotide conjugate comprising an enzymatically cleavable linker comprises a structure Nuc-L 1 -L2 -L 3 -Pol, where Nuc represents nucleotide, Pol represents polymerase, and L 1 -L 2 -L 3 In some embodiments, L 1 Indicates the connection of nucleotides to L 2 The region of the enzymatically cleavable linker, L 2 represents the cleavable portion of the enzymatically cleavable linker, L 3 Indicates that L 2 The region of the enzymatically cleavable linker attached to Pol.

[0123] In some embodiments, the enzymatically cleavable linker comprises an amino acid ester moiety. 2 In some embodiments, the ester group of the amino acid ester moiety can be cleaved by a protease comprising esterase activity. 2 The amino acid ester is connected to L 1 , which is in L 2 The ester cleavage may also be referred to as the spacer or nucleotide scar. 2 Also connected to the rest of the connector L 3 , comprising a linking chemistry for polymerase conjugation. In some embodiments, L 3 It may also contain a spacer or be referred to as a spacer. In some embodiments, L 2 Also contains an additional amino acid bound to the amine of the amino acid ester to serve as a protease substrate. 2 The stability of the ester is optimized to prevent spontaneous cleavage while retaining its ability to serve as a suitable substrate for the esterase activity of the protease comprising esterase activity.

[0124] In some embodiments, the linker is attached to the nucleotide at the nucleobase. In some embodiments, the linker is attached to the nucleotide at the sugar. In some embodiments, the linker is attached to the 5' phosphate group of the nucleotide, wherein the nucleotide is any polyphosphate nucleoside. In some embodiments, the linker is attached to the alpha phosphate. In some embodiments, the linker is attached to the gamma, beta, delta, epsilon, zeta, eta, or theta phosphate. In some embodiments, the linker is attached to the terminal phosphate. In some embodiments, the linker of the conjugate can be attached to the 7-position of deaza dGTP or the 5-position of dTTP or dUTP.

[0125] Additional tethered nucleotides can be found, for example, in PCT Publication WO 2017 / 223517, “Nucleic Acid Synthesis and Sequencing Using Tethered Nucleoside Triphosphates,” the entire contents of which are incorporated herein by reference.

[0126] In some embodiments, tethered nucleotides can be specifically linked to the cysteine residues of a polymerase using sulfhydryl-specific linking chemistry. Possible sulfhydryl-specific linking chemistry includes but is not limited to adjacent-pyridyl disulfide (OPSS), maleimide functional group, 3-aryl propiolate nitrile functional group, allenamide functional group, haloacetyl functional group (such as iodoacetyl or bromoacetyl), alkyl halide or perfluoroaryl, which can advantageously react with sulfhydryl groups surrounded by a specific amino acid sequence (Zhang, Chi, et al. Nature chemistry 8, (2015) 120-128.). Other linking chemistries for specific labeling of cysteine residues will be clear to those skilled in the art, or are described in related literature and text (e.g., Kim, Younggyu et al., Bioconjugate chemistry 19.3 (2008): 786-791).

[0127] Several TdT-dNTP conjugates have been prepared and tested, using different cleavable linkers to specifically tether the nucleotide to the polymerase. The use of a peptide bond in the linker results in a linker that can be cleaved by a protease. The cleavage of the peptide bond by the protease generates an amine and a carboxylic acid, both of which are charged under the typical buffer conditions for TdT activity. However, the continued presence of a charged functional group on the synthesized oligonucleotide can have a detrimental effect during the synthesis process.

[0128] In contrast, cleavage of an ester group produces an alcohol-charge-neutral cleavage product. Herein, it was demonstrated that nucleotides containing scars containing such alcohols do not hinder conjugate-based oligonucleotide synthesis (see Example 2). It was also demonstrated that linkers containing amino acid esters can be enzymatically cleaved by proteases containing esterase activity, such as proteinase K (see Example 2).

[0129] Therefore, in some embodiments, L 2 In some embodiments, the amino acid ester is the cleavage site of the linker, thereby facilitating the release of the polymerase from the nucleotide.

[0130] In addition, it was initially observed that the ester group of the glycine amino acid ester in the linker may be unstable, resulting in spontaneous cleavage of the conjugate and unwanted nucleotide insertion during the conjugate-based oligonucleotide synthesis process (see Examples 2 and 3). However, it was observed that the addition of aliphatic or bulky substituents on the α-carbon of the amino acid ester advantageously improved the stability of the adjacent ester (see Examples 4 and 6). In addition, atomic substitutions on the α-carbon of the amino acid ester can affect hyperconjugation, resulting in increased or decreased instability of the adjacent ester, as well as the cleavage rate of proteases containing esterase activity. Therefore, the selection of preferred substituents on the α-carbon of the amino acid ester can be used to achieve an acceptable balance between stability and linker cleavage kinetics (see Example 6).

[0131] Thus, in some embodiments, the amino acid ester comprises one or more substitutions on the α carbon, such as the addition of an aliphatic or bulky substituent. In some embodiments, the amino acid ester is represented by:

[0132]

[0133] where R 1 and R 1 ' are each independently selected from optionally substituted C 1-3 Alkyl, halogen, or optionally together with the atoms to which they are attached form an optionally substituted C3-C7 carbocycle.

[0134] Exemplary L having different substituents on the α carbon of the amino acid ester 2 The linker structure (e.g., to improve ester stability) is shown below:

[0135]

[0136] In addition, it was observed that the L-terminal amino acid ester 2 The moiety comprises one or more amino acids that improve the kinetics of cleavage of ester groups by proteases comprising esterase activity (see Example 5). 2 An amino acid ester comprises one or more amino acid residues adjacent to each other. In some embodiments, one or more amino acid residues are bound to the amine group of the amino acid ester.

[0137] In some embodiments, L 2 Contains or consists of:

[0138]

[0139] in

[0140] R 1 and R 1 ' are independently selected from hydrogen or optionally substituted C 1-3Alkyl, or together with the atoms to which they are attached, form an optionally substituted C3-C7 carbocycle;

[0141] Each R 3 is an optionally substituted group independently selected from the group consisting of hydrogen, C 1-6 Alkyl, benzyl, -OH, -O(C 1-6 alkyl) and -CN;

[0142] Each R c is hydrogen or optionally substituted C 1-6 alkyl; and

[0143] n is 1, 2 or 3.

[0144] In some embodiments, the one or more amino acids connected to the amine of amino acid ester include L-isomers or D-isomers of amino acid residues.Term " naturally occurring amino acid " refers to Ala, Asp, Cys, Glu, Phe, Gly, His, He, Lys, Leu, Met, Asn, Pro, Gin, Arg, Ser, Thr, Val, Trp, Tyr or citrulline. " D- " represents an amino acid with " D " (dextrorotatory) configuration, which is opposite to the configuration of naturally occurring (" L- ") amino acids. Amino acids as described herein can be purchased commercially (Sigma Chemical Co., Advanced Chemtech) or synthesized using methods known in the art. In some embodiments, amino acids with non-natural or artificial side chains are connected to the amine of amino acid ester.

[0145] As mentioned above, it was observed that the composition of the linker (i.e., L 2 The peptide sequence of L has a significant effect on the rate of protease-mediated deprotection. 2 Various permutations of the amino acids in the linker can produce conjugates with faster addition and deprotection kinetics. Such linkers can include variations in amino acid identity and the number of consecutive amino acids.

[0146] For example, you can choose to include L in the connector 2 One or more amino acids in the portion / combined with an amino acid ester to optimize protease binding and ester cleavage. Combinatorial libraries can be generated to test for optimal cleavage activity, and amino acids can be selected based on existing known peptide sequence targets of the protease. As disclosed herein, a protease comprising esterase activity can recognize the peptide portion of the linker and hydrolyze L 2 The ester group of the amino acid ester removes the polymerase attached to the nucleotide via the linker.

[0147] If desired, a spacer can be used between the nucleotides and the joint or between the joint and the label. Spacers of varying lengths can be used to increase the L2 utilization for proteases / esterases and to increase the efficiency and fidelity of the polymerase. Exemplary spacers include, for example, polyethylene glycol or other suitable spacers.

[0148] Contains L 2 Examples of linkers of structures comprising an amino acid ester bound to one or more amino acid residues are shown below:

[0149]

[0150]

[0151] For use in linker regions not related to the enzymatic cleavage disclosed herein (i.e., L 1 and L 3 ) There is considerable flexibility in the type of linker. Examples of suitable linker structures can include, but are not limited to, carbon chain linkers (e.g., C6, C12, C18, C24, etc.), peptide linkers (e.g., polyglycine or polyalanine ranging in length from about 1 residue to about 1,000 residues), or polyether linkers (e.g., PEG, PPG, PAG, PTMG ranging in length from about 1 polyether unit to about 1,000 polyether units).

[0152] In some embodiments, L 1 or L 3 is an atom chain selected from C, N, O, S, Si and P, preferably having 0 to 500 atoms, wherein L 1 Covalently linked to Nuc and L 2 , and where L 3 Covalently linked to L 2 and Pol. for the formation of L 1 or L 3 The atoms of the group may be combined in all chemically relevant ways, for example to form alkylene, alkenylene and alkynylene groups, carbamates, carbonates, ethers, polyoxyalkylenes, esters, amines, imines, polyamines, hydrazines, hydrazones, amides, ureas, semicarbazides, carbazides, alkoxyamines, alkoxyamines, carbamates, amino acids, peptides, acyloxyamines, hydroxamic acids or combinations thereof.

[0153] In some embodiments, in various embodiments, L 1 or L 3 comprises one or more carbon atoms, zero, one or more oxygen atoms, zero, one or more nitrogen atoms, zero, one or more sulfur atoms, or a combination thereof. 1 or L 347, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or any number or range of carbon atoms, oxygen atoms, nitrogen atoms, sulfur atoms, or a combination thereof.

[0154] In some embodiments, L 1 or L 3 comprises a polymer, such as a homopolymer or a heteropolymer. In some embodiments, L 1 or L 3 In some embodiments, the plurality of repeating units include the same repeating unit. In some embodiments, the plurality of repeating units include two or more different repeating units. The plurality of repeating units may include polyethers, such as paraformaldehyde, polyethylene glycol (PEG), polypropylene glycol (PPG), polyalkylene glycol (PAG), polytetramethylene glycol (PTMG) or a combination thereof. For example, the plurality of repeating units may include PEGix, PEG23, PEG24 or a combination thereof. The plurality of repeating units may include polyalkylenes, such as polyethylene, polypropylene, polybutylene or a combination thereof. In some embodiments, a repeating unit in the plurality of repeating units does not include an aromatic group. In some embodiments, a repeating unit in the plurality of repeating units includes one or more aromatic groups.

[0155] In some embodiments, L 1 or L 3 The invention can comprise any number of basic chemical starting blocks. For example, the linker can comprise a straight or branched alkyl, alkenyl or alkynyl chain or a combination thereof, which provides a linker between the nucleotide and the polymerase, between the nucleotide and the L 2 Between polymerase and L 2In some embodiments, the present invention provides a useful distance between the two ends of the nucleotide analog. For example, amino-alkyl linkers, such as amino-hexyl linkers, have been used to connect linkers to nucleotide analogs, and are usually hard enough to keep this type of distance. The longest chain of this type of linker can include up to 2 atoms, 3 atoms, 4 atoms, 5 atoms, 6 atoms, 7 atoms, 8 atoms, 9 atoms, 10 atoms or even 11-35 atoms or even 35-50 atoms. Straight or branched linkers can also contain heteroatoms other than carbon, including but not limited to oxygen, sulfur, phosphoric acid and nitrogen. Due to the hydrophilicity relevant to polyoxyethylene, polyoxyethylene chains (also commonly referred to as polyethylene glycol or PEG) are preferred linker components. Heteroatoms such as nitrogen and oxygen are inserted into the linker to affect the solubility and stability of the linker.

[0156] Including L 1 or L 3 The joint can be rigid or flexible in nature. The rigid structure comprises a plurality of chemical bonds (such as double bonds or triple bonds) between lateral rigid chemical groups (such as ring structures, such as aromatic compounds), adjacent groups, so as to prevent the groups from rotating relative to each other, and thus impart flexibility to the entire joint. Therefore, the required degree of rigidity can be changed according to the content of the joint or the number of bonds between the individual atoms constituting the joint. In addition, rigidity can be imparted by adding a ring structure along the joint. The ring structure can include an aromatic ring or a non-aromatic ring. The size of the ring can be 3 carbons to 4 carbons, to 5 carbons or even 6 carbons. The ring can also optionally include heteroatoms such as oxygen or nitrogen, and can also be aromatic or non-aromatic. The ring can additionally be optionally substituted by other alkyl and / or substituted alkyl groups.

[0157] Linkers comprising rings or aromatic structures may include, for example, aryl alkynes and aryl amides.Other examples of linkers disclosed herein include oligopeptide linkers, which may also optionally comprise ring structures in their structures.

[0158] In an embodiment, L 1 or L 3 It is C1-C 10 an alkylene chain in which 1-6 methylene units are optionally and independently replaced by -NH-, -O-, -C(O)-, -C(O)NH-, -NHC(O)-, -NHC(O)NH-, -C(O)O-, -OC(O)-, -SS-, optionally substituted cycloalkylene (e.g., C3-C8, C3-C6, or C5-C6), optionally substituted heterocycloalkylene (e.g., 3 to 8, 3 to 6, or 5 to 6 members), optionally substituted arylene (e.g., C6-C 10 、C 10 or phenylene) or a substituted or unsubstituted heteroarylene (e.g., 5- to 10-membered, 5- to 9-membered, or 5- to 6-membered).

[0159] In an embodiment, L 1 or L 3 is a bond, -NH-, -O-, -C(O)-, -C(O)NH-, -NHC(O)-, -NHC(O)NH-, -C(O)O-, -OC(O)-, -SS-, an optionally substituted alkylene group (e.g., C1-C 20 、C 10 -C 20 , C1-C8, C1-C6 or C1-C4), optionally substituted heteroalkylene (e.g., 2 to 20, 8 to 20, 2 to 10, 2 to 8, 2 to 6 or 2 to 4 members), optionally substituted cycloalkylene (e.g., C3-C8, C3-C6 or C5-C6), optionally substituted heterocycloalkylene (e.g., 3 to 8, 3 to 6 or 5 to 6 members), optionally substituted (e.g., C6-C 10 、C 10 or phenylene) or optionally substituted (e.g., 5 to 10, 5 to 9, or 5 to 6 membered). In an embodiment, L 1 or L 3 is an optionally substituted C1-C 20 In some embodiments, L 1 or L 3 Optionally substituted 2 to 20 membered heteroalkylene. In an embodiment, L 1 or L 3 L is an optionally substituted C3-C8 cycloalkylene. 1 or L 3 is an optionally substituted 3 to 8 membered heterocycloalkylene. 1 or L 3 is an optionally substituted C6-C 10 In an embodiment, L 1 or L 3 is an optionally substituted 5- to 10-membered heteroarylene group.

[0160] In some embodiments, L 1 R L In some embodiments, L 3 R L 1-6 instances of each R Lindependently selected from the group consisting of oxo, halogen, -CCl3, -CBr3, -CF3, -Cl3, -CN, -OH, -NH2, -COOH, -CONH2, -NO2, -SH, -SO3H, -SO4H, -SO2NH2, -NHNH2, -ONH2, -NHC(O)NHNH2, -NHC(O)NH2, -NHSO2H, -NHC(O)H, -NHC(O)OH, -NHOH, -OCCl3, -OCF3, -OCBr3, -OCI2, -OCHCI2, -OCHBr2, -OCHb, -OCHF2, -N3, optionally substituted alkyl (e.g., C1-C 20 、C 10 -C 20 , C1-C8, C1-C6 or C1-C4), optionally substituted heteroalkyl (e.g., 2 to 20, 8 to 20, 2 to 10, 2 to 8, 2 to 6 or 2 to 4 members), optionally substituted cycloalkyl (e.g., C3-C8, C3-C6 or C5-C6), optionally substituted heterocycloalkyl (e.g., 3 to 8, 3 to 6 or 5 to 6 members), optionally substituted aryl (e.g., C6-C 10 、C 10 or phenyl) and optionally substituted heteroaryl (e.g., 5- to 10-, 5- to 9-, or 5- to 6-membered).

[0161] In an embodiment, L 1 or L 3 is -(CH2CH2O) b -. In an embodiment, L 1 or L 3 Yes -CCCH2(OCH2CH2) a -NHC(O)-(CH2)c(OCH2CH2) b -. In an embodiment, L 1 or L 3 It is -CHCHCH2-NHC(O)-(CH2) c (0CH2CH2) b -. In an embodiment, L 1 or L 3 It is -CCCH2-NHC(O)-(CH2)c(OCH2CH2) b -. In an embodiment, L 1 or L 3is -CCCH2-. The symbol a is an integer from 0 to 8. In an embodiment, a is 1. In an embodiment, a is 0. The symbol b is an integer from 0 to 8. In an embodiment, b is 1 or 2. In an embodiment, b is an integer from 2 to 8. In an embodiment, b is 1. The symbol c is an integer from 0 to 8. In an embodiment, c is 3. In an embodiment, c is 1. In an embodiment, c is 2. In an embodiment, L1 or L3 is independently substituted or unsubstituted C1-C4 alkylene or substituted or unsubstituted 8- to 20-membered heteroalkylene.

[0162] L1 serves as the point of attachment to the nucleotide and includes a hydroxyl terminal group that binds to a portion of L2 during synthesis.

[0163] In some embodiments, L1 is a scar that can be cleaved enzymatically or chemically following cleavage of the polymerase-nucleotide linker / removal of the L2-L3-pol portion.

[0164] In some embodiments, L1 is selected from the group consisting of: a bond, an optionally substituted C 1-12 Alkylene chain, C4-C 20 Polyethylene glycol, optionally substituted C 2-12 Alkenylene chain and optionally substituted C 2-12 Alkyne chain, wherein 1-4 methylene units of L1 are optionally and independently replaced by -O-, -N(R b )-, -C(O)-, -S(O)-, -S(O)2-, phenylene, cyclopropylene; wherein each R b are independently hydrogen or optionally substituted C 1-6 alkyl.

[0165] In some embodiments, L1 comprises:

[0166]

[0167] Each R a are independently selected from the group consisting of halogen, hydroxy, cyano, optionally substituted C 1-6 Alkyl and optionally substituted C 1-6 Alkoxy.

[0168] In some embodiments, L 1 Selected from the group consisting of:

[0169]

[0170] In some embodiments, L3 comprises a bioconjugation group suitable for conjugating L3 to a polymerase.

[0171] In some embodiments, the bioconjugation group is an N-hydroxysuccinimide ester (NHS) group. In some embodiments, the bioconjugation group is a maleimide group. The linker can then be covalently attached to the polymerase by reacting the maleimide group with a cysteine residue of the polymerase.

[0172] In some embodiments, the polymerase can be operably linked to a linker moiety, including a covalent or non-covalent bond; an amino acid tag (e.g., a polyamino acid tag, a poly-His tag, a 6His tag); a chemical compound (e.g., polyethylene glycol); a protein-protein binding pair (e.g., biotin-avidin); an affinity coupling; a capture probe; or any combination of these. The linker moiety can be separated from the polymerase variant or can be part of the polymerase variant.

[0173] In some embodiments, the joint connecting the nucleotide and the polymerase comprises a saturated or unsaturated, substituted or unsubstituted, straight or branched carbon chain. In different embodiments, the length of the joint can be different. The length of the joint can vary according to the type of nucleotide and polymerase. In some embodiments, for each different nucleotide or nucleotide analog, the length of the joint in the enzyme-linked nucleotide is different. In some embodiments, the length of the joint is about 100-1500 Å. In some embodiments, the joint has a length of: Or any value or range between these two values. In some embodiments, the polymerase and the nucleotide are covalently linked, and the distance between the linked atom of the nucleotide and the polymerase is about to about In some embodiments, the distance between the linked atoms of the nucleotide and the polymerase is about to about In some embodiments, the distance between the linked atoms of the nucleotide and the polymerase is about to about In some embodiments, the distance between the linked atoms of the nucleotide and the polymerase is about to about In some embodiments, the distance between the linked atoms of the nucleotide and the polymerase is about to about

[0174] In some embodiments, the length of the linker will be defined as its persistence length, corresponding to the root mean square (RMS) distance between the ends of the linker, as characterized by dynamic simulations, 2-D capture experiments, or ab initio calculations based on the statistical distribution of polymers in the compact, collapsed, or fluid state required for the solution, suspension, or fluid conditions present. In some embodiments, the linker can have a persistence length of at least 0.1, at least 0.2, at least 0.4, at least 1, at least 2, at least 4, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 700, or at least 1,000 nm, or a persistence length within a range limited by or including any two or more of these values. In some embodiments, the linker used to attach the nucleotide to the enzyme can have a persistence length of about 0.1-1,000 nm, 0.5-500 nm, 0.5-400 nm, 0.5-300 nm, 0.5-200 nm, 0.5-100 nm, 0.5-50 nm, 1-500 nm, 1-400 nm, 1-300 nm, 1-200 nm, 1-100 nm, 1-50 nm, 1.5-500 nm, 1.5-400 nm, 1.5-300 nm, 1.5-200 nm, 1.5-100 nm, 1.5-50 nm, 5-500 nm, 5-400 nm, 5-300 nm, 5-200 nm, 5-100 nm, or 5-50 nm. In some embodiments, the linker can have a persistence length of less than about 5, 10, 20, 30, 40, 50, 60, 80, 100, 200, 300, 400, 500, 700, or 1,000 nm. In some embodiments, the linker provided for one nucleotide can be longer or shorter than the linker provided for another nucleotide. In some embodiments, the linker provided for one polymerase can be longer or shorter than the linker provided for another polymerase.

[0175] Conjugate

[0176] In some embodiments, the conjugate is represented by the formula:

[0177]

[0178] In some embodiments, the conjugate is represented by the structure of Formula (I) or (II):

[0179]

[0180] in

[0181] L 1 Selected from the group consisting of: optionally substituted C1-6 Alkylene chain, optionally substituted C 2-6 Alkenylene chain and optionally substituted C 1-6 Alkyne chain, wherein 1-4 methylene units are optionally and independently replaced by -O-, -N(R a )-, -C(O)-, -S-, -S(O)2- or phenylene;

[0182] L 2 is a cleavable linker;

[0183] L 3 To connect pol to L 2 Connectors

[0184] Each R a are independently hydrogen or C 1-6 alkyl;

[0185] R 2 is hydrogen or methyl;

[0186] R is ribose polyphosphate or deoxyribose polyphosphate; and

[0187] Pol is a polymerase.

[0188] In some embodiments, the conjugate is selected from the group consisting of:

[0189]

[0190]

[0191] When a conjugate comprising a polymerase and a nucleotide is incubated with a nucleic acid, it preferentially uses the nucleotide to which it is tethered (rather than using the nucleotide of another conjugate molecule) to extend the nucleic acid. As described above, the polymerase then remains attached to the nucleic acid via its tether to the added nucleotide until exposed to some stimulus that causes cleavage of the linkage to the added nucleotide. In this case, further extension of the polymerase-nucleotide conjugate is hindered due to "shadowing" when: 1) the attached polymerase molecule hinders access to the 3' OH of the extending DNA molecule, and 2) other nucleotides in the system are hindered from accessing the catalytic site of the polymerase, which remains attached to the 3' end of the extending nucleic acid. (The degree of shadowing can be described as the degree to which these two interactions are hindered.) To achieve subsequent extension, the linker tethering the incorporated nucleotide to the polymerase can be cleaved, thereby releasing the polymerase from the nucleic acid and thereby re-exposing its 3' OH group for subsequent extension.

[0192] The methods provided herein for nucleic acid synthesis that utilize shielding effects to achieve termination include an extension step in which the nucleic acid is preferentially exposed to the conjugate in the absence of free (i.e., untethered) nucleoside triphosphates, as shielded termination mechanisms may not prevent their incorporation into the nucleic acid.

[0193] In some embodiments, the termination that further prolongs can be " complete ", which means that after the nucleic acid molecule has been extended by the conjugate, further extension can not occur during the reaction. In other embodiments, the termination that further prolongs can be " incomplete ", which means that further extension can occur during the reaction, but compared with the initial extension, it occurs at a significantly reduced rate, for example, 100 times slower, or 1000 times slower, or 10,000 times slower or more. When the reaction stops after an appropriate amount of time, the conjugate that achieves incomplete termination can still be used to extend nucleic acid (for example, in a method for nucleic acid synthesis and sequencing) mainly by a single nucleotide. In some embodiments, the reagent containing the conjugate may also contain a polymerase that does not contain the tethered nucleotide, but those polymerases should not significantly affect the reaction because there are no free dNTPs in the mixture.

[0194] Reagents based on conjugates that utilize a shielding effect to achieve termination preferably contain only polymerase-nucleotide conjugates in which all polymerases remain folded in an active conformation. In some cases, if the polymerase portion of the conjugate is unfolded, its tethered nucleotides may become more accessible to the polymerase portions of other conjugate molecules. In these cases, the unshielded nucleotides may be more easily incorporated by other conjugate molecules, thereby bypassing the termination mechanism.

[0195] Polymerase-nucleotide conjugates that utilize shielding effects to realize termination are preferentially labeled with only a single nucleotide portion. Polymerase-nucleotide conjugates labeled with multiple nucleotides that can access the catalytic site can incorporate multiple nucleotides into the same nucleic acid in some cases. Therefore, the nucleotides of additional tethers can cause additional, undesirable nucleotides to be incorporated into the nucleic acid during the reaction. In addition, only one tethered nucleotide can occupy the (embedded) catalytic site of its polymerase at a time, so the nucleotides of one or more other tethers can have increased accessibility to the polymerase portion of other conjugate molecules, as discussed below.

[0196] Polymerase-nucleotide conjugates that utilize a shielding effect to achieve termination preferably include a linker that is as short as possible, which still enables the nucleotide to frequently approach the catalytic site of the polymerase molecule to which it is tethered in a productive conformation so that the nucleotide can be quickly incorporated into the nucleic acid. Such conjugates may also preferably utilize a linker that is connected to the polymerase at a position as close to the catalytic site as possible, thereby enabling the use of a shorter linker. The length of the linker will determine the maximum distance from the point of attachment that the tethered nucleotide or tethered nucleic acid can reach. Smaller distances can result in reduced accessibility of the tethered portion to other polymerase-nucleotide molecules, as discussed below. In some embodiments, the linker is approximately and Long. Shorter connectors (e.g. 8- connector) can increase shielding; longer connectors (e.g. or The shielding effect can be affected by a combination of factors, including but not limited to the structure of the polymerase, the length of the linker, the structure of the linker, the position of the linker to the polymerase, the binding affinity of the nucleotide to the catalytic site of the polymerase, the binding affinity of the nucleic acid to the polymerase, the preferred conformation of the polymerase, and the preferred conformation of the linker.

[0197] One contribution to shielding may be steric effects that prevent the 3'OH of a nucleic acid that has been extended by a conjugate from reaching the catalytic site of the polymerase portion of another conjugate. Steric effects may also hinder the tethered nucleotide from reaching the catalytic site of another polymerase-nucleotide conjugate molecule due to clashes between the conjugates that will occur during such methods. These steric effects may result in complete termination if they completely block the effective interaction between the tethered nucleotide (or extended nucleic acid) of one conjugate molecule and the other conjugate molecule, or incomplete termination if they only hinder such intermolecular interactions.

[0198] Another contribution to shielding comes from the binding affinity of the nucleotide of tethering and the catalytic site of polymerase. The tethering nucleotide of the conjugate will have a high effective concentration relative to the catalytic site of the polymerase of its tethering, so it may remain combined with the site for most of the time. When the catalytic site of the polymerase molecule of nucleotide and its tethering is combined, it can not be mixed by other polymerase molecules. Therefore, tethering reduces the effective concentration of nucleotides that can be used for intermolecular incorporation (i.e., incorporation catalyzed by the polymerase molecule of the nucleotide not tethering). This shielding effect can enhance termination by reducing the rate of extending nucleic acid by the polymerase part of another conjugate molecule using the nucleotide part of a conjugate molecule.

[0199] Another contribution to shielding comes from the binding affinity of the 3 ' zone of nucleic acid molecule and the catalytic site of polymerase molecule.After extending by conjugate, nucleic acid is tethered to conjugate by its 3 ' terminal nucleotide, and will have high effective concentration with respect to the catalytic site of the polymerase of its tether, so it may keep being combined with described site in most of the time.When nucleic acid is combined with the catalytic site of the polymerase molecule of its tether, it can not be extended by other conjugate molecule.This effect can strengthen termination by the speed that the nucleic acid that has been extended by the first conjugate is further extended by other conjugate molecule through reducing.

[0200] In some embodiments, the polymerase-nucleotide conjugate comprises an additional moiety that sterically hinders the tethered nucleotide (or the tethered nucleic acid after extension) from accessing the catalytic site of the other conjugate molecule. Such moieties include polypeptides or protein domains that can be inserted into the loops of the polymerase, as well as those that can be site-specifically linked to, for example, inserted non-natural amino acids or specific polypeptide tags, and other macromolecules such as polymers.

[0201] Tethered nucleotides can have high effective concentrations, enabling rapid incorporation kinetics. Depending on the length and geometry of the linker and its attachment site on the protein, the tethered nucleotide has a certain occupancy rate for the active site of the polymerase. This rate can be expressed as an effective concentration (the concentration of free nucleotide that produces an equivalent occupancy rate). By varying the nature of the linker and the attachment site, the effective concentration of the nucleotide can be controlled, thereby achieving a high effective concentration and therefore rapid incorporation. For example, very rough calculations show that a tethered nucleotide has a high effective concentration and therefore a high incorporation rate. The effective concentration of linker-tethered dNTPs will be approximately 50 mM (radius In this example, the local concentration of dNTPs can be increased by shortening the linker or decreased by lengthening the linker. Furthermore, tethering the nucleotide within the polymerase increases the effective concentration of the nucleotide because it cannot access the entire sphere due to limited interaction with the polymerase. For example, if a nucleotide is tethered to a completely flat surface, it can only occupy half of the sphere, so the effective concentration within that hemisphere is doubled relative to the untethered nucleotide.

[0202] polymerase

[0203] It is contemplated that any polymerase capable of extending a polynucleotide, incorporating a nucleotide into a polynucleotide, or incorporating a nucleotide analog into a polynucleotide is useful in the conjugates and methods described herein. In some embodiments, the polynucleotide is single-stranded. In some embodiments, the polynucleotide is double-stranded. In some embodiments, the polynucleotide is immobilized on a solid support.

[0204] Examples of DNA polymerases include polA, polB, polC, polD, polY, polX, reverse transcriptase (RT), and high-fidelity polymerases. In some cases, the polymerase is a modified polymerase. In some embodiments, the polymerase includes 29, B103, GA-1, PZA, 15, BS32, M2Y, Nf, Gl, Cp-1, PRD1, PZE, SF5, Cp-5, Cp-7, PR4, PR5, PR722, L17, 9°Nm TM 、Therminator TM DNA polymerase, Tne, Tma, Tfl, Tth, TIi, Stoffel fragment, Vent TM and Deep Vent TM DNA polymerase, KOD DNA polymerase, Tgo, JDF-3, Pfu, Taq, T7 DNA polymerase, T7 RNA polymerase, PGB-D, UlTma DNA polymerase, Escherichia coli DNA polymerase I, Escherichia coli DNA polymerase III, Archaea DP1I / DP2 DNA polymerase II, 9°N DNA polymerase, Taq DNA polymerase, DNA polymerase, Pfu DNA polymerase, SP6 RNA polymerase, RB69 DNA polymerase, avian myeloblastosis virus (AMV) reverse transcriptase, Moloney murine leukemia virus (MMLV) reverse transcriptase, II reverse transcriptase and III reverse transcriptase.

[0205] In some embodiments, the polymerase is DNA polymerase 1-KI enow fragment, Vent polymerase, DNA polymerase, KOD DNA polymerase, Taq polymerase, T7 DNA polymerase, T7 RNA polymerase, Therminator TM DNA polymerase, POLB polymerase, SP6 RNA polymerase, Escherichia coli DNA polymerase I, Escherichia coli DNA polymerase III, avian myeloblastosis virus (AMV) reverse transcriptase, Moloney murine leukemia virus (MMLV) reverse transcriptase, II reverse transcriptase or III reverse transcriptase.

[0206] The polymerase molecule used in the methods described herein can be polymerase θ, DNA polymerase or any enzyme that can extend nucleotide chain. In some embodiments, polymerase is tri29. In some embodiments, polymerase is a protein with a pocket that works around a terminal phosphate group (e.g., a triphosphate group). In some embodiments, the method uses TdT with 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid mutations to synthesize determined polynucleotides. In some embodiments, the method uses TdT with 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid mutations to surface accessible amino acid residues. In some embodiments, TdT is a TdT variant. In some embodiments, the TdT variant comprises a cysteine mutation. In some embodiments, the polymerase is mutated to increase the addition of the modified nucleotides combined with the polymerase to form a conjugate. In some cases, variant TdT and wild-type TdT comprise at least 70%, 80%, 90% or 95% sequence identity.

[0207] In some embodiments, the method uses a polymerase θ having 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid mutations to synthesize a defined polynucleotide. In some embodiments, the method uses a polymerase θ having 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid mutations to surface accessible amino acid residues. In some embodiments, the polymerase θ is a polymerase θ variant. In some cases, the variant polymerase θ comprises at least 70%, 80%, 90%, or 95% sequence identity to wild-type polymerase θ. In some embodiments, the polymerase θ is encoded by POLQ.

[0208] In some embodiments, the enzymes described herein (e.g., TdT) comprise one or more non-natural amino acids. In some cases, the non-natural amino acids include: lysine analogs; aromatic side chains; azido groups; alkynyl groups; or aldehyde groups or ketone groups. In some cases, the non-natural amino acids do not include aromatic side chains. In some embodiments, the non-natural amino acids are selected from N6-azidoethoxy-carbonyl-L-lysine (AzK), N6-propargyloxy-carbonyl-L-lysine (PraK), N6-(propargyloxy)-carbonyl-L-lysine (PrK), p-azido-phenylalanine, BCN-L-lysine, norbornene lysine, TCO-lysine, methyl tetrazine lysine, allyloxycarbonyl lysine, 2-amino-8- Oxonanoic acid, 2-amino-8-oxooctanoic acid, p-acetyl-L-phenylalanine, p-azidomethyl-L-phenylalanine (pAMF), p-iodo-L-phenylalanine, o-acetylphenylalanine, 2-amino-8-oxononanoic acid, p-propynyloxyphenylalanine, p-propargyl-phenylalanine, 3-methyl-phenylalanine, L-DOPA, fluorinated phenylalanine, isopropyl-L-phenylalanine, p-azido-L-phenylalanine, p- -Acyl-L-phenylalanine, p-benzoyl-L-phenylalanine, p-bromophenylalanine, p-amino-L-phenylalanine, isopropyl-L-phenylalanine, O-allyl-tyrosine, O-methyl-L-tyrosine, O-4-allyl-L-tyrosine, 4-propyl-L-tyrosine, phosphotyrosine, tri-O-acetyl-GlcNAcp-serine, L-phosphoserine, phosphoserine, L-3-(2-naphthyl)-1,2-dione (((3-(benzyloxy)-3-oxopropyl)amino)ethyl)seleno)propanoic acid, 2-amino-3-(phenylseleno)propanoic acid, selenocysteine, N6-(((2-azidobenzyl)oxy)carbonyl)-L-lysine, N6-(((3-azidobenzyl)oxy)carbonyl)-L-lysine, and N6-(((4-azidobenzyl)oxy)carbonyl)-L-lysine.

[0209] In some embodiments, the polymerase is a fusion protein. In some embodiments of the method, the fusion protein comprises maltose binding protein (MBP). In some embodiments, TdT is fused to other enzymes such as a helicase.

[0210] In some embodiments, the polymerase comprises a template-independent polymerase. In some embodiments, the polymerase comprises a Pol-X family polymerase. In some embodiments, the polymerase comprises a terminal deoxynucleotidyl transferase (TdT) or a variant thereof. In some embodiments, the template-independent polymerase comprises TdT or a variant thereof. In some embodiments, TdT or a variant thereof comprises a sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identity with SEQ ID NO: 1. In some embodiments of the method, TdT comprises a sequence identical to SEQ ID NO: 1. In some embodiments, the TdT variant comprises one or more amino acid substitutions, insertions or deletions relative to SEQ ID NO: 1.

[0211] >Terminal deoxynucleotidyl transferase (TdT)

[0212] MGGRDIVDGSEFSSPSPPGSQNVPAPAVKKISQYACQRRTTLNNYNQLFTDALDILAENDELRENEGSALAFMRASSVLKSLPFPITSMKDTEGIPSLGDKVKSIIEGIIEDGESSEAKAVLNDERYKSFKLFTSVFGVGLKTAEKWFRMGFRTLSKIQSDKSLRFTQMQKAGFLYYEDLVSCVNRPEAEAVSMLVK EAVVTFLPDALVTMTGGFRRGKMTGHDVDFLITSPEATEDEEQQLLHKVTDFWKQQGLLLYADILESTFEKFKQPSRKVDALDHFQKCFLILKLDHGRVHSEKSGQQEGKGWKAIRVDLVMSPYDRRAFALLGWTGSRQFERDLRRYATHERKMMLDNHALYDRTKRVFLEAESEEEIFAHGLDYIEPWERNA(SEQ ID NO:1)

[0213] In some embodiments, a template-independent polymerase having activity as described for EC class 2.7.7.31 is used. In some embodiments, the template-independent polymerase is a deoxynucleotidyl transferase or a DNA nucleotidyl transferase. Descriptions of such enzymes can be found in Bollum, FJ Deoxynucleotide-polymerizing enzymes of calfthymus gland. V. Homogeneous terminal deoxynucleotidyltransferase. J. Biol. Chem. 246 (1971) 909-916; Gottesman, ME and Canellakis, ES The terminal nucleotidyltransferases of calf thymus nuclei. J. Biol. Chem. 241 (1966) 4339-4352; and Krakow, JS, Coutsogeorgopoulos, C. and Canellakis, ES Studies on the incorporation of deoxyribonucleic acid. Biochim. Biophys. Acta 55 (1962) 639-650, among others.

[0214] Other polymerases that can be used that have the ability to extend single-stranded nucleic acids in the absence of a template include, but are not limited to, polymerase theta (Kent et al., eLife 5 (2016): e13740.), polymerase μ (Juarez et al., Nucleic acids research 34.16 (2006): 4572-4582.; or McElhinny et al., Molecular cell 19.3 (2005): 357-366.), or polymerases in which template-independent activity is induced, for example, by insertion of elements of a template-independent polymerase (Juarez et al., Nucleic acids research 34.16 (2006): 4572-4582).

[0215] In other DNA synthesis applications, the polymerase can be a template-dependent polymerase, ie, a DNA-directed DNA polymerase (eg, an enzyme having activity 2.7.7.7 using IUBMB nomenclature) or an RNA-directed DNA polymerase. Descriptions of such enzymes can be found in Richard son, A. Enzymatic synthesis of deoxyribonucleic acid. deoxythymidylate. J. Biol. Chem. 235 (1960) 3242-3249; and Zimmerman, BKP Purification and properties of deoxyribonucleic acid polymerase from Micrococcus lysodeikticus. J. Biol. Chem. 241 (1966) 2035-2041.

[0216] In some embodiments, polymerase comprises RNA polymerase.In these embodiments, RNA specific nucleotidyl transferase can be adopted, such as Escherichia coli Poly (A) polymerase (IUBMB EC 2.7.7.19) or Poly (U) polymerase etc. RNA nucleotidyl transferase can contain modification (such as single point mutation), and the modification affects the substrate specificity (Lunde et al., Nucleic acids research 40.19 (2012): 9815-9824.) for specific rNTP. In some embodiments, the very short tether between RNA nucleotidyl transferase and ribonucleotide can be used for inducing the high effective concentration of nucleotide, thereby forcing the incorporation of rNTP that may not be the natural substrate of nucleotidyl transferase.

[0217] polynucleotides

[0218] The term "nucleotide" and related terms refer to molecules comprising an aromatic base, a five-carbon sugar (e.g., ribose or deoxyribose) and at least one phosphate group. Typical nucleotides or atypical nucleotides are consistent with the use of the term. In some embodiments, the phosphate comprises monophosphate, diphosphate, or triphosphate, or a corresponding phosphate analog.

[0219] Nucleotides (and nucleosides) typically contain heterocyclic bases, including substituted or unsubstituted nitrogen-containing parent heteroaromatic rings commonly found in nucleic acids, including naturally occurring, substituted, modified or engineered variants or analogs thereof. Exemplary bases include, but are not limited to, purines and pyrimidines, for example: 2-aminopurine, 2,6-diaminopurine, adenine (A), ethyleneadenine, N6-Δ2-isopentenyladenine (6iA), N6-Δ2-isopentenyl-2-methylthioadenine (2ms6iA), N6-methyladenine, guanine (G), isoguanine, N2-dimethylguanine (dmG), 7-methylguanine (7mG), 2-thiopyrimidine, 6-thioguanine (6sG), hypoxanthine and O6-methylguanine; 7-deazapurines, such as 7-deazaadenine (7-deaza- A) and 7-deazaguanine (7-deaza-G); pyrimidines such as cytosine (C), 5-propynylcytosine, isocytosine, thymine (T), 4-thiothymine (4sT), 5,6-dihydrothymine, O4-methylthymine, uracil (U), 4-thiouracil (4sU) and 5,6-dihydrouracil (dihydrouracil; D); indoles such as nitroindole and 4-methylindole; pyrroles such as nitropyrrole; muscimol; inosine; hydroxymethylcytosine; 5-methylcytosine; base (Y); and methylated, glycosylated and acylated base moieties; etc. Other exemplary bases can be found in Fasman, 1989, in "Practical Handbook of Biochemistry and Molecular Biology", pages 385-394, CRC Press, Boca Raton, Fla.

[0220] Nucleotides (and nucleosides) typically contain a sugar moiety, such as a carbocyclic moiety (Ferraro and Gotor 2000 Chem. Rev. 100:4319-48), an acyclic moiety (Martinez et al., 1999 Nucleic Acids Research 27:1271-1274; Martinez et al., 1997 Bioorganic & Medicinal Chemistry Letters Vol. 7:3013-3016), and other sugar moieties (Joeng et al., 1993 J. Med. Chem. 36:2627-2638; Kim et al., 1993 J. Med. Chem. 36:30-7; Eschenmosser 1999 Science 284:2118-2124; and U.S. Pat. No. 5,558,991). Sugar moieties include: ribosyl; 2′-deoxyribosyl; 3′-deoxyribosyl; 2′,3′-dideoxyribosyl; 2′,3′-didehydrodideoxyribosyl; 2′-alkoxyribosyl; 2′-azidoribosyl; 2′-aminoribosyl; 2′-fluororibosyl; 2′-thiolribosyl; 2′-alkylthioribosyl; 3′-alkoxyribosyl; 3′-azidoribosyl; 3′-aminoribosyl; 3′-fluororibosyl; 3′-thiolribosyl; 3′-alkylthioribosyl carbocyclic; acyclic sugars or other modified sugars.

[0221] In some embodiments, nucleotide comprises a chain of one, two or three phosphorus atoms, wherein the chain is connected to the 5' carbon of the sugar moiety via an ester or phosphoramidite bond. In some embodiments, nucleotide is an analog with a phosphorus chain, wherein the phosphorus atom is linked together by the O, S, NH, methylene or ethylene inserted. In some embodiments, the phosphorus atom in the chain comprises a substituted side group, including O, S or BH3. In some embodiments, the chain comprises a phosphate group substituted by an analog, and the analog comprises phosphoramidate, phosphorothioate, phosphorodithioate and O-methylphosphoramidite groups.

[0222] In some embodiments, the polymerase of the conjugate can be covalently linked to the oligonucleotide or nucleotide via a nucleotide base. For example, the nucleotide or oligonucleotide can have a polymerase linked to the C5 position of a pyrimidine base or the C7 position of a 7-deazapurine base through a linker moiety.

[0223] The nucleotide used in the present disclosure can also include natural or non-natural bases. In this respect, natural deoxyribonucleic acid can have one or more bases selected from the group consisting of adenine, thymine, cytosine or guanine, and ribonucleic acid can have one or more bases selected from the group consisting of uracil, adenine, cytosine or guanine. The exemplary non-natural bases (no matter with natural backbone or analog structure) that can be included in nucleic acid include but are not limited to inosine, xanthine, hypoxanthine, isocytosine, isoguanine, 5-methylcytosine, 5-hydroxymethylcytosine, 2-aminoadenine, 6-methyladenine, 6-methylguanine, 2-propylguanine, 2-propyladenine, 2-thiouracil (thioLiracil), 2-thiothymine, 2-thiocytosine, 15-halouracil, 15-halocytosine, 5-propynyluracil, 5-propynyl cytosine, 6-azauracil, 6-azacytosine, 6-azathymine, 5-uracil, 4-thiouracil, 8-haloadenine or guanine, 8-aminoadenine or guanine, 8-thioadenine or guanine, 8-thioalkyladenine or guanine, 8-hydroxyadenine or guanine, 5-halogenated uracil or cytosine, 7-methylguanine, 7-methyladenine, 8-azaguanine, 8-azaadenine, 7-deazaguanine, 7-deazaadenine, 3-deazaguanine, 3-deazaadenine, etc.

[0224] In some embodiments, the phosphorylated nucleoside (e.g., nucleotide) to be tethered to the polymerase is a nucleoside comprising at least one phosphate group. In some embodiments, the nucleoside comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9 or more than 9 phosphate groups. In some embodiments, the nucleoside comprises at least 3 phosphate groups. In some embodiments, the phosphorylated nucleoside is adenosine, cytidine, uridine or guanosine, each of which comprises at least one phosphate group. In some embodiments, the phosphorylated nucleoside is a deoxynucleoside comprising at least one phosphate group. In some embodiments, the phosphorylated nucleoside is a deoxynucleoside comprising at least 3 phosphate groups. In some embodiments, the deoxynucleoside comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9 or more than 9 phosphate groups. In some embodiments, the phosphorylated nucleoside is deoxyadenosine, deoxycytidine, deoxythymidine or deoxyguanosine, each of which comprises at least one phosphate group. In some embodiments, the phosphorylated nucleoside is a nucleoside triphosphate, such as dNTP. In some embodiments, the phosphorylated nucleoside is a nucleoside tetraphosphate, a nucleoside pentaphosphate, a nucleoside hexaphosphate, a nucleoside heptaphosphate, a nucleoside octaphosphate or a nucleoside ninephosphate. In some embodiments, the phosphorylated nucleoside is a nucleoside hexaphosphate. In some embodiments, the phosphorylated nucleoside is a nucleoside triphosphate. In some embodiments, the phosphorylated nucleoside is selected from the group consisting of: deoxyadenosine triphosphate (dATP), deoxyguanosine triphosphate (dGTP), deoxycytidine triphosphate (dCTP), deoxythymidine triphosphate (dTTP), deoxyadenosine tetraphosphate, deoxyguanosine tetraphosphate, deoxycytidine tetraphosphate, deoxythymidine tetraphosphate, deoxyadenosine pentaphosphate, deoxyguanosine pentaphosphate, deoxycytidine pentaphosphate, deoxythymidine pentaphosphate, deoxyadenosine hexaphosphate, deoxyguanosine hexaphosphate, deoxythymidine hexaphosphate, deoxythymidine hexaphosphate and any combination thereof.

[0225] In some embodiments, the nucleotide analogs described herein comprise a reversible terminator group, such as an O-azidomethyl or O-NH2 group at the 3' position of the sugar or a (α-tert-butyl-2-nitrobenzyl)oxymethyl group at the 5' position of a pyrimidine or the 7' position of a 7-deazapurine (for an overview, see, e.g., Chen et al., Genomics, Proteomics & Bioinformatics 201311:34-40). In these embodiments, the nucleotide analogs, once incorporated into the nucleic acid, prevent or hinder further extension to achieve controlled termination of synthesis. In some embodiments, when used as part of a conjugate, the RTdNTP-polymerase conjugate does not rely on a shielding effect to achieve termination, e.g., when the 3' modified RTdNTP is tethered to the polymerase, the length of the linker used may exceed 100%. or

[0226] method

[0227] The conjugates described above can be used in nucleic acid synthesis methods. Nucleic acid synthesis can refer to the synthesis or generation of a product as a nucleic acid molecule (e.g., polynucleotide). Nucleic acid synthesis methods can include stepwise synthesis, wherein nucleotides are progressively inserted into nucleic acid polymers or polynucleotides. As a non-limiting example, a typical method for the progressive synthesis of polynucleotides includes progressively adding nucleotides to a starting molecule (e.g., an initial oligonucleotide) via the following cyclic steps: under conditions suitable for covalently binding nucleotides to the ends of the oligonucleotides catalyzed by a polymerase, a polymerase-nucleotide conjugate is added to the oligonucleotide. Successful incorporation of the nucleotides of the conjugate into the oligonucleotide can be referred to as "extension" or "extension reaction," which produces "extension products."

[0228] In some embodiments, the method includes: under the condition that the nucleotide of the first conjugate of polymerase catalysis is covalently added to the 3' hydroxyl of nucleic acid, the nucleic acid is hatched with the first conjugate to form an extension product. The reaction can be carried out using a nucleic acid connected to a solid support or in a solution (for example, not tethered to a solid support). Adding the conjugate to the nucleic acid causes the nucleic acid added with the 3' group to be shielded by the connected polymerase, thereby suppressing the subsequent addition of another nucleotide when the polymerase is connected. After the nucleic acid is extended by the first desired nucleotide, the method includes a deprotection (deshielding) step, in which the cleavable bond of the joint is cleaved, thereby releasing the polymerase from the extension product. The cracking of the joint removes the polymerase to produce a deprotected extension product. Deprotection enables the nucleic acid to be subsequently extended, thereby allowing these steps to be repeated cyclically to produce an extension product of a determined sequence. Specifically, in some embodiments, the method may also include after deprotection: under the condition that the nucleotide of the second conjugate of polymerase catalysis is covalently added to the 3' end of the extension product, the deprotected extension product is hatched together with the second conjugate.

[0229] In some embodiments, the method may include: (a) incubating the nucleic acid with a first conjugate under conditions where a polymerase catalyzes the covalent addition of a nucleotide (i.e., a single nucleotide) of the first conjugate to the 3' hydroxyl group of the nucleic acid to produce an extension product; (b) cleaving the cleavable bond of the linker, thereby releasing the polymerase from the extension product and deprotecting the extension product; (c) incubating the deprotected extension product with a second conjugate as described in claim 1 under conditions where a polymerase catalyzes the covalent addition of a nucleotide of the second conjugate to the 3' end of the extension product to produce a second extension product; (d) repeating steps (b)-(c) multiple times (e.g., 2 to 100 times or more) on the second extension product to produce an extension oligonucleotide of a defined sequence. Steps (b)-(c) can be repeated as many times as desired until an extension product of a defined sequence and length is synthesized. The final product can be 2-100 bases in length, although in theory, the method can be used to produce products of any length, including greater than 200 bases or greater than 500 bases.

[0230] In some embodiments, as provided herein, the method for nucleic acid synthesis is implemented in a reaction buffer composition. In some embodiments, the reaction buffer composition is an aqueous solution. In some embodiments, the reaction buffer composition comprises a group of components suitable for the stability of any surface or matrix of a polymerase, nucleotides, polymerase-nucleotide conjugates, starting molecules, nucleic acid molecule products, and methods disclosed herein. In some such embodiments, the reaction buffer composition comprises a group of components suitable for implementing the catalytic step (e.g., the polynucleotide polymerization performed by a polymerase) described in the method for nucleic acid synthesis of the present disclosure.

[0231] The conditions under which nucleic acid synthesis is performed can be varied. For example, the number of times each step in the stepwise nucleotide addition cycle is performed can be varied to improve the purity of the various products produced by the nucleic acid synthesis methods described herein.

[0232] In some embodiments, nucleic acid molecule products (i.e., polynucleotide products) are produced according to the nucleic acid synthesis method of the present disclosure. In some embodiments, nucleic acid molecule products (i.e., polynucleotide products) have a target sequence (i.e., predetermined sequence). "Target" sequence or "predetermined" sequence refers to the desired polynucleotide sequence intentionally produced by the nucleic acid synthesis method. The predetermined sequence can include any number of nucleotides comprising core bases (e.g., adenine, thymine, guanine, cytosine and / or uracil). In some embodiments, nucleotides are modified nucleotides (i.e., nucleotide analogs). In some embodiments, core bases are modified core bases. In some embodiments, the predetermined sequence includes one or more designated positions, which can be random core bases. Comprise that the position with random core bases can be useful, for example, in the polynucleotide product, introduce random mutations.

[0233] In some embodiments, the disclosure includes a method for synthesizing polynucleotides, the method comprising contacting a precursor polynucleotide with a conjugate comprising a nucleotide covalently linked to a polymerase via a cleavable joint, wherein the nucleotide comprises the protected core base. In some embodiments, the method for synthesizing polynucleotides is included in cracking the cleavable joint after adding the nucleotide to the precursor polynucleotide. In some embodiments, the method for synthesizing polynucleotides includes repeating contact as described herein, addition, and optional cleavage step once or many. In some embodiments, the removal of one or more blocking groups as described herein includes contacting the polynucleotides with an enzyme that can remove the one or more blocking groups from the protected core base. In some embodiments, the method for synthesizing polypeptides includes contacting the polynucleotides with two or more enzymes that can remove the one or more blocking groups from the protected core base.

[0234] In some embodiments, the synthesis of polynucleotides comprises the stepwise addition of nucleotides to a starting molecule (e.g., an initial oligonucleotide) via the following cyclic steps: adding a polymerase-nucleotide conjugate to the oligonucleotide, binding the nucleotide to the 3' end of the oligonucleotide catalyzed by the polymerase, and cleaving the polymerase from the added nucleotides. These steps can be repeated until the desired polynucleotide is synthesized. As described herein, the use of nucleotides comprising protected nucleobases during polynucleotide synthesis helps to improve the efficiency and accuracy of synthesis by inhibiting the formation of secondary structures that interfere with the addition of introduced nucleotides by the polymerase during synthesis.

[0235] Although synthesis can be completed entirely with protected nucleotides, synthesis with a combination of unmodified nucleotides and protected nucleotides can also be effectively used to improve polynucleotide synthesis. In some embodiments, only one of the four nucleotides added (e.g., from G or T) is protected during synthesis. In some embodiments, protected nucleotides are added only at target positions where secondary or tertiary structures are predicted, which may interfere with synthesis. Such structures can be predicted in various ways based on the presence of complementary DNA regions, and corresponding tools exist, such as the NUPACK algorithm ( http: / / www.nupack.org / home / model). Therefore, in some embodiments, the synthesis of improved complete polynucleotides can include the use of only 1 or 2 protected nucleotides. In some embodiments, about 5%, about 10%, about 20%, about 30%, about 50%, substantially all, or 100% of a specific nucleotide is incorporated into the polynucleotide in its protected form. In some embodiments, less than 5%, less than 10%, less than 20%, less than 30%, or less than 50% of a specific nucleotide is incorporated into the polynucleotide in its protected form. In some embodiments, more than 5%, more than 10%, more than 20%, more than 30%, or more than 50% of a specific nucleotide is incorporated into the polynucleotide in its protected form. In some embodiments, only protected guanine nucleotides are used in the nucleotide synthesis reaction. Removing protecting groups in the terminal positions of nucleic acids may be more challenging than removing them from internal DNA positions. Therefore, in some embodiments, nucleotide synthesis is performed so that the last and first 1, 2, or 3 positions of the synthesized nucleic acid do not contain protected nucleotides.

[0236] The nucleic acid molecule product or polynucleotide product produced by the methods described herein can contain multiple products. In some embodiments, multiple products include nucleic acid molecules containing target sequences (i.e., predetermined sequences). In some embodiments, multiple products include nucleic acid molecules containing sequences that are not target sequences. In some embodiments, multiple products include nucleic acid molecule products containing target sequences and nucleic acid molecule products that are not target sequences. The "purity" of multiple products can refer to the abundance of nucleic acid molecule products with target sequences and the ratio of the abundance of nucleic acid molecule products without target sequences. The purity of product can be assessed by many methods known in the art for determining nucleic acid sequences. Any suitable nucleic acid sequencing method can be used. For example, it can be not limited to but can be assessed by Sanger sequencing, next generation sequencing (e.g., Illumina sequencing) or long read sequencing (e.g., small molecules, real-time sequencing (SMRT) and nanopore sequencing).

[0237] In some embodiments, the nucleic acid synthesis methods according to the present disclosure produce products with a purity between about 10% and about 99.99%. In some embodiments, the nucleic acid synthesis methods produce products with a purity of at least 10%. In some embodiments, the nucleic acid synthesis methods produce products with a purity of at least 10%. In some embodiments, the nucleic acid synthesis methods produce products with a purity of at least 20%. In some embodiments, the nucleic acid synthesis methods produce products with a purity of at least 30%. In some embodiments, the nucleic acid synthesis methods produce products with a purity of at least 40%. In some embodiments, the nucleic acid synthesis methods produce products with a purity of at least 50%. In some embodiments, the nucleic acid synthesis methods produce products with a purity of at least 60%. In some embodiments, the nucleic acid synthesis methods produce products with a purity of at least 70%. In some embodiments, the nucleic acid synthesis methods produce products with a purity of at least 80%. In some embodiments, the nucleic acid synthesis methods produce products with a purity of at least 90%. In some embodiments, the nucleic acid synthesis methods produce products with a purity of at least 95%. In some embodiments, the nucleic acid synthesis method produces a product that is at least 99% pure.

[0238] In any of the embodiments outlined above, the nucleoside triphosphates can be deoxyribonucleoside triphosphates or ribonucleoside triphosphates. In some embodiments, the conjugate can include an RNA polymerase that is connected to the ribonucleoside triphosphates. In these embodiments, the nucleotides added to the nucleic acid can be ribonucleotides. In other embodiments, the conjugate includes a DNA polymerase that is connected to deoxyribonucleoside triphosphates. In these embodiments, the nucleotides added to the nucleic acid can be deoxyribonucleotides.

[0239] In some embodiments, the nucleotide is a nucleotide analog. In some embodiments, the nucleotide analog is a reversible terminator. Reversible terminators are terminators known in the art for use in nucleic acid synthesis. The use of reversible terminators in nucleic acid synthesis has been previously described; see, for example, WO 2021 / 122539 A1, WO 2018 / 215803 A1, WO 2021 / 094251 A1, and WO 2020 / 081985 A1.

[0240] In some embodiments, the nucleotides may contain reversible terminators ("RTdNTPs"), and the deprotection step of the method further comprises removing a blocking group (e.g., removing a terminator group) from the added nucleotides to produce a deprotected extension product. Deprotection enables subsequent extension of the nucleic acid, thereby allowing cyclic repetition of these steps to produce extension products of a defined sequence.

[0241] Sequencing methods are also provided. These methods can include incubating a duplex comprising a primer and a template with a composition comprising a set of conjugates, wherein the conjugates correspond to G, A, T, and C and are distinguishably labeled, for example, fluorescently labeled; detecting which nucleotide has been added to the primer by detecting a label tethered to a polymerase that has added the nucleotide to the primer; deprotecting the extension product by cleaving a linker; and repeating the incubation, detection, and deprotection steps to obtain the sequence of at least a portion of the template.

[0242] Conjugate synthesis

[0243] In some examples, the polymerase-nucleotide conjugate is prepared by first synthesizing an intermediate compound comprising a linker and a nucleotide (referred to herein as a "linker-nucleotide"), and then attaching the intermediate compound to the polymerase.

[0244] Those skilled in the art will appreciate that the conjugates and nucleotides disclosed herein can be prepared in a manner similar to the reaction schemes shown below. The synthetic methods outlined in these reaction schemes can be described for specific nucleotides. Similar synthetic methods can be applied to related nucleotide analogs.

[0245] There are several known reactions and functional groups suitable for producing nucleotide polymerase conjugates with cleavable linkers consistent with those described herein. For example, linkage of the conjugate components can be achieved by disulfide formation (forming readily cleavable linkages), amide formation, ester formation, protein-ligand bonds (e.g., biotin-streptavidin bonds), by alkylation (e.g., using substituted iodoacetamide reagents), or adduct formation using aldehydes and amines or hydrazines.

[0246] In some embodiments, the individual components of the conjugate comprise sites suitable for conjugation to facilitate conjugate synthesis (i.e., conjugation groups). Examples of such conjugation groups include, but are not limited to, hydroxyl, ester, amine, carbonate, acetal, aldehyde, aldehyde hydrate, alkenyl, acrylate, methacrylate, acrylamide, active sulfone, hydrazide, thiol, alkanoic acid, acyl halide, isocyanate, isothiocyanate, maleimide, vinyl sulfone, dithiopyridine, vinylpyridine, iodoacetamide, epoxide, glyoxal, diketone, mesylate, tosylate, and tresylate.

[0247] Other examples of conjugated groups include -NH2, -COOH, -COOCH3, -N-hydroxysuccinimide, and -maleimide. In some embodiments, the bioconjugate reactive group can be protected (e.g., with a protecting group). Additional examples of bioconjugate reactive groups and resulting bioconjugate reactive linkers can be found, for example, in PCT Publication WO2021 / 226327, the entire contents of which are incorporated herein by reference.

[0248] Linker-nucleotide synthesis

[0249] Described herein are exemplary reaction schemes for preparing linker-nucleotide conjugates having linkers comprising amino acid esters, including linkers described herein and variants thereof. In some embodiments, the desired nucleotides can be commercially obtained, and the hydroxyl groups are protected by TBS before the L1-OH group is conjugated to the exocyclic oxygen or amine on the nucleobase. Alternatively, modified nucleotides can be obtained that contain an L1-OH group bound to a nucleotide, such as an L1-OH group bound to the C5 of a pyrimidine or an L1-OH group bound to the C7 of a 7-deazapurine. An Fmoc-protected amino acid ester is then coupled to the hydroxyl group of the L1-OH, followed by deprotection of the hydroxyl group and triphosphorylation of the nucleoside.

[0250] To complete the linker, the Fmoc is removed from the amino acid ester amine group and coupled to the remainder of the linker (including L3), which is capable of binding to the polymerase.

[0251] Nucleotide ligation

[0252] In some embodiments, the linker binds to a portion of a nucleotide at an atom that does not participate in base pairing. In other embodiments, the linker binds to a nucleobase of a nucleotide at an atom that participates in base pairing. In some embodiments, the linker is considered to be at least an atom that connects the polymerase to any atom in a monocyclic or polycyclic ring system that is bonded to the 1′ position of a sugar (e.g., a pyrimidine or purine or a 7-deazapurine or an 8-aza-7-deazapurine).

[0253] Certain polymerases have high tolerance for modification of certain parts of nucleotides, for example, some polymerases tolerate modification of the 5-position of pyrimidines and the 7-position of purines well (He and Seela., Nucleic Acids Research 30.24(2002):5485-5496.; or Hottin et al., Chemistry. 2017 Feb 10;23(9):2109-2118). In some embodiments, the linker is attached to the 5-position of a pyrimidine or the 7-position of a 7-deazapurine. In other embodiments, the linker can be attached to the exocyclic amine of a nucleobase, for example, by N-alkylating the exocyclic amine of cytosine with a nitrobenzyl moiety as discussed below.

[0254] In other embodiments, the linker is connected to the sugar or α-phosphate of the nucleotide. In some embodiments, the linker is connected to the terminal phosphate of the nucleotide. In all embodiments, the linker used should be long enough to allow the nucleotide to approach the active site of the polymerase to which it is tethered. As will be described in more detail below, the polymerase of the conjugate is capable of catalyzing the addition of the nucleotide to which it is connected to the 3' end of the nucleic acid.

[0255] The conjugation of nucleotides or other base pairing moieties to the linker can be achieved by any means known in the art of chemical conjugation methods. The nucleotide bases can be obtained or modified to comprise the L1 portion of the linker. The remainder of the linker can be connected to L1 using the methods exemplified herein. Those skilled in the art will know or be able to determine the appropriate method for connecting the linker based on the reactivity of these bases.

[0256] In some embodiments, it is contemplated that base-modified nucleotides containing added free amine groups are used for conjugation with linkers described herein. For example, primary amines can be attached to bases in such a way that they can react with heterobifunctional polyethylene glycol (PEG) linkers to produce nucleotides containing variable length PEG linkers. Examples of such amine-containing nucleotides include 5-propargylamino-dNTPs, 5-propargylamino-NTPs, aminoallyl-dNTPs, and aminoallyl-NTPs.

[0257] polymerase ligation

[0258] In some embodiments, tethered nucleotides can be specifically linked to the cysteine residues of a polymerase using sulfhydryl-specific linking chemistry. Possible sulfhydryl-specific linking chemistry includes but is not limited to adjacent-pyridyl disulfide (OPSS), maleimide functional group, 3-aryl propiolate nitrile functional group, allenamide functional group, haloacetyl functional group (such as iodoacetyl or bromoacetyl), alkyl halide or perfluoroaryl, which can advantageously react with sulfhydryl groups surrounded by a specific amino acid sequence (Zhang, Chi, et al. Nature chemistry 8, (2015) 120-128.). Other linking chemistries for specific labeling of cysteine residues will be clear to those skilled in the art, or are described in related literature and text (e.g., Kim, Younggyu et al., Bioconjugate chemistry 19.3 (2008): 786-791).

[0259] In other embodiments, the linker can be linked to a lysine residue via an amine-reactive functional group, such as an NHS ester, a sulfo-NHS ester, a tetrafluorophenyl ester or a pentafluorophenyl ester, an isothiocyanate, a sulfonyl chloride, etc. In other embodiments, the linker can be linked to the polymerase via linkage to a genetically inserted unnatural amino acid, such as p-propargyloxyphenylalanine or p-azidophenylalanine, which can undergo an azide-alkyne Huisgen cycloaddition, although many suitable unnatural amino acids suitable for site-specific labeling exist and are described in the literature (e.g., Lang and Chin., Chemical reviews 114.9 (2014): 4764-4806).

[0260] In other embodiments, the linker can be specifically connected to the N-terminal of the polymerase. In some embodiments, the polymerase is mutated to have an N-terminal serine or threonine residue, which can be specifically oxidized to produce an N-terminal aldehyde for subsequent coupling to, for example, hydrazides. In other embodiments, the polymerase is mutated to have an N-terminal cysteine residue that can be specifically labeled with an aldehyde to form a thiazolidine. In other embodiments, the N-terminal cysteine residue can be labeled with a peptide linker by native chemical connection.

[0261] In other embodiments, the peptide tag sequence can be inserted into a polymerase that can be specifically labeled with a synthetic group by an enzyme, for example, as demonstrated in the literature using biotin ligase, transglutaminase, lipoic acid ligase, bacterial taxonomic enzyme, and phosphopantetheinyltransferase (e.g., as described in references 74-78 of Stephanopoulos & Francis Nat. Chem. Biol. 7, (2011) 876-884).

[0262] In other embodiments, the linker is connected to a tag domain fused to a polymerase. For example, a linker with a corresponding reactive portion can be used to covalently tag a SNAP tag, a CLIP tag, a HaloTag tag, and an acyl carrier protein domain (e.g., as described in references 79-82 of Stephanopoulos & Francis Nat. Chem. Biol. 7, (2011) 876-884).

[0263] In other embodiments, the linker is linked to an aldehyde specifically generated within the polymerase, as described in Carrico et al. (Nat. Chem. Biol. 3, (2007) 321-322). For example, after an amino acid sequence recognized by formylglycine generating enzyme (FGE) is inserted into the polymerase, it can be exposed to FGE, which will specifically convert the cysteine residues in the recognition sequence into formylglycine (i.e., generate an aldehyde). This aldehyde can then be specifically labeled with, for example, a hydrazide or aminooxy moiety of the linker.

[0264] In some embodiments, the linker can be connected to the polymerase by non-covalent binding of the portion of the linker to the portion fused to the polymerase. Examples of such connection strategies include fusing the polymerase to a streptavidin that can bind to the biotin portion of the linker, or fusing the polymerase to an anti-digoxigenin that can bind to the digoxigenin portion of the linker. In some embodiments, site-specific labeling can result in the linker being connected to the polymerase, and the connection can be easily reversed (e.g., orthopyridyl disulfide (OPSS) groups forming disulfide bonds with cysteine, which disulfide bonds can be cleaved using reducing agents, such as TCEP), and other connection chemistries will produce permanent connections.

[0265] In any embodiment, as will be apparent to those skilled in the art, the polymerase can be mutated to ensure specific attachment of the tethered nucleotide to a particular position of the polymerase. For example, using sulfhydryl-specific attachment chemistries such as maleimide or orthopyridyl disulfide, accessible cysteine residues in the wild-type polymerase can be mutated to non-cysteine residues to prevent labeling at those positions. In this "unreactive cysteine" context, cysteine residues can be introduced at the desired attachment position by mutation. These mutations preferably do not interfere with the activity of the polymerase.

[0266] In some embodiments, the linker is specifically connected to the amino acid of polymerase.In these cases, it is preferred that the linker is connected to the position where amino acid can be mutated without losing polymerase activity, such as position 180,188,253 or 302 of mouse TdT (as numbered in crystal structure PDBID:4127). It is preferred that the linker is not connected to the amino acid participating in polymerase catalytic activity, to avoid interfering with catalysis. It is known that the residues participating in catalysis and the method for determining whether the residue participates in catalysis (such as by site-specific mutagenesis) are clear to those skilled in the art, and are reviewed in the literature (such as Joyce et al. (Journal of Bacteriology 177.22 (1995): 6321.) and Jara and Martinez (The Journal of Physical Chemistry B120.27 (2016): 6504-6514.))

[0267] Other strategies for site-specific attachment of synthetic groups to proteins will be clear to the skilled person and reviewed in the literature (eg Stephanopoulos & Francis Nat. Chem. Biol. 7, (2011) 876-884).

[0268] Equivalents and scope

[0269] Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the disclosure described herein.The scope of the present disclosure is not intended to be limited to the above description, but is set forth in the appended claims.

[0270] In the claims, unless indicated to the contrary or otherwise obvious from the context, articles such as "a," "an," and "the" may mean one or more than one. A claim or description including "or" between one or more members of a group is deemed satisfied if one, more than one, or all of the group members are present in, used in, or otherwise associated with a given product or process, unless indicated to the contrary or otherwise obvious from the context. The present disclosure includes embodiments in which exactly one group member is present in, employed in, or otherwise associated with a given product or process. The present disclosure includes embodiments in which more than one or all of the group members are present in, employed in, or otherwise associated with a given product or process.

[0271] It should also be noted that the term "comprising" is intended to be open ended and allows but does not require the inclusion of additional elements or steps. When the term "comprising" is used herein, the term "consisting of" is also encompassed and disclosed.

[0272] When ranges are given, the endpoints are included. Further, it should be understood that, unless otherwise indicated or otherwise obvious from the context and understanding of one of ordinary skill in the art, in various embodiments of the present disclosure, values expressed as ranges can be assumed to be any specific value or sub-range within the stated range, to one-tenth of the unit of the lower limit of the stated range, unless the context clearly dictates otherwise.

[0273] All cited sources, such as references, publications, databases, database entries, and techniques cited herein, are incorporated herein by reference, even if not explicitly stated in the citation. If statements in a cited source conflict with statements in this application, the statements in this application control.

[0274] Section and table headings are not intended to be limiting.

[0275] Example

[0276] The following are examples of specific embodiments of the present disclosure. The examples are provided for illustrative purposes only and are not intended to limit the scope of the present disclosure in any way. Efforts have been made to ensure accuracy with respect to the values used (e.g., amounts, temperatures, etc.), but of course some experimental errors and deviations should be allowed for.

[0277] Unless otherwise indicated, the implementation of the present disclosure will adopt conventional protein chemistry, biochemistry, recombinant DNA technology and pharmacological methods that are within the skill of the art. Such techniques are fully described in the literature. See, for example, T. Creighton, Proteins: Structures and Molecular Properties (W. H. Freeman and Company, 1993); A. Lehninger, Biochemistry (Worth Publishers, Inc., current addition); Sambrook, et al., Molecular Cloning: A Laboratory Manual (2nd edition, 1989); Methods In Enzymology (S. Colowick and N. Kaplan, ed., Academic Press, Inc.); Remington's Pharmaceutical Sciences, 18th edition (East on, Pennsylvania: Mack Publishing Company, 1990); Carey and Sundberg Advanced Organic Chemistry 3rd edition (Plenum Press) Volumes A and B (1992).

[0278] General Synthesis Scheme of the Disclosed Conjugates

[0279] In some embodiments, the conjugates of the present disclosure can be prepared as outlined in Scheme 1

[0280] Solution 1

[0281]

[0282] wherein B is a nucleobase; PG is a protecting group; X is -N- or -O- of the nucleobase; and L 1 、L 3 、R 1 、R 1’ 、R 2 、R 3 , n and Pol are defined herein.

[0283] The nucleotide portion of the polymerase-nucleotide conjugates of the present disclosure can be prepared as outlined below:

[0284] Protection of hydroxyl groups on nucleotides

[0285] In some embodiments, the 3' and 5' hydroxyl groups of the nucleotides can be protected with TBS. This protection reaction can be achieved using TBS-Cl (tert-butyldimethylsilyl chloride; TBDMS) in imidazole to form TBS esters at the 3' and 5' positions. For example, the protection of the 3' and 5' hydroxyl groups of the nucleotides can be achieved as outlined below:

[0286]

[0287] The protecting group of the protected hydroxyl group is not particularly limited, and any protecting group described in, for example, GREENE'SPROTECTIVE GROUPS IN ORGANIC SYNTHESIS, 5th edition, JOHN WILLY & SONS (2014) can be mentioned, the entire contents of which are incorporated herein by reference. Specifically, methyl, benzyl, p-methoxybenzyl, tert-butyl, methoxyethyl, ethoxyethyl, cyanoethyl, cyanoethoxymethyl, phenylcarbamoyl, 1,1-dioxothiomorpholine-4-thiocarbamoyl, acetyl, pivaloyl, benzoyl, triethylsilyl, triisopropylsilyl, tert-butyldimethylsilyl, [(triisopropylsilyl)oxy]methyl (Tom) group, 1-(4-chlorophenyl)-4-ethoxypiperidin-4-yl (Cpep) group, etc. can be mentioned. From the viewpoint of economic efficiency and easy availability, the hydroxy protecting group is preferably triethylsilyl, triisopropylsilyl or tert-butyldimethylsilyl, more preferably tert-butyldimethylsilyl. The protection and deprotection of hydroxy groups are well known and can be carried out by, for example, the method described in GREENE'S PROTECTIVE GROUPS IN ORGANIC SYNTHESIS mentioned above.

[0288] L 1 -OH ("scar") added to the nucleotide

[0289] Among them L 1 An exemplary reaction scheme identified as "scar" is shown below:

[0290]

[0291] Guanine

[0292]

[0293] BOP and DBU are added to the solution of the TBS-protected nucleoside derivative in acetonitrile, and the mixture is stirred at room temperature. The mixture is then diluted with EtOAc and washed with deionized water, followed by brine washing. The organic layer is dried with anhydrous Na2SO4, filtered, and evaporated under reduced pressure. The gained residue is dissolved in DCM and then added to hexane. An oily viscous layer is formed at the bottom, and a turbid hexane layer is formed on the top. The hexane layer is decanted on sodium sulfate. The same hexane washing procedure is repeated more than twice with an oily layer. The hexane layer is collected, and evaporated under reduced pressure to obtain the intermediate in a white foam shape. This product is used for the next reaction as it is.

[0294]

[0295] The intermediate was dissolved in anhydrous THF. Cs2CO3 and E-butene-1,4-diol were added and the mixture was stirred at 60°C. The reaction mixture was diluted with EtOAc and washed with deionized water and then with brine. The organic layer was dried over anhydrous Na2SO4, filtered, and evaporated under reduced pressure. The crude material was purified by silica gel column chromatography using a 20-50% EtOAC / hexane gradient to obtain the TBS-protected modified nucleoside containing scar as a white foam.

[0296] The above reaction is specifically for the addition of E-butene-1,4-diol, however, many other L1-OH groups can be added by one of ordinary skill in the art using the guidance of the above reaction.

[0297] Thymine

[0298] The reaction scheme for adding L1-OH to O4 of thymine and aligning with the O6 conjugate of guanine is provided below. Examples of alternative L1-OH groups are also shown.

[0299]

[0300] Uracil

[0301] Similar to thymine, the reaction scheme for adding L1-OH to O4 of uracil can be performed as follows:

[0302]

[0303] Adenine

[0304] The following is a reaction scheme for adding the L1-OH group to the N6 group of adenine. Note that this scheme also includes amino acid esters that have been bonded to the L1 group.

[0305]

[0306] Cytosine

[0307] The following provides a method for converting L 1 Reaction scheme for the addition of an -OH group to the N4 group of cytosine. Note that this scheme also includes amino acid esters already bound to the L1 group.

[0308]

[0309] These nucleotides are also commercially available as deoxyribonucleoside triphosphates.

[0310] Coupling amino acid esters

[0311] An exemplary reaction scheme for coupling an amino acid ester to the hydroxyl group of the L1-OH group bound to a nucleotide is shown and described below. Although an amino acid ester having a cyclopropyl R group on the α-carbon is shown, any natural or synthetic amino acid ester having a different R group can be used (e.g., as shown in Example 6). Similarly, this reaction can be used to direct the addition of an amino acid ester to the L1-OH group attached to any nucleotide.

[0312]

[0313] Under an inert atmosphere and ambient temperature, DMAP (0.6 equivalent) and Fmoc amino acid (1.2 equivalent) are added to a solution of the nucleotide L1-OH compound shown above in dry DCM (0.2M). The reaction mixture is cooled to 0°C and EDC·HCl is then slowly added. The reaction mixture is stirred at ambient temperature for 16 hours. The reaction solution is then diluted with additional DCM and extracted with water. The aqueous layer is then washed with more DCM (2x). The combined organic layers are then washed with saturated ammonium bicarbonate, dried with Na2SO4, and concentrated. The crude material is purified by silica gel column chromatography using a 10-35% EtOAC / hexane gradient.

[0314] OH deprotection

[0315] In some embodiments, TBS-protected hydroxyl groups can be deprotected using the following protocol.

[0316]

[0317] 3HF-TEA (10 equivalents) was added dropwise to a solution of 0.1M protected nucleosides in anhydrous THF at 0°C. The mixture was stirred for 16 hours while warming to ambient temperature. The reaction was quenched by adding a few drops of MeOH and the solvent was removed under reduced pressure. The residue was diluted with DCM and washed with water. The aqueous layer was extracted with DCM (1x) again. The combined organic layers were dried over anhydrous Na2SO4, filtered and concentrated. Compounds were purified by silica gel column chromatography using a 1-7% MeOH / DCM gradient. All compounds were obtained as white foamy solids.

[0318] Triphosphorylation and Fmoc deprotection

[0319] The reaction scheme for triphosphorylating nucleotide analogs and deprotecting amino acid esters is shown below:

[0320]

[0321] For the triphosphorylation reaction, the nucleoside analog is placed in a 10mL round-bottom flask equipped with a stir bar, and tetrabutylammonium pyrophosphate is placed in a separate 5mL conical tube. The two flasks are placed in a vacuum desiccator with PO and dried under vacuum for at least 16 hours. In addition, the molecular sieves and three small round-bottom flasks are placed in a drying oven for at least 16 hours. Molecular sieves are added to the two small flasks in the oven and flame activated under vacuum. When they cool, another small flask is connected to a Hickman distillation apparatus and flame dried. After cooling, the first two flasks are backfilled with nitrogen. Trimethyl phosphate and tributylamine are then placed over the molecular sieves in the first two flasks for drying. POCl is then freshly distilled using a Hickman distillation apparatus. The vacuum desiccator is purged with N gas, and the interior of the flask is then transferred to a nitrogen balloon or Schlenk line. Trimethyl phosphate (40 equivalents) is added to the nucleoside, and the mixture is cooled to -5°C. To this mixture of nucleosides was added anhydrous tributylamine (3 equiv.) followed by the slow addition of POCl3 (2.1, 1.3, 1.5, 1.8 equiv., respectively) via a microsyringe. The combined mixture was stirred at -5°C for 45 minutes. After 45 minutes, the reaction mixture was treated with a mixture of tributylammonium pyrophosphate (5 equiv., 0.5 M in anhydrous acetonitrile) and tributylamine (6 equiv.). After 1 hour, the mixture was treated with triethylammonium bicarbonate (0.5 M, 1:2 of the total reaction volume) and stirred at ambient temperature for 1 hour. The reaction mixture was then further treated with N-methylpiperidine (1 / 5 of the total reaction volume) and stirred at ambient temperature for 90 minutes before extraction with dichloromethane (2X). The product was then purified by reverse phase HPLC (0.1 M triethylammonium acetate buffer / acetonitrile, 4–47%, 0–15 minutes, flow rate 5 ml min -1) Purify the aqueous layer. Product-containing fractions were combined and lyophilized to provide the desired product as a triethylammonium salt. The resulting solid was reconstituted in RNase-DI water and used for further experiments.

[0322] Adding amino acids / L3 to amino acid esters

[0323] The following exemplary reaction scheme can be used to add the L3 portion of the linker that binds to the polymerase, along with any amino acids adjacent to the amino acid ester and a portion of the L2 portion of the linker:

[0324]

[0325] OPSS-Gly-NHS can be synthesized according to the following reaction scheme.

[0326]

[0327] This reaction scheme can also be used to add more amino acids or to include alternative amino acids in L2, such as alanine:

[0328]

[0329] In order to add one or more amino acids to the joint, peptide synthesis can also be performed using standard solid phase or liquid phase chemistry as needed. Methods for peptide synthesis are well known to those skilled in the art (Fodor et al., Science 251:767 (1991); Gallop et al., J.Med.Chem.37:1233-1251 (1994); Gordon et al., J.Med.Chem.37:1385-1401 (1994)). It should be understood that the peptide joint can be synthesized and then added to the NTP as a peptide, or can be synthesized by sequentially adding amino acids.

[0330] Scheme 2: Synthesis of OPSS-Gly-ACC-EtS-dGTP

[0331] A complete exemplary reaction scheme for the synthesis of a nucleotide conjugated to a linker (L1-L2-L3) comprising an amino acid ester as part of the L2 moiety is shown below:

[0332]

[0333] Example 1: Preparation of polymerase-nucleotide conjugates

[0334] Synthesis of linker-nucleotides with different R groups on amino acid esters

[0335] Modified nucleotides having an amino acid ester moiety of L1 and L2 with various substitutions on the alpha carbon of the amino acid were synthesized according to the following reaction scheme:

[0336]

[0337] Conjugation of L1 to O6 of guanine

[0338] To a solution of nucleoside derivative 1 (4g, 8.08mmol) in acetonitrile (16mL), BOP (7.13g, 16.14mmol) and DBU (2.408mL, 16.14mmol) were added, and the mixture was stirred at room temperature for 1.5 hours. The mixture was then diluted with EtOAc (100mL), washed with deionized water (3 × 50mL), and then washed with brine (30mL). The organic layer was dried with anhydrous Na2SO4, filtered, and evaporated under reduced pressure. The obtained residue was dissolved in 3 to 4mL DCM, and then added to 50mL hexanes. Two layers were formed, with an oily viscous layer at the bottom and a turbid hexane layer at the top. The hexane layer was decanted on sodium sulfate. The same hexane washing procedure was repeated more than twice with oily residue. The hexane layer was collected and evaporated under reduced pressure to obtain compound 2 (3.63g, 73%) in white foam. This product was used for the next reaction as it is.

[0339] Compound 2 (3.62 g, 5.89 mmol) was dissolved in anhydrous THF (30 mL), CsCO (3.84 g, 11.79 mmol) and E-butene-1,4-diol (1.30 g, 14.72 mmol) were added, and the mixture was stirred at 60 ° C for 1.5 hours. The reaction mixture was then diluted with EtOAc (100 mL), washed with deionized water (2 × 50 mL) and brine (25 mL). The organic layer was dried over anhydrous NaSO, filtered, and evaporated under reduced pressure. The crude material was purified by silica gel column chromatography using a 20-50% EtOAC / hexane gradient to obtain 3 (1.884 g, 56%) as a white foam.

[0340] Add amino acid ester

[0341] At ambient temperature and under inert atmosphere, DMAP (0.6 equivalent) and Fmoc amino acid (1.2 equivalent) (corresponding to dimethyl, cyclopropyl, cyclobutyl, cyclopentyl or cyclohexyl R groups) are added to a solution of compound 3 (1 equivalent) in anhydrous DCM (0.2M). The reaction mixture is cooled to 0 ° C and EDC HCl is then slowly added. The reaction mixture is stirred for 16 hours at ambient temperature. The reaction solution is then diluted with additional DCM and extracted with water. The water layer is then washed with more DCM (2x). The combined organic layer is then washed with saturated ammonium bicarbonate, dried with Na2SO4, and concentrated. The crude material is purified by silica gel column chromatography using a 10-35% EtOAC / hexane gradient.

[0342] TBS deprotection

[0343] At 0 ° C, 3HF-TEA (10 equivalents) was added dropwise to a 0.1M solution of protected nucleoside (4-8) in anhydrous THF. The mixture was stirred for 16 hours while warming to ambient temperature. The reaction was quenched by adding a few drops of MeOH and the solvent was removed under reduced pressure. The residue was diluted with DCM and washed with water. The aqueous layer was extracted with DCM (1x) again. The combined organic layers were dried over anhydrous Na2SO4, filtered and concentrated. The compounds were purified by silica gel column chromatography using a 1-7% MeOH / DCM gradient. All compounds were obtained as white foamy solids.

[0344] Synthesis of triphosphates and Fmoc deprotection

[0345] Nucleoside analogs (9 / 11 / 12 / 13, 1 equivalent) were placed in a 10 mL round-bottom flask equipped with a stirring bar, and tetrabutylammonium pyrophosphate was placed in a separate 5 mL conical tube. The two flasks were placed in a vacuum desiccator with PO and dried under vacuum for at least 16 hours. In addition, molecular sieves and three small round-bottom flasks were placed in a drying oven for at least 16 hours. Molecular sieves were loaded into the two small flasks in the oven and flame activated under vacuum. When they cooled, another small flask was connected to a Hickman distillation apparatus and flame dried. After cooling, the first two flasks were backfilled with nitrogen. Trimethyl phosphate and tributylamine were then placed on the molecular sieves in the first two flasks for drying. POCl was then freshly distilled using a Hickman distillation apparatus. The vacuum desiccator was purged with N gas, and the interior of the flask was then transferred to a nitrogen balloon or Schlenk line. Trimethyl phosphate (40 equivalents) was added to the nucleoside, and the mixture was cooled to -5°C. To this mixture of nucleosides was added anhydrous tributylamine (3 equiv.) followed by the slow addition of POCl3 (2.1, 1.3, 1.5, 1.8 equiv., respectively) via a microsyringe. The combined mixture was stirred at -5°C for 45 minutes. After 45 minutes, the reaction mixture was treated with a mixture of tributylammonium pyrophosphate (5 equiv., 0.5 M in anhydrous acetonitrile) and tributylamine (6 equiv.). After 1 hour, the mixture was treated with triethylammonium bicarbonate (0.5 M, 1:2 of the total reaction volume) and stirred at ambient temperature for 1 hour. The reaction mixture was then further treated with N-methylpiperidine (1 / 5 of the total reaction volume) and stirred at ambient temperature for 90 minutes before extraction with dichloromethane (2X). The product was then purified by reverse phase HPLC (0.1 M triethylammonium acetate buffer / acetonitrile, 4–47%, 0–15 minutes, flow rate 5 ml min -1 ) Purify the aqueous layer. Product-containing fractions were combined and lyophilized to provide the desired product as a triethylammonium salt. The resulting solid was reconstituted in RNase-DI water and used for further experiments.

[0346] Addition of the glycine-L3 portion of the linker

[0347] The glycine amino acid portion of L2 and the L3 portion of the linker that connects L2 to the polymerase are then added to the compound synthesized above according to the following reactions:

[0348]

[0349] Specifically, a 20 μL reaction was set up with 2 mM nucleotide, 6 mM (3 equiv.) OPSS-Gly-NHS ester, and 0.1 M sodium bicarbonate (50 equiv.), and the amide bond was formed by reaction with OPSS-Gly-NHS ester [2,5-dioxopyrrolidin-1-yl(3-(pyridin-2-yldisulfanyl)propionyl)glycine ester)].

[0350] To compare the reactivity of the primary amines (14-18) in these nucleotides, the amide bond formation rate was determined by injecting 10 μL of the reaction volume onto an HPLC RP C-18 column at 10 and 50 min (0.1 M triethylammonium acetate buffer / acetonitrile, 4–47%, 0–20 min, flow rate 1 ml min). -1 ). The percentage of conversion to product at each time point is listed in Table 1. As shown, for each amino acid ester R group, complete linker-nucleotide synthesis was successful. As described in Example 1, conjugates with polymerase were formed.

[0351] Table 1. Conversion of amines to amides at reaction times of 10 and 50 minutes.

[0352]

[0353] Generation of polymerase (TdT) mutants with different adaptor ligation positions

[0354] An inducible plasmid expressing murine TdT with a single cysteine at position 182 was generated (for a complete protocol, see Palluk et al., Nature Biotechnology, 2018). For more details on the preparation of polymerase-nucleotide conjugates, see also U.S. Patent Publication No. 2019 / 0112627, "Nucleic Acid Synthesis and Sequencing Using Tethered Nucleoside Triphosphates," which is incorporated herein by reference in its entirety.

[0355] Protein expression and purification of mutants

[0356] BL21 (DE3) Gold cells (Agilent) are used to carry out TdT expression in TB culture medium, and the TB culture medium contains the antibiotic for the resistance marker of plasmid. 50mL of overnight culture is used to inoculate 400mL expression culture with 1 / 20vol. Cells are grown under 37 ℃ and 200rpm shaking until they reach OD 0.6. IPTG is added to a final concentration of 0.5mM, and expression is continued for 16-20 hours at 16 ℃. By harvesting cells by centrifugation for 10 minutes at 8000G and resuspending them in 20mL buffer A (20mM Tris-HCl, 0.5M NaCl, pH 8)+5mM imidazoles. Cell lysis is carried out using ultrasonic treatment, and then centrifuged at 30,000G for 20 minutes. Supernatant is applied to a gravity column containing 1mL Ni-NTA agarose (Qiagen). The column was washed with 20 volumes of buffer A + 40 mM imidazole and bound protein was eluted using 4 mL of buffer A + 500 mM imidazole. The protein was concentrated to approximately 0.15 mL using a Vivaspin 20 column (MWCO 10 kDa, Sartorius) and then purified using a Pur-A-Lyzer TM Dialysis kit Mini 12000 tubes (Sigma) were dialyzed overnight against 200 mL of TdT storage buffer (100 mM NaCl, 200 mM K2HPO4, pH 6.5).

[0357] Ni-purified samples were applied to a HiTrap Q HP anion column. Proteins were eluted with a linear gradient from 100% Q buffer A (100 mM NaCl, 20 mM K2HPO4, pH 6.5) to 100% Q buffer B (1 M NaCl, 20 mM K2HPO4, pH 6.5). Fractions containing TdT were identified using SDS-PAGE analysis and pooled and concentrated.

[0358] Linkage of tethered nucleoside triphosphates to polymerase

[0359] To prepare the TdT-nucleotide conjugate, a cleavable linker-nucleotide was first synthesized that had a moiety capable of site-specific conjugation to cysteine (i.e., maleimide). Equimolar amounts of TdT and linker-nucleotide were then incubated overnight at 4° C. in 500 mM NaCl, 20 mM K HPO (pH 6.5). The TdT conjugate was separated from unreacted linker-nucleotide using an S200 size exclusion column (Cytiva) pre-equilibrated in 20 mM Tris acetate, 50 mM potassium acetate (pH 7.9).

[0360] Example 2: Amino Acid Ester Cleavable Linker

[0361] Several TdT-dNTP conjugates have been prepared and tested, using different cleavable linkers to specifically tether the nucleotide to the polymerase. The use of a peptide bond in the linker results in a linker that can be cleaved by a protease. The cleavage of the peptide bond by the protease generates an amine and a carboxylic acid, both of which are charged under the typical buffer conditions for TdT activity. However, the continued presence of a charged functional group on the synthesized oligonucleotide can have a detrimental effect during the synthesis process.

[0362] As an alternative, the use of amino acid esters in the linker to create a cleavable linkage between the nucleotide and TdT was explored. Cleavage of the ester yields an alcohol as a charge-neutral cleavage product.

[0363] Conjugates containing ester linkers

[0364] Initially, two classes of ester-containing linkers were designed that could be cleaved by esterases to leave an alcohol-containing scar on the nucleobase. One class was based on a hydroxypropargyl scar (Linker 1) and contained a cleavable leucine amino acid ester, and the other class was based on a smaller hydroxymethyl scar (Linker 2) and contained a cleavable glycine amino acid ester. Two amino acid ester dTTP analogs (Linkers 1 and 2, Figure 1 ; synthesized by Jena Bioscience) was linked to a cysteine-reactive cross-linker and conjugated to TdT.

[0365] First, the effect of incorporating nucleotides with alcohol scars left by linker cleavage on processivity during polynucleotide synthesis was tested.

[0366] Tolerance of hydroxymethyl scars during synthesis

[0367] To compare the extension kinetics of the addition of the conjugate to the 3' end of an unscarred nucleotide compared to a hydroxymethyl-scarred nucleotide, two cycles of synthesis were performed, first adding the dTTP conjugate to a native DNA primer, then cleaving the polymerase from the dTTP, leaving the hydroxymethyl scar, and then adding the dTTP conjugate to a DNA primer containing a hydroxymethyl-scarred nucleotide at the 3' end.

[0368] like Figure 2 -A, exposure of a native DNA primer to a dTTP conjugate for 1 second resulted in approximately 35% extension yield. The reaction was allowed to proceed to completion ( Figure 2 -B), followed by cleavage of the linker, generating a DNA primer containing a hydroxymethyl scarred nucleotide at the 3' end. Exposure of the scarred primer to a dTTP conjugate for 1 second again resulted in an extension yield of approximately 35%, similar to that observed with native DNA ( Figure 2 -C).

[0369] Because the conjugate-induced extension rate of hydroxymethyl-scarred DNA is comparable to that of native DNA, ester-containing linkers that leave an alcohol scar on the nucleotide are acceptable for DNA synthesis. The alcohol generated by ester cleavage is a charge-neutral cleavage product, allowing for undisturbed nucleotide addition and providing an improvement over protease cleavage, which leaves a charged nucleotide scar that can adversely affect oligomer synthesis.

[0370] Containing ester connector

[0371] Next, we tested whether TdT-dTTP conjugates with linkers containing peptides that leave a hydroxypropargyl scar (Linker 1 - Leucine amino acid ester) and a hydroxymethyl scar (Linker 2 - Glycine amino acid ester) could successfully incorporate dTTP onto the 3' end of oligonucleotides.

[0372] The ssDNA primer was extended for 60 seconds with 1) Linker 1 conjugate, 2) Linker 2 conjugate, 3) Linker 2 conjugate (repeat), and 4) no conjugate. Upon incubation with the ssDNA primer, the TdT-dTTP conjugate containing either linker formed a covalent primer extension complex with >95% yield in less than one minute, as measured by gel shift assay on SDS-PAGE ( Figure 3 ). T / P: TdT / DNA complex. P: ssDNA primer (bound). The band above the complex is a product with more than one base added to the primer.

[0373] As shown, the conjugate containing amino acid ester cleavable linkers 1 and 2 was successfully added to the 3' end of the oligonucleotide and was suitable for oligonucleotide synthesis.

[0374] Esterase screening and enzymatic cleavage

[0375] Several commercially available enzymes were screened to identify an esterase to cleave the conjugates prepared with linkers 1 and 2. The serine protease proteinase K (Barthel et al., Enhancing terminal deoxynucleotidyl transferase activity on substrates with 3'terminal structures for enzymatic De Novo DNA synthesis. Genes (Base I). 2020; 11: 1-9.), which is known to have esterase activity, cleanly cleaves the two linkers into alcohol products. The oxypropargyl ester linker 1 was completely cleaved within 2 minutes. The shorter oxymethyl ester linker 2 required 7.5 minutes at 40 ° C for complete cleavage. Therefore, amino acid ester linkers in polymerase-nucleotide conjugates, which have been shown to have good oligonucleotide incorporation kinetics, can also be successfully cleaved using proteases containing esterase activity.

[0376] Synthesis of linker 2dT(10)

[0377] As shown above, it was observed that the minimal hydroxymethyl scar left by Linker 2 was well tolerated during oligonucleotide synthesis. To explore this further, a 10-mer poly-T (dT(10)) was synthesized using a dTTP conjugate with Linker 2 at the end of the starting oligomer.

[0378] The synthesis of dT (10) using the dTTP conjugate with linker 2 had excellent coupling yields, with deletions well below the detection limit of 1% / step. However, the insertion rate was approximately 2.5% / step, which is undesirably high. This high insertion rate is likely due to spontaneous cleavage of the ester. Because some esters are known to be unstable, the stability of linkers containing amino acid esters has been previously tested, and both linkers 1 and 2 were found to decompose upon heating. When the ester is hydrolyzed, a free nucleotide is released and can be incorporated by the TdT portion of the primer-TdT complex or by initial extension of the free nucleotide by a free polymerase, resulting in the addition of the second nucleotide during the coupling step. Therefore, the high insertion rate is likely due to spontaneous cleavage of the ester.

[0379] High insertion rate: 2.5% at pH 6.5; 0.5% at pH 7.9.

[0380] Joint instability determination

[0381] Since cleavage may be due to base-catalyzed hydrolysis, the stability of the amino acid ester linker conjugates was tested over a pH range of 8.5 to 6.5.

[0382] Specifically, dTTP conjugates containing linker 2 were stored overnight at pH 6.5, pH 7.0, pH 7.5, pH 8.0, pH 8.5, or stored overnight in buffer alone as a control. Primer extension was performed and the incubated dTTP conjugate was incorporated into the primer. The resulting extension products were then run on a capillary electrophoresis pattern. The results are shown in FIG. Figure 4 , where the insertion is represented by a second peak at approximately 59. The insertion indicates the presence of free dNTPs in the conjugate solution. When stored at pH>6.5, the linker 2 conjugate releases dNTPs. However, at pH 6.5, a significant decrease in nucleotide release from the conjugate was observed.

[0383] Synthesis of dT(10), dT(100), and dT(200) at pH 6.5

[0384] Next, a new dT(10) synthesis was performed at pH 6.5 using the Linker 2 conjugate. The insertion rate of this synthesis dropped to less than 1% / step, which is below the detection limit of the assay.

[0385] To more accurately measure low error rates, the synthesis length was increased to dT 100- and 200-mers to improve the detection limit of synthesis errors. dT 100- and 200-mers were synthesized onto the 3' end of a 60 pb primer using a linker 2dTTP conjugate at pH 6.5. The resulting dT 100- and 200-mers included a hydroxymethyl scar. For both syntheses, the synthesis proceeded with a low deletion rate of <0.1% / step and an insertion rate of approximately 0.5% / step ( Figure 5 , part A).

[0386] For comparison with chemical synthesis, two dT 100mers were ordered from IDT (synthesized without capping to enable comparison of deletions) and the chemically synthesized oligomers were assayed via capillary electrophoresis. A magnified comparison between the enzymatically synthesized dT 100 oligomer above and the 100mer dT produced via chemical synthesis as measured by capillary electrophoresis is shown in Figure 5 In Part B, the most abundant species in both syntheses was +100 nt. The main error types were insertions (+101 nt) in the enzymatic product and deletions (+99 nt) in the chemical product. Enzymatic extension showed a lower deletion rate than chemical synthesis.

[0387] These results demonstrate that enzymatic synthesis using the conjugate exhibits exceptionally low deletion rates. Chemical oligonucleotide synthesis has never achieved deletion rates below 0.1% per step, so by this criterion, the enzymatic synthesis method is already superior. Furthermore, the stepwise yields for 100-mers and 200-mers were essentially identical, indicating that yield did not decline as the synthesis proceeded. In contrast, chemical oligomer syntheses typically exhibit increasing deletions with longer synthesis times. These results confirm that a fully enzymatic oligonucleotide system will be able to synthesize longer oligomers than is possible using chemical synthesis.

[0388] Example 3: A, C, T and G conjugates with protease ester cleavable linkers

[0389] Next, a panel of A, C, T, and G conjugates were tested, each with a protease ester cleavable linker to allow for the synthesis of oligonucleotides with a full set of nucleotides. Figure 1 The amino acid ester linker (Linker 6) is shown.

[0390] The dNTPs were synthesized by coupling an easily synthesized ester-containing linker -COOH to commercially available aminopropargyl dNTPs (Linker 6; Figure 1 ) to assemble linker 6 into linker-nucleotides. An ester linker-COOH was purchased from WuXi AppTec (using non-SBIR funds) and contracted with MyChem LLC (San Diego) to couple it to aminopropargyl dNTPs, thereby generating a complete set of linker 6 nucleotides (A, C, G, T).

[0391] Then, conjugates were prepared from each linker nucleotide and their extension and cleavage kinetics were tested. Specifically, the oligonucleotide primer was exposed to TdT-dATP, TdT-dCTP, TdT-dGTP, and TdT-dTTP conjugates containing linker 6 for 1 minute, and then the linker was cleaved with proteinase K for 4 minutes. The cleaved extension products were then measured by capillary electrophoresis. Figure 6 As shown in Figure A, all four linker 6-dNTP conjugates showed >99% coupling yield in less than 1 minute.

[0392] Next, a cleavage time course was performed on a TdT-dTTP conjugate containing linker 6 and incorporated into a primer oligonucleotide. Specifically, the conjugate was coupled for 1 second, and the complex was exposed to proteinase K for 30 seconds, 60 seconds, 120 seconds, or 240 seconds. The extension products were then determined by capillary electrophoresis. Figure 6 As shown in Figure B, the amino acid ester linker (Linker 6) was cleaved >99% by Proteinase K in less than 4 minutes.

[0393] Thus, successful oligonucleotide synthesis has been demonstrated using all four nucleotide conjugates containing amino acid ester linkers, and efficient cleavage of the linker with a protease containing esterase activity (proteinase K) to separate the polymerase from the incorporated nucleotides.

[0394] However, a peak indicating spontaneous nucleotide insertion during the coupling reaction was still observed, which can be attributed to the instability of the ester linker as demonstrated above.

[0395] dT(10) synthesis

[0396] To further test the performance of these conjugates, dT(10) oligomers were first synthesized and the products analyzed by capillary electrophoresis. Coupling yields were excellent, with deletions <1% / step, below the limit of detection. As expected based on the insertions observed above, conjugates containing the linker 6 ester showed high levels of insertion, approximately 3.5% / step. Therefore, it is desirable to improve ester stability to mitigate spontaneous cleavage of the linker. In addition, improving ester stability may allow for extended coupling times to further reduce deletion rates.

[0397] Example 4: ACC ester improves linker stability

[0398] As shown above, amino acid esters can be successfully cleaved by proteases containing esterase activity (proteinase K) to promote polymerase cleavage from nucleotides after incorporation, thereby leaving a neutral alcohol scar on the nucleotide that does not hinder oligonucleotide synthesis. However, the amino acid esters in linkers 2 and 6 are unstable and spontaneously cleave, resulting in unwanted insertions during oligonucleotide synthesis. Linker 2 amino acid ester is a glycine analog that is not substituted on the α carbon. Here, whether adding aliphatic or bulky substituents to the α carbon of the amino acid ester will improve the stability of the ester was tested. Exemplary side group substitutions that improve stability are shown in Figure G.

[0399] The stability of conjugates with different cleavable linkers (glycine ester vs. ACC ester) was determined using the following protocol:

[0400] Prepare master mix (MM) with 20mM tris acetate pH 7.9, 50mM potassium acetate, 100 μM cobalt acetate (II) and 100nM DNA oligomer substrate. In order to trigger the addition reaction, MM is mixed with corresponding TdT-dNTP conjugate solution (2 μM solution in 20mM tris acetate pH 7.9, 50mM potassium acetate and 0.1% Tween-20) in 1:1. Then the mixture is incubated at room temperature for 5 minutes, then quenched with EDTA (to the final EDTA concentration of 32mM). Now, the sample is separated for incubation at different temperatures. After incubation for 4 hours, the sample is diluted 10 times in the HiDi containing 20mM DTT. The samples of these dilutions are then analyzed by capillary electrophoresis.

[0401] The stability of the glycine ester bond and the ACC ester bond was examined after exposure to 45°C for 60 minutes ( Figure 7 ) after addition of 1% cyclopropyl, it is clear that the hydrolysis of the ACC ester is not as significant as that of the corresponding glycine ester. It is suspected that this stabilization is the result of the hyperconjugation effect of the cyclopropyl group.

[0402] Example 5: Peptides adjacent to ACC esters improve enzymatic cleavage

[0403] Next, the rate of ProK-mediated linker cleavage towards the ACC ester bond was determined.

[0404] First, Figure 8A The conjugates shown with an ACC ester bond were incorporated into the following oligomer substrates: a master mix (MM) was prepared with 20 mM tris acetate pH 7.9, 50 mM potassium acetate, 100 μM cobalt (II) acetate, and 100 nM DNA oligomer substrate. To initiate the addition reaction, the MM was mixed 1:1 with the corresponding TdT-dNTP conjugate solution (2 μM solution in 20 mM tris acetate pH 7.9, 50 mM potassium acetate, and 0.1% Tween-20). The reaction was allowed to incubate for 5 minutes before being quenched by adding EDTA (40 mM final concentration).

[0405] Next, the ACC ester bond was cleaved as follows: The quenched spiked reaction was then mixed with a solution of ProK (40 U / mL final concentration) and EDTA (40 mM final concentration) in a 1: 1 ratio. Aliquots of the ProK reaction were removed and quenched at different time points by dilution in HiDi with Pefabloc and subsequently analyzed by capillary electrophoresis.

[0406] Figure 8A The cleavage products observed by capillary electrophoresis after 60 seconds of ProK treatment are shown. As shown, ProK treatment of the ACC ester linker in the OPSS-ACC-OEt-dATP conjugate produces very few complete cleavage products, indicating that the ACC ester linker structure is a poor substrate for ProK. Because linker cleavage is a key step in conjugate-based oligonucleotide synthesis, various modifications were explored to enhance the cleavage activity of the ACC ester by ProK.

[0407] Among the various possible changes in the linker structure, the effect of introducing one or more amino acids in the linker near the amine of the ACC ester was investigated. Figure 8B and Figure 8CAs shown, linkers containing one or two glycine residues bound to the amine of the ACC ester (OPSS-Gly-ACC-OEt-dATP and OPSS-2×Gly-ACC-OEt-dATP, respectively) were added to dATP. After preparing polymerase-nucleotide conjugates using these linkers, the conjugates were incorporated into oligomeric substrates, and the ACC ester bond of the ACC ester conjugates without glycine residues was cleaved as described above. The resulting products were then analyzed by capillary electrophoresis. Figure 8B and Figure 8C The cleavage products observed by capillary electrophoresis after 60 seconds of ProK treatment are shown. As shown, the ACC ester linker with a single glycine residue was completely cleaved by ProK after 60 seconds ( Figure 8B ), while the ACC ester linker with two adjacent glycine residues showed almost complete cleavage after 60 seconds ( Figure 8C Thus, incorporation of one or more amino acid residues near a stable protease ester significantly improves the kinetics of ProK cleavage that removes the polymerase from incorporated nucleotides during oligonucleotide synthesis.

[0408] In summary, it was shown that the addition of one or more amino acids to the L2 structure adjacent to the amino acid ester improves the kinetics of enzymatic cleavage of the linker by proteases comprising esterase activity.

[0409] Additional kinetics of ProK cleavage of the ACC ester-glycine linker

[0410] The kinetics of ProK-mediated linker cleavage of 1×Gly and 2×Gly ACC ester linkers were further explored.

[0411] Short ssDNA oligomers labeled with FAM are extended by 1 nucleotide, and dATP is conjugated to terminal deoxynucleotidyl transferase (TdT) via a linker with an aminocyclopropylcarboxyethyl group and one (1XG) or two (2XG) glycines. After the extension reaction, the extended DNA is incubated with ProK, or not incubated with ProK as a negative control, and quenched with pefabloc after 15 seconds (s), 30s, 60s, 4 minutes (m), 8m or 16m. DTT is added to the assay solution to remove any protein that is not cleaved from the linker. Capillary electrophoresis (CE) is used to analyze the cleaved and uncleaved DNA fragments. The fragment displacement in the electropherogram is observed to determine the size of the fragments and thus the extent of linker cleavage. Fragments that shift to the left in the electropherogram are smaller fragments and therefore indicate the presence of a cleaved linker. The cleaved and uncleaved linker peaks are marked on Figure 9 middle.

[0412] Figure 9The electropherograms shown in Figure 3 show that cleavage of the 1XG linker is complete within 60 seconds, while cleavage of the 2XG linker is complete within 4 minutes. This data indicates that ProK is capable of cleaving linkers having aminocyclopropylcarboxyethyl (ACC amino acid ester) and one (1XG) or two (2XG) glycines as substrates. Furthermore, the data show that the rate at which ProK cleaves a linker can be tuned by varying the number of amino acids adjacent to the amino acid ester within the linker.

[0413] In summary, a single glycine residue adjacent to an amino acid ester had improved cleavage kinetics compared to two glycine residues adjacent to a stable amino acid ester in the L2 group. However, both linkers had significantly improved ProK cleavage kinetics compared to linkers without an amino acid adjacent to a stable amino acid ester in the linker.

[0414] Nucleotide addition / oligomer synthesis

[0415] Next, the effect of each of the three L2 groups of the linker determined above (ACC, Gly-ACC, and 2xGly-ACC) on the nucleotide incorporation kinetics of the corresponding TdT-nucleotide conjugates was investigated.

[0416] A master mix (MM) was prepared using 20 mM tris acetate pH 7.9, 50 mM potassium acetate, 100 μM cobalt (II) acetate, and 100 nM DNA oligomer substrate with a CCC 3' end. To initiate the addition reaction, the MM was mixed 1:1 with the corresponding TdT-dNTP conjugate solution (2 μM solution in 20 mM tris acetate pH 7.9, 50 mM potassium acetate, and 0.1% Tween-20). The reaction was terminated at different time points by adding EDTA and ProK. The resulting mixture was then diluted in HiDi and fragment analysis was performed by capillary electrophoresis.

[0417] Figure 10 Results are shown for each TdT-nucleotide conjugate, adding the conjugate to the primer 3.8 seconds after conjugate addition. The data demonstrate that progressively removing individual glycines, from 2 to 1 to 0 glycines in L2, gradually increases the kinetics of nucleotide incorporation. While the conjugate incorporation reactions proceeded relatively quickly for all three conjugates considering the 3.8 second time point, a single amino acid in the linker adjacent to the amino acid ester may be preferred to optimize linker cleavage and conjugate incorporation speed.

[0418] In summary, the L2 component of the cleavable linker has been engineered for rapid protease activity and nucleotide addition. It has been found that additional amino acids attached to the amine of ACC promote these rapid reactions, with a single additional glycine exhibiting the fastest cleavage kinetics. The additional peptide bond does not result in a free amine that could be cleaved by non-target proteases. Furthermore, the linker nucleotides we developed exhibit very high conjugation efficiency.

[0419] These improvements resulted in TdT-dNTP conjugates that support the synthesis of long, high-quality oligonucleotides with very short cycle times via rapid nucleotide addition, rapid deprotection, and benign deprotection conditions.

[0420] Example 6: Ring Expansion / ACC Variants - Ester Stability

[0421] As shown above, the use of cyclopropyl R groups on amino acid esters increases the stability of the ester group in the linker, thereby minimizing unwanted spontaneous cleavage that could lead to insertion during oligomer synthesis. Here, other amino acid ester R groups were explored to determine whether they also impart ester bond stability in the linker prior to cleavage by proK and whether they are also suitable for addition to oligonucleotides via TdT conjugation.

[0422] Tested connector

[0423] To test the stability of the esters imparted by ring expansions and other substitutions at the R groups of the amino acid esters in the L2 portion of the linker, modified nucleotides containing L1 and L2 groups with different R groups on the amino acid esters were synthesized and subjected to high pH conditions. Specifically, as described in Example 1, the following nucleotide linker compounds were synthesized with dimethyl, cyclopropyl, cyclobutyl, cyclopentyl, and cyclohexyl R groups:

[0424]

[0425] Ester stability of compounds 14-18

[0426] To compare the stability of the ester bond in five analogs (compounds 14-18), the hydrolysis rate of these analogs was studied in TP8 buffer at 50 ° C over a 20-hour time course. A 3 mM solution of the nucleotide analogs (14-18) in 1x TP8 (24 μL, pH 8) buffer was incubated at 50 ° C. 4 μL aliquots were taken at 1 minute, 10 minutes, 30 minutes, 1 hour, 3 hours and 20 hours, and neutralized with 8 μL KP 6.5 buffer and frozen at -80 ° C. The samples were then thawed and analyzed by analytical RP-HPLC. (0.1 M triethylammonium acetate buffer / acetonitrile, 4–90%, 0–14 minutes, flow rate 1 ml min -1The hydrolysis rate was determined using the percentage of the hydrolysis product peak area relative to the total area of the starting product and hydrolysis product peaks.

[0427] Figure 11 Shown is a graph of the ester hydrolysis of each nucleotide analog (14-18) over time at 50°C.

[0428] Ester stability of the linker after incorporation into the starting DNA

[0429] Next, the ester stability of the expanded series conjugates after incorporation into oligonucleotides was determined by incubating the extended oligonucleotides at 50°C for varying amounts of time.

[0430] The polymerase nucleotide conjugates of each ring expansion series (compounds 14-18) are synthesized as described above. The concentrated TdT conjugate stock solution is diluted to 0.2 mg / mL with 1x TP8+0.1% P20. The diluted conjugate is then mixed with 50 nM starting DNA and a solution of 100 μM Co in 1xTP8 (10 μL+10 μL) at 1:1 to initiate an extension reaction at ambient temperature with the following final concentrations: 25 nM DNA, 50 μM Co, 0.05% P20, 1x TP8. After 4 minutes, 40 μL 20mM EDTA are added to terminate the reaction. The merged mixture is then incubated at 50°C. At 1 hour, 4 hours and 16 hours, 10 μL aliquots are taken out from the reaction mixture and frozen in -80°C refrigerators until they are thawed for analysis simultaneously. 1 μL of each aliquot was diluted with 9 μL assay solution (75% HiDi with DNA ladder and 20 mM DTT) and analyzed by capillary electrophoresis. For comparison, a control sample (allyl G) was used in which the linker of the extended DNA product had been completely removed by the protease.

[0431] The results are Figure 12 The reference dotted line marks the position of the Allyl G control sample.

[0432] in conclusion

[0433] After incubation at 50°C for 16 hours, the linkers on the expanded series (ACC, AiB, AC4C, AC5C, and AC6C) were minimally hydrolyzed and showed only minor differences across the entire series, consistent with the stability of the ACC amino acid ester. Since the conjugates are typically incubated at 37°C for only a few minutes for standard synthesis, the linkers across the expanded series have acceptable hydrolytic stability for single nucleotide extension.

[0434] In addition, based on the teachings herein, optimized R groups can be used to optimize the balance between ester stability and the cleavage rate of proteases containing esterase activity. This can be accomplished using one of the R groups shown in Figure G or similarly substituted amino acid esters to achieve acceptable stability while increasing the kinetics of linker cleavage.

[0435] Other implementation plans

[0436] It is to be understood that the words which have been used are words of description rather than limitation and that changes may be made within the purview of the following claims without departing from the true scope and spirit of the disclosure in its broader aspects.

[0437] While the disclosure has been described at length and with a certain particularity with respect to several described embodiments, it is not intended that the disclosure should be limited to any such details or embodiments or any particular embodiment, but rather that the disclosure should be construed with reference to the appended claims so as to provide the broadest possible interpretation of such claims in light of the prior art and thereby effectively encompass the intended scope of the disclosure.

[0438] All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety. In the event of a conflict, the present specification, including definitions, will control. In addition, section headings, materials, methods, and examples are illustrative only and are not intended to be limiting.

Claims

1. A conjugate comprising a polymerase, a nucleotide, and a cleavable linker attached to the polymerase and the nucleotide, wherein the cleavable linker comprises an amino acid ester.

2. The conjugate of claim 1, wherein the amino acid ester is linked to an amino acid.

3. The conjugate of claim 2, wherein the amine group of the amino acid ester is bound to the amino acid.

4. The conjugate of claim 3, comprising a peptide of at least 2, at least 3, at least 4 or at least 5 amino acids bound to the amine group of the amino acid ester.

5. The conjugate of any one of claims 1 to 4, wherein the one or more amino acids are selected from the group consisting of alanine, arginine, asparagine, aspartic acid, cysteine, glutamic acid, glutamine, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, and valine.

6. The conjugate of any one of claims 1 to 4, wherein the amino acid is or comprises glycine.

7. The conjugate of any one of claims 1 to 4, wherein the amino acid is a non-naturally occurring amino acid, or the amino acid comprises a non-naturally occurring amino acid.

8. The conjugate of any one of the preceding claims, wherein the cleavable linker is bound to the α-phosphate, sugar or nucleobase of the nucleotide.

9. The conjugate of any one of the preceding claims, wherein the amino acid ester is represented by the formula: where R 1 and R 1' are each independently selected from hydrogen and optionally substituted C 1-6 Alkyl, or optionally together with the atoms to which they are attached, form an optionally substituted C3-C7 carbocycle.

10. The conjugate of claim 9, wherein the amino acid ester is represented by a compound selected from the group consisting of:

11. The conjugate of any one of claims 1 to 8, wherein the linker comprises the following structure: where R 1 and R 1' are each independently selected from hydrogen and optionally substituted C 1-6 Alkyl or optionally together with the atoms to which they are attached form an optionally substituted C3-C7 carbocycle; Each R 2 is an optionally substituted group independently selected from the group consisting of hydrogen, C 1-6 Alkyl, phenyl, C1-C6 carbocycle and 3-7 heterocycle; Each R 3 is hydrogen or optionally substituted C 1-6 alkyl; and n is 1, 2, 3, 4 or 5.

12. The conjugate of claim 11, wherein R 3 For hydrogen.

13. The conjugate of claim 11 or 12, wherein R 2 For hydrogen.

14. The conjugate of claim 11 or 12, wherein R 2 Selected from the group consisting of: hydrogen, -Me, -iso-Pr, -tert-butyl, isobutyl, -CH2Ph, -CH2OH, -CH2SH, -CH2CH2SCH3, -CH2COOH, -CH2CH2COOH, -CH2CONH2, -CH2CH2CONH2, -CH2CH2, CH2CH2NH2, 15. The conjugate of any one of claims 11 to 14, wherein n is 1.

16. The conjugate of any one of claims 11 to 15, wherein R 1 and R 1' Together they form an optionally substituted C3-C7 carbocycle.

17. The conjugate of claim 16, wherein R 1 and R 1' Together they form an optionally substituted C3 carbocycle.

18. The conjugate of any one of claims 1 to 8, wherein the linker comprises the following structure:

19. The conjugate of claim 1, comprising the following structure: Nuc—L1—L2—L3—Pol in: Nuc is nucleotide; Pol is polymerase; L1 is the first part of the linker that connects the nucleotide to L2; L2 is the second part of the linker represented by the following formula: where R 1 and R 1' are each independently selected from optionally substituted C 1-6 alkyl, halogen, or optionally together with the atoms to which they are attached form an optionally substituted C3-C7 carbocycle; Each R 2 is an optionally substituted group independently selected from the group consisting of hydrogen, C 1-6 Alkyl, phenyl, C1-C6 carbocycle and 3-7 heterocycle; Each R 3 is hydrogen or optionally substituted C 1-6 alkyl; n is 0, 1, 2, 3, 4 or 5; Where * denotes the connection point of L2 and L1; and ** denotes the connection point of L2 and L3; Among them L 2 is cleavable; and L 3 To connect pol to L 2 connector.

20. The conjugate of claim 19, wherein L 1 Selected from the group consisting of: a bond, an optionally substituted C 1-12 Alkylene chain, C4-C 20 Polyethylene glycol, optionally substituted C 2-12 Alkenylene chain and C 2-12 Alkyne chain, where L 1 1-6 methylene units are optionally and independently replaced by -O-, -N(R b )-, -N=C(H)-, -C(O)-, -S-, -S(O)-, -S(O)2-, optionally substituted phenylene or optionally substituted cyclopropylene.

21. The conjugate of claim 20, wherein L 1 Include: Each R a are independently selected from the group consisting of halogen, hydroxy, cyano, optionally substituted C 1-6 Alkyl and optionally substituted C 1-6 Alkoxy.

22. The conjugate of any one of claims 19-21, wherein L 2 comprising an amino acid ester selected from the group consisting of:

23. The conjugate of claim 22, wherein L 2 It is represented by:

24. The conjugate of any one of claims 19-23, wherein L 1 Binds to the nucleobase of the nucleotide.

25. The conjugate of claim 24, wherein L 1 Binds to the nucleobase at the oxygen or nitrogen that participates in base pairing.

26. The conjugate of claim 25, wherein the nucleobase is selected from the group consisting of:

27. The conjugate of any one of claims 19-23, wherein L 1 bound to the sugar of the nucleotide.

28. The conjugate of any one of claims 19-23, wherein L 1 Binds to the phosphate of the nucleotide.

29. The conjugate of claim 28, wherein the phosphate is an alpha phosphate.

30. The conjugate of any one of the above claims, wherein the nucleotide is a ribonucleotide polyphosphate or a deoxyribonucleotide polyphosphate.

31. The conjugate of any one of the above claims, wherein the nucleotide is selected from the group consisting of adenine, guanine, cytosine, uracil, and thymine.

32. The conjugate of any one of the above claims, wherein the polymerase is a template-independent polymerase.

33. The conjugate of claim 32, wherein the polymerase is TdT.

34. The conjugate of any one of the above claims, wherein the linker is cleavable by a protease comprising esterase activity.

35. The conjugate of claim 34, wherein the linker is cleavable by proteinase K.

36. The conjugate of claim 34 or 35, wherein the linker is capable of being cleaved at the ester group on L2, leaving a compound represented by Nuc-L1-OH after said cleavage.

37. A method for synthesizing a polynucleotide, comprising: The polynucleotide is incubated with the conjugate of any one of claims 1-36.

38. The method of claim 37, further comprising extending the polynucleotide by adding the nucleotide bound to the conjugate to the 3' OH of the polynucleotide.

39. The method of claim 38, further comprising cleaving the cleavable linker after adding the nucleotide to the precursor polynucleotide.

40. The method of claim 39, further comprising repeating the incubating, extending, and lysing steps one or more times.

41. The method of claim 39 or 40, wherein the cleaving comprises contacting the extended polynucleotide with an enzyme comprising esterase activity under conditions sufficient to cleave the linker, thereby releasing the polymerase from the extension product.

42. The method of claim 41, wherein the enzyme is a protease comprising esterase activity.

43. The method of any one of claims 39-42, further comprising removing a scar remaining after said cleavage of said linker attached to said nucleotide.

44. The method of claim 43, wherein scar removal is performed after polynucleotide synthesis is complete.

45. The method of claim 43, wherein the scar removal is performed after a portion of the polynucleotide is synthesized.

46. The method of claim 43, wherein the scar removal is performed after cleavage of the linker and before addition of the next nucleotide during polynucleotide synthesis.

47. A method for synthesizing a polynucleotide, comprising: (a) incubating the nucleic acid with the first conjugate under conditions where a polymerase catalyzes the covalent addition of a nucleotide of the first conjugate to the 3' hydroxyl group of the nucleic acid to produce a first extension product; (b) cleaving the cleavable bond of the linker, thereby releasing the polymerase from the extension product to unmask the 3' hydroxyl terminus of the first extension product; (c) incubating the extension product with the second conjugate under conditions where the polymerase catalyzes the covalent addition of a nucleotide of the second conjugate of any one of claims 1 to 36 to the 3' end of the first extension product to produce a second extension product; (d) repeating steps (b) to (c) multiple times on the second extension product to generate an extended nucleic acid of a defined sequence.

48. A sequencing method comprising: incubating a duplex comprising a primer and a template with a composition comprising a set of conjugates according to any one of claims 1 to 36, wherein the conjugates correspond to G, A, T (or U), and C and are distinguishably labeled; detecting which nucleotide has been added to the primer by detecting a signal from the distinguishable label; cleaving the cleavable bond of the linker, thereby releasing the polymerase from the extension product to unmask the 3' hydroxyl terminus of the first extension product; as well as The incubation, detection, and cleavage steps are repeated to determine the sequence of the template.

49. A modified nucleotide comprising a cleavable linker, wherein the cleavable linker comprises an amino acid ester.

50. The modified nucleotide of claim 49, wherein the amino acid ester is linked to an amino acid.

51. The modified nucleotide of claim 50, wherein the amine group of the amino acid ester is bound to the amino acid.

52. The modified nucleotide of claim 51, comprising a peptide of at least 2, at least 3, at least 4, or at least 5 amino acids bound to the amine group of the amino acid ester.

53. The modified nucleotide of any one of claims 49-52, wherein the one or more amino acids are selected from the group consisting of alanine, arginine, asparagine, aspartic acid, cysteine, glutamic acid, glutamine, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, and valine.

54. The modified nucleotide of any one of claims 49-52, wherein the amino acid is or comprises glycine.

55. The modified nucleotide of any one of claims 49-52, wherein the amino acid is or comprises a non-naturally occurring amino acid.

56. The modified nucleotide of any one of claims 49-55, wherein the cleavable linker is bound to the α-phosphate, sugar or nucleobase of the nucleotide.

57. The modified nucleotide of any one of claims 49-56, wherein the amino acid ester is represented by: where R 1 and R 1' are each independently selected from hydrogen and optionally substituted C 1-6 Alkyl, or optionally together with the atoms to which they are attached, form an optionally substituted C3-C7 carbocycle.

58. The modified nucleotide of claim 57, wherein the amino acid ester is represented by a compound selected from the group consisting of:

59. The modified nucleotide of any one of claims 49-57, wherein the linker comprises the structure: where R 1 and R 1' are each independently selected from hydrogen and optionally substituted C 1-6 Alkyl or optionally together with the atoms to which they are attached form an optionally substituted C3-C7 carbocycle; Each R 2 is an optionally substituted group independently selected from the group consisting of hydrogen, C 1-6 Alkyl, phenyl, C1-C6 carbocycle and 3-7 heterocycle; Each R 3 is hydrogen or optionally substituted C 1-6 alkyl; and n is 1, 2, 3, 4 or 5.

60. The modified nucleotide of claim 59, wherein R 3 For hydrogen.

61. The modified nucleotide of claim 59 or 60, wherein R 2 For hydrogen.

62. The modified nucleotide of claim 59 or 60, wherein R 2 Selected from the group consisting of: hydrogen, -Me, -iso-Pr, -tert-butyl, isobutyl, -CH2Ph, -CH2OH, -CH2SH, -CH2CH2SCH3, -CH2COOH, -CH2CH2COOH, -CH2CONH2, -CH2CH2CONH2, -CH2CH2, CH2CH2NH2, 63. The modified nucleotide of any one of claims 59-62, wherein n is 1.

64. A modified nucleotide as described in any one of claims 59-63, wherein R 1 and R 1' Together they form an optionally substituted C3-C7 carbocycle.

65. The modified nucleotide of claim 64, wherein R 1 and R 1' Together they form an optionally substituted C3 carbocycle.

66. The modified nucleotide of any one of claims 49-65, wherein the linker comprises the structure:

67. The modified nucleotide of claim 1, comprising the structure: Nuc—L1—L2 in: Nuc is nucleotide; L1 is the first part of the linker that connects the nucleotide to L2 L2 is the second part of the linker represented by the following formula: where R 1 and R 1' are each independently selected from optionally substituted C 1-6 alkyl, halogen, or optionally together with the atoms to which they are attached form an optionally substituted C3-C7 carbocycle; Each R 2 is an optionally substituted group independently selected from the group consisting of hydrogen, C 1-6 Alkyl, phenyl, C1-C6 carbocycle and 3-7 heterocycle; Each R 3 is hydrogen or optionally substituted C 1-6 alkyl; and n is 0, 1, 2, 3, 4 or 5.

68. The modified nucleotide of claim 67, wherein L 1 Selected from the group consisting of: a bond, an optionally substituted C 1-12 Alkylene chain, C4-C 20 Polyethylene glycol, optionally substituted C 2-12 Alkenylene chain and C 2-12 Alkyne chain, where L 1 1-6 methylene units are optionally and independently replaced by -O-, -N(R b )-, -N=C(H)-, -C(O)-, -S-, -S(O)-, -S(O)2-, optionally substituted phenylene or optionally substituted cyclopropylene.

69. The modified nucleotide of claim 68, wherein L 1 Include: Each R a are independently selected from the group consisting of halogen, hydroxy, cyano, optionally substituted C 1-6 Alkyl and optionally substituted C 1-6 Alkoxy.

70. The modified nucleotide of any one of claims 67-69, wherein L 2 comprising an amino acid ester selected from the group consisting of:

71. The modified nucleotide of any one of claims 67-69, wherein L 2 It is represented by:

72. The modified nucleotide of any one of claims 67-71, wherein L 1 Binds to the nucleobase of the nucleotide.

73. The modified nucleotide of claim 72, wherein L 1 Binds to the nucleobase at the oxygen or nitrogen that participates in base pairing.

74. A modified nucleotide as described in any one of claims 67-71, wherein L 1 bound to the sugar of the nucleotide.

75. The modified nucleotide of any one of claims 67-71, wherein L 1 Binds to the phosphate of the nucleotide.

76. The modified nucleotide of any one of claims 67-75, wherein the phosphate is an alpha phosphate.

77. The modified nucleotide of any one of claims 67-76, wherein the nucleotide is a ribonucleotide polyphosphate or a deoxyribonucleotide polyphosphate.

78. The modified nucleotide of any one of claims 67-77, wherein the nucleotide is selected from the group consisting of adenine, guanine, cytosine, uracil, and thymine.

79. The modified nucleotide of any one of claims 67-78, wherein the linker is capable of being cleaved by a protease comprising esterase activity.

80. The modified nucleotide of claim 79, wherein the linker is cleavable by proteinase K.

81. The modified nucleotide of claim 79 or 80, wherein the linker is capable of being cleaved at the ester group on L2, leaving a compound represented by Nuc-L1-OH after said cleavage.

Citation Information

Patent Citations

  • Nucleic Acid Synthesis and Sequencing Using Tethered Nucleoside Triphosphates

    US20190112627A1

  • DNA sequencing method using acyclonucleoside triphosphates

    US5558991A

  • Nucleic acid synthesis and sequencing using tethered nucleoside triphosphates

    WO2017223517A1

  • Use of terminal transferase enzyme in nucleic acid synthesis

    WO2018215803A1

  • Modified template-independent enzymes for polydeoxynucleotide synthesis

    WO2020081985A1