Mirror image crispr-cas compositions
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- DXOME CO LTD
- Filing Date
- 2025-09-17
- Publication Date
- 2026-05-28
AI Technical Summary
The lack of efficient and reliable technologies to synthesize large mirror-image biomolecules, such as D-form proteins and nucleic acids, has constrained the development of mirror-image biology systems.
Development of novel mirror image CRISPR systems and compositions, including D-form DNA endonucleases and ligases, which are mirror image forms of their L-form counterparts, for site-specific cleaving and editing of L-form DNA and RNA.
Enables efficient synthesis and utilization of mirror-image biomolecules, facilitating the development of mirror-image biology systems for various applications.
Abstract
Description
Attorney Ref.: 39339-64210 Client Ref.: 003WO MIRROR IMAGE CRISPR-CAS COMPOSITIONS 1. SEQUENCE LISTING
[0001] The instant application contains a Sequence Listing. 2. BACKGROUND
[0002] Proteins composed entirely of D-amino acids and the achiral amino acid glycine are mirror image forms of their native L-protein counterparts. Mirror image proteins (also referred to as D-proteins or D-form proteins) have a wide range of applications in structural biology, peptide / protein drug design, and mechanistic studies of biological processes. D- proteins can be prepared using chemical protein synthesis methods including native chemical ligation.
[0003] Phosphoramidate chemistry has provided oligonucleotide synthesis of DNA and RNA sequences, including mirror image DNA and RNA. Researchers have generated mirror-image forms of some biomolecules as an attempt to create a mirror-image artificial system based on D-proteins such as D-form polymerases, and L-form nucleic acids. The lack of efficient and reliable technologies to synthesize large mirror-image biomolecules have prohibitively constrained the development of such systems.
[0004] Mirror-image biology systems and compositions, and novel applications for utilizing them are of great interest. 3. SUMMARY
[0005] The present disclosure provides novel mirror image CRISPR systems and compositions that include components such as mirror image CRISPR-associated endonucleases or mirror image guide RNA (e.g., mirror image gRNA or sgRNA) for processing of mirror image nucleic acids. The compositions find use in a variety of applications where site-specific cleaving or editing of mirror image nucleic acids, e.g., L- form DNA and / or RNA, can be utilized.
[0006] Accordingly, in a first aspect, the present disclosure provides a D-form DNA endonuclease. In some embodiments, the D-form DNA endonuclease is a mirror image form of a L-form DNA endonuclease or a variant thereof.
[0007] In some embodiments, the D-form DNA endonuclease is a mirror image form of a L- form Cas protein.Attorney Ref.: 39339-64210 Client Ref.: 003WO
[0008] In some embodiments, the D-form DNA endonuclease is a mirror image isomer of a class 1, type I, III, or IV Cas endonuclease or a functional derivative thereof. In some embodiments, the D-form DNA endonuclease is a mirror image isomer of a class 2, type II, V, or VI Cas endonuclease or a functional derivative thereof.
[0009] In some embodiments, the D-form DNA endonuclease is a mirror image isomer of a L-Cas protein selected from Cas3, Cas9, Cas10, Cas12, Cas13, Cas14 and functional derivatives thereof.
[0010] In some embodiments, the D-form DNA endonuclease has a sequence of a L-Cas protein with one or more variations, wherein the one or more variations comprise a substitution from alanine to cysteine. In some embodiments, the D-form DNA endonuclease has a sequence of a L-Cas protein with one or more variations, wherein the one or more variations comprise a substitution from isoleucine to alanine, leucine, phenylalanine, valine, threonine, or tyrosine.
[0011] In some embodiments, the D-form DNA endonuclease comprises: i) a sequence selected from SEQ ID NOs: 1-19; or ii) a sequence having at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to the sequence set forth in i).
[0012] In some embodiments, the the D-form DNA endonuclease comprises: a. a sequence having at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to the sequence of SEQ ID NO: 10; and b. one or more variations selected from: A62C, A105C, A157C, A188C, A228C, A284C, A349C, A388C, A449C, A498C, A541C, A602C, A660C, A712C, A801C, A799C, A844C, A894C, A954C, A1010C, A1067C, A1135C, A1179C, A1131C, and A1267C.
[0013] In some embodiments, the D-form DNA endonuclease comprises: a. a sequence having at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to the sequence of SEQ ID NO: 10; and b. one or more variations selected from: I37L; I53V; I57Y; A62C; I96A; A105C; I111T; I121F; I128A; A157C; A188C; I207A; I212V; A228C; I229A; I231A; I237L; I265V; A284C; I285A; I295L; I303A; I330L; A349C; I359A; I374V; A388C; I394A; I401A; I418F; I423L; I424A; I442V; A449C; I466A; A498C; I503A; A541C; A602C; I626Y; I633A; A660C; I693Y; A712C; I726V; I731V; A751C; A799C; A844C; I850V; A894C; I915L; I917A; I935V; A954C; I984V;Attorney Ref.: 39339-64210 Client Ref.: 003WO A1010C; A1067C; A1135C; I1138V; I1164A; A1179C; I1183A; I1199V; I1212F; I1219F; A1231C; A1267C; and I1302V.
[0014] In some embodiments, the D-form DNA endonuclease comprises: a. a sequence having at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to the sequence of SEQ ID NO: 16; and b. one or more variations selected from: A38C; A83C; A111C; A162C; A226C; A264C; A364C; A428C; and A476C.
[0015] In some embodiments, thethe D-form DNA endonuclease comprises: a. a sequence having at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to the sequence of SEQ ID NO: 16; and b. one or more variations selected from: I15V; I24V; I28V; A38C; I76A, A83C, I92V, I103L, I106A, A111C, I124L, A162C, A226C, I248V, A264C, I272A, I281A, I289A, I297V, I345A, A364C, A428C, A476C, and I484L.
[0016] In some embodiments, the D-form DNA endonuclease comprises: a. a sequence having at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to the sequence of SEQ ID NO: 16; and b. one or more variations selected from: I15V; I24V; I28V; I76A; I92V; I103L; I106A; I124L; I248V; I272A; I281A; I289A; V294A; I297V; I345A; and I484L.
[0017] In a second aspect, the present disclosure provides a D-form DNA ligase. In some embodiments, the D-form DNA ligase is a mirror image form of an L-form DNA ligase or a variant thereof.
[0018] In some embodiments, the D-form DNA ligase is a mirror image form of a L-form T4 ligase.
[0019] In some embodiments, the D-form DNA ligase has a sequence of a L-T4 ligase with one or more variations, wherein the one or more variations comprise a substitution from alanine to cysteine. In some embodiments, the D-form DNA ligase has a sequence of a L-T4 ligase with one or more variations, wherein the one or more variations comprise a substitution from isoleucine to alanine, leucine, or valine. In some embodiments, the D-form DNA ligase has a sequence of a L-T4 ligase with one or more variations, wherein the one or more variations comprise a substitution from valine to alanine.
[0020] In some embodiments, the D-form DNA ligase comprises: i) a sequence selected from SEQ ID NOs: 20-22; orAttorney Ref.: 39339-64210 Client Ref.: 003WO ii) a sequence having at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to the sequence set forth in i).
[0021] In some embodiments, the D-form DNA ligase comprises: a. a sequence having at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to the sequence of SEQ ID NO: 21; and b. a substitution of V317A.
[0022] In a third aspect, the present disclosure provides a composition. In some embodiments, the composition is a mirror image DNA-cleaving and / or modifying composition.
[0023] In some embodiments, the composition comprises a D-form DNA endonuclease (e.g., as described herein) and an L-form guide RNA, optionally wherein the L-form guide RNA comprises a sequence complementary to a target sequence of a target L-DNA. In some embodiments, the L-form guide RNA comprises a sequence complementary to a target sequence of a target L-DNA.
[0024] In some embodiments, the composition further comprises a target L-DNA. In some embodiments, the target L-DNA comprises a duplex segment that includes the target sequence to which the L-guide RNA is directed. In some embodiments, the target L-DNA is comprised within a folded nanostructure.
[0025] In some embodiments, the composition further comprises a D-form DNA ligase (e.g., as described herein).
[0026] In a fourth aspect, the present disclosure provides a D-form endonuclease having 90% or greater sequence identity to SEQ ID NOs: 9-12, wherein the D-form endonuclease is prepared via native chemical ligation (NCL) of two or more precursor fragments between one or more positions defined by the following pairs of amino acid residues: i) between tyr61and cys62; ii) between asn104and cys105; iii) between asn156and cys157; iv) between thr187and cys188; v) between lys227and cys228; vi) between leu283and cys284; vii) between thr348and cys349; viii) between asn387and cys388; ix) between ala448and cys449;Attorney Ref.: 39339-64210 Client Ref.: 003WO x) between ser497and cys498; xi) between leu540and cys541; xii) between ala601and cys602; xiii) between tyr659and cys660; xiv) between tyr711and cys712; xv) between phe750and cys751; xvi) between met798and cys799; xvii) between arg843and cys844; xviii) between asn893and cys894; xix) between ala953and cys954; xx) between lys1009and cys1010; xxi) between pro1066and cys1067; xxii) between pro1134and cys1135; xxiii) between pro1178and cys1179; xxiv) between ala1230and cys1231; and xxv) between Gly1266and cys1267.
[0027] In some embodiments, the D-form endonuclease is prepared via desulfurization of one or more cysteine residues of a precursor fragment to alanine residues at one or more positions selected from 62, 105, 157, 188, 228, 284, 349, 388, 449, 498, 541, 602, 660, 712, 751, 799, 844, 894, 954, 1010, 1067, 1135, 1179, 1231, and 1267.
[0028] In some embodiments, the D-form endonuclease is prepared via NCL coupling of two precursor fragments between positions tyr659and cys660.
[0029] In some embodiments, the D-form endonuclease is prepared via NCL coupling of: i) met1-tyr659-NHNH2 (cys(Acm)65, 205, 334, 379, 608); and ii) cys660-asn1307-OH (cys(Acm)674, 1025, 1248).
[0030] In some embodiments, the D-form endonuclease has at least 95% (e.g., 95% or greater, 96% or greater, 97% or greater, 98% or greater, or 99% or greater) sequence identity to the D-Cas12 protein of SEQ ID NOs: 9-12.
[0031] In a fifth aspect, the present disclosure provides a D-polypeptide hydrazide having a structure selected from a group consisting of: a. met1-tyr61-NHNH2 ; b. cys62-asn104-NHNH2(cys(Acm)65); c. cys105-asn156-NHNH2 ;Attorney Ref.: 39339-64210 Client Ref.: 003WO d. thz157-thr187-NHNH2 ; e. cys188-lys227-NHNH2(cys(Acm)205); f. cys228-leu283-NHNH2 ; g. thz284-thr348-NHNH2(cys(Acm)334); h. cys349-asn387-NHNH2 (cys(Acm)379); i. cys388-ala448-NHNH2; j. thz449-ser497-NHNH2 ; k. cys498-leu540-NHNH2; l. thz541-ala601-NHNH2 ; m. cys602-tyr659-NHNH2 (cys(Acm)608); n. thz660-tyr711-NHNH2(cys(Acm)674); o. cys712-phe750-NHNH2 ; p. cys751-met798-NHNH2; q. thz799-arg843-NHNH2 ; r. cys844-asn893-NHNH2; s. cys894-ala953-NHNH2 ; t. thz954-lys1009-NHNH2; u. cys1010-pro1066-NHNH2 (cys(Acm)1025) ; v. cys1067-pro1134-NHNH2; w. thz1135-pro1178-NHNH2 ; x. cys1179-ala1230-NHNH2; y. thz1231-Gly1266-NHNH2 (cys(Acm)1248); and z. cys1267-asn1307-OH and having a sequence with at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to any one of SEQ ID NOs: 9-12 between two amino acid residues corresponding to the positions identified in the structure.
[0032] In a sixth aspect, the present disclosure provides a D-form endonuclease having at least 90% sequence identity to a sequence selected from SEQ ID NOs: 15-19, wherein the D- form endonuclease is prepared via native chemical ligation (NCL) of two or more precursor fragments at one or more positions defined by the following pairs of amino acid residues: i) between gln37and cys38; ii) between his82and cys83; iii) between arg110and cys111;Attorney Ref.: 39339-64210 Client Ref.: 003WO iv) between Gly161and cys162; v) between trp225and cys226; vi) between phe263and cys264; vii) between phe319and cys320; viii) between trp363and cys364; ix) between gln427and cys428; and x) between ala475and cys476.
[0033] In some embodiments, the D-form endonuclease is prepared via desulfurization of one or more cysteine residues of a precursor fragment to alanine residues at one or more positions selected from 38, 83, 111, 162, 226, 264, 320, 364, 428, and 476.
[0034] In some embodiments, the D-form endonuclease is prepared via NCL coupling of two precursor fragments between positions phe263and cys264.
[0035] In some embodiments, the D-form endonuclease is prepared via NCL coupling of: i) met1-phe263-NHNH2 (cys(Acm)125) ; and ii) cys264-lys500-OH (cys(Acm)295, 435, 440, 462, 465).
[0036] In some embodiments, the D-form endonuclease has at least 95% (e.g., 95% or greater, 96% or greater, 97% or greater, 98% or greater, or 99% or greater) sequence identity to the D-Cas14 protein of SEQ ID NOs: 15-19.
[0037] In a seventh aspect, the present disclosure provides a D-polypeptide hydrazide having a structure selected from a group consisting of: a. met1-gln37-NHNH2; b. cys38-his82-NHNH2 ; c. thz83-arg110-NHNH2; d. cys111-Gly161-NHNH2 (Cys(Acm)125) ; e. thz162-trp225-NHNH2; f. cys226-phe263-NHNH2 ; g. thz264-phe319-NHNH2(Cys(Acm)295) ; h. cys320-trp363-NHNH2 ; i. thz364-gln427-NHNH2 ; j. thz428-ala475-NHNH2 (cys(Acm)435, 440, 462, 465) ; and k. cys476-lys500-OHAttorney Ref.: 39339-64210 Client Ref.: 003WO and having a sequence with at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to any one of SEQ ID NOs: 15-19 between two amino acid residues corresponding to the positions identified in the structure.
[0038] In an eighth aspect, the present disclosure provides a D-form endonuclease having at least 90% sequence identity to a sequence selected from SEQ ID NOs: 15-19, wherein the D- form endonuclease is prepared via native chemical ligation (NCL) of two or more precursor fragments at one or more positions defined by the following pairs of amino acid residues: i) between ala76and cys77; ii) between leu124and cys125; iii) between Gly161and cys162; iv) between trp225and cys226; v) between ala294and cys295; vi) between trp363and cys364; and vii) between leu434and cys435.
[0039] In some embodiments, the D-form endonuclease is prepared via desulfurization of one or more cysteine residues of a precursor fragment to alanine residues at one or more positions selected from 77, 125, 162, 226, 295, 364, and 435.
[0040] In some embodiments, the D-form endonuclease is prepared via NCL coupling of two precursor fragments between positions ala294and cys295.
[0041] In some embodiments, the D-form endonuclease is prepared via NCL coupling of: i) met1-ala294-NHNH2; and ii) cys295-lys500-OH.
[0042] In some embodiments, the D-form endonuclease has at least 95% (e.g., 95% or greater, 96% or greater, 97% or greater, 98% or greater, or 99% or greater) sequence identity to the D-Cas14 protein of SEQ ID NOs: 15-19.
[0043] In a ninth aspect, the present disclosure provides a D-polypeptide hydrazide having a structure selected from a group consisting of: a. met1-ala76-NHNH2; b. cys77-leu124-NHNH2; c. thz125-Gly161-NHNH2; d. thz162-trp225-NHNH2; e. cys226-ala294-NHNH2; f. thz295-trp363-NHNH2;Attorney Ref.: 39339-64210 Client Ref.: 003WO g. cys364-leu434-NHNH2; and h. cys435-lys500-OH and having a sequence with at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to any one of SEQ ID NOs: 15-19 between two amino acid residues corresponding to the positions identified in the structure.
[0044] In a tenth aspect, the present disclosure provides D-form T4 ligase having a sequence of a L-T4 ligase with one or more variations. In some embodiments, the one or more variations comprise a substitution from valine to alanine. In some embodiments, the one or more variations comprise a substitution from alanine to cysteine. In some embodiments, the one or more variations comprise a substitution from isoleucine to alanine, leucine, or valine.
[0045] In some embodiments, the D-form T4 ligase comprises: i) a sequence selected from SEQ ID Nos: 20-22; or ii) a sequence having at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to the sequence set forth in i).
[0046] In some embodiments, the D-form T4 ligase having at least 90% sequence identity to a sequence selected from SEQ ID NOs: 20-22, wherein the D-form T4 ligase is prepared via native chemical ligation (NCL) of two or more precursor fragments at one or more positions defined by the following pairs of amino acid residues: i) between leu76and cys77; ii) between Gly116and cys117; iii) between arg164and cys165; iv) between thr201and cys202; v) between thr255and cys256; vi) between lys316and cys317; vii) between lys388and cys389; and viii) between lys412and cys413.
[0047] In some embodiments, the D-form T4 ligase is prepared via desulfurization of one or more cysteine residues of a precursor fragment to alanine residues at one or more positions selected from 77, 117, 165, 202, 256, 317, 389, and 413.
[0048] In some embodiments, the D-form T4 ligase is prepared via NCL coupling of two precursor fragments between positions arg164and cys165.
[0049] In some embodiments, the D-form T4 ligase is prepared via NCL coupling of: i) met1-arg164-NHNH2 (Cys(Acm)115); andAttorney Ref.: 39339-64210 Client Ref.: 003WO ii) cys165-leu487-OH.
[0050] In some embodiments, the D-form T4 ligase has at least 95% (e.g., 95% or greater, 96% or greater, 97% or greater, 98% or greater, or 99% or greater) sequence identity to the D-form T4 ligase of SEQ ID NOs: 20-22.
[0051] In an eleventh aspect, the present disclosure provides D-polypeptide hydrazide having a structure selected from a group consisting of: a. met1- leu76-NHNH2 ; b. cys77-Gly116-NHNH2(Cys(Acm)115); c. cys117-arg164-NHNH2 ; d. thz165-thr201-NHNH2 ; e. cys202-thr255-NHNH2; f. thz256-lys316-NHNH2 (Cys(Acm)276); g. cys317-lys388-NHNH2; h. thz389-lys412-NHNH2 (Cys(Acm)398); and i. cys413-leu487-OH (Cys(Acm)439, 441) and having a sequence with at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to any one of SEQ ID NOs: 20-22 between two amino acid residues corresponding to the positions identified in the structure.
[0052] In a twelfth aspect, the present disclosure provides a method of making a D-form Cas12 having 90% or greater sequence identity to SEQ ID NOs: 10-12, comprising coupling of two precursor fragments between positions tyr659and cys660via native chemical ligation (NCL).
[0053] In some embodiments, the two precursor fragments are: met1-tyr659-NHNH2 (cys(Acm)65, 205, 334, 379, 608); and cys660-asn1307-OH (cys(Acm)674, 1025, 1248).
[0054] In some embodiments, the the D-form Cas12 is prepared via native chemical ligation (NCL) couplings of precursor fragments between positions defined by the following pairs of amino acid residues: i) between tyr61and cys62; ii) between asn104and cys105; iii) between asn156and cys157; iv) between thr187and cys188; v) between lys227and cys228;Attorney Ref.: 39339-64210 Client Ref.: 003WO vi) between leu283and cys284; vii) between thr348and cys349; viii) between asn387and cys388; ix) between ala448and cys449; x) between ser497and cys498; xi) between leu540and cys541; xii) between ala601and cys602; xiii) between tyr659and cys660; xiv) between tyr711and cys712; xv) between phe750and cys751; xvi) between met798and cys799; xvii) between arg843and cys844; xviii) between asn893and cys894; xix) between ala953and cys954; xx) between lys1009and cys1010; xxi) between pro1066and cys1067; xxii) between pro1134and cys1135; xxiii) between pro1178and cys1179; xxiv) between ala1230and cys1231; and xxv) between Gly1266and cys1267.
[0055] In some embodiments, the precursor fragments are derived from one or more of the following fragments: a. met1-tyr61-NHNH2; b. cys62-asn104-NHNH2 (cys(Acm)65); c. cys105-asn156-NHNH2; d. thz157-thr187-NHNH2; e. cys188-lys227-NHNH2(cys(Acm)205); f. cys228-leu283-NHNH2; g. thz284-thr348-NHNH2 (cys(Acm)334); h. cys349-asn387-NHNH2 (cys(Acm)379); i. cys388-ala448-NHNH2; j. thz449-ser497-NHNH2; k. cys498-leu540-NHNH2;Attorney Ref.: 39339-64210 Client Ref.: 003WO l. thz541-ala601-NHNH2; m. cys602-tyr659-NHNH2(cys(Acm)608); n. thz660-tyr711-NHNH2 (cys(Acm)674); o. cys712-phe750-NHNH2; p. cys751-met798-NHNH2; q. thz799-arg843-NHNH2; r. cys844-asn893-NHNH2; s. cys894-ala953-NHNH2; t. thz954-lys1009-NHNH2; u. cys1010-pro1066-NHNH2 (cys(Acm)1025); v. cys1067-pro1134-NHNH2; w. thz1135-pro1178-NHNH2; x. cys1179-ala1230-NHNH2; y. thz1231-Gly1266-NHNH2 (cys(Acm)1248); and z. cys1267-asn1307-OH.
[0056] In some embodiments, the method further comprises desulfurizing one or more cysteine residues of a precursor fragment to alanine residues at one or more positions selected from 62, 105, 157, 188, 228, 284, 349, 388, 449, 498, 541, 602, 660, 712, 751, 799, 844, 894, 954, 1010, 1067, 1135, 1179, 1231, and 1267.
[0057] In a thirteenth aspect, the present disclosure provides a method of making a D-form Cas14 having 90% or greater sequence identity to SEQ ID NOs: 16-19, comprising coupling of two precursor fragments between positions phe263and cys264via native chemical ligation (NCL).
[0058] In some embodiments, the two precursor fragments are: met1-phe263-NHNH2(cys(Acm)125); and cys264-lys500-OH (cys(Acm)295, 435, 440, 462, 465).
[0059] In some embodiments, the D-form Cas14 is prepared via native chemical ligation (NCL) couplings of precursor fragments between positions defined by the following pairs of amino acid residues: i) between gln37and cys38; ii) between his82and cys83; iii) between arg110and cys111; iv) between Gly161and cys162;Attorney Ref.: 39339-64210 Client Ref.: 003WO v) between trp225and cys226; vi) between phe263and cys264; vii) between phe319and cys320; viii) between trp363and cys364; ix) between gln427and cys428; and x) between ala475and cys476.
[0060] In some embodiments, the precursor fragments are derived from one or more of the following fragments: a. met1-gln37-NHNH2; b. cys38-his82-NHNH2; c. thz83-arg110-NHNH2; d. cys111-Gly161-NHNH2 (Cys(Acm)125); e. thz162-trp225-NHNH2; f. cys226-phe263-NHNH2; g. thz264-phe319-NHNH2(Cys(Acm)295); h. cys320-trp363-NHNH2; i. thz364-gln427-NHNH2; j. thz428-ala475-NHNH2 (cys(Acm)435, 440, 462, 465); and k. cys476-lys500-OH.
[0061] In some embodiments, the method further comprises desulfurizing one or more cysteine residues of a precursor fragment to alanine residues at one or more positions selected from 38, 83, 111, 162, 226, 264, 320, 364, 428, and 476.
[0062] In a fourteenth aspect, the present disclosure provides a method of making a D-form Cas14 having 90% or greater sequence identity to SEQ ID NOs: 16-19, comprising coupling of two precursor fragments between positions ala294and cys295via native chemical ligation (NCL).
[0063] In some embodiments, the two precursor fragments are: met1-ala294-NHNH2; and cys295-lys500-OH.
[0064] In some embodiments, the D-form Cas14 is prepared via native chemical ligation (NCL) couplings of precursor fragments between positions defined by the following pairs of amino acid residues: i) between ala76and cys77;Attorney Ref.: 39339-64210 Client Ref.: 003WO ii) between leu124and cys125; iii) between Gly161and cys162; iv) between trp225and cys226; v) between ala294and cys295; vi) between trp363and cys364; and vii) between leu434and cys435.
[0065] In some embodiments, the precursor fragments are derived from one or more of the following fragments: a. met1- ala76-NHNH2 ; b. cys77-leu124-NHNH2 ; c. thz125-Gly161-NHNH2; d. thz162-trp225-NHNH2 ; e. cys226-ala294-NHNH2; f. thz295-trp363-NHNH2 ; g. cys364-leu434-NHNH2; and h. cys435-lys500-OH.
[0066] In some embodiments, the method further comprises desulfurizing one or more cysteine residues of a precursor fragment to alanine residues at one or more positions selected from 77, 125, 162, 226, 295, 364, and 435.
[0067] In a fifteenth aspect, the present disclosure provides a method of making a D-form T4 ligase having 90% or greater sequence identity to SEQ ID NO 21 or SEQ ID NO 22, comprising coupling of two precursor fragments between positions arg164and cys165via native chemical ligation (NCL).
[0068] In some embodiments, the two precursor fragments are: met1-arg164-NHNH2(Cys(Acm)115); and cys165-leu487-OH.
[0069] In some embodiments, the D-form T4 ligase is prepared via native chemical ligation (NCL) couplings of precursor fragments between positions defined by the following pairs of amino acid residues: i) between leu76and cys77; ii) between Gly116and cys117; iii) between arg164and cys165; iv) between thr201and cys202;Attorney Ref.: 39339-64210 Client Ref.: 003WO v) between thr255and cys256; vi) between lys316and cys317; vii) between lys388and cys389; and viii) between lys412and cys413.
[0070] In some embodiments, the precursor fragments are derived from one or more of the following fragments: a. met1- leu76-NHNH2 ; b. cys77-Gly116-NHNH2(Cys(Acm)115); c. cys117-arg164-NHNH2 ; d. thz165-thr201-NHNH2 ; e. cys202-thr255-NHNH2; f. thz256-lys316-NHNH2 (Cys(Acm)276); g. cys317-lys388-NHNH2; h. thz389-lys412-NHNH2 (Cys(Acm)398); and i. cys413-leu487-OH (Cys(Acm)439, 441).
[0071] In some embodiments, the method further comprises desulfurizing one or more cysteine residues of a precursor fragment to alanine residues at one or more positions selected from 77, 117, 165, 202, 256, 317, 389, and 413.
[0072] In a sixteenth aspect, the present disclosure provides a method of cleaving and / or modifying a target L-DNA. In some embodiments, the method comprises: contacting a composition as described herein with a target L-DNA; wherein a guide RNA of the composition hybridizes to the target L-DNA sequence, thereby directing a D-form DNA endonuclease to bind to said target L-DNA sequence and cleave and / or modify the target L- DNA sequence.
[0073] In a seventeenth aspect, the present disclosure provides a method of modifying a target L-DNA. In some embodiments, the method comprises: contacting a composition as described herein with a target L-DNA; wherein a guide RNA of the composition hybridizes to the target L-DNA sequence, thereby directing a D-form DNA endonuclease to bind to said target L-DNA sequence and cleave and / or modify the target L-DNA sequence, and wherein a D-form DNA ligase ligates the cleaved and / or modified target L-DNA sequence.
[0074] In an eighteenth aspect, the present disclosure provides a method of editing a target L- DNA. In some embodiments, the method comprises: interacting a composition as described herein and a target L-DNA in a mixture; and incubating the mixture thereby inducingAttorney Ref.: 39339-64210 Client Ref.: 003WO cleavage and / or modification of the target L-DNA. In some embodiments, the method comprises: interacting a composition as described herein and a target L-DNA in a mixture; and incubating the mixture thereby inducing modification and ligation of the target L-DNA. 4. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] These and other features, aspects, and advantages of the present invention will become better understood with regard to the following description, and accompanying drawings.
[0076] FIGs.1A-1G illustrate an exemplary method of preparing D-form Cas12 via fragment couplings using native chemical ligation (NCL) according to the retrosynthetic scheme and synthetic precursors shown in FIGs.1A-1G.
[0077] FIGs.2A-2C illustrate an exemplary method of preparing D-form Cas14a.1 via fragment couplings using native chemical ligation (NCL) according to the retrosynthetic scheme and synthetic precursors shown in FIGs.2A-2C.
[0078] FIG.3 shows a sequence alignment of Cas12 with variants having alanine to cysteine mutations (e.g., at positions selected for fragment condensation via native chemical ligation) and variants having substitutions of isoleucine for e.g., alanine or leucine. It is understood that although the sequences are depicted in L-form, that mirror image D-forms of these sequences are meant to be encompassed by this disclosure.
[0079] FIG.4 shows a sequence alignment of Cas14a.1 with variants having alanine to cysteine mutations (e.g., at positions selected for fragment condensation via native chemical ligation) and variants having substitutions of isoleucine for e.g., alanine or leucine. It is understood that although the sequences are depicted in L-form, that mirror image D-forms of these sequences are meant to be encompassed by this disclosure.
[0080] FIG.5 illustrates an exemplary method of preparing D-form Cas14a.1 mutant (Cas14a.1m) via fragment couplings using native chemical ligation (NCL) according to the retrosynthetic scheme and synthetic precursors shown in FIG.5.
[0081] FIG.6 shows an analytical HPLC chromatogram of the crude D-Cas14a.1m-2 (λ = 214 nm). Column: Welch XB-C4, 4.6 x 250 mm, 5 µm. Buffers: A: 0.1% TFA / H2O, B: 0.1% TFA / CH3CN. Gradient: 20 to 70% of Buffer B over 30 min.Attorney Ref.: 39339-64210 Client Ref.: 003WO
[0082] FIG.7 shows an analytical HPLC chromatogram of the crude D-Cas14a.1m-3 (λ = 214 nm). Column: Welch XB-C4, 4.6 x 250 mm, 5 µm. Buffers: A: 0.1% TFA / H2O, B: 0.1% TFA / CH3CN. Gradient: 20 to 70% of Buffer B over 30 min.
[0083] FIG.8 shows an analytical HPLC chromatogram of the crude D-Cas14a.1m-4 (λ = 214 nm). Column: Welch XB-C4, 4.6 x 250mm, 5 µm. Buffers: A: 0.1% TFA / H2O, B: 0.1% TFA / CH3CN. Gradient: 20 to 70% of Buffer B over 30 min.
[0084] FIG.9 shows an analytical HPLC chromatogram of the crude D-Cas14a.1m-5 (λ = 214 nm). Column: Welch XB-C4, 4.6 x 250 mm, 5 µm. Buffers: A: 0.1% TFA / H2O, B: 0.1% TFA / CH3CN. Gradient: 20 to 70% of Buffer B over 30 min.
[0085] FIG.10 shows an analytical HPLC chromatogram of the crude D-Cas14a.1m-6 (λ = 214 nm). Column: Welch XB-C4, 4.6 x 250mm, 5 µm. Buffers: A: 0.1% TFA / H2O, B: 0.1% TFA / CH3CN. Gradient: 20 to 70% of Buffer B over 30 min.
[0086] FIG.11 shows an analytical HPLC chromatogram of the crude D-Cas14a.1m-7 (λ = 214 nm). Column: Welch XB-C4, 4.6 x 250 mm, 5 µm. Buffers: A: 0.1% TFA / H2O, B: 0.1% TFA / CH3CN. Gradient: 20 to 70% of Buffer B over 30 min.
[0087] FIG.12 shows an analytical HPLC chromatogram of the crude D-Cas14a.1m-8 (λ = 214 nm). Column: Welch XB-C4, 4.6 x 250 mm, 5 µm. Buffers: A: 0.1% TFA / H2O, B: 0.1% TFA / CH3CN. Gradient: 20 to 70% of Buffer B over 30 min.
[0088] FIG.13 illustrates an exemplary method of preparing D-form T4 ligase via fragment couplings using native chemical ligation (NCL) according to the retrosynthetic scheme and synthetic precursors shown in FIG.13.
[0089] FIG.14 shows an analytical HPLC chromatogram of the crude D-T4-1 (λ = 214 nm). Column: Welch XB-C4, 4.6 x 250 mm, 5 µm. Buffers: A: 0.1% TFA / H2O, B: 0.1% TFA / CH3CN. Gradient: 20 to 70% of Buffer B over 30 min.
[0090] FIG.15 shows an analytical HPLC chromatogram of the crude D-T4-7 (λ = 214 nm). Column: Welch XB-C4, 4.6 x 250 mm, 5 µm. Buffers: A: 0.1% TFA / H2O, B: 0.1% TFA / CH3CN. Gradient: 20 to 70% of Buffer B over 30 min.
[0091] FIG.16 illustrates the synthesis of Fmoc-2-Cl-(Trt)-NHNH2 as detailed in exemplary general methods. 5. DETAILED DESCRIPTION 5.1. DefinitionsAttorney Ref.: 39339-64210 Client Ref.: 003WO
[0092] Unless otherwise defined herein, scientific and technical terms used in connection with the present invention shall have the meanings that are commonly understood by those of ordinary skill in the art. Further, unless otherwise required by context, singular terms shall include pluralities and plural terms shall include the singular. Generally, nomenclatures used in connection with, and techniques of, cell and tissue culture, molecular biology, immunology, microbiology, genetics and protein and nucleic acid chemistry and hybridization described herein are those well-known and commonly used in the art. The methods and techniques of the present invention are generally performed according to conventional methods well known in the art and as described in various general and more specific references that are cited and discussed throughout the present specification unless otherwise indicated. Enzymatic reactions and purification techniques are performed according to manufacturer’s specifications, as commonly accomplished in the art or as described herein. The terminology used in connection with, and the laboratory procedures and techniques of, analytical chemistry, synthetic organic chemistry, and medicinal and pharmaceutical chemistry described herein are those well-known and commonly used in the art. Standard techniques can be used for chemical syntheses, chemical analyses, pharmaceutical preparation, formulation, and delivery, and treatment of patients.
[0093] The following terms, unless otherwise indicated, shall be understood to have the following meanings:
[0094] With regard to the binding of an antibody to a target molecule, the terms “bind,” “specific binding,” “specifically binds to,” “specific for,” “selectively binds,” and “selective for” a particular antigen (e.g., a polypeptide target) or an epitope on a particular antigen mean binding that is measurably different from a non-specific or non-selective interaction (e.g., with a non-target molecule). Specific binding can be measured, for example, by measuring binding to a target molecule and comparing it to binding to a non-target molecule. Specific binding can also be determined by competition with a control molecule that mimics the epitope recognized on the target molecule. In that case, specific binding is indicated if the binding of the antibody to the target molecule is competitively inhibited by the control molecule.
[0095] An “isolated nucleic acid” refers to a nucleic acid molecule that has been separated from a component of its natural environment. An isolated nucleic acid includes a nucleic acid molecule contained in cells that ordinarily contain the nucleic acid molecule, but the nucleicAttorney Ref.: 39339-64210 Client Ref.: 003WO acid molecule is present extrachromosomally or at a chromosomal location that is different from its natural chromosomal location.
[0096] The term “pharmaceutical composition” refers to a preparation which is in such form as to permit the biological activity of an active ingredient contained therein to be effective, and which contains no additional components which are unacceptably toxic to a subject to which the formulation would be administered.
[0097] A “pharmaceutically acceptable carrier” refers to an ingredient in a pharmaceutical formulation, other than an active ingredient, which is nontoxic to a subject. A pharmaceutically acceptable carrier includes, but is not limited to, a buffer, excipient, stabilizer, or preservative.
[0098] Terms such as "associated," "connected," "attached," "linked," and "conjugated" are used interchangeably herein and encompass direct as well as indirect connection, attachment, linkage or conjugation unless the context clearly dictates otherwise.
[0099] "Complementary" or "substantially complementary" refers to the ability to hybridize or base pair between nucleotides or nucleic acids, such as, for instance, between the two strands of a double stranded DNA molecule or between a polynucleotide primer and a primer binding site on a single stranded nucleic acid to be sequenced or amplified. Complementary nucleotides are, generally, A and T (or A and U), or C and G. Two single-stranded RNA or DNA molecules are said to be substantially complementary when the nucleotides of one strand, optimally aligned and compared and with appropriate nucleotide insertions or deletions, pair with at least about 80% of the nucleotides of the other strand, usually at least about 90% to 95%, and more preferably from about 98 to 100%.
[0100] Alternatively, substantial complementarity exists when an RNA or DNA strand will hybridize under selective hybridization conditions to its complement. Typically, selective hybridization will occur when there is at least about 65% complementary over a stretch of at least 14 to 25 nucleotides, preferably at least about 75%, more preferably at least about 90% complementary.
[0101] "Preferential binding" or "preferential hybridization" refers to the increased propensity of one polynucleotide to bind to a complementary polynucleotide in a sample as compared to noncomplementary polynucleotides in the sample or as compared to the propensity of the one polynucleotide to form an internal secondary structure such as a hairpin or stem-loop structure under at least one set of hybridization conditions.Attorney Ref.: 39339-64210 Client Ref.: 003WO
[0102] Stringent hybridization conditions will typically include salt concentrations of less than about 1 M, more usually less than about 500 mM and preferably less than about 200 mM. Hybridization temperatures can be as low as 5.degree.C., but are typically greater than 22.degree.C., more typically greater than about 30.degree.C., and preferably in excess of about 37.degree.C. Longer fragments may require higher hybridization temperatures for specific hybridization. Other factors may affect the stringency of hybridization, including base composition and length of the complementary strands, presence of organic solvents and extent of base mismatching, and the combination of parameters used is more important than the absolute measure of any one alone. Other hybridization conditions which may be controlled include buffer type and concentration, solution pH, presence and concentration of blocking reagents to decrease background binding such as repeat sequences or blocking protein solutions, detergent type(s) and concentrations, molecules such as polymers which increase the relative concentration of the polynucleotides, metal ion(s) and their concentration(s), chelator(s) and their concentrations, and other conditions known in the art. Less stringent, and / or more physiological, hybridization conditions can be used.
[0103] The terms "polynucleotide," "oligonucleotide," "nucleic acid" and "nucleic acid molecule" are used interchangeably herein to refer to a polymeric form of nucleotides of any length, and may comprise ribonucleotides, deoxyribonucleotides, analogs thereof, mirror image isomers thereof, or mixtures thereof. This term refers only to the primary structure of the molecule. Thus, the term includes triple-, double- and single-stranded deoxyribonucleic acid ("DNA"), as well as triple-, double- and single-stranded ribonucleic acid ("RNA"). It also includes modified, for example by alkylation, and / or by capping, and unmodified forms of the polynucleotide. More particularly, the terms "polynucleotide," "oligonucleotide," "nucleic acid" and "nucleic acid molecule" include polydeoxyribonucleotides (containing 2- deoxy-D-ribose), polyribonucleotides (containing D-ribose), including tRNA, rRNA, hRNA, and mRNA, whether spliced or unspliced, any other type of polynucleotide which is an N- or C-glycoside of a purine or pyrimidine base, and other polymers containing nonnucleotidic backbones, for example, polyamide (e.g., peptide nucleic acids ("PNAs")) and polymorpholino (commercially available from the Anti-Virals, Inc., Corvallis, Oreg., as Neugene) polymers, and other synthetic sequence-specific nucleic acid polymers providing that the polymers contain nucleobases in a configuration which allows for base pairing and base stacking, such as is found in DNA and RNA. There is no intended distinction in length between the terms "polynucleotide," "oligonucleotide," "nucleic acid" and "nucleic acidAttorney Ref.: 39339-64210 Client Ref.: 003WO molecule," and these terms are used interchangeably herein. These terms refer only to the primary structure of the molecule. Thus, these terms include, for example, 3'-deoxy-2',5'- DNA, oligodeoxyribonucleotide N3'P5' phosphoramidates, 2'-O-alkyl-substituted RNA, double- and single-stranded DNA, as well as double- and single-stranded RNA, and hybrids thereof including for example hybrids between DNA and RNA or between PNAs and DNA or RNA, and also include known types of modifications, for example, labels, alkylation, "caps," substitution of one or more of the nucleotides with an analog, intemucleotide modifications such as, for example, those with uncharged linkages (e.g., methyl phosphonates, phosphotriesters, phosphoramidates, carbamates, etc.), with negatively charged linkages (e.g., phosphorothioates, phosphorodithioates, etc.), and with positively charged linkages (e.g., aminoalkylphosphoramidates, aminoalkylphosphotriesters), those containing pendant moieties, such as, for example, proteins (including enzymes (e.g. nucleases), toxins, antibodies, signal peptides, poly-L-lysine, etc.), those with intercalators (e.g., acridine, psoralen, etc.), those containing chelates (of, e.g., metals, radioactive metals, boron, oxidative metals, etc.), those containing alkylators, those with modified linkages (e.g., alpha anomeric nucleic acids, etc.), as well as unmodified forms of the polynucleotide or oligonucleotide.
[0104] It will be appreciated that, as used herein, the terms "nucleoside" and "nucleotide" will include those moieties which contain not only the known purine and pyrimidine bases, but also other heterocyclic bases which have been modified. Such modifications include methylated purines or pyrimidines, acylated purines or pyrimidines, halogenated purines or pyrimidines, or other heterocycles. Modified nucleosides or nucleotides can also include modifications on the sugar moiety, e.g., wherein one or more of the hydroxyl groups are replaced with halogen, aliphatic groups, or are functionalized as ethers, amines, or the like. The term "nucleotidic unit" is intended to encompass nucleosides and nucleotides.
[0105] Furthermore, modifications to nucleotidic units include rearranging, appending, substituting for or otherwise altering functional groups on the purine or pyrimidine base, which form hydrogen bonds to a respective complementary pyrimidine or purine. The resultant modified nucleotidic unit optionally may form a base pair with other such modified nucleotidic units but not with A, T, C, G or U. Abasic sites may be incorporated which do not prevent the function of the polynucleotide. Some or all of the residues in the polynucleotide can optionally be modified in one or more ways.Attorney Ref.: 39339-64210 Client Ref.: 003WO
[0106] Standard A-T and G-C base pairs form under conditions which allow the formation of hydrogen bonds between the N3-H and C4-oxy of thymidine and the N1 and C6-NH2, respectively, of adenosine and between the C2-oxy, N3 and C4-NH2, of cytidine and the C2- NH2, N'-H and C6-oxy, respectively, of guanosine. Thus, for example, guanosine (2-amino- 6-oxy-9-.beta.-D-ribo-furanosyl-purine) may be modified to form isoguanosine (2-oxy-6- amino-9-.beta.-D-ribofuranosyl-purine). Such modification results in a nucleoside base which will no longer effectively form a standard base pair with cytosine. However, modification of cytosine (1-.beta.-ribofuranosyl-2-oxy-4-amino-pyrimidine) to form isocytosine (1-.beta.-D- ribofuranosyl-2-amino-4-oxy-pyrimidine) results in a modified nucleotide which will not effectively base pair with guanosine but will form a base pair with isoguanosine (U.S. Pat. No.5,681,702 to Collins et al.). Isocytosine is available from Sigma Chemical Co. (St. Louis, Mo.); isocytidine may be prepared by the method described by Switzer et al. (1993) Biochem.32:10489-10496 and references cited therein; and isoguanine nucleotides may be prepared using the method described by Switzer et al. (1993), supra, and Mantsch et al. (1993) Biochem.14:5593-5601, or by the method described in U.S. Pat. No.5,780,610 to Collins et al. Other nonnatural base pairs may be synthesized by the method described in Piccirilli et al. (1990) Nature 343:33-37 for the synthesis of 2,6-diaminopyrimidine and its complement (1-methylpyrazolo-[4,3]pyrimidine-5,7-(4H,6H)-dione. Other such modified nucleotidic units which form unique base pairs are known, such as those described in Leach et al. (1992) J. Am. Chem. Soc.114:3675-3683 and Switzer et al. (1993), supra.
[0107] Aptamers are nucleic acid-based affinity agents. These affinity reagents typically bind targets with nanomolar or better affinities and have specificities comparable to monoclonal antibodies. However, unlike antibodies, aptamers are smaller, more chemically-defined, and can be prepared by chemical synthesis, facilitating downstream conjugation to drugs and peptides such as the cytotoxins and endosome disruptors proposed here. In vivo, aptamers can be safe, non-immunogenic, and efficacious.
[0108] Aptamers are generated by iterative rounds of in vitro selection. Briefly, randomized pools of RNA or ssDNA are incubated with target molecules (or mirror image target molecules) under carefully chosen selection conditions. Binding species are partitioned away from non-binders, amplified to generate a new pool, and the process is repeated until a desired ‘phenotype’ is achieved or until sequence diversity is significantly diminished. Multiple cycles of selection and amplification tend to winnow a pool of upwards of 1015molecules to only those few species that have the highest affinities and specificities for aAttorney Ref.: 39339-64210 Client Ref.: 003WO target. Whereas aptamers have traditionally been generated against soluble proteins or small molecule targets, aptamer selection techniques have more recently expanded to include cell surface targets, whole cells, and even in vivo selection procedures. Some aptamers selected against specific cell surface receptors have been successfully used for the delivery of cargoes into cells and are being investigated for in vivo targeting and delivery.
[0109] CRISPR refers to clustered regularly interspaced short palindromic repeats and includes a family of nucleic acid, e.g., DNA, sequences.
[0110] Ranges recited herein are understood to be shorthand for all of the values within the range, inclusive of the recited endpoints. For example, a range of 1 to 50 is understood to include any number, combination of numbers, or sub-range from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, and 50. 5.2. Mirror image DNA-cleaving and / or modifying compositions
[0111] As summarized above, the present disclosure provides mirror image nucleic acid- cleaving and / or modifying compositions including, e.g., mirror image CRISPR-Cas components. In some embodiments, the composition includes a mirror image CRISPR- associated endonuclease (e.g., DNA endonuclease) and a mirror image guide RNA that together provide for processing of target mirror image nucleic acids. In some embodiments, the composition further comprises a D-form DNA ligase (e.g., T4 ligase). Exemplary components of the compositions of this disclosure are now described in greater detail. 5.3. D-form endonuclease
[0112] As summarized above, the present disclosure provides compositions including mirror image endonucleases that can cleave target L-DNA or L-RNA. In some embodiments, the endonuclease is a mirror image DNA endonuclease. It is understood the mirror image DNA endonuclease is a D-protein.
[0113] In some embodiments, the D-form DNA endonuclease is a mirror image form of a L- form DNA endonuclease or a variant thereof. In some embodiments, the D-protein DNA endonuclease is a D-form Cas protein. Cas proteins generally include (1) at least one RNA recognition or RNA binding domain and (2) at least one nuclease domain. The RNA recognition or RNA binding domain can interact with guide RNAs (gRNAs, described in more detail below). The nuclease domains can possess catalytic activity for cleaving a nucleic acid molecule. Cleaving includes breaking the covalent bonds between nucleotides in aAttorney Ref.: 39339-64210 Client Ref.: 003WO nucleic acid molecule. Cleaving can produce blunt ends or staggered ends, and the Cas proteins can cleave single-stranded or double-stranded nucleic acid molecules. The nuclease domain can have endonuclease, exonuclease, or both endonuclease and exonuclease activity. Additionally, the nuclease domain can have 3’-5’ nuclease activity, 5’-3’ nuclease activity, or both. Cas proteins can additionally comprise DNA binding domains, helicase domains, protein-protein interaction domains, dimerization domains, and other domains. A mirror image Cas can interact with mirror image guide RNAs and can cleave target L-nucleic acid.
[0114] In some embodiments, the D-form Cas protein is a mirror image isomer of a class 1, type I, III, or IV Cas endonuclease or a functional derivative thereof.
[0115] In some embodiments, the D-form Cas protein is a mirror image isomer of a class 2, type II, V, or VI Cas endonuclease or a functional derivative thereof.
[0116] In some embodiments, the D-form Cas protein is a mirror image isomer of a L-Cas protein selected from Cas3, Cas9, Cas10, Cas12, Cas13, Cas14 and functional derivatives thereof. Examples of D-form Cas proteins which can be prepared in mirror image form include, but are not limited to, Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5e (CasD), Cas6, Cas6e, Cas6f, Cas7, Cas8a1, Cas8a2, Cas8b, Cas8c, Cas9 (Csn1 or Csx12), Cas10, Cas10d, CasF, CasG, CasH, Csy1, Csy2, Csy3, Cse1 (CasA), Cse2 (CasB), Cse3 (CasE), Cse4 (CasC), Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Csn2, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, and Cu1966, and homologs or modified versions or functional derivatives thereof.
[0117] In some embodiments, the D-protein DNA endonuclease recognizes a protospacer adjacent motif (PAM) having the L-form sequence NGG or NNGG, wherein N is any L-form nucleotide, or a functional derivative thereof. In some embodiments, the D-protein DNA endonuclease is a type II Cas endonuclease or a functional derivative thereof. In some embodiments, the DNA endonuclease is Cas9. In some embodiments, the Cas9 is from Streptococcus pyogenes (spCas9). In some embodiments, the Cas9 is from Staphylococcus lugdunensis (SluCas9).
[0118] In some embodiments, the D-form Cas protein comprises: i) a sequence selected from SEQ ID NOs: 1-19; or ii) a sequence having at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to the sequence set forth in i).
[0119] In some embodiments, the D-form Cas protein comprises:Attorney Ref.: 39339-64210 Client Ref.: 003WO i) a sequence selected from SEQ ID NOs: 2, 4, 6, 8, 10-12, 14, and 16-19; or ii) a sequence having at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to the sequence set forth in i).
[0120] In some embodiments, the D-form Cas protein has a sequence corresponding to a L- Cas protein with one or more variations, wherein the one or more variations comprise a substitution from isoleucine to a suitable neutral amino acid, such as alanine, leucine or valine.
[0121] Described herein in the experimental section are exemplary D-form Cas proteins and methods and synthetic precursors for preparing the same. Native chemical ligation (NCL) is used to prepare large polypeptides by the assembling of two or more unprotected peptides segments. In native chemical ligation, the thiol group of an N-terminal cysteine residue of an unprotected peptide attacks the C-terminal thioester of a second unprotected peptide. This reversible transthioesterification step is chemoselective and regioselective and leads to form a thioester intermediate. This intermediate rearranges by an intramolecular S,N-acyl shift that results in the formation of a native amide (peptide) bond at the ligation site. Native chemical ligation is performed using fragments having C-terminal hydrazides, e.g., according to the procedure described by Zheng et al. (Chemical synthesis of proteins using peptide hydrazides as thioester surrogates. Nat Protoc 8, 2483–2495 (2013)) or Zuo et al. (Total Chemical Synthesis of Proteins (2021): 87-118). 5.3.1. D-form Cas precursors and synthesis
[0122] This disclosure provides synthetic precursors of D-form Cas proteins and methods and technology for preparing such proteins.
[0123] Also provided are D-form Cas proteins having a mirror image sequence corresponding to a L-Cas protein, or synthetic precursor thereof, but with one or more sequence variations. In some embodiments, the variations are related to the particular synthetic strategy involved with preparing such D-proteins via native chemical ligation (NCL). In some embodiments, the one or more variations include a substitution from D- alanine to D-cysteine.
[0124] In some embodiments, the one or more variations include a substitution from D- isoleucine to a neutral D-amino acid residue. In some embodiments, the one or more variations include a substitution from D-isoleucine to D-alanine, D-leucine, D-phenylalanine, D-valine, D-threonine, or D-tyrosine. In some embodiments, the one or more variations include a substitution from D-isoleucine to alanine, leucine, or valine.Attorney Ref.: 39339-64210 Client Ref.: 003WO
[0125] In some embodiments, the D-Cas protein is a class 1, type I, III, or IV Cas endonuclease or a functional derivative thereof. In some embodiments, the D-form Cas protein is a class 2, type II, V, or VI Cas endonuclease or a functional derivative thereof. In some embodiments, the D-Cas protein is selected from Cas3, Cas9, Cas10, Cas12, Cas13, Cas14 and functional derivatives thereof.
[0126] In some embodiments, the D-form Cas protein comprises: i) a sequence selected from SEQ ID NOs: 1-19 including one or more variations which include a substitution from D-isoleucine to a neutral D-amino acid residue; or ii) a sequence having at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to the sequence set forth in i) including one or more variations which include a substitution from D-isoleucine to a neutral D-amino acid residue.
[0127] In some embodiments, the neutral D-amino acid residue is D-alanine, D-leucine, D- phenylalanine, D-valine, D-threonine, or D-tyrosine.
[0128] In some embodiments, the D-form Cas protein comprises: i) a sequence selected from SEQ ID NOs: 2, 4, 6, 8, 10-12, 14, and 16-19 including one or more variations which include a substitution from D-isoleucine to a neutral D-amino acid residue; or ii) a sequence having at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to the sequence set forth in i) including one or more variations which include a substitution from D-isoleucine to a neutral D-amino acid residue.
[0129] In some embodiments, the neutral D-amino acid residue is D-alanine, D-leucine, D- phenylalanine, D-valine, D-threonine, or D-tyrosine. 5.3.1. D-form Cas12
[0130] In some embodiments, the D-form Cas protein is a D-form Cas12 (also referred to as D-Cas12 or mirror image Cas12).
[0131] In some embodiments, the D-form Cas12 has 90% or greater (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to one of sequences SEQ ID NOs: 9-12. In some embodiments, the D-form polypeptide can be prepared via native chemical ligation (NCL) of two or more precursor fragments between one or more positions defined by the following pairs of amino acid residues, e.g., corresponding to a mirror image Cas12 sequence such as one of sequences SEQ ID NOs: 9-12: i) between tyr61and cys62; ii) between asn104and cys105;Attorney Ref.: 39339-64210 Client Ref.: 003WO iii) between asn156and cys157; iv) between thr187and cys188; v) between lys227and cys228; vi) between leu283and cys284; vii) between thr348and cys349; viii) between asn387and cys388; ix) between ala448and cys449; x) between ser497and cys498; xi) between leu540and cys541; xii) between ala601and cys602; xiii) between tyr659and cys660; xiv) between tyr711and cys712; xv) between phe750and cys751; xvi) between met798and cys799; xvii) between arg843and cys844; xviii) between asn893and cys894; xix) between ala953and cys954; xx) between lys1009and cys1010; xxi) between pro1066and cys1067; xxii) between pro1134and cys1135; xxiii) between pro1178and cys1179; xxiv) between ala1230and cys1231; and xxv) between Gly1266and cys1267.
[0132] Accordingly, this disclosure provides synthetic precursors of a D-form Cas12 that are D-polypeptide hydrazide fragments selected from one of the following segments of a D- Cas12, e.g., where the underlying sequence of the fragment can correspond to any one of sequences SEQ ID NOs: 9-12: a. met1-tyr61-NHNH2 b. cys62-asn104-NHNH2 (cys(Acm)65) c. cys105-asn156-NHNH2 d. thz157-thr187-NHNH2 e. cys188-lys227-NHNH2(cys(Acm)205) f. cys228-leu283-NHNH2Attorney Ref.: 39339-64210 Client Ref.: 003WO g. thz284-thr348-NHNH2 (cys(Acm)334) h. cys349-asn387-NHNH2(cys(Acm)379) i. cys388-ala448-NHNH2 j. thz449-ser497-NHNH2k. cys498-leu540-NHNH2 l. thz541-ala601-NHNH2m. cys602-tyr659-NHNH2 (cys(Acm)608) n. thz660-tyr711-NHNH2(cys (Acm)674) o. cys712-phe750-NHNH2 p. cys751-met798-NHNH2 q. thz799-arg843-NHNH2r. cys844-asn893-NHNH2 s. cys894-ala953-NHNH2t. thz954-lys1009-NHNH2 u. cys1010-pro1066-NHNH2(cys (Acm)1025) v. cys1067-pro1134-NHNH2 w. thz1135-pro1178-NHNH2x. cys1179-ala1230-NHNH2 y. thz1231-Gly1266-NHNH2(cys(Acm)1248), and z. cys1267-asn1307-OH.
[0133] In some embodiments, the D-polypeptide hydrazides having an underlying sequence corresponding to a segment of SEQ ID NOs: 9-12. FIGs.1A-1G illustrate an exemplary convergent fragment condensation strategy using NCL for preparing a D-Cas12. Example 1 of the experimental section provides further details of exemplary procedures for accomplishing the same.
[0134] In some embodiments, the D-Cas12 polypeptide is prepared via NCL coupling of two precursor fragments between positions tyr659and cys660. Dialysis and refolding of the D- Cas12 polypeptide after synthesis can provide a functional endonuclease.
[0135] In some embodiments, the D-polypeptide is prepared via NCL coupling of: i) met1-tyr659-NHNH2 (cys(Acm)65, 205, 334, 379, 608); and ii) cys660-asn1307-OH (cys
[0136] In some embodiments,further includes desulfurizing one or more cysteine residues of a precursor D-polypeptide to alanine residuesAttorney Ref.: 39339-64210 Client Ref.: 003WO at one or more positions selected from 62, 105, 157, 188, 228, 284, 349, 388, 449, 498, 541, 602, 660, 712, 751, 799, 844, 894, 954, 1010, 1067, 1135, 1179, 1231, and 1267. See e.g., Example 1.
[0137] In some embodiments, the D-form Cas12 comprises: a. a sequence having at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to the sequence of SEQ ID NO: 10; and b. one or more variations selected from: A62C, A105C, A157C, A188C, A228C, A284C, A349C, A388C, A449C, A498C, A541C, A602C, A660C, A712C, A801C, A799C, A844C, A894C, A954C, A1010C, A1067C, A1135C, A1179C, A1131C, and A1267C.
[0138] In some embodiments, the D-form Cas12 comprises: a. a sequence having at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to the sequence of SEQ ID NO: 10; and b. one or more variations selected from: I37L; I53V; I57Y; A62C; I96A; A105C; I111T; I121F; I128A; A157C; A188C; I207A; I212V; A228C; I229A; I231A; I237L; I265V; A284C; I285A; I295L; I303A; I330L; A349C; I359A; I374V; A388C; I394A; I401A; I418F; I423L; I424A; I442V; A449C; I466A; A498C; I503A; A541C; A602C; I626Y; I633A; A660C; I693Y; A712C; I726V; I731V; A751C; A799C; A844C; I850V; A894C; I915L; I917A; I935V; A954C; I984V; A1010C; A1067C; A1135C; I1138V; I1164A; A1179C; I1183A; I1199V; I1212F; I1219F; A1231C; A1267C; and I1302V. 5.3.2. D-form Cas14
[0139] In some embodiments, the D-form Cas protein is a D-form Cas14 (also referred to as D-Cas14 or mirror image Cas14).
[0140] In some embodiments, the D-form Cas14 has 90% or greater (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to one of sequences SEQ ID NOs: 15- 19. In some embodiments, the D-form polypeptide can be prepared via native chemical ligation (NCL) of two or more precursor fragments between one or more positions defined by the following pairs of amino acid residues, e.g., corresponding to a mirror image Cas14 sequence such as one of sequences SEQ ID NOs: 15-19: i) between gln37and cys38; ii) between his82and cys83; iii) between arg110and cys111;Attorney Ref.: 39339-64210 Client Ref.: 003WO iv) between Gly161and cys162; v) between trp225and cys226; vi) between phe263and cys264; vii) between phe319and cys320; viii) between trp363and cys364; ix) between gln427and cys428; and x) between ala475and cys476.
[0141] Accordingly, this disclosure provides synthetic precursors of a D-form Cas14 that are D-polypeptide hydrazide fragments selected from one of the following segments of a D- Cas14, e.g., where the underlying sequence of the fragment can correspond to any one of sequences SEQ ID NOs: 15-19: a. met1-gln37-NHNH2; b. cys38-his82-NHNH2; c. thz83-arg110-NHNH2; d. cys111-Gly161-NHNH2(Cys(Acm)125); e. thz162-trp225-NHNH2; f. cys226-phe263-NHNH2; g. thz264-phe319-NHNH2 (Cys(Acm)295); h. cys320-trp363-NHNH2; i. thz364-gln427-NHNH2; j. thz428-ala475-NHNH2(cys(Acm)435, 440, 462, 465); and k. cys476-lys500-OH.
[0142] In some embodiments, the D-polypeptide hydrazides having an underlying sequence corresponding to a segment of SEQ ID NOs: 15-19. FIGs.2A-2C illustrate an exemplary convergent fragment condensation strategy using NCL for preparing a D-Cas14. Example 2 of the experimental section provides further details of exemplary procedures for accomplishing the same.
[0143] In some embodiments, the D-Cas14 polypeptide is prepared via NCL coupling of two precursor fragments between positions phe263and cys264. Dialysis and refolding of the D- Cas14 polypeptide after synthesis can provide a functional endonuclease.
[0144] In some embodiments, the D-Cas14 polypeptide is prepared via NCL coupling of: met1-phe263-NHNH2(cys(Acm)125); and cys264-lys500-OH (cys(Acm)295, 435, 440, 462, 465).Attorney Ref.: 39339-64210 Client Ref.: 003WO
[0145] In some embodiments, preparing the D-Cas14 polypeptide further includes desulfurizing one or more cysteine residues of a precursor D-polypeptide fragment to alanine residues at one or more positions selected from 38, 83, 111, 162, 226, 264, 320, 364, 428, and 476. See e.g., Example 2.
[0146] In some embodiments, the D-form Cas14 has 90% or greater (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to one of sequences SEQ ID NOs: 15- 19. In some embodiments, the D-form polypeptide can be prepared via native chemical ligation (NCL) of two or more precursor fragments between one or more positions defined by the following pairs of amino acid residues, e.g., corresponding to a mirror image Cas14 sequence such as one of sequences SEQ ID NOs: 15-19: i) between ala76and cys77; ii) between leu124and cys125; iii) between Gly161and cys162; iv) between trp225and cys226; v) between ala294and cys295; vi) between trp363and cys364; and vii) between leu434and cys435.
[0147] Accordingly, this disclosure provides synthetic precursors of a D-form Cas14 that are D-polypeptide hydrazide fragments selected from one of the following segments of a D- Cas14, e.g., where the underlying sequence of the fragment can correspond to any one of sequences SEQ ID NOs: 15-19: a. met1-ala76-NHNH2; b. cys77-leu124-NHNH2; c. thz125-Gly161-NHNH2; d. thz162-trp225-NHNH2; e. cys226-ala294-NHNH2; f. thz295-trp363-NHNH2; g. cys364-leu434-NHNH2; and h. cys435-lys500-OH.
[0148] In some embodiments, the D-polypeptide hydrazides having an underlying sequence corresponding to a segment of SEQ ID NOs: 15-19. FIG.5 illustrates an exemplary convergent fragment condensation strategy using NCL for preparing a D-Cas14. Example 3Attorney Ref.: 39339-64210 Client Ref.: 003WO of the experimental section provides further details of exemplary procedures for accomplishing the same.
[0149] In some embodiments, the D-Cas14 polypeptide is prepared via NCL coupling of two precursor fragments between positions ala294and cys295. Dialysis and refolding of the D- Cas14 polypeptide after synthesis can provide a functional endonuclease.
[0150] In some embodiments, the D-Cas14 polypeptide is prepared via NCL coupling of: met1-ala294-NHNH2; and cys295-lys500-OH.
[0151] In some embodiments, preparing the D-Cas14 polypeptide further includes desulfurizing one or more cysteine residues of a precursor D-polypeptide fragment to alanine residues at one or more positions selected from 77, 125, 162, 226, 295, 364, and 435. See e.g., Example 3.
[0152] In some embodiments, the D-form Cas14 comprises: a. a sequence having at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to the sequence of SEQ ID NO: 16; and b. one or more variations selected from: A38C; A83C; A111C; A162C; A226C; A264C; A364C; A428C; and A476C.
[0153] In some embodiments, the D-form Cas14 comprises: a. a sequence having at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to the sequence of SEQ ID NO: 16; and b. one or more variations selected from: I15V; I24V; I28V; A38C; I76A, A83C, I92V, I103L, I106A, A111C, I124L, A162C, A226C, I248V, A264C, I272A, I281A, I289A, I297V, I345A, A364C, A428C, A476C, and I484L.
[0154] In some embodiments, the D-form Cas14 comprises: a. a sequence having at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to the sequence of SEQ ID NO: 16; and b. one or more variations selected from: I15V; I24V; I28V; I76A; I92V; I103L; I106A; I124L; I248V; I272A; I281A; I289A; V294A; I297V; I345A; and I484L. 5.4. L-form guide RNA
[0155] In some embodiments, the mirror image nucleic acid-cleaving and / or modifying composition includes a mirror image guide RNA (also referred to as L-form guide RNA or L- guide RNA).Attorney Ref.: 39339-64210 Client Ref.: 003WO
[0156] A “guide RNA” or “gRNA” includes an RNA molecule that binds to a Cas protein and targets the Cas protein to a specific location within a target nucleic acid molecule. Guide RNAs can comprise two segments: a “nucleic acid-targeting segment” (e.g., crRNA) and a “protein-binding segment” (e.g., tracrRNA). “Segment” includes a segment, section, or region of a molecule, such as a contiguous stretch of nucleotides in an RNA. The nucleic acid-targeting segment of the gRNA is the segment that hybridizes to a target sequence in the target nucleic acid molecule. The nucleic acid-targeting segment can be partially or completely complementarily to the target sequence in the target nucleic acid molecule. The protein-binding segment of the gRNA is the segment that binds to the Cas protein. In some cases, the nucleic acid-targeting segment and the protein-binding segment are located in one gRNA molecule. In some cases, the nucleic acid-targeting segment and the protein-binding segment are located in two separate gRNA molecules. The two separate gRNA molecules can each have a segment of homology that allow the two separate gRNA molecules to hybridize to each other and guide the Cas protein to the target sequence. The two separate gRNA molecules can include: an “activator-RNA” and a “targeter-RNA.” The targeter-RNA can have a nucleic acid-targeting segment. The activator-RNA can have a protein-binding segment. The activator-RNA and the targeter-RNA can each have a segment of homology that allows that activator-RNA and the targeter-RNA to hybridize together and guide the Cas protein to the target sequence. Other gRNAs are a single RNA molecule (single RNA polynucleotide), which can also be called a “single-molecule gRNA,” a “single-guide RNA,” or an “sgRNA.” See, e.g., WO / 2013 / 176772A1, WO / 2014 / 065596A1, WO / 2014 / 089290A1, WO / 2014 / 093622A2, WO / 2014 / 099750A2, WO / 2013142578A1, and WO 2014 / 131833A1, each of which is herein incorporated by reference in its entirety for all purposes. The terms “guide RNA” and “gRNA” are inclusive, including both double-molecule gRNAs and single- molecule gRNAs. This disclosure is meant to include mirror image forms (also referred to as L-forms) of any of these guide RNAs which are readily available.
[0157] The L-form guide RNA can direct an associated D-form DNA endonuclease to a particular target nucleotide sequence of interest through hybridization of the guide RNA to the target L-DNA sequence. A target L-DNA sequence can include L-form DNA, L-form RNA, or a combination of both and can be single-stranded or double-stranded.
[0158] A targeting segment of a L-form gRNA interacts with a target L-form DNA in a sequence-specific manner via hybridization (i.e., base pairing). As such, the nucleotide sequence of the L-form DNA-targeting segment may vary and determines the location withinAttorney Ref.: 39339-64210 Client Ref.: 003WO the target L-form DNA with which the gRNA and the target L-form DNA will interact (e.g., hybridize). The L-DNA-targeting segment of a subject L-form gRNA can be modified to hybridize to any desired sequence within a target L-form DNA. Naturally occurring crRNAs differ depending on the CRISPR-Cas9 system and organism but often contain a targeting segment of between 21 to 72 nucleotides length, flanked by two direct repeats (DR) of a length of between 21 to 46 nucleotides (see, e.g., WO2014 / 131833, herein incorporated by reference in its entirety for all purposes). In the case of S. pyogenes, the DRs are 36 nucleotides long and the targeting segment is 30 nucleotides long. The DR located 3′ of the targeting segment is complementary to and hybridizes with the corresponding tracrRNA, which in turn binds to the Cas9 protein. Mirror image forms of such sequences are meant to be encompassed by this disclosure.
[0159] The L-DNA-targeting segment of a L-form gRNA can have a length of from about 12 nucleotides to about 100 nucleotides. For example, the L-DNA-targeting segment can have a length of from about 12 nucleotides (nt) to about 80 nt, from about 12 nt to about 50 nt, from about 12 nt to about 40 nt, from about 12 nt to about 30 nt, from about 12 nt to about 25 nt, from about 12 nt to about 20 nt, or from about 12 nt to about 19 nt. Alternatively, the DNA- targeting segment can have a length of from about 19 nt to about 20 nt, from about 19 nt to about 25 nt, from about 19 nt to about 30 nt, from about 19 nt to about 35 nt, from about 19 nt to about 40 nt, from about 19 nt to about 45 nt, from about 19 nt to about 50 nt, from about 19 nt to about 60 nt, from about 19 nt to about 70 nt, from about 19 nt to about 80 nt, from about 19 nt to about 90 nt, from about 19 nt to about 100 nt, from about 20 nt to about 25 nt, from about 20 nt to about 30 nt, from about 20 nt to about 35 nt, from about 20 nt to about 40 nt, from about 20 nt to about 45 nt, from about 20 nt to about 50 nt, from about 20 nt to about 60 nt, from about 20 nt to about 70 nt, from about 20 nt to about 80 nt, from about 20 nt to about 90 nt, or from about 20 nt to about 100 nt.
[0160] A L-form guide RNA can have at least a spacer sequence that hybridizes to a target L- nucleic acid sequence of interest and a mirror image CRISPR repeat sequence. In Type II CRISPR-Cas systems, the L-form gRNA also has a second L-form RNA called the tracrRNA sequence. In the Type II CRISPR-Cas systems L-form guide RNA (L-form gRNA), the mirror image CRISPR repeat sequence in the crRNA and L-form tracrRNA sequence hybridize to each other to form a duplex. In the Type V CRISPR-Cas systems L-form guide RNA (gRNA), the L-form crRNA forms a duplex with the target nucleic acid molecule. InAttorney Ref.: 39339-64210 Client Ref.: 003WO both systems, the L-form duplex binds a site-directed mirror image polypeptide such that the L-form guide RNA and site-direct polypeptide form a complex.
[0161] Site-specific cleavage of target L-DNA by D-form Cas can occur at locations determined by both (i) base-pairing complementarity between the L-form gRNA and the target L-DNA and (ii) the presence of a short motif, called the protospacer adjacent motif (PAM), in the target L-DNA. The Cas protein may only cleave the target nucleic acid molecule when the gRNA is hybridized to the target sequence and the Cas protein recognizes the PAM sequence. The PAM can flank a CRISPR L-form RNA recognition sequence. In some cases, the PAM sequence is 5’ of the target sequence in the target nucleic acid molecule. In some cases, the PAM sequence is 3’ of the target sequence in the target nucleic acid molecule. Optionally, the CRISPR L-form RNA recognition sequence can be flanked on the 3′ end by the PAM. The PAM sequence recognized by the Cas protein is specific for each Cas protein. For example, the PAM sequence of Cas9 is 5’-NGG-3’; N can refer to any nucleotide base).
[0162] In some embodiments, the D-protein DNA endonuclease recognizes a protospacer adjacent motif (PAM) having the L-form sequence NGG or NNGG, wherein N is any L-form nucleotide, or a functional derivative thereof. In some embodiments, the D-protein DNA endonuclease is a type II Cas endonuclease or a functional derivative thereof. In some embodiments, the DNA endonuclease is Cas9. In some embodiments, the Cas9 is from Streptococcus pyogenes (spCas9). In some embodiments, the Cas9 is from Staphylococcus lugdunensis (SluCas9).
[0163] Guide RNA, CRISPR components, and target sequences that can be adapted for use in the nanostructures and compositions of this disclosure include those sequences and materials described in US 10,457,960; WO2020081843; Jinek et al. (Science, 337, 816-821 (2012)); and Deltcheva et al. (Nature.471, 602- 607 (2011)), and mirror image versions of the same.
[0164] The L-form guide RNA can be chemically synthesized using methods of solid phase oligonucleotide synthesis readily available in the art.
[0165] In some embodiments, the target L-DNA is single stranded. In some embodiments, the target L-DNA is double stranded. In some embodiments, the target L-DNA can include a duplex segment that includes the target sequence to which the L-guide RNA is directed.
[0166] In some embodiments, the target L-DNA is comprised within a folded L-DNA nanostructure. In some embodiments, folded L-DNA nanostructure includes the duplex sequence that comprises: a first trigger L-DNA strand comprising a target cleavage site,Attorney Ref.: 39339-64210 Client Ref.: 003WO wherein the first trigger strand is linked to a first L-DNA domain; and a complementary second trigger L-DNA strand linked to a second L-DNA domain; wherein L-DNA cleavage at the target cleavage site results in a conformational rearrangement of the first and second L- DNA domains. 5.5. D-form DNA ligase
[0167] The present disclosure further provides a mirror image DNA ligase (also referred to as D-form DNA ligase or D-DNA ligase). The D-form DNA ligase can be used by itself or together with the D-form endonuclease disclosed herein. It is understood the mirror image DNA ligase is a D-protein.
[0168] In some embodiments, the D-form DNA ligase is a mirror image protein of T4 ligase (also referred to as D-form T4 ligase or D-T4 ligase).
[0169] In some embodiments, the D-form DNA ligase has a sequence of a L-T4 ligase with one or more variations. In some embodiments, the one or more variations comprise a substitution from alanine to cysteine. In some embodiments, the one or more variations comprise a substitution from isoleucine to alanine, leucine, or valine. In some embodiments, the one or more variations comprise a substitution from valine to alanine.
[0170] In some embodiments, the D-form DNA ligase comprises: i) a sequence selected from SEQ ID NOs: 20-22; or ii) a sequence having at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to the sequence set forth in i).
[0171] In some embodiments, the D-form DNA ligase comprises: i) a sequence having at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to the sequence of SEQ ID NO: 21; and ii) a substitution of V317A.
[0172] In some embodiments, the D-form T4 ligase has 90% or greater (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to one of sequences SEQ ID NOs: 20- 22. In some embodiments, the D-form T4 ligase can be prepared via native chemical ligation (NCL) of two or more precursor fragments between one or more positions defined by the following pairs of amino acid residues, e.g., corresponding to a mirror image T4 ligase sequence such as one of sequences SEQ ID NOs: 20-22: i) between leu76and cys77; ii) between Gly116and cys117; iii) between arg164and cys165;Attorney Ref.: 39339-64210 Client Ref.: 003WO iv) between thr201and cys202; v) between thr255and cys256; vi) between lys316and cys317; vii) between lys388and cys389; and viii) between lys412and cys413.
[0173] Accordingly, this disclosure provides synthetic precursors of a D-form T4 ligase that are D-polypeptide hydrazide fragments selected from one of the following segments of a D- form T4 ligase, e.g., where the underlying sequence of the fragment can correspond to any one of sequences SEQ ID NOs: 20-22: a. met1- leu76-NHNH2 ; b. cys77-Gly116-NHNH2(Cys(Acm)115); c. cys117-arg164-NHNH2 ; d. thz165-thr201-NHNH2; e. cys202-thr255-NHNH2 ; f. thz256-lys316-NHNH2(Cys(Acm)276); g. cys317-lys388-NHNH2 ; h. thz389-lys412-NHNH2(Cys(Acm)398); and i. cys413-leu487-OH (Cys(Acm)439, 441).
[0174] In some embodiments, the D-polypeptide hydrazides having an underlying sequence corresponding to a segment of SEQ ID NOs: 20-22. FIG.13 illustrates an exemplary convergent fragment condensation strategy using NCL for preparing a D-form T4 ligase. Example 4 of the experimental section provides further details of exemplary procedures for accomplishing the same.
[0175] In some embodiments, the D-form T4 ligase is prepared via NCL coupling of two precursor fragments between positions arg164and cys165. Dialysis and refolding of the D- form T4 ligase after synthesis can provide a functional endonuclease.
[0176] In some embodiments, the D-form T4 ligase is prepared via NCL coupling of: met1-arg164-NHNH2 (Cys(Acm)115); and cys165-leu487-OH.
[0177] In some embodiments, preparing the D-form T4 ligase further includes desulfurizing one or more cysteine residues of a precursor fragment to alanine residues at one or more positions selected from 77, 117, 165, 202, 256, 317, 389, and 413. See e.g., Example 4. 5.6. Methods of useAttorney Ref.: 39339-64210 Client Ref.: 003WO
[0178] Provided herein are methods of using the mirror image DNA-cleaving and / or modifying compositions to site specifically cleave a target L-DNA. As described herein, the mirror image DNA-cleaving and / or modifying composition can include a L-form guide RNA that directs the D-form DNA endonuclease to a particular site of a target L-DNA sequence.
[0179] In some embodiments, the method is a method of detecting a target L-DNA in a sample. In some embodiments, the method is a method of cleaving and / or modifying a target L-DNA. The method can include: contacting the mirror image DNA-cleaving and / or modifying composition (e.g., as described herein) with a target L-DNA; wherein a guide RNA of the composition hybridizes to the target L-DNA sequence, thereby directing a D- form DNA endonuclease to bind to said target L-DNA sequence and cleave and / or modify the target L-DNA sequence.
[0180] In some embodiments, the method is a method of modifying a target L-DNA. The method can include: contacting the mirror image DNA-cleaving and / or modifying composition (e.g., as described herein) with a target L-DNA; wherein a guide RNA of the composition hybridizes to the target L-DNA sequence, thereby directing a D-form DNA endonuclease to bind to said target L-DNA sequence and cleave and / or modify the target L- DNA sequence; and wherein a D-form DNA ligase ligates the cleaved and / or modified target L-DNA sequence.
[0181] In some embodiments, the method is a method of editing a target L-DNA sequence. The method of editing can include: interacting the mirror image DNA-cleaving and / or modifying composition (e.g., as described herein) and a target L-DNA in a mixture; and incubating the mixture thereby inducing cleavage and / or modification of the target L-DNA.
[0182] In some embodiments, the method is a method of editing a target L-DNA sequence. The method of editing can include: interacting the mirror image DNA-cleaving and / or modifying composition (e.g., as described herein) and a target L-DNA in a mixture; and incubating the mixture thereby inducing modification and ligation of the target L-DNA.
[0183] The L-form guide RNA of the compositions and methods can be a single guide RNA or a dual-guide RNA.
[0184] In some embodiments, the target L-DNA sequence is adjacent to a protospacer adjacent motif (PAM). In certain embodiments, cleavage of a double-stranded targets L-DNA sequence is dependent upon the presence of a PAM. In certain embodiments, cleavage of a single-stranded target sequence is PAM-independent. A protospacer adjacent motif is generally within about 1 to about 10 nucleotides from the target L-DNA sequence, includingAttorney Ref.: 39339-64210 Client Ref.: 003WO about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, or about 10 nucleotides from the target nucleotide sequence. The PAM can be 5′ or 3′ of the target L- DNA sequence. In some embodiments, the PAM is 5′ of the target L-DNA sequence. Generally, the PAM is a consensus sequence of about 3-4 nucleotides, but in particular embodiments, can be 1, 2, 3, 4, 5, 6, 7, 8, 9, or more nucleotides in length. In some embodiments, the PAM is 5′ of the target sequence and is T-rich. 5.7. Method of making
[0185] In yet another aspect, the present disclosure provides a method of making a D-form endonuclease, L-form guide RNA, or D-form DNA ligase. In some embodiments, the D- form endonuclease is a D-form DNA endonuclease. In some embodiments, the D-form DNA endonuclease is a D-form Cas12 or a D-form Cas14. In some embodiments, the D-form DNA ligase is a D-form T4 ligase.
[0186] The present disclosure also provides a D-form endonuclease, L-form guide RNA, or D-form DNA ligase generated by a method described herein.
[0187] In some embodiments, the method is a method of making a D-form Cas12 having 90% or greater sequence identity to SEQ ID NOs: 10-12. In some embodiments, the method comprises coupling of two precursor fragments between positions tyr659and cys660via native chemical ligation (NCL). In some embodiments, the two precursor fragments are met1-tyr659- NHNH2(cys(Acm)65, 205, 334, 379, 608) and cys660-asn1307-OH (cys(Acm)674, 1025, 1248). In some embodiments, the D-form Cas12 is prepared via native chemical ligation (NCL) couplings of precursor fragments between positions defined by the following pairs of amino acid residues: i) between tyr61and cys62; ii) between asn104and cys105; iii) between asn156and cys157; iv) between thr187and cys188; v) between lys227and cys228; vi) between leu283and cys284; vii) between thr348and cys349; viii) between asn387and cys388; ix) between ala448and cys449; x) between ser497and cys498; xi) between leu540and cys541; xii) between ala601and cys602;Attorney Ref.: 39339-64210 Client Ref.: 003WO xiii) between tyr659and cys660; xiv) between tyr711and cys712; xv) between phe750and cys751; xvi) between met798and cys799; xvii) between arg843and cys844; xviii) between asn893and cys894; xix) between ala953and cys954; xx) between lys1009and cys1010; xxi) between pro1066and cys1067; xxii) between pro1134and cys1135; xxiii) between pro1178and cys1179; xxiv) between ala1230and cys1231; and xxv) between Gly1266and cys1267.
[0188] In some embodiments, the precursor fragments are derived from one or more of the following fragments: a. met1-tyr61-NHNH2 ; b. cys62-asn104-NHNH2(cys(Acm)65); c. cys105-asn156-NHNH2; d. thz157-thr187-NHNH2; e. cys188-lys227-NHNH2 (cys(Acm)205); f. cys228-leu283-NHNH2; g. thz284-thr348-NHNH2 (cys(Acm)334); h. cys349-asn387-NHNH2(cys(Acm)379); i. cys388-ala448-NHNH2; j. thz449-ser497-NHNH2; k. cys498-leu540-NHNH2; l. thz541-ala601-NHNH2; m. cys602-tyr659-NHNH2 (cys(Acm)608); n. thz660-tyr711-NHNH2 (cys(Acm)674); o. cys712-phe750-NHNH2; p. cys751-met798-NHNH2; q. thz799-arg843-NHNH2; r. cys844-asn893-NHNH2;Attorney Ref.: 39339-64210 Client Ref.: 003WO s. cys894-ala953-NHNH2; t. thz954-lys1009-NHNH2; u. cys1010-pro1066-NHNH2 (cys(Acm)1025); v. cys1067-pro1134-NHNH2; w. thz1135-pro1178-NHNH2; x. cys1179-ala1230-NHNH2; y. thz1231-Gly1266-NHNH2 (cys(Acm)1248); and z. cys1267-asn1307-OH.
[0189] In some embodiments, the method of making the D-form Cas12 having 90% or greater sequence identity to SEQ ID NOs: 10-12 further comprises desulfurizing one or more cysteine residues of a precursor fragment to alanine residues at one or more positions selected from 62, 105, 157, 188, 228, 284, 349, 388, 449, 498, 541, 602, 660, 712, 751, 799, 844, 894, 954, 1010, 1067, 1135, 1179, 1231, and 1267.
[0190] In some embodiments, the method is a method of making a D-form Cas14 having 90% or greater sequence identity to SEQ ID NOs: 16-19. In some embodiments, the method comprises coupling of two precursor fragments between positions phe263and cys264via native chemical ligation (NCL). In some embodiments, the two precursor fragments are met1-phe263- NHNH2 (cys(Acm)125) and cys264-lys500-OH (cys(Acm)295, 435, 440, 462, 465). In some embodiments, the D-form Cas14 is prepared via native chemical ligation (NCL) couplings of precursor fragments between positions defined by the following pairs of amino acid residues: i) between gln37and cys38; ii) between his82and cys83; iii) between arg110and cys111; iv) between Gly161and cys162; v) between trp225and cys226; vi) between phe263and cys264; vii) between phe319and cys320; viii) between trp363and cys364; ix) between gln427and cys428; and x) between ala475and cys476.
[0191] In some embodiments, the precursor fragments are derived from one or more of the following fragments: a. met1-gln37-NHNH2;Attorney Ref.: 39339-64210 Client Ref.: 003WO b. cys38-his82-NHNH2; c. thz83-arg110-NHNH2; d. cys111-Gly161-NHNH2 (Cys(Acm)125); e. thz162-trp225-NHNH2; f. cys226-phe263-NHNH2; g. thz264-phe319-NHNH2(Cys(Acm)295); h. cys320-trp363-NHNH2; i. thz364-gln427-NHNH2; j. thz428-ala475-NHNH2 (cys(Acm)435, 440, 462, 465); and a. cys476-lys500-OH.
[0192] In some embodiments, the method of making the D-form Cas14 having 90% or greater sequence identity to SEQ ID NOs: 16-19 further comprises desulfurizing one or more cysteine residues of a precursor fragment to alanine residues at one or more positions selected from 38, 83, 111, 162, 226, 264, 320, 364, 428, and 476.
[0193] In some embodiments, the method is a method of making a D-form Cas14 having 90% or greater sequence identity to SEQ ID NOs: 16-19. In some embodiments, the method comprises coupling of two precursor fragments between positions ala294and cys295via native chemical ligation (NCL). In some embodiments, the two precursor fragments are met1-ala294- NHNH2and cys295-lys500-OH. In some embodiments, the D-form Cas14 is prepared via native chemical ligation (NCL) couplings of precursor fragments between positions defined by the following pairs of amino acid residues: i) between ala76and cys77; ii) between leu124and cys125; iii) between Gly161and cys162; iv) between trp225and cys226; v) between ala294and cys295; vi) between trp363and cys364; and vii) between leu434and cys435.
[0194] In some embodiments, the precursor fragments are derived from one or more of the following fragments: a. met1- ala76-NHNH2 ; b. cys77-leu124-NHNH2; c. thz125-Gly161-NHNH2 ;Attorney Ref.: 39339-64210 Client Ref.: 003WO d. thz162-trp225-NHNH2 ; e. cys226-ala294-NHNH2; f. thz295-trp363-NHNH2 ; g. cys364-leu434-NHNH2; and h. cys435-lys500-OH.
[0195] In some embodiments, the method of making the D-form Cas14 having 90% or greater sequence identity to SEQ ID NOs: 16-19 further comprises desulfurizing one or more cysteine residues of a precursor fragment to alanine residues at one or more positions selected from 77, 125, 162, 226, 295, 364, and 435.
[0196] In some embodiments, the method is a method of making a D-form T4 ligase having 90% or greater sequence identity to SEQ ID NO 21 or SEQ ID NO 22. In some embodiments, the method comprises coupling of two precursor fragments between positions arg164and cys165via native chemical ligation (NCL). In some embodiments, the two precursor fragments are met1-arg164-NHNH2 (Cys(Acm)115) and cys165-leu487-OH. In some embodiments, the D-form T4 ligase is prepared via native chemical ligation (NCL) couplings of precursor fragments between positions defined by the following pairs of amino acid residues: i) between leu76and cys77; ii) between Gly116and cys117; iii) between arg164and cys165; iv) between thr201and cys202; v) between thr255and cys256; vi) between lys316and cys317; vii) between lys388and cys389; and viii) between lys412and cys413.
[0197] In some embodiments, the precursor fragments are derived from one or more of the following fragments: a. met1- leu76-NHNH2 ; b. cys77-Gly116-NHNH2 (Cys(Acm)115); c. cys117-arg164-NHNH2 ; d. thz165-thr201-NHNH2 ; e. cys202-thr255-NHNH2; f. thz256-lys316-NHNH2 (Cys(Acm)276);Attorney Ref.: 39339-64210 Client Ref.: 003WO g. cys317-lys388-NHNH2 ; h. thz389-lys412-NHNH2(Cys(Acm)398); and i. cys413-leu487-OH (Cys(Acm)439, 441).
[0198] In some embodiments, the method of making the D-form T4 ligase having 90% or greater sequence identity to SEQ ID NO 21 or SEQ ID NO 22 further comprises desulfurizing one or more cysteine residues of a precursor fragment to alanine residues at one or more positions selected from 77, 117, 165, 202, 256, 317, 389, and 413. 5.8. Kits and Compositions
[0199] In another aspect, a kit containing materials useful for the use of the mirror image DNA-cleaving and / or modifying compositions described above is provided. The kit may include one or more components of the compositions (e.g., as disclosed herein), such as a D- form DNA endonuclease, an L-form guide RNA, and a target L-DNA. In some embodiments, the kit comprises a D-form DNA ligase. In some embodiments, the kit comprises a D-form DNA endonuclease, an L-form guide RNA and a D-form DNA ligase.
[0200] Also provided are compositions including two or more components of the mirror image DNA-cleaving and / or modifying compositions (e.g., as disclosed herein), such as a D- form DNA endonuclease or a L-form guide RNA, and a target L-DNA. In some embodiments, the composition includes a D-form DNA endonuclease, an L-form guide RNA, and a target L-DNA, wherein the L-guide RNA comprises a sequence complementary to a target sequence of a target L-DNA. In some embodiments the kit further comprises a D-form DNA ligase.
[0201] In some embodiments, the kits and compositions are useful for detecting, cleaving and / or modifying a target L-DNA site of a L-DNA molecule in a sample.
[0202] Elements of the kit may be provided individually or in combinations, and may be provided in any suitable container. In some embodiments, the kit includes instructions in one or more languages. In some embodiments, a kit comprises one or more reagents for use in a process utilizing one or more of the elements described herein. Reagents may be provided in any suitable container. For example, a kit may provide one or more reaction or storage buffers. Reagents may be provided in a form that is usable in a particular assay, or in a form that requires addition of one or more other components before use (e.g. in concentrate or lyophilized form).
[0203] The kit can include a container and a label or package insert on or associated with the container. Suitable containers include, for example, bottles, vials, tubes, syringes, IV solutionAttorney Ref.: 39339-64210 Client Ref.: 003WO bags, etc. The containers may be formed from a variety of materials such as glass or plastic. The container holds a composition which is by itself or combined with another composition effective for treating, preventing and / or diagnosing the condition and may have a sterile access port (for example the container may be an intravenous solution bag or a vial having a stopper pierceable by a hypodermic injection needle). At least one component of a composition disclosed herein is included. The label or package insert indicates that the composition is used for treating the condition of choice. The kit in this embodiment disclosed herein may further comprise a package insert indicating that the compositions can be used to treat a particular condition. Alternatively, or additionally, the kit may further comprise a second (or third) container comprising a pharmaceutically acceptable buffer, such as bacteriostatic water for injection (BWFI), phosphate-buffered saline, Ringer's solution and dextrose solution. It may further include other materials desirable from a commercial and user standpoint, including other buffers, diluents, filters, needles, and syringes. 6. EXAMPLES 6.1. Exemplary General Methods
[0204] Materials. All solvents and reagents were reagent grade, purchased commercially, and used without further purification unless specified. All chemicals were purchased from Sigma-Aldrich, Fisher Scientific, TCI, etc.
[0205] HO-TCP(Cl)-ProTide was purchased from CEM. Fmoc-Hydrazine was purchased from TCI.2-Chlorotrityl chloride resin (loading=0.98 mmol g−1) was purchased from Purepep. Fmoc-D-amino acids, D-4-thiazolidinecarboxylic acid, hydrazine hydrate, ethyl cyanoglyoxylate-2-oxime (Oxyma), N,N′-diisopropylcarbodiimide (DIC), trifluoroacetic acid, N,N-dimethylformamide (DMF), dichloromethane, piperidine, thioanisole, triisopropylsilane, 1,2-ethanedithiol, and trifluoroacetic acid for peptide synthesis were purchased commercially from Chempep, Sigma-Aldrich, Alfa Aesar, TCI, etc. The reagents for NCL reaction, i.e, guanidine hydrochloride (Gn^HCl), Na2HPO4^12H2O, NaH2PO4^2H2O, sodium nitrite (NaNO2), sodium hydroxide (NaOH), hydrochloric acid sodium 2- mercaptoethanesulfonate, 4-mercaptophenylacetic acid (MPAA), tris(2- carboxyethyl)phosphine hydrochloride (TCEP^HCl), DL-1,4-dithiothreitol (DTT), 2,2′- azobis[2-(2-imidazolin-2-yl) propane] dihydrochloride (VA-044), glutathione (reduced form), and palladium chloride (PdCl2) were purchased commercially from Sigma-Aldrich, Alfa Aesar, TCI, Duksan, etc.Attorney Ref.: 39339-64210 Client Ref.: 003WO
[0206] Fmoc-Hydrazine Loading Onto the 2-Cl-trt Resin From ProTide: Standard protocol was followed for loading Fmoc-NHNH2on HO-TCP(Cl)-ProTide resin. Briefly, 1 g (0.3 mmol) of resin was activated with thionyl chloride (1 mL in 8 mL of DCM) under N2 for overnight. Then, the resin was washed thoroughly with 3 x 5 mL of DCM, DMF, DCM. Following this, Fmoc-hydrazine (1.5 mmol, 5 eq) in 10 mL of DMF and 5 equiv. of DIPEA (1.5 mmol) was added to the resin and mixed for 4 hours. This was followed by washing the resin thoroughly with 3 x 5 mL of DMF, DCM, diethyl ether, and drying under vacuum for at least 3 hours before used for peptide synthesis.
[0207] Fmoc-Hydrazine Loading Onto the 2-Cl-trt Resin From Aapptec / Chempep: The above-mentioned protocol was followed to load Fmoc-NHNH2, except resin activation with thionyl chloride was omitted because of high loading capacity of these resins.
[0208] Fmoc-Based Solid-Phase Peptide Synthesis (Fmoc-SPPS): All peptides were synthesized by Fmoc-based SPPS on a Liberty Blue automated microwave peptide synthesizer (CEM) and PurePep® Chorus automated peptide synthesizer. All the peptide hydrazides were synthesized on Fmoc-hydrazine 2-chlorotrityl chloride resin. The peptide hydrazides (D-TdT-WT1 to D-TdT-WT5) were synthesized on Fmoc-hydrazine loaded ProTide resin and the peptide hydrazide (D-TdT-WT6) was synthesized on Fmoc-hydrazine- 2-chlorotrityl chloride resin from Chempep or Aapptec. The D-TdT-WT7 was synthesized on Rink amide resin to give C-Terminal “-OH”. For each peptide hydrazide, the first residue was attached to the hydrazine-2-chlorotrityl chloride resin by a double coupling method using 5 equiv. amino acid, 10 equiv. DIC, and 5 equiv. Oxymapure at 50oC for 10 min. All resins were swelled in DMF for 30 min before coupling. The Fmoc groups of assembled amino acids were removed by treatment with 20% piperidine and 0.1 M Oxyma in DMF at 85-90 °C. Coupling of amino acids except Fmoc-Cys(Trt)-OH, Fmoc-Arg(pbf)-OH, and Fmoc- His(Trt)-OH was carried out at 85 °C using 5 equiv. amino acid, 5 equiv. Oxymapure, and 10 equiv. DIC for 2-3 min. The coupling reactions for Fmoc-Cys(Trt)-OH, Fmoc-Arg(pbf)-OH, and Fmoc-His(Trt)-OH were carried out at 50 °C for 10 min to avoid side reactions at high temperature. Trifluoroacetyl thiazolidine-4-carboxylic acid-OH was coupled twice using 5 equiv. Oxymapure and 10 equiv. DIC at 50 °C for 10 min. Double coupling strategy was used for the peptides beyond 20 amino acids. After the completion of peptide chain assembly, peptides were cleaved from resin using H2O / thioanisole / triisopropylsilane / 1,2- ethanedithiol / trifluoroacetic acid (0.5 / 0.5 / 0.5 / 0.25 / 8.25) (vol / vol). The cleavage reaction took 2.5 h under agitation at 27 °C. Cold ether was added to precipitate the crude peptide. AfterAttorney Ref.: 39339-64210 Client Ref.: 003WO centrifugation, the supernatant was discarded and the precipitates were washed twice with ether. The crude peptides were dissolved in CH3CN / H2O, analyzed by RP-HPLC, and purified by semi-preparative HPLC. Collected peptide fractions were analyzed by electrospray ionization mass spectrometry (ESI-MS).
[0209] Native Chemical Ligation (NCL), Method A: The C-terminal peptide hydrazide segment was dissolved in acidified ligation buffer (aqueous solution of 6 M Gn^HCl and 0.1 M NaH2PO4, pH 3.0). The mixture was cooled in an ice-salt bath (-10 to -15 ° C), and 10 equiv. NaNO2in acidified ligation buffer (pH 3.0) was added. The activation reaction system was kept in an ice-salt bath under stirring for 20-25 min, after which 40 equiv. MPAA in ligation buffer and 1 equiv. N-terminal cysteine peptide were added, and the pH of the solution was adjusted to 6.6-6.8 at room temperature. After overnight reaction, 150 mM TCEP in ligation buffer (pH adjusted to 7.0) was added to dilute the system twice and the reaction system was kept at room temperature for 30 min with stirring. Finally, the ligation product was analyzed by HPLC and purified by semi-preparative HPLC. Purified ligation fractions were analyzed by ESI-MS.
[0210] Native Chemical Ligation (NCL), Method B: General approach. Peptide hydrazides were dissolved at a specified concentration (usually 1-20 mg / mL) in 6 M Gn^HCl containing 200 mM MPAA in a 1.5-2 mL Eppendorf tube. This will form a heterogeneous suspension and mixing sonication / vortexing should break any large pieces of solid MPAA. The pH can be adjusted to pH 3 if needed. A stock solution of acac is made in water (10x-20x) and 1 equiv. to 5 equiv. acac are added to the peptide mixture. Finally, a small stir bar is added to the tube and the mixture is allowed to stir for 2-4 hours.
[0211] After 2-4 hours, an equimolar amount of the Cys-fragment peptide (the C-terminal fragment of the ligated product) is dissolved in 6 M Gn^HCl with 200 mM Na2HPO4 and 50 mM TCEP (TCEP is used to reduce disulfides between the Cys fragments and is not necessary) (pH 8.5) to a volume equal to that which the thioesterification reaction is occurring. These two solutions are mixed together, upon which the resultant pH will be around 5-7 and the MPAA emulsion dissolves. The solution should be clear or slightly yellow. The pH of the combined solution is then adjusted to pH 7-7.4 by the addition of 1 M NaOH. The ligation reaction can then be left to stir for 4-18 hours.
[0212] Desulfurization: Cys-containing peptide (3 mg / mL) is dissolved in desulfurization buffer (0.1 M aqueous phosphate buffer containing 6 M Gn^HCl, 200 mM TCEP, 40 mM reduced L-glutathione, and 20 mM 2,2′-azobis [2-(2-imidazolin-2-yl)propane]Attorney Ref.: 39339-64210 Client Ref.: 003WO dihydrochloride, pH 6.8). The mixture is stirred at 37° C overnight, and the desulfurization product is analyzed by HPLC and ESI-MS, and purified by semi-preparative HPLC.
[0213] Acm Deprotection: Acetamidomethyl (Acm) group is removed by the Pd-assisted deprotection strategy. Acm-protected peptide is dissolved in Acm deprotection buffer (aqueous solution of 6 M Gn^HCl, 0.1 M phosphate, and 40 mM TCEP, pH 7.0) to a final concentration of 1 mM, after which 20 equiv. PdCl2is added. The reaction mixture is incubated with agitation at 25° C overnight. DTT is added to 50 mM final concentration to quench the reaction. The reaction mixture is stirred for 1 hour and purified by semi- preparative HPLC. 6.2. Example 1: Preparation of D-Cas12
[0214] D-form Cas12 is prepared via fragment couplings using native chemical ligation (NCL) according to the retrosynthetic scheme shown in FIGs.1A-1G.
[0215] D-Amino-acid sequence corresponding to L-form Cas12 sequence (accession number: U2UMQ6 (SEQ ID NO: 10)). The sequence includes many possible native chemical ligation junction sites including sites N-terminal to alanine.
[0216] In this exemplary synthesis, particular pairs of amino acids that are underlined and bolded are selected as coupling sites for native chemical ligation, and thus define the size and sequence of precursor fragments that are synthesized. The coupling sites are selected adjacent to particular D-alanine residues of the final sequence. It is understood that for each coupling site, a N-terminal D-cysteine residue is installed in the precursor fragments to be coupled (Table 1) instead of the D-alanine. See e.g., FIG.3, variant sequence “Cas12_Alanine to Cysteine”. After NCL coupling(s) the D-cysteine residue can be converted back to a D- alanine via a desulfurization procedure to arrive at the final sequence.
[0217] The parent sequence (SEQ ID NO: 10) includes isoleucine amino acid residues (total 15), at least several of which can be substituted with a suitable replacement D-amino acid residue (e.g., D-alanine, D-leucine, D-valine, glycine, etc). 1 mtqfeGftnl yqvsktlrfe lipqGktlkh iqeqGfieed karndhykel 51 kpiidriykt yadqclqlvq ldwenlsaai dsyrkektee trnalieeqa 101 tyrnaihdyf iGrtdnltda inkrhaeiyk GlfkaelfnG kvlkqlGtvt 151 ttehenallr sfdkfttyfs Gfyenrknvf saedistaip hrivqdnfpk 201 fkenchiftr litavpslre hfenvkkaiG ifvstsieev fsfpfynqll 251Attorney Ref.: 39339-64210 Client Ref.: 003WO tqtqidlynq llGGisreaG tekikGlnev lnlaiqknde tahiiaslph 301 rfiplfkqil sdrntlsfil eefksdeevi qsfckyktll rnenvletae 351 alfnelnsid lthifishkk letissalcd hwdtlrnaly erriseltGk 401 itksakekvq rslkhedinl qeiisaaGke lseafkqkts eilshahaal 451 dqplpttlkk qeekeilksq ldsllGlyhl ldwfavdesn evdpefsarl 501 tGiklemeps lsfynkarny atkkpysvek fklnfqmptl asGwdvnkek 551 nnGailfvkn GlyylGimpk qkGrykalsf eptektseGf dkmyydyfpd 601 aakmipkcst qlkavtahfq thttpillsn nfiepleitk eiydlnnpek 651 epkkfqtaya kktGdqkGyr ealckwidft rdflskytkt tsidlsslrp 701 ssqykdlGey yaelnpllyh isfqriaeke imdavetGkl ylfqiynkdf 751 akGhhGkpnl htlywtGlfs penlaktsik lnGqaelfyr pksrmkrmah 801 rlGekmlnkk lkdqktpipd tlyqelydyv nhrlshdlsd earallpnvi 851 tkevsheiik drrftsdkff fhvpitlnyq aanspskfnq rvnaylkehp 901 etpiiGidrG ernliyitvi dstGkileqr slntiqqfdy qkkldnreke 951 rvaarqawsv vGtikdlkqG ylsqviheiv dlmihyqavv vlenlnfGfk 1001 skrtGiaeka vyqqfekmli dklnclvlkd ypaekvGGvl npyqltdqft 1051 sfakmGtqsG flfyvpapyt skidpltGfv dpfvwktikn hesrkhfleG 1101 fdflhydvkt Gdfilhfkmn rnlsfqrGlp Gfmpawdivf eknetqfdak 1151 GtpfiaGkri vpvienhrft Gryrdlypan elialleekG ivfrdGsnil 1201 pkllenddsh aidtmvalir svlqmrnsna atGedyinsp vrdlnGvcfd 1251 srfqnpewpm dadanGayhi alkGqlllnh lkeskdlklq nGisnqdwla 1301 yiqelrn (SEQ ID NO: 10)
[0218] SEQ ID NO: 11 is a sequence of D-form Cas12 variant including D-alanine to D- cysteine substitutions at the coupling sites. The substitutions include: A62C, A105C, A157C, A188C, A228C, A284C, A349C, A388C, A449C, A498C, A541C, A602C, A660C, A712C, A801C, A799C, A844C, A894C, A954C, A1010C, A1067C, A1135C, A1179C, A1131C, and A1267C. 1Attorney Ref.: 39339-64210 Client Ref.: 003WO mtqfeGftnl yqvsktlrfe lipqGktlkh iqeqGfieed karndhykel 51 kpiidriykt ycdqclqlvq ldwenlsaai dsyrkektee trnalieeqa 101 tyrncihdyf iGrtdnltda inkrhaeiyk GlfkaelfnG kvlkqlGtvt 151 ttehencllr sfdkfttyfs Gfyenrknvf saedistcip hrivqdnfpk 201 fkenchiftr litavpslre hfenvkkciG ifvstsieev fsfpfynqll 251 tqtqidlynq llGGisreaG tekikGlnev lnlciqknde tahiiaslph 301 rfiplfkqil sdrntlsfil eefksdeevi qsfckyktll rnenvletce 351 alfnelnsid lthifishkk letissalcd hwdtlrncly erriseltGk 401 itksakekvq rslkhedinl qeiisaaGke lseafkqkts eilshahacl 451 dqplpttlkk qeekeilksq ldsllGlyhl ldwfavdesn evdpefscrl 501 tGiklemeps lsfynkarny atkkpysvek fklnfqmptl csGwdvnkek 551 nnGailfvkn GlyylGimpk qkGrykalsf eptektseGf dkmyydyfpd 601 ackmipkcst qlkavtahfq thttpillsn nfiepleitk eiydlnnpek 651 epkkfqtayc kktGdqkGyr ealckwidft rdflskytkt tsidlsslrp 701 ssqykdlGey ycelnpllyh isfqriaeke imdavetGkl ylfqiynkdf 751 ckGhhGkpnl htlywtGlfs penlaktsik lnGqaelfyr pksrmkrmch 801 rlGekmlnkk lkdqktpipd tlyqelydyv nhrlshdlsd earcllpnvi 851 tkevsheiik drrftsdkff fhvpitlnyq aanspskfnq rvncylkehp 901 etpiiGidrG ernliyitvi dstGkileqr slntiqqfdy qkkldnreke 951 rvacrqawsv vGtikdlkqG ylsqviheiv dlmihyqavv vlenlnfGfk 1001 skrtGiaekc vyqqfekmli dklnclvlkd ypaekvGGvl npyqltdqft 1051 sfakmGtqsG flfyvpcpyt skidpltGfv dpfvwktikn hesrkhfleG 1101 fdflhydvkt Gdfilhfkmn rnlsfqrGlp Gfmpcwdivf eknetqfdak 1151 GtpfiaGkri vpvienhrft Gryrdlypcn elialleekG ivfrdGsnil 1201 pkllenddsh aidtmvalir svlqmrnsna ctGedyinsp vrdlnGvcfd 1251 srfqnpewpm dadanGcyhi alkGqlllnh lkeskdlklq nGisnqdwlaAttorney Ref.: 39339-64210 Client Ref.: 003WO 1301 yiqelrn (SEQ ID NO: 11)
[0219] SEQ ID NO: 12 is a sequence of D-form Cas12 variant with selected D-isoleucine substitutions (15 total) to D-alanine, D-leucine, D-phenylalanine, D-valine, D-threonine, or D-tyrosine (see bold residues). The positions of the D-isoleucine substitutions were based in part on the alignment and consensus sequence of Cas12 family members from different species of bacteria compared to Acidaminococcus.
[0220] The substitutions include: I37L; I53V; I57Y; A62C; I96A; A105C; I111T; I121F; I128A; A157C; A188C; I207A; I212V; A228C; I229A; I231A; I237L; I265V; A284C; I285A; I295L; I303A; I330L; A349C; I359A; I374V; A388C; I394A; I401A; I418F; I423L; I424A; I442V; A449C; I466A; A498C; I503A; A541C; A602C; I626Y; I633A; A660C; I693Y; A712C; I726V; I731V; A751C; A799C; A844C; I850V; A894C; I915L; I917A; I935V; A954C; I984V; A1010C; A1067C; A1135C; I1138V; I1164A; A1179C; I1183A; I1199V; I1212F; I1219F; A1231C; A1267C; and I1302V. 1 mtqfeGftnl yqvsktlrfe lipqGktlkh iqeqGfleed karndhykel 51 kpvidryykt ycdqclqlvq ldwenlsaai dsyrkektee trnalaeeqa 101 tyrncihdyf tGrtdnltda fnkrhaeayk GlfkaelfnG kvlkqlGtvt 151 ttehencllr sfdkfttyfs Gfyenrknvf saedistcip hrivqdnfpk 201 fkenchaftr lvtavpslre hfenvkkcaG afvstsleev fsfpfynqll 251 tqtqidlynq llGGvsreaG tekikGlnev lnlcaqknde tahilaslph 301 rfaplfkqil sdrntlsfil eefksdeevl qsfckyktll rnenvletce 351 alfnelnsad lthifishkk letvssalcd hwdtlrncly erraseltGk 401 atksakekvq rslkhedfnl qelasaaGke lseafkqkts evlshahacl 451 dqplpttlkk qeekealksq ldsllGlyhl ldwfavdesn evdpefscrl 501 tGaklemeps lsfynkarny atkkpysvek fklnfqmptl csGwdvnkek 551 nnGailfvkn GlyylGimpk qkGrykalsf eptektseGf dkmyydyfpd 601 ackmipkcst qlkavtahfq thttpyllsn nfaepleitk eiydlnnpekAttorney Ref.: 39339-64210 Client Ref.: 003WO epkkfqtayc kktGdqkGyr ealckwidft rdflskytkt tsydlsslrp 701 ssqykdlGey ycelnpllyh isfqrvaeke vmdavetGkl ylfqiynkdf 751 ckGhhGkpnl htlywtGlfs penlaktsik lnGqaelfyr pksrmkrmch 801 rlGekmlnkk lkdqktpipd tlyqelydyv nhrlshdlsd earcllpnvv 851 tkevsheiik drrftsdkff fhvpitlnyq aanspskfnq rvncylkehp 901 etpiiGidrG ernllyatvi dstGkileqr slntvqqfdy qkkldnreke 951 rvacrqawsv vGtikdlkqG ylsqviheiv dlmvhyqavv vlenlnfGfk 1001 skrtGiaekc vyqqfekmli dklnclvlkd ypaekvGGvl npyqltdqft 1051 sfakmGtqsG flfyvpcpyt skidpltGfv dpfvwktikn hesrkhfleG 1101 fdflhydvkt Gdfilhfkmn rnlsfqrGlp Gfmpcwdvvf eknetqfdak 1151 GtpfiaGkri vpvaenhrft Gryrdlypcn elaalleekG ivfrdGsnvl 1201 pkllenddsh afdtmvalfr svlqmrnsna ctGedyinsp vrdlnGvcfd 1251 srfqnpewpm dadanGcyhi alkGqlllnh lkeskdlklq nGisnqdwla 1301 yvqelrn (SEQ ID NO: 12)
[0221] A series of synthetic precursor fragments (Table 1) are prepared using solid phase peptide synthesis (SPPS) via Fmoc protecting group strategy. Alternatively, a Boc SPPS methodology is used.
[0222] Table 1: precursor fragments for NCL preparation of exemplary D-form Cas12. Name, length SEQ. ID. NO: Structure Cas12N1 61 AA 40 66 met1-t r61-NHNHAttorney Ref.: 39339-64210 Client Ref.: 003WO Cas12.C1, 52 AA 53, 76 thz660-tyr711-NHNH (cys(Acm)67) Cas12.C2, 39 AA 54, 7712 750247 cys 2751-phe798-NHNH2 C 1 C348 AA 55 t NHNH- form Cas12. SEQ. ID. NO: Sequence v - v s - e yAttorney Ref.: 39339-64210 Client Ref.: 003WO cllpnvvtkevsheiikdrrftsdkfffhvpitlnyqaanspskfnqrvn- NHNH2c lkeh et iiGidrGernll atvidstGkile rslntv fd kkldn t m - i - v s e y n -Attorney Ref.: 39339-64210 Client Ref.: 003WO
[0224] FIGs.1A-1G illustrate the convergent fragment assembly strategy for construction of D-Cas12 via NCL. Protected N-terminal D-thiazolidine (thz) residues are converted to N- terminal D-cysteine residues prior to NCL fragment couplings.
[0225] Selected D-cysteine residues that are not located at the N-terminal of a fragment are maintained within the sequence by utilizing an orthogonal acetamidomethyl (Acm) protecting group strategy during SPPS, NCL couplings and desulfurization.
[0226] D-Cysteine residues that are incorporated into the synthetic precursor fragments for purposes of performing the NCL fragment condensation strategy are converted to D-alanine residues using a metal-free radical-based desulfurization procedure after selected couplings, as shown in FIGs.1A-1G. 6.3. Example 2: Preparation of D-Cas14
[0227] D-form Cas14 is prepared via fragment couplings using native chemical ligation (NCL) according to the retrosynthetic scheme shown in FIGs.2A-2C.
[0228] D-Amino-acid sequence corresponding to L-form Cas14 sequence (accession number: QBM01093.1) (SEQ ID NO: 16). The sequence includes several possible native chemical ligation junction sites including sites N-terminal to alanine.
[0229] In this exemplary synthesis, particular pairs of amino acids that are underlined and bolded are selected as coupling sites for native chemical ligation, and thus define the size and sequence of precursor fragments that are synthesized. D-Cysteine residues that are not located at the N-terminal of a fragment are maintained as cysteine by utilizing an orthogonal acetamidomethyl (Acm) protecting group strategy during synthesis. After NCL coupling(s) the D-cysteine residue can be converted back to a D-alanine via a desulfurization procedure to arrive at the final sequence.
[0230] The parent sequence (SEQ ID NO: 16) includes isoleucine amino acid residues (total 35), at least several of which can be substituted with a suitable replacement D-amino acid residue (e.g., D-alanine, D-leucine, D-valine, glycine, etc). See e.g., FIG.4, variant sequence “Cas14 Isoleucine substitution”. 1 mevqktvmkt lslrilrply sqeiekeike ekerrkqaGG tGeldGGfyk 51 klekkhsemf sfdrlnllln qlqreiakvy nhaiselyia tiaqGnksnk 101 hyissivynr ayGyfynayi alGicskvea nfrsnelltq qsalptaksd 151 nfpivlhkqk GaeGedGGfr isteGsdlif eipipfyeyn Genrkepykw 201Attorney Ref.: 39339-64210 Client Ref.: 003WO vkkGGqkpvl klilstfrrq rnkGwakdeG tdaeirkvte Gkyqvsqiei 251 nrGkklGehq kwfanfsieq piyerkpnrs ivGGldvGir splvcainns 301 fsrysvdsnd vfkfskqvfa frrrllskns lkrkGhGaah klepitemte 351 kndkfrkkii erwakevtnf fvknqvGivq iedlstmkdr edhffnqylr 401 Gfwpyyqmqt lienklkeyG ievkrvqaky tsqlcsnpnc rywnnyfnfe 451 yrkvnkfpkf kcekcnleis adynaarnls tpdiekfvak atkGinlpek (SEQ ID NO: 16)
[0231] SEQ ID NO: 17 is a sequence of D-form Cas14a.1 variant including D-alanine to D- cysteine substitutions at the coupling sites. The substitutions include: A38C; A83C; A111C; A162C; A226C; A264C; A364C; A428C; and A476C. 1 mevqktvmkt lslrilrply sqeiekeike ekerrkqcGG tGeldGGfyk 51 klekkhsemf sfdrlnllln qlqreiakvy nhciselyia tiaqGnksnk 101 hyissivynr cyGyfynayi alGicskvea nfrsnelltq qsalptaksd 151 nfpivlhkqk GceGedGGfr isteGsdlif eipipfyeyn Genrkepykw 201 vkkGGqkpvl klilstfrrq rnkGwckdeG tdaeirkvte Gkyqvsqiei 251 nrGkklGehq kwfcnfsieq piyerkpnrs ivGGldvGir splvcainns 301 fsrysvdsnd vfkfskqvfa frrrllskns lkrkGhGaah klepitemte 351 kndkfrkkii erwckevtnf fvknqvGivq iedlstmkdr edhffnqylr 401 Gfwpyyqmqt lienklkeyG ievkrvqcky tsqlcsnpnc rywnnyfnfe 451 yrkvnkfpkf kcekcnleis adynacrnls tpdiekfvak atkGinlpek (SEQ ID NO: 17)
[0232] SEQ ID NO: 18 is a sequence of D-form Cas14a.1 variant with selected D-isoleucine substitutions (15 total) to either D-alanine, D-leucine or D-valine (see bold residues). The substitutions include: I15V; I24V; I28V; A38C; I76A, A83C, I92V, I103L, I106A, A111C, I124L, A162C, A226C, I248V, A264C, I272A, I281A, I289A, I297V, I345A, A364C, A428C, A476C, and I484L. 1 mevqktvmkt lslrvlrply sqevekevke ekerrkqcGG tGeldGGfyk 51Attorney Ref.: 39339-64210 Client Ref.: 003WO klekkhsemf sfdrlnllln qlqreaakvy nhciselyia tvaqGnksnk 101 hylssavynr cyGyfynayi alGlcskvea nfrsnelltq qsalptaksd 151 nfpivlhkqk GceGedGGfr isteGsdlif eipipfyeyn Genrkepykw 201 vkkGGqkpvl klilstfrrq rnkGwckdeG tdaeirkvte Gkyqvsqvei 251 nrGkklGehq kwfcnfsieq payerkpnrs avGGldvGar splvcavnns 301 fsrysvdsnd vfkfskqvfa frrrllskns lkrkGhGaah klepatemte 351 kndkfrkkii erwckevtnf fvknqvGivq iedlstmkdr edhffnqylr 401 Gfwpyyqmqt lienklkeyG ievkrvqcky tsqlcsnpnc rywnnyfnfe 451 yrkvnkfpkf kcekcnleis adynacrnls tpdlekfvak atkGinlpek (SEQ ID NO: 18)synthetic precursor fragments (Table 2) are prepared using solid phase peptide synthesis (SPPS) via Fmoc protecting group strategy. Alternatively, a Boc SPPS methodology is used.
[0234] Table 3: precursor fragments for NCL preparation of exemplary D-form Cas14. Name, length SED. ID. NO: Structure Cas14a.1.1, 37 AA 84, 95 met1-gln37-NHNH2
[0235] Table 4: sequences of the precursor fragments for NCL preparation of exemplary D- form Cas14. SEQ. Sequence k kAttorney Ref.: 39339-64210 Client Ref.: 003WO lilstfrrqrnkGw-NHNH289ckdeGtdaeirkvteGkyqvsqveinrGkklGehqkwf-NHNH2Thz-nf i rk nr vGGldvG r lv (A m) vnn f r vds e ) k s. of D-Cas14 via NCL. N-terminal D-thiazolidine (thz) residues protect N-terminals of fragments during a coupling at a C-terminal of the fragment. The protected N-terminal D-thiazolidine (thz) residues are then converted into N-terminal D-cysteine residues prior to NCL fragment couplings at the N-terminal position.
[0237] Selected D-cysteine residues that are not located at the N-terminal of a fragment are maintained within the sequence by utilizing an orthogonal acetamidomethyl (Acm) protecting group strategy during SPPS, NCL couplings and desulfurization.
[0238] D-Cysteine residues that are incorporated into the synthetic precursor fragments for purposes of performing the NCL fragment condensation strategy are converted to D-alanine residues using a metal-free radical-based desulfurization procedure after selected couplings, as shown in FIGs.2A-2C. 6.4. Example 3: Preparation of D-Cas14a.1 Mutant
[0239] D-form Cas14a.1 mutant was prepared via fragment couplings using native chemical ligation (NCL) according to the retrosynthetic scheme shown in FIG.5. The D-form Cas14a.1 mutant was designed as eight synthetic peptides (D-Cas14a.1m-1 to D-Cas14a.1m- 8) via solid phase peptide synthesis which will ligate at the certain cysteine residue as shown in FIG.5.Attorney Ref.: 39339-64210 Client Ref.: 003WO
[0240] D-Amino-acid sequence corresponding to L-form Cas14 sequence (accession number: QBM01093.1) (SEQ ID NO: 16). The sequence includes several possible native chemical ligation junction sites including sites N-terminal to alanine.
[0241] In this exemplary synthesis, particular pairs of amino acids that are underlined and bolded were selected as coupling sites for native chemical ligation, and thus define the size and sequence of precursor fragments that were synthesized. D-Cysteine residues that are not located at the N-terminal of a fragment were maintained as cysteine by utilizing an orthogonal acetamidomethyl (Acm) protecting group strategy during synthesis. After NCL coupling(s) the D-cysteine residue can be converted back to a D-alanine via a desulfurization procedure to arrive at the final sequence.
[0242] The parent sequence (SEQ ID NO: 16) includes isoleucine amino acid residues (total 35), at least several of which can be substituted with a suitable replacement D-amino acid residue (e.g., D-alanine, D-leucine, D-valine, glycine, etc). See e.g., FIG.4, variant sequence “Cas14 Isoleucine substitution”.
[0243] SEQ ID No: 19 is a sequence of D-form Cas14a.1 mutant including selected D- isoleucine substitutions (15 total) to either D-alanine, D-leucine or D-valine (see bold residues). The substitutions include: I15V; I24V; I28V; I76A; I92V; I103L; I106A; I124L; I248V; I272A; I281A; I289A; V294A; I297V; I345A; and I484L. 1 mevqktvmkt lslrvlrply sqevekevke ekerrkqaGG tGeldGGfyk 51 klekkhsemf sfdrlnllln qlqreaakvy nhaiselyia tvaqGnksnk 101 hylssavynr ayGyfynayi alGlcskvea nfrsnelltq qsalptaksd 151 nfpivlhkqk GaeGedGGfr isteGsdlif eipipfyeyn Genrkepykw 201 vkkGGqkpvl klilstfrrq rnkGwakdeG tdaeirkvte Gkyqvsqvei 251 nrGkklGehq kwfanfsieq payerkpnrs avGGldvGar splacavnns 301 fsrysvdsnd vfkfskqvfa frrrllskns lkrkGhGaah klepatemte 351 kndkfrkkii erwakevtnf fvknqvGivq iedlstmkdr edhffnqylr 401 Gfwpyyqmqt lienklkeyG ievkrvqaky tsqlcsnpnc rywnnyfnfe 451 yrkvnkfpkf kcekcnleis adynaarnls tpdlekfvak atkGinlpek (SEQ ID NO: 19)Attorney Ref.: 39339-64210 Client Ref.: 003WO
[0244] A series of synthetic precursor fragments (Table 3) were prepared using solid phase peptide synthesis (SPPS) via Fmoc protecting group strategy. Alternatively, a Boc SPPS methodology was used.
[0245] Table 5: precursor fragments for NCL preparation of exemplary D-form Cas14a.1 mutant. ESI-MS: ESI-MS: Name, length SEQ. ID. NO: Structure m / z m / z d- form Cas14a.1 mutant. SEQ. ID. NO: Sequence e k n p l f
[0247] FIG.5 illustrates the convergent fragment assembly strategy for construction of D- Cas14a.1 mutant via NCL. N-terminal D-thiazolidine (thz) residues protect N-terminals of fragments during a coupling at a C-terminal of the fragment. The protected N-terminal D- thiazolidine (thz) residues are then converted into N-terminal D-cysteine residues prior to NCL fragment couplings at the N-terminal position.Attorney Ref.: 39339-64210 Client Ref.: 003WO
[0248] Selected D-cysteine residues that are not located at the N-terminal of a fragment are maintained within the sequence by utilizing an orthogonal acetamidomethyl (Acm) protecting group strategy during SPPS, NCL couplings and desulfurization.
[0249] D-Cysteine residues that are incorporated into the synthetic precursor fragments for purposes of performing the NCL fragment condensation strategy are converted to D-alanine residues using a metal-free radical-based desulfurization procedure after selected couplings, as shown in FIG.5.
[0250] D-Cas14a.1m-2@48-mer, D-Cas14a.1m-3@37-mer, D-Cas14a.1m-4@64-mer, D- Cas14a.1m-5@69-mer, D-Cas14a.1m-6@69-mer, D-Cas14a.1m-7@71-mer, and D- Cas14a.1m-8@66-mer were synthesized on a Liberty Blue automated peptide synthesizer by following the conditions mentioned in the exemplary general methods, were purified by HPLC, and were analyzed by ESI-MS. The results are provided in FIGs.6-12, Table 3, and Table 4. 6.5. Example 4: Preparation of D-T4 Ligase
[0251] D-form T4 ligase is prepared via fragment couplings using native chemical ligation (NCL) according to the retrosynthetic scheme shown in FIG.13. The D-form T4 ligase was designed as nine synthetic peptides (D-T4-1 to D-T4-9) via solid phase peptide synthesis which will ligate at the certain cysteine residue as shown in FIG.13.
[0252] D-Amino-acid sequence corresponding to L-form T4 ligase sequence (accession number: P00970) (SEQ ID NO: 21). The sequence includes several possible native chemical ligation junction sites including sites N-terminal to alanine.
[0253] In this exemplary synthesis, particular pairs of amino acids that are underlined and bolded are selected as coupling sites for native chemical ligation, and thus define the size and sequence of precursor fragments that were synthesized. D-Cysteine residues that are not located at the N-terminal of a fragment are maintained as cysteine by utilizing an orthogonal acetamidomethyl (Acm) protecting group strategy during synthesis. After NCL coupling(s) the D-cysteine residue can be converted back to a D-alanine via a desulfurization procedure to arrive at the final sequence.
[0254] The parent sequence (SEQ ID NO: 21) includes isoleucine amino acid residues (total 37), at least several of which can be substituted with a suitable replacement D-amino acid residue (e.g., D-alanine, D-leucine, D-valine, glycine, etc).
[0255] SEQ ID NO: 22 is a sequence of D-form T4 ligase including selected substitution (see bold residues). The substitution includes V317A.Attorney Ref.: 39339-64210 Client Ref.: 003WO 1 milkilneia siGstkqkqa ileknkdnel lkrvyrltys rGlqyyikkw 51 pkpGiatqsf Gmltltdmld fieftlatrk ltGnaaieel tGyitdGkkd 101 dvevlrrvmm rdlecGasvs iankvwpGli peqpqmlass ydekGinkni 151 kfpafaqlka dGarcfaevr Gdelddvrll sraGneylGl dllkeelikm 201 taearqihpe GvlidGelvy heqvkkepeG ldflfdaype nskakefaev 251 aesrtasnGi ankslkGtis ekeaqcmkfq vwdyvplvei yslpafrlky 301 dvrfskleqm tsGydkaili enqvvnnlde akviykkyid qGleGiilkn 351 idGlwenars knlykfkevi dvdlkivGiy phrkdptkaG GfilesecGk 401 ikvnaGsGlk dkaGvkshel drtrimenqn yyiGkilece cnGwlksdGr 451 tdyvklflpi airlredktk antfedvfGd fhevtGl (SEQ ID NO: 22)
[0256] A series of synthetic precursor fragments (Table 4) were prepared using solid phase peptide synthesis (SPPS) via Fmoc protecting group strategy. Alternatively, a Boc SPPS methodology was used.
[0257] Table 7: precursor fragments for NCL preparation of exemplary D-form T4 ligase. Name, length SEQ. ID. ESI-MS: ESI-MS: NO: Structure m / z Calculated m / z Observed
[0258] Table 8: sequence of the precursor fragments for NCL preparation of exemplary D- form T4 ligase. SEQ. Sequence iAttorney Ref.: 39339-64210 Client Ref.: 003WO 32ctrkltGnaaieeltGyitdGkkddvevlrrvmmrdlec(Acm)G-NHNH233 csvsiankvwpGlipeqpqmlassydekGinknikfpafaqlkadGar- NHNH r y v k- T4 ligase via NCL. N-terminal D-thiazolidine (thz) residues protect N-terminals of fragments during a coupling at a C-terminal of the fragment. The protected N-terminal D-thiazolidine (thz) residues are then converted into N-terminal D-cysteine residues prior to NCL fragment couplings at the N-terminal position.
[0260] Selected D-cysteine residues that are not located at the N-terminal of a fragment are maintained within the sequence by utilizing an orthogonal acetamidomethyl (Acm) protecting group strategy during SPPS, NCL couplings and desulfurization.
[0261] D-Cysteine residues that are incorporated into the synthetic precursor fragments for purposes of performing the NCL fragment condensation strategy are converted to D-alanine residues using a metal-free radical-based desulfurization procedure after selected couplings, as shown in FIG.13.
[0262] D- T4-1@His6+76-mer and D-T4-7@72-mer were synthesized on a Liberty Blue automated peptide synthesizer by following the conditions mentioned in the exemplary general methods, were purified by HPLC, and were analyzed by ESI-MS. The results are provided in FIGs.14-15, Table 5, and Table 6. 7. EQUIVALENTS AND INCORPORATION BY REFERENCE
[0263] While the invention has been particularly shown and described with reference to a preferred embodiment and various alternate embodiments, it will be understood by persons skilled in the relevant art that various changes in form and details can be made therein without departing from the spirit and scope of the invention.Attorney Ref.: 39339-64210 Client Ref.: 003WO
[0264] All references, issued patents, and patent applications cited within the body of the instant specification, including U.S. Provisional Appl. No.63 / 695,718, are hereby incorporated by reference in their entirety, for all purposes. 8. SEQUENCES SEQ. ID. NO: Sequence AE G D N I A H E L I G L R H L S K F A K E A e G d n i a h e l i G l r h l s k f a k e a R N AAttorney Ref.: 39339-64210 Client Ref.: 003WO (Cas9 KQLLKVQKAYHQLDQSFIDTYIDLLETRRTYYEGPGEGSPFGWKDIKEWYEMLMGHCTYF Streptococcus PEELRSVKYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQKKKPTLKQIA aureus) KEILVNEEDIKGYRVTSTGKPEFTNLKVYHDIKDITARKEIIENAELLDQIAKILTIYQS R R A S L K N L S A I r n a f a s r r a s l k n l s a i H P D L K L L Y F G K N K K N F T I Y T E N K D F R AAttorney Ref.: 39339-64210 Client Ref.: 003WO 6 mnfkilpiaidlGvkntGvfsafyqkGtslerldnknGkvyelskdsytllmnnrtarrh qrrGidrkqlvkrlfkliwteqlnlewdkdtqqaisflfnrrGfsfitdGyspeylnivp (D-form Cas9 eqvkailmdifddynGeddldsylklateqeskiseiynklmqkilefklmklctdikdd l k l l y f G k n k k n f t i y t e n k d f r a R R E S K T E L S K L S R K D F r r e s k t e l s k l s r k d fAttorney Ref.: 39339-64210 Client Ref.: 003WO ekyivsalGevtkaefrqredfkk MTQFEGFTNLYQVSKTLRFELIPQGKTLKHIQEQGFIEEDKARNDHYKELKPIIDRIYKT YADQCLQLVQLDWENLSAAIDSYRKEKTEETRNALIEEQATYRNAIHDYFIGRTDNLTDA F V H D L L L D A H K D P V I V F L M t a f v h d l l l d a h k d p v i v f l mAttorney Ref.: 39339-64210 Client Ref.: 003WO rlGekmlnkklkdqktpipdtlyqelydyvnhrlshdlsdearcllpnvi tkevsheiikdrrftsdkfffhvpitlnyqaanspskfnqrvncylkehp etpiiGidrGernliyitvidstGkileqrslntiqqfdyqkkldnreke N A N T K E K T N I S E C I L I F S F H G S TAttorney Ref.: 39339-64210 Client Ref.: 003WO KIENTNDTL mGnlfGhkrwyevrdkkdfkikrkvkvkrnydGnkyilninennnkekidnnkfirkyin ykkndnilkeftrkfhaGnilfklkGkeGiiriennddfleteevvlyieayGkseklka n t k e k t n i s e c i l i f s f h G s tAttorney Ref.: 39339-64210 Client Ref.: 003WO (D-form hylssavynrcyGyfynayialGlcskveanfrsnelltqqsalptaksd Cas14a.1 nfpivlhkqkGceGedGGfristeGsdlifeipipfyeynGenrkepykw variant with vkkGGqkpvlklilstfrrqrnkGwckdeGtdaeirkvteGkyqvsqveiAttorney Ref.: 39339-64210 Client Ref.: 003WO 25 Thz-skveanfrsnelltqqsalptaksdnfpivlhkqkG-NHNH (D- Cas14a1m-3) l h lAttorney Ref.: 39339-64210 Client Ref.: 003WO (Cas12.N7)eevlqsfc(Acm)kyktllrnenvlet-NHNH247 (C 12N8) cealfnelnsadlthifishkkletvssalc(Acm)dhwdtlrn-NHNH2q n n k k h s 2 l h lAttorney Ref.: 39339-64210 Client Ref.: 003WO 71 Thz-iqkndetahiiaslphrfiplfkqilsdrntlsfileefksd (Cas12.N7) eeviqsfc(Acm)kyktllrnenvlet-NHNH2722q n n k s v 2Attorney Ref.: 39339-64210 Client Ref.: 003WO (Cas14a.1.1) 96 cGGtGeldGGfykklekkhsemfsfdrlnlllnqlqreiakvynh- (C 14 12) NHNH v 2
Claims
Attorney Ref.: 39339-64210 Client Ref.: 003WO WHAT IS CLAIMED IS:
1. A D-form DNA endonuclease, wherein the D-form DNA endonuclease is a mirror image form of a L-form DNA endonuclease or a variant thereof.
2. The D-form DNA endonuclease of claim 1, wherein the D-form DNA endonuclease is a mirror image form of a L-form Cas protein.
3. The D-form DNA endonuclease of claim 2, wherein D-form DNA endonuclease is a mirror image isomer of a class 1, type I, III, or IV Cas endonuclease or a functional derivative thereof.
4. The D-form DNA endonuclease of claim 2, wherein D-form DNA endonuclease is a mirror image isomer of a class 2, type II, V, or VI Cas endonuclease or a functional derivative thereof.
5. The D-form DNA endonuclease of any one of claims 1 to 4, wherein the D-form DNA endonuclease is a mirror image isomer of a L-Cas protein selected from Cas3, Cas9, Cas10, Cas12, Cas13, Cas14 and functional derivatives thereof.
6. The D-form DNA endonuclease of claim 5, wherein the D-form DNA endonuclease has a sequence of a L-Cas protein with one or more variations, wherein the one or more variations comprise a substitution from alanine to cysteine.
7. The D-form DNA endonuclease of claim 5, wherein the D-form DNA endonuclease has a sequence of a L-Cas protein with one or more variations, wherein the one or more variations comprise a substitution from isoleucine to alanine, leucine, phenylalanine, valine, threonine, or tyrosine.
8. The D-form DNA endonuclease of any one of claims 1 to 7, wherein the D-form DNA endonuclease comprises: i) a sequence selected from SEQ ID NOs: 1-19; or ii) a sequence having at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to the sequence set forth in i).Attorney Ref.: 39339-64210 Client Ref.: 003WO 9. The D-form DNA endonuclease of claim 8, wherein the D-form DNA endonuclease comprises: a. a sequence having at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to the sequence of SEQ ID NO: 10; and b. one or more variations selected from: A62C, A105C, A157C, A188C, A228C, A284C, A349C, A388C, A449C, A498C, A541C, A602C, A660C, A712C, A801C, A799C, A844C, A894C, A954C, A1010C, A1067C, A1135C, A1179C, A1131C, and A1267C.
10. The D-form DNA endonuclease of claim 8 or 9, wherein the D-form DNA endonuclease comprises: a. a sequence having at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to the sequence of SEQ ID NO: 10; and b. one or more variations selected from: I37L; I53V; I57Y; A62C; I96A; A105C; I111T; I121F; I128A; A157C; A188C; I207A; I212V; A228C; I229A; I231A; I237L; I265V; A284C; I285A; I295L; I303A; I330L; A349C; I359A; I374V; A388C; I394A; I401A; I418F; I423L; I424A; I442V; A449C; I466A; A498C; I503A; A541C; A602C; I626Y; I633A; A660C; I693Y; A712C; I726V; I731V; A751C; A799C; A844C; I850V; A894C; I915L; I917A; I935V; A954C; I984V; A1010C; A1067C; A1135C; I1138V; I1164A; A1179C; I1183A; I1199V; I1212F; I1219F; A1231C; A1267C; and I1302V.
11. The D-form DNA endonuclease of claim 8, wherein the D-form DNA endonuclease comprises: a. a sequence having at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to the sequence of SEQ ID NO: 16; and b. one or more variations selected from: A38C; A83C; A111C; A162C; A226C; A264C; A364C; A428C; and A476C.
12. The D-form DNA endonuclease of claim 8 or 11, wherein the D-form DNA endonuclease comprises: a. a sequence having at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to the sequence of SEQ ID NO: 16; andAttorney Ref.: 39339-64210 Client Ref.: 003WO b. one or more variations selected from: I15V; I24V; I28V; A38C; I76A, A83C, I92V, I103L, I106A, A111C, I124L, A162C, A226C, I248V, A264C, I272A, I281A, I289A, I297V, I345A, A364C, A428C, A476C, and I484L.
13. The D-form DNA endonuclease of claim 8, wherein the D-form DNA endonuclease comprises: a. a sequence having at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to the sequence of SEQ ID NO: 16; and b. one or more variations selected from: I15V; I24V; I28V; I76A; I92V; I103L; I106A; I124L; I248V; I272A; I281A; I289A; V294A; I297V; I345A; and I484L.
14. A composition comprising the D-form DNA endonuclease of any one of claims 1 to 13 and an L-form guide RNA, optionally wherein the L-form guide RNA comprises a sequence complementary to a target sequence of a target L-DNA.
15. The composition of claim 14, further comprising: a target L-DNA comprising the target sequence.
16. The composition of claim 15, wherein the target L-DNA comprises a duplex segment that comprises the target sequence to which the L-guide RNA is directed.
17. The composition of any one of claims 14 to 16, wherein the target L-DNA is comprised within a folded nanostructure.
18. A D-form DNA ligase, wherein the D-form DNA ligase is a mirror image form of an L-form DNA ligase or a variant thereof.
19. The D-form DNA ligase of claim Error! Reference source not found., wherein the D-form DNA ligase is a mirror image form of a L-form T4 ligase.Attorney Ref.: 39339-64210 Client Ref.: 003WO 20. The D-form DNA ligase of claim Error! Reference source not found. or 19, wherein the D-form DNA ligase has a sequence of a L-T4 ligase with one or more variations, wherein the one or more variations comprise a substitution from alanine to cysteine.
21. The D-form DNA ligase of any one of claims Error! Reference source not found. to 20, wherein the D-form DNA ligase has a sequence of a L-T4 ligase with one or more variations, wherein the one or more variations comprise a substitution from isoleucine to alanine, leucine, or valine.
22. The D-form DNA ligase of any one of claims Error! Reference source not found. to 21, wherein the D-form DNA ligase has a sequence of a L-T4 ligase with one or more variations, wherein the one or more variations comprise a substitution from valine to alanine.
23. The D-form DNA ligase of any one of claims Error! Reference source not found. to 22, wherein the D-form DNA ligase comprises: i) a sequence selected from SEQ ID NOs: 20-22; or ii) a sequence having at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to the sequence set forth in i).
24. The D-form DNA ligase of claim 23, wherein the D-form DNA ligase comprises: a. a sequence having at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to the sequence of SEQ ID NO: 21; and b. a substitution of V317A.
25. The composition of any one of claims 14 to 17, further comprising the D-form DNA ligase of any one of claims Error! Reference source not found. to 24.
26. A D-polypeptide hydrazide having a structure selected from a group consisting of: a. met1-tyr61-NHNH2 ; b. cys62-asn104-NHNH2 (cys(Acm)65); c. cys105-asn156-NHNH2 ; d. thz157-thr187-NHNH2; e. cys188-lys227-NHNH2 (cys(Acm)205);Attorney Ref.: 39339-64210 Client Ref.: 003WO f. cys228-leu283-NHNH2 ; g. thz284-thr348-NHNH2(cys(Acm)334); h. cys349-asn387-NHNH2 (cys(Acm)379); i. cys388-ala448-NHNH2; j. thz449-ser497-NHNH2 ; k. cys498-leu540-NHNH2; l. thz541-ala601-NHNH2 ; m. cys602-tyr659-NHNH2(cys(Acm)608); n. thz660-tyr711-NHNH2 (cys(Acm)674); o. cys712-phe750-NHNH2 ; p. cys751-met798-NHNH2; q. thz799-arg843-NHNH2 ; r. cys844-asn893-NHNH2; s. cys894-ala953-NHNH2 ; t. thz954-lys1009-NHNH2; u. cys1010-pro1066-NHNH2 (cys(Acm)1025) ; v. cys1067-pro1134-NHNH2; w. thz1135-pro1178-NHNH2 ; x. cys1179-ala1230-NHNH2; y. thz1231-Gly1266-NHNH2 (cys(Acm)1248); and z. cys1267-asn1307-OH and having a sequence with at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to any one of SEQ ID NOs: 9-12 between two amino acid residues corresponding to the positions identified in the structure.
27. A D-form endonuclease having 90% or greater sequence identity to SEQ ID NOs: 9- 12, wherein the D-form endonuclease is prepared via native chemical ligation (NCL) of two or more precursor fragments between one or more positions defined by the following pairs of amino acid residues: i) between tyr61and cys62; ii) between asn104and cys105; iii) between asn156and cys157; iv) between thr187and cys188;Attorney Ref.: 39339-64210 Client Ref.: 003WO v) between lys227and cys228; vi) between leu283and cys284; vii) between thr348and cys349; viii) between asn387and cys388; ix) between ala448and cys449; x) between ser497and cys498; xi) between leu540and cys541; xii) between ala601and cys602; xiii) between tyr659and cys660; xiv) between tyr711and cys712; xv) between phe750and cys751; xvi) between met798and cys799; xvii) between arg843and cys844; xviii) between asn893and cys894; xix) between ala953and cys954; xx) between lys1009and cys1010; xxi) between pro1066and cys1067; xxii) between pro1134and cys1135; xxiii) between pro1178and cys1179; xxiv) between ala1230and cys1231; and xxv) between Gly1266and cys1267.
28. The D-form endonuclease of claim 27, wherein the D-form endonuclease is prepared via desulfurization of one or more cysteine residues of a precursor fragment to alanine residues at one or more positions selected from 62, 105, 157, 188, 228, 284, 349, 388, 449, 498, 541, 602, 660, 712, 751, 799, 844, 894, 954, 1010, 1067, 1135, 1179, 1231, and 1267.
29. The D-form endonuclease of claim 27 or 28, wherein the D-form endonuclease is prepared via NCL coupling of two precursor fragments between positions tyr659and cys660.
30. The D-form endonuclease of claim 29, wherein the D-form endonuclease is prepared via NCL coupling of: i) met1-tyr659-NHNH2 (cys(Acm)65, 205, 334, 379, 608); andAttorney Ref.: 39339-64210 Client Ref.: 003WO ii) cys660-asn1307-OH (cys(Acm)674, 1025, 1248).
31. The D-form endonuclease of any one of claims 27 to 30, having at least 95% (e.g., 95% or greater, 96% or greater, 97% or greater, 98% or greater, or 99% or greater) sequence identity to the D-Cas12 protein of SEQ ID NOs: 9-12.
32. A D-polypeptide hydrazide having a structure selected from a group consisting of: a. met1-gln37-NHNH2; b. cys38-his82-NHNH2 ; c. thz83-arg110-NHNH2 ; d. cys111-Gly161-NHNH2(Cys(Acm)125) ; e. thz162-trp225-NHNH2 ; f. cys226-phe263-NHNH2; g. thz264-phe319-NHNH2 (Cys(Acm)295) ; h. cys320-trp363-NHNH2; i. thz364-gln427-NHNH2 ; j. thz428-ala475-NHNH2(cys(Acm)435, 440, 462, 465) ; and k. cys476-lys500-OH and having a sequence with at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to any one of SEQ ID NOs: 15-19 between two amino acid residues corresponding to the positions identified in the structure.
33. A D-form endonuclease having at least 90% sequence identity to a sequence selected from SEQ ID NOs: 15-19, wherein the D-form endonuclease is prepared via native chemical ligation (NCL) of two or more precursor fragments at one or more positions defined by the following pairs of amino acid residues: i) between gln37and cys38; ii) between his82and cys83; iii) between arg110and cys111; iv) between Gly161and cys162; v) between trp225and cys226; vi) between phe263and cys264; vii) between phe319and cys320;Attorney Ref.: 39339-64210 Client Ref.: 003WO viii) between trp363and cys364; ix) between gln427and cys428; and x) between ala475and cys476.
34. The D-form endonuclease of claim 33, wherein the D-form endonuclease is prepared via desulfurization of one or more cysteine residues of a precursor fragment to alanine residues at one or more positions selected from 38, 83, 111, 162, 226, 264, 320, 364, 428, and 476.
35. The D-form endonuclease of claim 33 or 34, wherein the D-form endonuclease is prepared via NCL coupling of two precursor fragments between positions phe263and cys264.
36. The D-form endonuclease of claim 35, wherein the D-form endonuclease is prepared via NCL coupling of: i) met1-phe263-NHNH2(cys(Acm)125) ; and ii) cys264-lys500-OH (cys(Acm)295, 435, 440, 462, 465).
37. The D-form endonuclease of any one of claims 33 to 36, having at least 95% (e.g., 95% or greater, 96% or greater, 97% or greater, 98% or greater, or 99% or greater) sequence identity to the D-Cas14 protein of SEQ ID NOs: 15-19.
38. A D-polypeptide hydrazide having a structure selected from a group consisting of: a. met1-ala76-NHNH2; b. cys77-leu124-NHNH2; c. thz125-Gly161-NHNH2; d. thz162-trp225-NHNH2; e. cys226-ala294-NHNH2; f. thz295-trp363-NHNH2; g. cys364-leu434-NHNH2; and h. cys435-lys500-OH and having a sequence with at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to any one of SEQ ID NOs: 15-19 between two amino acid residues corresponding to the positions identified in the structure.Attorney Ref.: 39339-64210 Client Ref.: 003WO 39. A D-form endonuclease having at least 90% sequence identity to a sequence selected from SEQ ID NOs: 15-19, wherein the D-form endonuclease is prepared via native chemical ligation (NCL) of two or more precursor fragments at one or more positions defined by the following pairs of amino acid residues: i) between ala76and cys77; ii) between leu124and cys125; iii) between Gly161and cys162; iv) between trp225and cys226; v) between ala294and cys295; vi) between trp363and cys364; and vii) between leu434and cys435.
40. The D-form endonuclease of claim 39, wherein the D-form endonuclease is prepared via desulfurization of one or more cysteine residues of a precursor fragment to alanine residues at one or more positions selected from 77, 125, 162, 226, 295, 364, and 435.
41. The D-form endonuclease of claim 39 or 40, wherein the D-form endonuclease is prepared via NCL coupling of two precursor fragments between positions ala294and cys295.
42. The D-form endonuclease of claim 41, wherein the D-form endonuclease is prepared via NCL coupling of: i) met1-ala294-NHNH2; and ii) cys295-lys500-OH.
43. The D-form endonuclease of any one of claims 39 to 42, having at least 95% (e.g., 95% or greater, 96% or greater, 97% or greater, 98% or greater, or 99% or greater) sequence identity to the D-Cas14 protein of SEQ ID NOs: 15-19.
44. A D-form T4 ligase having a sequence of a L-T4 ligase with one or more variations, wherein the one or more variations comprise a substitution from valine to alanine.
45. The D-form T4 ligase of claim 44, wherein the one or more variations comprise: i) a substitution from alanine to cysteine; orAttorney Ref.: 39339-64210 Client Ref.: 003WO ii) a substitution from isoleucine to alanine, leucine, or valine.
46. The D-form T4 ligase of claim 44 or 45, comprising: i) a sequence selected from SEQ ID Nos: 20-22; or ii) a sequence having at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to the sequence set forth in i).
47. A D-polypeptide hydrazide having a structure selected from a group consisting of: a. met1- leu76-NHNH2 ; b. cys77-Gly116-NHNH2 (Cys(Acm)115); c. cys117-arg164-NHNH2; d. thz165-thr201-NHNH2 ; e. cys202-thr255-NHNH2; f. thz256-lys316-NHNH2 (Cys(Acm)276); g. cys317-lys388-NHNH2; h. thz389-lys412-NHNH2 (Cys(Acm)398); and i. cys413-leu487-OH (Cys(Acm)439, 441) and having a sequence with at least 90% (e.g., 95% or greater, 98% or greater, or 99% or greater) sequence identity to any one of SEQ ID NOs: 20-22 between two amino acid residues corresponding to the positions identified in the structure.
48. A D-form T4 ligase having at least 90% sequence identity to a sequence selected from SEQ ID NOs: 20-22, wherein the D-form T4 ligase is prepared via native chemical ligation (NCL) of two or more precursor fragments at one or more positions defined by the following pairs of amino acid residues: i) between leu76and cys77; ii) between Gly116and cys117; iii) between arg164and cys165; iv) between thr201and cys202; v) between thr255and cys256; vi) between lys316and cys317; vii) between lys388and cys389; and viii) between lys412and cys413.Attorney Ref.: 39339-64210 Client Ref.: 003WO 49. The D-form T4 ligase of claim 48, wherein the D-form T4 ligase is prepared via desulfurization of one or more cysteine residues of a precursor fragment to alanine residues at one or more positions selected from 77, 117, 165, 202, 256, 317, 389, and 413.
50. The D-form T4 ligase of claim 48 or 49, wherein the D-form T4 ligase is prepared via NCL coupling of two precursor fragments between positions arg164and cys165.
51. The D-form T4 ligase of claim 50, wherein the D-form T4 ligase is prepared via NCL coupling of: i) met1-arg164-NHNH2(Cys(Acm)115); and ii) cys165-leu487-OH.
52. The D-form T4 ligase of any one of claims 48 to 51, having at least 95% (e.g., 95% or greater, 96% or greater, 97% or greater, 98% or greater, or 99% or greater) sequence identity to the D-form T4 ligase of SEQ ID NOs: 20-22.
53. A method of making a D-form Cas12 having 90% or greater sequence identity to SEQ ID NOs: 10-12, comprising coupling of two precursor fragments between positions tyr659and cys660via native chemical ligation (NCL).
54. The method of claim 53, wherein the two precursor fragments are: met1-tyr659-NHNH2(cys(Acm)65, 205, 334, 379, 608); and cys660-asn1307-OH (cys(Acm)674, 1025, 1248).
55. The method of claim 53 or 54, wherein the D-form Cas12 is prepared via native chemical ligation (NCL) couplings of precursor fragments between positions defined by the following pairs of amino acid residues: i) between tyr61and cys62; ii) between asn104and cys105; iii) between asn156and cys157; iv) between thr187and cys188; v) between lys227and cys228;Attorney Ref.: 39339-64210 Client Ref.: 003WO vi) between leu283and cys284; vii) between thr348and cys349; viii) between asn387and cys388; ix) between ala448and cys449; x) between ser497and cys498; xi) between leu540and cys541; xii) between ala601and cys602; xiii) between tyr659and cys660; xiv) between tyr711and cys712; xv) between phe750and cys751; xvi) between met798and cys799; xvii) between arg843and cys844; xviii) between asn893and cys894; xix) between ala953and cys954; xx) between lys1009and cys1010; xxi) between pro1066and cys1067; xxii) between pro1134and cys1135; xxiii) between pro1178and cys1179; xxiv) between ala1230and cys1231; and xxv) between Gly1266and cys1267.
56. The method of claim 54 or 55, wherein the precursor fragments are derived from one or more of the following fragments: a. met1-tyr61-NHNH2 ; b. cys62-asn104-NHNH2(cys(Acm)65); c. cys105-asn156-NHNH2; d. thz157-thr187-NHNH2; e. cys188-lys227-NHNH2 (cys(Acm)205); f. cys228-leu283-NHNH2; g. thz284-thr348-NHNH2 (cys(Acm)334); h. cys349-asn387-NHNH2 (cys(Acm)379); i. cys388-ala448-NHNH2; j. thz449-ser497-NHNH2 ;Attorney Ref.: 39339-64210 Client Ref.: 003WO k. cys498-leu540-NHNH2; l. thz541-ala601-NHNH2; m. cys602-tyr659-NHNH2 (cys(Acm)608); n. thz660-tyr711-NHNH2(cys(Acm)674); o. cys712-phe750-NHNH2; p. cys751-met798-NHNH2; q. thz799-arg843-NHNH2; r. cys844-asn893-NHNH2; s. cys894-ala953-NHNH2; t. thz954-lys1009-NHNH2; u. cys1010-pro1066-NHNH2(cys(Acm)1025); v. cys1067-pro1134-NHNH2; w. thz1135-pro1178-NHNH2; x. cys1179-ala1230-NHNH2; y. thz1231-Gly1266-NHNH2(cys(Acm)1248); and z. cys1267-asn1307-OH.
57. The method of any one of claims 53 to 56, further comprising desulfurizing one or more cysteine residues of a precursor fragment to alanine residues at one or more positions selected from 62, 105, 157, 188, 228, 284, 349, 388, 449, 498, 541, 602, 660, 712, 751, 799, 844, 894, 954, 1010, 1067, 1135, 1179, 1231, and 1267.
58. A method of making a D-form Cas14 having 90% or greater sequence identity to SEQ ID NOs: 16-19, comprising coupling of two precursor fragments between positions phe263and cys264via native chemical ligation (NCL).
59. The method of claim 58, wherein the two precursor fragments are: met1-phe263-NHNH2 (cys(Acm)125); and cys264-lys500-OH (cys(Acm)295, 435, 440, 462, 465).
60. The method of claim 58 or 59, wherein the D-form Cas14 is prepared via native chemical ligation (NCL) couplings of precursor fragments between positions defined by the following pairs of amino acid residues:Attorney Ref.: 39339-64210 Client Ref.: 003WO i) between gln37and cys38; ii) between his82and cys83;between arg110and cys111;between Gly161and cys162;between trp225and cys226; vi) between phe263and cys264; vii) between phe319and cys320; viii) between trp363and cys364; ix) between gln427and cys428; and x) between ala475and cys476.
61. The method of claim 59 or 60, wherein the precursor fragments are derived from one or more of the following fragments: a. met1-gln37-NHNH2; b. cys38-his82-NHNH2; c. thz83-arg110-NHNH2; d. cys111-Gly161-NHNH2(Cys(Acm)125); e. thz162-trp225-NHNH2; f. cys226-phe263-NHNH2; g. thz264-phe319-NHNH2 (Cys(Acm)295); h. cys320-trp363-NHNH2; i. thz364-gln427-NHNH2; j. thz428-ala475-NHNH2(cys(Acm)435, 440, 462, 465); and k. cys476-lys500-OH.
62. The method of any one of claims 58 to 61, further comprising desulfurizing one or more cysteine residues of a precursor fragment to alanine residues at one or more positions selected from 38, 83, 111, 162, 226, 264, 320, 364, 428, and 476.
63. A method of making a D-form Cas14 having 90% or greater sequence identity to SEQ ID NOs: 16-19, comprising coupling of two precursor fragments between positions ala294and cys295via native chemical ligation (NCL).Attorney Ref.: 39339-64210 Client Ref.: 003WO 64. The method of claim 63, wherein the two precursor fragments are: met1-ala294-NHNH2; and cys295-lys500-OH.
65. The method of claim 63 or 64, wherein the D-form Cas14 is prepared via native chemical ligation (NCL) couplings of precursor fragments between positions defined by the following pairs of amino acid residues: i) between ala76and cys77; ii) between leu124and cys125; iii) between Gly161and cys162; iv) between trp225and cys226; v) between ala294and cys295; vi) between trp363and cys364; and vii) between leu434and cys435.
66. The method of claim 64 or 65, wherein the precursor fragments are derived from one or more of the following fragments: a. met1- ala76-NHNH2 ; b. cys77-leu124-NHNH2; c. thz125-Gly161-NHNH2 ; d. thz162-trp225-NHNH2; e. cys226-ala294-NHNH2 ; f. thz295-trp363-NHNH2; g. cys364-leu434-NHNH2 ; and h. cys435-lys500-OH.
67. The method of any one of claims 63 to 66, further comprising desulfurizing one or more cysteine residues of a precursor fragment to alanine residues at one or more positions selected from 77, 125, 162, 226, 295, 364, and 435.
68. A method of making a D-form T4 ligase having 90% or greater sequence identity to SEQ ID NO 21 or SEQ ID NO 22, comprising coupling of two precursor fragments between positions arg164and cys165via native chemical ligation (NCL).Attorney Ref.: 39339-64210 Client Ref.: 003WO 69. The method of claim 68, wherein the two precursor fragments are: met1-arg164-NHNH2 (Cys(Acm)115); and cys165-leu487-OH.
70. The method of claim 68 or 69, wherein the D-form T4 ligase is prepared via native chemical ligation (NCL) couplings of precursor fragments between positions defined by the following pairs of amino acid residues: i) between leu76and cys77; ii) between Gly116and cys117; iii) between arg164and cys165; iv) between thr201and cys202; v) between thr255and cys256; vi) between lys316and cys317; vii) between lys388and cys389; and viii) between lys412and cys413.
71. The method of claim 69 or 70, wherein the precursor fragments are derived from one or more of the following fragments: a. met1- leu76-NHNH2 ; b. cys77-Gly116-NHNH2(Cys(Acm)115); c. cys117-arg164-NHNH2 ; d. thz165-thr201-NHNH2; e. cys202-thr255-NHNH2 ; f. thz256-lys316-NHNH2(Cys(Acm)276); g. cys317-lys388-NHNH2 ; h. thz389-lys412-NHNH2(Cys(Acm)398); and i. cys413-leu487-OH (Cys(Acm)439, 441).
72. The method of any one of claims 68 to 71, further comprising desulfurizing one or more cysteine residues of a precursor fragment to alanine residues at one or more positions selected from 77, 117, 165, 202, 256, 317, 389, and 413.Attorney Ref.: 39339-64210 Client Ref.: 003WO 73. A method of cleaving and / or modifying a target L-DNA, comprising: contacting the composition of claim 14 with a target L-DNA; wherein a guide RNA of the composition hybridizes to the target L-DNA sequence, thereby directing a D-form DNA endonuclease to bind to said target L-DNA sequence and cleave and / or modify the target L-DNA sequence.
74. A method of modifying a target L-DNA, comprising: contacting the composition of claim 25 with a target L-DNA; wherein a guide RNA of the composition hybridizes to the target L-DNA sequence, thereby directing a D-form DNA endonuclease to bind to said target L-DNA sequence and cleave and / or modify the target L-DNA sequence, and wherein a D-form DNA ligase ligates the cleaved and / or modified target L-DNA sequence.
75. A method of editing a target L-DNA, comprising: interacting the composition of claim 14 and a target L-DNA in a mixture; and incubating the mixture thereby inducing cleavage and / or modification of the target L- DNA.
76. A method of editing a target L-DNA, comprising: interacting the composition of claim 25 and a target L-DNA in a mixture; and incubating the mixture thereby inducing modification and ligation of the target L- DNA.